Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
(CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.
| Date | Stars |
|---|---|
| 2026-07-31 | 366 |
| 2026-08-01 | 366 |
| 2026-08-06 | 366 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# SEED-Bench: Benchmarking Multimodal Large Language Models [SEED-Bench-H](https://github.com/AILab-CVC/SEED-Bench/blob/main/SEED-Bench-H/SEED-Bench-H.pdf) [SEED-Bench-2-Plus Arxiv](https://arxiv.org/abs/2404.16790) [SEED-Bench-2 Arxiv](https://arxiv.org/abs/2311.17092) [SEED-Bench-1 Arxiv](https://arxiv.org/abs/2307.16125) <img src="https://github.com/AILab-CVC/SEED-Bench/blob/main/figs/seed-bench-2.jpg" width = "600" alt="图片名称" align=center /> SEED-Bench-H is a comprehensive integration of previous SEED-Bench series (SEED-Bench, SEED-Bench-2, SEED-Bench-2-Plus), with additional evaluation dimension. It consists of 28K multiple-choice questions with precise human annotations, spanning 34 dimensions, including the evaluation of both text and image generation. SEED-Bench-2-Plus comprises 2.3K multiple-choice questions with precise human annotations, spanning three broad categories: Charts, Maps, and Webs, each of which covers a wide spectrum of textrich scenarios in the real world. SEED-Bench-2 comprises 24K multiple-choice questions with accurate human annotations, which spans 27 dimensions, including the evaluation of both text and image generation. SEED-Bench-1 consists of 19K multiple-choice questions with accurate human annotations, covering 12 evaluation dimensions including both the spatial and temporal understanding. ## News **[2025.1.14]** [SEED-Bench](https://arxiv.org/abs/2307.16125) has been included in the [OpenCompass-Dataset Community](https://hub.opencompass.org.cn/dataset-detail/SEED-Bench), thanks to [OpenCompass](https://opencompass.org.cn/home). **[2024.7.11]** [SEED-Bench-H](https://github.com/AILab-CVC/SEED-Bench/blob/main/SEED-Bench-H/SEED-Bench-H.pdf), [SEED-Bench-2-Plus](https://arxiv.org/abs/2404.16790), [SEED-Bench-2](https://arxiv.org/abs/2311.17092), and [SEED-Bench-1](https://arxiv.org/abs/2307.16125) data is released on [ModelScope](https://modelscope.cn/organization/TencentARC?tab=dataset), thanks to [ModelScope Community](https://modelscope.cn). **[2024.6.18]** [SEED-Bench-2](https://arxiv.org/abs/2311.17092) can be evaluated on [VLMEvalKit](https://github.com/open-compass/VLMEvalKit), thanks to [kennymckormick](https://github.com/kennymckormick). **[2024.5.30]** We released [SEED-Bench-H](https://github.com/AILab-CVC/SEED-Bench/blob/main/SEED-Bench-H/SEED-Bench-H.pdf), which is a comprehensive integration of previous SEED-Bench series ([SEED-Bench](https://arxiv.org/abs/2311.17092), [SEED-Bench-2](https://arxiv.org/abs/2311.17092), [SEED-Bench-2-Plus](https://arxiv.org/abs/2404.16790)), with additional evaluation dimension. The additional evaluation dimension including Image to Latex, Visual Story Comprehension, Few-shot Segmentation, Few-shot Keypoint, Few-shot Depth, and Few-shot Object Detection. Please refer [SEED-Bench-H](https://github.com/AILab-CVC/SEED-Bench/blob/main/SEED-Bench-H/SEED-Bench-H.pdf) for detailed. Corresponding dataset is released on [SEED-Bench-H](https://huggingface.co/datasets/AILab-CVC/SEED-Bench-H). **[2024.5.25]** [SEED-Bench-2-Plus](https://arxiv.org/abs/2404.16790) can be evaluated on [VLMEvalKit](https://github.com/open-compass/VLMEvalKit), thanks to [kennymckormick](https://github.com/kennymckormick). **[2024.4.26]** We are excited to announce the release of [SEED-Bench-2-Plus](https://arxiv.org/abs/2404.16790), a benchmark specifically designed for text-rich visual comprehension. The accompanying dataset is released on [SEED-Bench-2-Plus](https://huggingface.co/datasets/AILab-CVC/SEED-Bench-2-plus). **[2024.4.23]** We are pleased to share the comprehensive evaluation results for [Gemini-Vision-Pro](https://gemini.google.com/) and [Claude-3-Opus](https://www.anthropic.com/news/claude-3-family) on [SEED-Bench-1](https://arxiv.org/abs/2307.16125) and [SEED-Bench-2](https://arxiv.org/abs/2311.17092). You can access detailed performance on the [SEED-Bench Leaderboard](https://huggingface.co/spaces/AILab-CVC/SEE
Excerpt of 13,516 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:ff5c0a3a2f7f5d31, desc:multimodal