Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A collection of benchmarks and datasets for evaluating LLM.
| Date | Stars |
|---|---|
| 2026-07-31 | 575 |
| 2026-08-06 | 575 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# llm_benchmarks A collection of benchmarks and datasets for evaluating LLM. ## Knowledge and Language Understanding ### Massive Multitask Language Understanding (MMLU) * **Description:** Measures general knowledge across 57 different subjects, ranging from STEM to social sciences. * **Purpose:** To assess the LLM's understanding and reasoning in a wide range of subject areas. * **Relevance:** Ideal for multifaceted AI systems that require extensive world knowledge and problem solving ability. * **Source:** [Measuring Massive Multitask Language Understanding](https://arxiv.org/abs/2009.03300) * **Resources:** * [MMLU GitHub](https://github.com/hendrycks/test) * [MMLU Dataset](https://people.eecs.berkeley.edu/~hendrycks/data.tar) ### AI2 Reasoning Challenge (ARC) * **Description:** Tests LLMs on grade-school science questions, requiring both deep general knowledge and reasoning abilities. * **Purpose:** To evaluate the ability to answer complex science questions that require logical reasoning. * **Relevance:** Useful for educational AI applications, automated tutoring systems, and general knowledge assessments. * **Source:** [Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge](https://arxiv.org/abs/1803.05457) * **Resources:** * [ARC Dataset: HuggingFace](https://huggingface.co/datasets/ai2_arc) * [ARC Dataset: Allen Institute](https://allenai.org/data/arc) ### General Language Understanding Evaluation (GLUE) * **Description:** A collection of various language tasks from multiple datasets, designed to measure overall language understanding. * **Purpose:** To provide a comprehensive assessment of language understanding abilities in different contexts. * **Relevance:** Crucial for applications requiring advanced language processing, such as chatbots and content analysis. * **Source:** [GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding](https://arxiv.org/abs/1804.07461) * **Resources:** * [GLUE Homepage](https://gluebenchmark.com/) * [GLUE Dataset](https://huggingface.co/datasets/glue) ### Natural Questions * **Description:** A collection of real-world questions people have Googled, paired with relevant Wikipedia pages to extract answers. * **Purpose:** To test the ability to find accurate short and long answers from web-based sources. * **Relevance:** Essential for search engines, information retrieval systems, and AI-driven question-answering tools. * **Source:** [Natural Questions: A Benchmark for Question Answering Research](https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00276/43518/Natural-Questions-A-Benchmark-for-Question) * **Resources:** * [Natural Questions Homepage](https://ai.google.com/research/NaturalQuestions) * [Natural Questions Dataset: Github](https://github.com/google-research-datasets/natural-questions) ### LAnguage Modelling Broadened to Account for Discourse Aspects (LAMBADA) * **Description:** A collection of passages testing the ability of language models to understand and predict text based on long-range context. * **Purpose:** To assess the models' comprehension of narratives and their predictive abilities in text generation. * **Relevance:** Important for AI applications in narrative analysis, content creation, and long-form text understanding. * **Source:** [The LAMBADA Dataset: Word prediction requiring a broad discourse context](https://arxiv.org/abs/1606.06031) * **Resources:** * [LAMBADA Dataset: HuggingFace](https://huggingface.co/datasets/lambada) ### HellaSwag * **Description:** Tests natural language inference by requiring LLMs to complete passages in a way that requires understanding intricate details. * **Purpose:** To evaluate the model's ability to generate contextually appropriate text continuations. * **Relevance:** Useful in content creation, dialogue systems, and applications requiring advanced text generation capabilities. * **Source:** [HellaSwag: Can a Machine Really Finish Your Sen
Excerpt of 28,148 characters
Read on GitHub3
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:398b9d1f9a00acec, desc:datasets