Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
| Date | Stars |
|---|---|
| 2026-07-31 | 2775 |
| 2026-08-03 | 2774 |
| 2026-08-06 | 2774 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Chain-of-Thought Hub: Measuring LLMs' Reasoning Performance  "A fantasy graph illustrating a chain of stars in a dark night with blue sky, digital art, super resolution". Midjourney V5 ---- By [Yao Fu](https://franxyao.github.io/), [Litu Ou](https://github.com/Leonard907), [Mingyu Chen](https://github.com/Spehhhhh), [Yuhao Wan](https://github.com/Yuhao-Wan), [Hao Peng](https://haopeng-nlp.github.io/), [Tushar Khot](https://allenai.org/team/tushark), [Wenhu Chen](https://wenhuchen.github.io/) From University of Edinburgh, University of Washington, Allen Institute for AI, University of Waterloo [[paper](https://arxiv.org/abs/2305.17306)] [[blog](https://yaofu.notion.site/Towards-Complex-Reasoning-the-Polaris-of-Large-Language-Models-c2b4a51355b44764975f88e6a42d4e75)] [[twitter](https://twitter.com/Francis_YAO_/status/1663472109299937280)] Recently, there are a lot of progress in LLMs. Many claim that a small model less than 10B can achieve comparable performance to GPT-3.5. Really? > In a casual conversation, the distinction between GPT-3.5 and GPT-4 can be subtle. The difference comes out when **\*the complexity of the task reaches a sufficient threshold\*** — GPT-4 is more reliable, creative, and able to handle much more nuanced instructions than GPT-3.5. -- *GPT-4 release blog* The key differentiator is whether a model can do **complex tasks**, like the old saying: "chit-chat is cheap, show me the reasoning." This is why we compile a list of complex reasoning tasks including math (GSM8K), science (MATH, TheoremQA), symbolic (BBH), knowledge (MMLU, C-Eval), coding (HumanEval), factual (SummEdits), and long-context (RepoBench, Qspr, QALT, BkSS) to measure the models' performance on challenging tasks. More importantly, we envisage large language models to become the next-generation computational platform and foster an ecosystem of LLM-based new applications. When this comes, chain-of-thought prompt engineering will be the next-generation system calls and shell scripts. The credibility of chain-of-thought hub comes from the very carefully mediculously picked datasets and models that can clearly help the development of LLMs. The resutls and scripts from Chain-of-thought Hub is being used and referred by leading industrial and academic organizations in the space of large language models. We devide the tasks into three categories: main, experimental, and long-context. * Main: datasets that are stable and consistently referred by places where LLMs are built. * Experimental: datasets that has the potential to test future LLM capabilities. * Long-context: datasets that require reasoning over very long context, an important direction of future LLMs. <details> <summary>[List of datasets we consider]</summary> | Section | Dataset | Description | | ------- | ------- | ----------- | | Main | GSM8K | Grade-level math word problems | | Main | MATH | Competition-level math and science problems | | Main | MMLU | Multi-discipline knowledge | | Main | BBH | Challenging language and symbolic reasoning | | Main | HumanEval | Python coding | | Main | C-Eval | Chineses multi-discipline knowledge | | Experimental | TheoremQA | Theorem proving | | Experimental | SummEdits | Factual reasoning | | Long Ctx | Qspr | Question answering over research papers | | Long Ctx | QALT | Multiple-choice questions over long articles and stories | | Long Ctx | BkSS | Reordering of summaries of parts of novels | </details> **[Call for contribution]**: would love to invite community members to: * Send a PR to fill in a missing number in the table * Raise an issue to suggest / brainstorm a new task / benchmark that measures **reasoning over very long context** * Raise an issue to suggest / brainstorm a new task / benchmark that measures **complex API calls and tool usage** * Raise an issue to suggest other good ta
Excerpt of 22,841 characters
Read on GitHub29
Diane Wan
28
ipruning · China
19
morecry
14
Litu Ou · [email protected] · United Kingdom
11
2
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:f99dd3feed9109e1, llm:description: 'Benchmarking large language models\' complex reasoning ability with chain-of-thought prompting'
matched fp:f99dd3feed9109e1, llm:description: 'Benchmarking large language models\' complex reasoning ability with chain-of-thought prompting'