Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A prize for finding tasks that cause large language models to show inverse scaling
| Date | Stars |
|---|---|
| 2026-07-31 | 622 |
| 2026-08-03 | 622 |
| 2026-08-06 | 622 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<p align="center"> <img src="docs/promo-image.png" alt="Two graphs, one with regular scaling marked 'Many tasks like this', and one with inverse scaling marked 'Any tasks like this?'" width=500px/> </p> # Inverse Scaling Prize **TL;DR: Win up to $100,000 for finding an important task where larger language models do worse.** _~~Submissions due August 27, 2022 (Round 1) and October 27, 2022 (Round 2).~~_ The contest has ended! Results: [Round 1](https://irmckenzie.co.uk/round1), [Round 2](https://irmckenzie.co.uk/round2). ## Recent changes ### 11 October, 2023 * Added Modus Tollens caveat to [data-release README](data-release/README.md) ### 21 March, 2023 * Updated prize pool info ### 1 March, 2023 * Released [data for all winning tasks](data-release/README.md) ### 17 December, 2022 * Updated prize eligibility for FAR employees ### 12 December, 2022 * Added prize terms update to the ‘Prize information’ section * Updated ‘About us’ ### 9 October, 2022 * Added Huggingface Hub evaluation setup description to the tips section ### 4 October, 2022 * BUG FIX: Reported total probabilities should now be more accurate for all classification tasks ### 26 September, 2022 * Demonstrating positive scaling on the ‘incorrect’ answer is now allowed * Added stronger recommendation to aim for roughly 1000 examples * Added requirement to name the task * Added request to include data for control experiments * Added field to specify how the data was generated * Added field for links to dataset sources * Added field for code that generated dataset * Added requirement that submissions in multiple parts should upload all .csv files together in one .zip * Added request to make file names anonymous * Added option to use a variable number of classes in classification datasets * Added print out to colabs of the total probability given to class labels * Added reminder that the submitted plot should be from our official colab * Added request that people edit their form submission rather than resubmit to update * Added reminder to specify correct behavior on the task * Added field to specify whether the task is zero-shot or few-shot * Updated terms and conditions ## Motivation As language models get larger, they seem to only get better. Larger language models score better on benchmarks and unlock new capabilities like arithmetic [\[1\]](#ref1), few-shot learning [\[1\]](#ref1), and multi-step reasoning [\[2\]](#ref2). However, language models are not without flaws, exhibiting many biases [\[3\]](#ref3) and producing plausible misinformation [\[4\]](#ref4). The purpose of this contest is to find evidence for a stronger failure mode: tasks where language models get **worse** as they become better at language modeling (next word prediction). The standard paradigm in natural language processing today is to pretrain large language models to autocomplete text corpora. The resulting models are then either frozen and used directly for other tasks (zero-shot or using few-shot learning), or additionally trained on other tasks (fine-tuning). We focus on the case of zero-shot/few-shot evaluation on downstream tasks without task-specific gradient optimization: it's typically easier to use in practice and to study. Scaling laws [\[5\]](#ref5)[\[6\]](#ref6) show that language models get predictably better (in terms of test loss and downstream performance [\[7\]](#ref7)) as the number of parameters, amount of compute used, and dataset size increase. The improvement follows a power law in each of parameters, compute, and dataset size. We hypothesize that there are tasks with trends in the opposite direction: task performance gets monotonically, predictably worse as the overall test loss of the language model improves. We call this phenomenon *inverse scaling*, in contrast with the standard scaling laws. There are some tasks that appear to show inverse scaling under some conditions [\[4\]](#ref4)[\[8\]](#ref8)[\[10\]](#ref10), but such tasks appear to be rare. This
Excerpt of 48,834 characters
Read on GitHub35
8
Alexander Lyzhov
6
Erjan K · Netherlands
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:4c5c2652c0976a61, llm:description: 'A prize for finding tasks that cause large language models to show inverse scaling'
matched fp:4c5c2652c0976a61, llm:description: 'A prize for finding tasks that cause large language models to show inverse scaling'