Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A curated list of awesome instruction tuning datasets, models, papers and repositories.
| Date | Stars |
|---|---|
| 2026-07-31 | 346 |
| 2026-08-06 | 345 |
Today
-1 stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome-instruction-tuning A curated list of open-source instruction tuning datasets, models, papers, repositories. ## Datasets and Models ### Modified from Traditional NLP Following [Longpre et al.](https://arxiv.org/pdf/2301.13688.pdf), we list all existing instruction tuning datasets modified from traditional NLP tasks. |Release| Datasets| Number of Tasks | Number of Instances | Model_name | Base | Model_Size| | ---- | ---- | ---- | ---- | ---- | ---- | ---- | | 2020-05 | [UnifiedQA](https://github.com/allenai/unifiedqa) |46 | 750k | UnifiedQA | RoBerta | 110-340 M | | 2021-04 | [CrossFit](https://github.com/INK-USC/CrossFit) |159 | 71.M | BART-CrossFit | BART | 140 M | | 2021-04 | [Natural Inst v1.0](https://instructions.apps.allenai.org/) |61 | 620 k | Gen. BART | BART |140 M | | 2021-09 | [Flan 2021](https://github.com/google-research/FLAN/tree/main#flan-2021) |62 | 4.4M | Flan-LaMDA | LaMDA | 137B | | 2021-10 | [P3](https://github.com/bigscience-workshop/promptsource) | 62 | 12M |TO, TO+, TO++ | T5-LM| 3-11B | | 2021-10 | [MetalCL](https://github.com/facebookresearch/MetaICL) |142 | 3.5M |MetalCL | GPT-2 | 770 M | | 2021-11 | [ExMix](https://github.com/google-research/text-to-text-transfer-transformer) | 107 | 500 k | ExT5 | T5 | 220M-11B | | 2022-04 | [Super-Natural Inst.](https://github.com/allenai/natural-instructions) |1613 | 5M | Tk-Instruct | T5-LM, mT5 | 17-13B | | 2022-10 | [GLM](https://github.com/THUDM/GLM-130B) | 77 | 12M | GLM-130B | GLM | 130 B | | 2022-10 | [Flan 2022](https://github.com/google-research/FLAN/tree/main/flan/v2) |1836 | 15M | Flan-T5, Flan-PaLM | T5-LM, PaLM | 10 M-540 B | | 2022-11 | [xP3](https://huggingface.co/datasets/bigscience/xP3) | 71 | 81M | BLOOMz, mTO | BLOOM, mT5 | 13-176B | | 2022-12 | [Unnatural Inst.](https://github.com/orhonovich/unnatural-instructions) | 117 | 64 k | T5-LM-Unnat. Inst. | T5-LM | 11B | ### Generated by LLMs |Release| Model_name | Base | Model_Size| Datasets | Number of Instances | Language| | ---- | ---- | ---- | ---- | ---- | ---- | ---- | | 2022-12 | GPT-3 Self Inst. | GPT-3 | 175B | Self-Instruct | 82 k |En | | 2023-03-03|[alpaca](https://github.com/tatsu-lab/stanford_alpaca)| LLaMA | 7B |[alpaca_data](https://github.com/tatsu-lab/stanford_alpaca/blob/main/alpaca_data.json)| 52 k | En | | 2023-03-19|[alpaca-lora](https://github.com/tloen/alpaca-lora/commits/main) | LLaMA | 7B 13B 30B|[alpaca_data](https://github.com/tatsu-lab/stanford_alpaca/blob/main/alpaca_data.json)、[alpaca_data_cleaned](https://github.com/tloen/alpaca-lora/blob/main/alpaca_data_cleaned.json) |52 k | En| | 2023-03-23| [Chinese-Vicuna](https://github.com/Facico/Chinese-Vicuna) | LLaMA | 7B 13B | [BELLE](https://github.com/LianjiaTech/BELLE)、[GuanacoDataset](https://huggingface.co/datasets/JosephusCheung/GuanacoDataset) | 1M | Zh | | 2023-03-24| [Alpaca-CoT](https://github.com/PhoebusSi/Alpaca-CoT) | LLaMA | 7B | [dataset](https://github.com/PhoebusSi/Alpaca-CoT#statistics) | ---- | En Zh | | 2023-03-25|[dolly](https://github.com/databrickslabs/dolly)| dolly | 6B |[alpaca_data](https://github.com/tatsu-lab/stanford_alpaca/blob/main/alpaca_data.json)| 52 k|En | | 2023-03-25|[guanaco](https://huggingface.co/KBlueLeaf/guanaco-7B-leh)| LLaMA | 7B |[GuanacoDataset](https://huggingface.co/datasets/JosephusCheung/GuanacoDataset)| 534 k | En Zh Ja De| | 2023-03-28| [Chinese-LLaMA-Alpaca](https://github.com/ymcui/Chinese-LLaMA-Alpaca) | LLaMA | 7B | [alpaca_data_zh](https://github.com/ymcui/Chinese-LLaMA-Alpaca/tree/main/data)、[pCLUE](https://github.com/CLUEbenchmark/pCLUE)、[translation2019zh](https://github.com/brightmart/nlp_chinese_corpus#5%E7%BF%BB%E8%AF%91%E8%AF%AD%E6%96%99translation2019zh)、[alpaca_data](https://github.com/tatsu-lab/stanford_alpaca/blob/main/alpaca_data.json)、Self-Instruct | 2M | Zh | |2023-03-29|[ColossalChat](https://github.com/hpcaitech/ColossalAI)| LLaMA |7B 13B |[InstructionWild](https://github.com/XueFuzhao/Instr
Excerpt of 8,610 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:fa849ace5cafba9c, topic:awesome, topic:awesome-list, desc:curated list
matched fp:fa849ace5cafba9c, topic:instruction-tuning, name:instruction tuning, desc:instruction tuning
matched fp:fa849ace5cafba9c, topic:llm, topic:gpt