Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A collection of open-source dataset to train instruction-following LLMs (ChatGPT,LLaMA,Alpaca)
| Date | Stars |
|---|---|
| 2026-07-31 | 1154 |
| 2026-08-05 | 1154 |
| 2026-08-06 | 1154 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# awesome-text/visual-instruction-tuning-dataset A collection of open-source instruction tuning datasets to train (text and multi-modal) chat-based LLMs (GPT-4, ChatGPT,LLaMA,Alpaca). We currently include three types of dataset: 1. visual-instruction-tuning (e.g. image-instruction-answer) 2. text-instruction-tuning datasets. 3. red-teaming | Reinforcement Learning from Human Feedback (RLHF) Datasets Instruction Tuning / Reinforcement Learning from Human Feedback (RLHF) Dataset is a key component of instruction-following LLMs such as ChatGPT. This repo is dedicated to providing a comprehensive list of datasets used for instruction tuning in various LLMs, making it easier for researchers and developers to access and utilize these resources. Lists of codebse to train your LLMs: - [nichtdax/awesome-totally-open-chatgpt](https://github.com/nichtdax/awesome-totally-open-chatgpt): A codebase of totally open alternatives to ChatGPT Size: The number of instruction tuning pairs Lingual-Tags: - EN: Instruction datasets in English - CN: Instruction datasets in Chinese - ML: [Multi-lingual] Instruction datasets in multiple languages Task-Tags: - MT: [Multi-task] Datasets containing multiple tasks - TS: [Task-specific] Datasets tailored for specific tasks Generation-method: - HG: [Human Generated Dataset] Datasets created by humans - SI: [Self-Instruct] Datasets generated using self-instruct methods - MIX: [Mixed Dataset] Dataset contains both human and machine generated data - COL: [Collection of Dataset] Dataset made from a collection of other datasets # Table of Contents 1. [The template](#the-template) 2. [The Multi-modal Instruction Dataset](#the-multi-modal-instruction-datasets) - [(Vision-CAIR/MiniGPT-4)|5K|EN|MT|MIX](https://minigpt-4.github.io/) - [(haotian-liu/LLaVA)|150K|EN|MT|MIX](https://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K) 3. [The Instruction tuning Dataset](#the-instruction-following-datasets) - [(tatsu-lab/Alpaca)|52K|EN|MT|SI](https://github.com/tatsu-lab/stanford_alpaca) - [(gururise/Cleaned Alpaca)|52K|EN|MT|SI](https://github.com/gururise/AlpacaDataCleaned) - [(XueFuzhao/InstructionWild)|52K|EN|CN|MT|SI](https://github.com/XueFuzhao/InstructionWild) - [(JosephusCheung/GuanacoDataset)|534K|ML|MT|SI](https://huggingface.co/datasets/JosephusCheung/GuanacoDataset) - [(Hello-SimpleAI/HC3)|24K|EN|MT|MIX](https://huggingface.co/datasets/Hello-SimpleAI/HC3) - [(Hello-SimpleAI/HC3-Chinese)|13K|CN|MT|MIX](https://huggingface.co/datasets/Hello-SimpleAI/HC3) - [(allenai/prosocial-dialog)|58K|EN|MT|MIX](https://huggingface.co/datasets/allenai/prosocial-dialog) - [(allenai/natural-instructions)|1.6K|ML|MT|HG](https://github.com/allenai/natural-instructions) - [(bigscience/xP3)|N/A|ML|MT|MIX](https://huggingface.co/datasets/bigscience/xP3) - [(nomic-ai/gpt4all)|437k|EN|MT|COL](https://github.com/nomic-ai/gpt4all) - [(PhoebusSi/Alpaca-CoT)|500k|ML|MT|COL](https://huggingface.co/datasets/QingyiSi/Alpaca-CoT) - [(google-research/FLAN)|N/A|EN|MT|MIX](https://github.com/google-research/FLAN/tree/main/flan/v2) - [(thunlp/UltraChat)|280k|EN|TS|MIX](https://github.com/thunlp/UltraChat) - [(cascip/ChatAlpaca)|10k|EN|MT|MIX](https://github.com/cascip/ChatAlpaca) - [(YeungNLP/firefly-train-1.1M)|1100k|CN|MT|COL](https://huggingface.co/datasets/YeungNLP/firefly-train-1.1M) - [(orhonovich/unnatural-instructions)|240K|EN|MT|MIX](https://github.com/orhonovich/unnatural-instructions) - [(Instruction-Tuning-with-GPT-4/GPT-4-LLM)|52K|EN|CN|MT|SI](https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM) - [(databrickslabs/dolly)|15K|EN|MT|HG](https://github.com/databrickslabs/dolly/tree/master/data) - [(OpenAssistant/oasst1)|161K|ML|MT|HG](https://huggingface.co/datasets/OpenAssistant/oasst1) - [(RyokoAI/ShareGPT52K)|90K|ML|MT|SI](https://huggingface.co/datasets/RyokoAI/ShareGPT52K) - [(zjunlp/Mol-Instructions)|2043K|ML|MT|MIX](https://huggingf
Excerpt of 22,174 characters
Read on GitHub33
3
Zhenran Xu · Alibaba Group
3
2
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:26b9e8734bd25660, topic:llama, topic:language-model
matched fp:26b9e8734bd25660, topic:instruction-tuning
matched fp:26b9e8734bd25660, topic:datasets, name:dataset, desc:dataset