Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A blueprint for creating Pretraining and Fine-Tuning datasets for Indic languages
| Date | Stars |
|---|---|
| 2026-07-31 | 412 |
| 2026-08-02 | 412 |
| 2026-08-06 | 412 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages
🏆 ACL 2024 Outstanding Paper Award 🏆
This repository contains the artifacts and resources for curating Pre-training and Fine-tuning datasets for Indic Languages.
<p align="left">
<a href="https://github.com/AI4Bharat/IndicLLMSuite/blob/master/LICENSE">
<img src="https://img.shields.io/badge/License-MIT-green">
</a>
<a href="https://aclanthology.org/2024.acl-long.843/">
<img src="https://img.shields.io/badge/ACL%20-2024-blue">
</a>
</p>

[📜 Paper](https://arxiv.org/abs/2403.06350) | [🌐 Blog](https://ai4bharat.iitm.ac.in) | [🤗 Data](https://huggingface.co/collections/ai4bharat/indicllmsuite-65ee7d225c337fcfa0991707)
IndicLLMSuite is the largest Pre-training and Instruction Fine-tuning dataset collection across 22 Indic languages. We open-source our pre-training dataset "Sangraha," the Instruction Fine-tuning dataset "IndicAlign-Instruct," and the Toxic alignment dataset "IndicAlign-Toxic." We also open-source all the code and other resources used for curating these datasets, including "Setu," a comprehensive data cleaning, filtering, and deduplication pipeline for Indic languages. We hope that this will advance the development of LLMs for Indian Languages.
We release the below Artifacts:
- [Sangraha](#sangraha)
- [IndicAlign](#indicalign)
- [IndicAlign-Instruct](#indicalign-instruct)
- [IndicAlign-Toxic](#indicalign-toxic)
- [Data Pipelines](#data-pipelines)
- [Setu](#setu)
- [Setu-translate](#setu-translate)
- [Setu-transliterate](#setu-transliterate)
- [Other Resources](#other-resources)
- [Portal for URL Verification](#portal-for-url-verification)
- [Portal for Human Data Audit](#portal-for-human-data-audit)
- [List of Toxic Words](#list-of-toxic-words)
- [Romanization Dictionary](#romanization-dictionary)
## Sangraha
Sangraha is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora, and large-scale translations. It has three broad components:
- **Sangraha Verified**: Contains scraped data from "human-verified" Websites, OCR-extracted data from high-quality Indic language PDFs, transcribed data from various Indic language videos, podcasts, movies, courses, etc. For scraping the data from the Web, we use the open-source framework [Webcorpus](https://github.com/AI4Bharat/webcorpus/tree/46af15a794fe101b0d9444f4a03dcd903be19fb7). For OCR, we collect PDFs from various sources with Internet Archive being the most prominent one. We release the code to download all the Indic Language PDFs from Internet Archive [here](https://github.com/AI4Bharat/sangraha-internet-archive-download/tree/94d32100b40f08450733895f9edb252d0590d9cb)
- **Sangraha Unverified**: High-quality Indic language data extracted from existing multilingual corpora. We employ perplexity filtering using n-gram language models trained on Sangraha Verified by extending the pipeline proposed by [CCNet](https://github.com/facebookresearch/cc_net/blob/main/cc_net/perplexity.py).
- **Sangraha Synthetic**: Wikimedia English translated to 14 Indic languages and further "romanized" from 14 languages by transliteration to English. We use the [Setu-translate](#setu-translate) and [Setu-transliterate](#setu-transliterate) pipelines for translation and transliteration, respectively.
Sangraha can be downloaded from [Huggingface 🤗](https://huggingface.co/datasets/ai4bharat/sangraha).
The list of languages in Sangraha:
<table>
<tbody>
<tr>
<td>Assamese (asm)</td>
<td>Konkani (gom)</td>
<td>Maithili (mai)</td>
<td>Oriya (ori)</td>
<td>Tamil (tam)</td>
</tr>
<tr>
<td>Bengali (ben)</td>
<td>Gujarati (guj)</td>
<td>Malayalam (mal)</td>
<td>Punjabi (pan)</tdExcerpt of 11,290 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:185c07a0e55e8e09, desc:fine-tuning, desc:fine tuning