Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Scalable data pre processing and curation toolkit for LLMs
| Date | Stars |
|---|---|
| 2026-07-31 | 1689 |
| 2026-08-04 | 1699 |
| 2026-08-06 | 1699 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
35.0
growth rate 0.00%/day
<div align="center"> <a href="https://github.com/NVIDIA-NeMo/Curator/blob/main/LICENSE"></a> <a href="https://codecov.io/github/NVIDIA-NeMo/Curator"></a> <a href="https://pypi.org/project/nemo-curator/"></a> <a href="https://github.com/NVIDIA-NeMo/Curator/graphs/contributors"></a> <a href="https://github.com/NVIDIA-NeMo/Curator/releases"></a> <a href="https://pypi.org/project/nemo-curator/"></a> </div> # NVIDIA NeMo Curator **NeMo Curator helps ML engineers and data teams build repeatable, GPU-accelerated pipelines that load, filter, deduplicate, and transform large text, image, video, and audio datasets for AI training.** Run the same pipeline on a laptop or across a multi-node Ray cluster. > *Part of the [NVIDIA NeMo](https://www.nvidia.com/en-us/ai-data-science/products/nemo/) software suite for managing the AI agent lifecycle.* ## What's Hot Don't miss the latest capabilities developers are picking up: | Feature | What it unlocks | Read this | |---------|-----------------|-----------| | **Curator on Slurm** | Run multi-node Ray pipelines on HPC clusters — text, image, video, and audio workloads at scale | [Slurm Deployment Guide](https://docs.nvidia.com/nemo/curator/latest/admin/deployment/slurm-multi-node-ray) | | **Audio Curation** | Build ALM and speech datasets with composite quality filtering, audio tagging, and speaker diarization | [Audio Guide](https://docs.nvidia.com/nemo/curator/latest/curate-audio) | | **Inference Server** | Spin up an OpenAI-compatible LLM endpoint inside your pipeline for SDG, classification, and synthetic data workflows | [Inference Server](https://docs.nvidia.com/nemo/curator/latest/curate-text/synthetic/inference-server) | > Want something featured here? Open an issue or ping `@nemo-curator-leads`. ## Updates - **2026-04** — NeMo Curator 26.04: Cosmos-Xenna 0.2.0 upgrade, simplified `Resources` API, Ray runtime upgrade. See the [release notes](https://docs.nvidia.com/nemo/curator/latest/about/release-notes). - **2026-02** — NeMo Curator 26.02: Ray-based pipeline architecture for all modalities — text, image, video, and audio. --- ## What You Can Build | Modality | Common Operations | Guide | |----------|-------------------|-------| | **Text** | Deduplication, classification, quality filtering, language detection | [Text Guide](https://docs.nvidia.com/nemo/curator/latest/get-started/text) | | **Image** | Aesthetic filtering, NSFW detection, embedding generation, deduplication | [Image Guide](https://docs.nvidia.com/nemo/curator/latest/get-started/image) | | **Video** | Scene detection, clip extraction, motion filtering, deduplication | [Video Guide](https://docs.nvidia.com/nemo/curator/latest/get-started/video) | | **Audio** | ASR transcription, quality assessment, WER filtering | [Audio Guide](https://docs.nvidia.com/nemo/curator/latest/get-started/audio) | ### Use NeMo Curator when… - You need **repeatable curation pipelines** — not one-off notebooks or ad-hoc scripts. - You need **GPU and distributed execution** for data-heavy stages (dedupe, classification, embedding, inference). - You need **modality-aware building blocks** for text, image, video, or audio. - You want **recipes that map to NVIDIA training workflows** like Nemotron and Nemotron-CC. --- ## Quick Start Three paths, depending on what you're trying to do. Each path is self-contained. NeMo Curator uses [`uv`](https://docs.astral.sh/uv/) for ins
Excerpt of 13,062 characters
Read on GitHub161
Praateek Mahajan · NVIDIA · United States
81
75
74
59
Lawrence Lane · NVIDIA · United States
58
oliver könig · Nvidia · Denmark
43
38
36
32
Ao Tang · NVIDIA · Canada
27
Abhinav Garg · Stanford University; IIT Kanpur
23
Charlie Truong
22
15
14
14
13
10
Onur Yilmaz · NVIDIA
9
9
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:b263163236fac64b, topic:large-language-models, topic:llm
matched fp:b263163236fac64b, topic:fine-tuning