Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Two levels: three broad groups, and a flat list of categories beneath them. Every repository gets exactly one primary category — the counts below are primary assignments, so they add up.
The plumbing that runs models in production: serving, storage, orchestration, observability and data movement.
Runtimes and servers that execute model inference at speed and scale.
Making models smaller and faster: quantization, distillation, pruning, sparsity and kernel-level work.
Storage and retrieval for embeddings: vector indexes, hybrid search and ANN libraries.
Experiment tracking, model registries, feature stores, workflow schedulers and deployment tooling.
Logging, tracing, cost tracking and production monitoring for LLM and ML systems.
Unified APIs, model routers, proxies, caching layers and rate limiting across providers.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Cluster scheduling, GPU sharing, serverless compute and the distributed execution layer.
Building the models themselves: training, fine-tuning, evaluation, and the modality-specific research stacks.
Core deep-learning frameworks and libraries for pretraining and distributed training.
Adapting pretrained models: PEFT/LoRA, instruction tuning, RLHF, DPO and preference alignment.
Released model weights, reference implementations and architecture research.
Detection, segmentation, tracking, OCR models, 3D reconstruction and classical vision.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Diffusion models, image editing, upscaling and the surrounding creative tooling.
Video synthesis and editing, talking heads, avatars, animation and 3D asset generation.
Vision-language models, document understanding and any-to-any architectures.
Text and multimodal embedding models, rerankers and representation learning.
Tokenization, parsing, classical NLP pipelines, translation and information extraction.
RL algorithms, environments, simulators and decision-making systems.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Published datasets, dataset tooling and corpora for training and evaluation.
What people build on top of models: agents, retrieval, developer tools, interfaces and end-user products.
Autonomous and multi-agent systems, agent frameworks, planning and tool use.
Retrieval-augmented generation, document Q&A, knowledge graphs and memory systems.
Coding assistants, autonomous software engineers, code review bots and IDE integrations.
Model Context Protocol servers and clients, plugin systems and third-party tool connectors.
Prompt management, templating, DSLs, constrained decoding and guaranteed-schema output.
Chat interfaces, desktop clients, playgrounds and UI component libraries for AI apps.
Visual builders, low-code AI workflows, pipeline orchestration and automation platforms.
Agents that drive browsers, desktops and GUIs, plus web automation and testing.
Robot learning, manipulation, autonomous driving, simulation and embodied agents.
Guardrails, PII redaction, prompt-injection defense, interpretability and offensive/defensive AI security.
AI applied to a specific vertical: healthcare, finance, legal, science, education and more.
Courses, books, tutorials, paper collections and curated awesome lists.