Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
| Date | Stars |
|---|---|
| 2026-07-31 | 765 |
| 2026-08-02 | 781 |
| 2026-08-06 | 781 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome Agent Evals [](https://awesome.re) > A curated, opinionated, **non-BS** library of the best resources for **building and evaluating AI agents** — papers, blog posts, talks, courses, tools, and benchmarks. Maintained by [**BenchFlow**](https://benchflow.ai) · Join our [**Discord**](https://discord.gg/mZ9Rc8q8W3) Most "awesome" lists are link dumps. This one is **annotated and verified**: every entry says *what it is and why it belongs*, URLs are checked, quotes are verbatim, and dead/abandoned tools are pruned (not silently listed). It was assembled by: - a **depth-4 recursive citation crawl** (11.6k papers, ranked by in-degree) to surface the academic canon, - **targeted practitioner-web discovery** for the industry sources citation graphs miss (Eugene Yan, Han-Chung Lee, Hamel Husain, Shreya Shankar, Nathan Lambert, …), - **47 talks & podcasts transcribed and deep-noted** (verbatim + timestamps), and - **per-section gap audits** with adversarial verification. **443+ curated links · 146 deep reading notes** (see [`notes/`](notes/)). Markers: 🆕 = released/updated 2025–2026 · ⚠️ = caveat. Contributions welcome — see [CONTRIBUTING](CONTRIBUTING.md). > 📘 **Playbook:** [**PATTERNS.md**](PATTERNS.md) — real, runnable code + worked examples for LLM-as-judge (aligned to humans), pass@k/pass^k, error analysis, trajectory & world-state grading, CI gating, verifiable rewards, and more. ## Contents - [📘 Playbook — real code & worked examples (PATTERNS.md)](PATTERNS.md) - [⭐ Must-read starter set (read these first)](#-must-read-starter-set-read-these-first) - [1 · Why we need evals](#1-why-we-need-evals) - [2 · "If you can eval it, you have built it" — eval ⇄ capability ⇄ RL environment](#2-if-you-can-eval-it-you-have-built-it-eval-capability-rl-environment) - [3 · The model / harness / skill decomposition](#3-the-model-harness-skill-decomposition) - [4 · Observability & the output / eval space (the surfaces you can grade)](#4-observability-the-output-eval-space-the-surfaces-you-can-grade) - [5 · Evaluation infrastructure (the eval stack: datasets, scorers, online/offline, tracing, CI)](#5-evaluation-infrastructure-the-eval-stack-datasets-scorers-onlineoffline-tracing-ci) - [6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)](#6-benchmark-vs-eval-and-benchmark-integrity-contamination-saturation-label-errors-leaderboard-gaming) - [7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)](#7-evals-rl-environments-verifiers-reward-design-difficulty-calibration-lifecycle) - [8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)](#8-llm-as-judge-verifiers-alignment-biases-verifiable-vs-judgeable) - [9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)](#9-agent-specific-evaluation-trajectories-tool-use-multi-turn-world-state-multi-agent-localization) - [10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)](#10-safety-adversarial-evaluation-prompt-injection-jailbreaks-action-authorization-benchmark-auditing) - [🎙 Talks, podcasts & slides (transcribed + noted)](#-talks-podcasts-slides-transcribed-noted) - [💬 Eval mentions](#-eval-mentions) - [Companies & landscape (eval / RL-environment market)](#companies-landscape-eval-rl-environment-market) - [Notes on provenance & gaps](#notes-on-provenance-gaps) - [Deep notes](#deep-notes) - [Contributing](#contributing) - [License](#license) --- ## ⭐ Must-read starter set (read these first) 1. **[The Second Half](https://ysymyth.github.io/The-Second-Half/)** — Shunyu Yao — <https://ysymyth.github.io/The-Second-Half/> · *blog* — "Evaluation becomes more important than training." The field-level *why*. 2. **[An LLM-as-Judge Won't Save the Product, Fixing Your Process Will](https://eugeneyan.com/writing/eval-process/)** — Eugen
Excerpt of 138,861 characters
Read on GitHubXiangyi Li · @benchflow-ai · United States
37
2
1
1
1
1
1
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:07231160d248677c, topic:awesome, topic:awesome-list
matched fp:07231160d248677c, topic:llm
matched fp:07231160d248677c, topic:llm-evaluation
matched fp:07231160d248677c, topic:ai-agents, desc:ai agents