Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Agents' Last Exam
| Date | Stars |
|---|---|
| 2026-07-31 | 916 |
| 2026-08-06 | 936 |
Today
+20 stars today
This week
— stars this week
This month
— stars this month
Momentum
80.0
growth rate 0.00%/day
<div align="center">
<h1>
<img src="assets/logo.png" alt="" width="43" valign="middle" />
Agents' Last Exam
</h1>
*Challenge and measure AI agents on economically valuable, real-world tasks.*
[](https://agents-last-exam.org/)
[](https://arxiv.org/abs/2606.05405)
[](https://huggingface.co/agents-last-exam)
[](https://agenthle.org/leaderboard)
[](LICENSE)
[](LICENSE-DATA)
[](https://groups.google.com/g/agenthle-news)
Led by **[UC Berkeley RDI](https://rdi.berkeley.edu/)** × **RDI Foundation**
<br/>
<a href="assets/teaser.pdf"><img src="assets/teaser.png" alt="ALE benchmark: domains and example workflows" width="100%" /></a>
</div>
---
Agents' Last Exam aims to build the **broadest-coverage agent evaluation
benchmark to date**, measuring performance on long-horizon, economically
valuable tasks with verifiable outcomes. Co-led by Berkeley RDI and built with
hundreds of industry experts, ALE organizes real professional work into 55
subdomains across 13 industry clusters, with reference to O\*NET / SOC 2018
(the U.S. federal occupational taxonomy).
<table align="center">
<tr>
<td align="center" width="25%"><b>Broadest Coverage</b><br/><sub>55 subdomains<br/>across 13 clusters</sub></td>
<td align="center" width="25%"><b>Verifiable Outcomes</b><br/><sub>Hidden references<br/>+ deterministic graders</sub></td>
<td align="center" width="25%"><b>Long-Horizon</b><br/><sub>Multi-step workflows<br/>on real OS sandboxes</sub></td>
<td align="center" width="25%"><b>Economically Valuable</b><br/><sub>Sourced and validated<br/>by industry experts</sub></td>
</tr>
</table>
This repository is the **open evaluation framework**: the `ale_run` toolkit that
provisions sandboxes, runs agents, and grades them, plus around **150 public
tasks** across all 55 subdomains and reference integrations for several agent
harnesses.
---
## Quick start
Choose where ALE should create or attach each task sandbox:
| Provider | Best for | Guide |
|---|---|---|
| **Google Cloud VMs** | Elastic batch runs on published Ubuntu, Windows, and GPU images | [Cloud quick start](docs/quickstart.md) |
| **AWS (EC2 + S3)** | Elastic batch runs on the published Ubuntu and Windows images | [AWS setup guide](https://agents-last-exam.org/docs?p=pages/aws.html) |
| **Alibaba Cloud (ECS + OSS)** | Elastic batch runs on the published Ubuntu and Windows images | [Alibaba setup guide](https://agents-last-exam.org/docs?p=pages/aliyun.html) |
| **QEMU/KVM VMs** | CPU-compatible Ubuntu and Windows tasks on a Linux host with KVM | [QEMU/KVM guide](https://agents-last-exam.org/docs?p=pages/local.html) |
| **Local containers (Docker)** | The lighter supported Ubuntu subset | [Local container guide](https://agents-last-exam.org/docs?p=pages/local-docker.html) |
| **Existing sandbox** | Debugging against a CUA-enabled machine you already operate | [Static provider guide](https://agents-last-exam.org/docs?p=pages/static.html) |
Google Cloud is the recommended production path. The quick start covers the
one-time project setup, image copy, credentials, demo run, and grading flow.
### Roadmap
Where the frameworkExcerpt of 10,296 characters
Read on GitHub110
14
3
2
Wanyi Chen
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:19d19d32ad1e4d49, llm:Repository name and description: 'Agents' Last Exam' (no topics, no README provided). Python project likely related to agents. Limited information; could be about agents (research/exam) but unclear. Classified under 'agents' based on name.