Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
| Date | Stars |
|---|---|
| 2026-07-31 | 445 |
| 2026-08-06 | 449 |
Today
+4 stars today
This week
— stars this week
This month
— stars this month
Momentum
16.0
growth rate 0.00%/day
<div align="center">
<p align="center">
<img src="./assets/toolathlon.svg" alt="Logo" width="500" height="200"/>
</p>
# The Tool Decathlon: Benchmarking Language Agents for <br>Diverse, Realistic, and Long-Horizon Task Execution
[](https://toolathlon.xyz/)
[](https://discord.gg/Da3AaW4rVs)
[](https://arxiv.org/abs/2510.25726)
[](https://huggingface.co/datasets/hkust-nlp/Toolathlon-Verified_Trajectories/tree/main)
[](https://github.com/hkust-nlp/Toolathlon)
</div>
## Introduction
Toolathlon is a benchmark to assess language agents' general tool use in realistic environments. It features 600+ diverse tools based on real-world software environments. Each task requires long-horizon tool calls to complete. This repository corresponds to the **Toolathlon-Verified** final release. Below we show a demo task where the agent needs to automatically check assignments in the email box, and grade them on Canvas.
<div align="center">
<img src="assets/demo.gif" width="100%" alt="Demo">
</div>
## News
[2026.07.09] 📣 We have also uploaded trajectories for Muse Spark 1.1 to the [**Toolathlon-Verified** trajectories](https://huggingface.co/datasets/hkust-nlp/Toolathlon-Verified_Trajectories/tree/main) dataset on Hugging Face.
[2026.06.30] 🎉 **Toolathlon-Verified** is released. This release marks the verified final version of Toolathlon, with task prompts, ground truths, and evaluators reviewed and aligned for the final benchmark release. We have also uploaded trajectories for 8 models to the [**Toolathlon-Verified** trajectories](https://huggingface.co/datasets/hkust-nlp/Toolathlon-Verified_Trajectories/tree/main) dataset on Hugging Face.
[2025.12.12] 📣 We have set up a new documentation page for common issues and updates, please refer to [UpdateLogs_CommonIssues.md](UpdateLogs_CommonIssues.md) for more details. The document will be updated regularly to track the latest changes and issues.
[2025.12.05] 📣 We have updated the trajectories for 4 new models (gemini-3-pro, claude-4.5-opus, gpt-5.1, deepseek-v3.2-thinking) on [Hugging Face](https://huggingface.co/datasets/hkust-nlp/Toolathlon-Trajectories), take a look if you are interested.
[2025.11.28] 🚀 We have provided a ready-to-use public eval service (you do not need to set anything up!), please refer to [EVAL_SERVICE_README.md](EVAL_SERVICE_README.md) for more details.
## Sandbox Version
We also provide a sandboxed version of Toolathlon in the [`toolathlon-sandbox`](https://github.com/hkust-nlp/Toolathlon/tree/toolathlon-sandbox) branch.
## Quick Start
> There are four ways to run Toolathlon evaluation:
> 1. Using our public evaluation service: Check [EVAL_SERVICE_README.md](EVAL_SERVICE_README.md) for more details.
> 2. Set up your own Toolathlon evaluation service on your own machine as detailed below.
> 3. If you are a major user that will use Toolathlon evaluation a lot, you can also contact us ([email protected]), we may be able to provide a dedicated evaluation service for you (for free).
> 4. If you have an API endpoint and just want to test your model, you can contact us ([email protected]) and we are happy to help you run evaluation on Toolathlon with your given API endpoint.
### Using Our Public Evaluation Service
We provide Toolathlon evaluation as a service on public servers, where we have set up all the required MCP accounts and you don't need to worry about the setup -- you don't even need to install any MCP-related dependenciExcerpt of 16,873 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:7da93a315766b8dd, llm:Repository title and description: 'The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution' (ICLR 2026). Focus is benchmarking language agents.
matched fp:7da93a315766b8dd, llm:Repository title and description: 'The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution' (ICLR 2026). Focus is benchmarking language agents.
matched fp:7da93a315766b8dd, llm:Repository title and description: 'The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution' (ICLR 2026). Focus is benchmarking language agents.