Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.
| Date | Stars |
|---|---|
| 2026-07-31 | 654 |
| 2026-08-02 | 654 |
| 2026-08-06 | 654 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<div align="center">
<h1>Awesome LLM Eval</h1>
<a href="https://awesome.re"><img src="https://awesome.re/badge.svg"/></a>
</div>
[English](README_EN.md) | [中文](README_CN.md)
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on Large Language Models and exploring the boundaries and limits of Generative AI.
The is the official project of our survey: [Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap](arxiv.org/abs/2508.18646).
**NOTE:** As we cannot update the arXiv paper in real time, please refer to this repo for the latest updates and the paper may be updated later. We also welcome any pull request or issues to help us improve this work. Your contributions will be acknowledged in <a href="#acknowledgements">acknowledgements</a>.
If you find our survey useful, please kindly cite our paper:
```bibtex
@misc{wang2025llmevalroadmap,
title={Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap},
author={Jun Wang and Ninglun Gu and Kailai Zhang and Zijiao Zhang and Yelun Bao and Jin Yang and Xu Yin and Liwei Liu and Yihuan Liu and Pengyong Li and Gary G. Yen and Junchi Yan},
year={2025},
eprint={2508.18646},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2508.18646},
}
```

## Table of Contents
- [News](#News)
- [Tools](#Tools)
- [Datasets / Benchmark](#Datasets-or-Benchmark)
- [General](#General)
- [Domain](#Domain)
- [RAG-Evaluation](#RAG-Evaluation)
- [Agent-Capabilities](#Agent-Capabilities)
- [Coding-Capabilities](#Coding-Capabilities)
- [Multimodal/Cross-modal](#Multimodal-Cross-modal)
- [Long-Context](#Long-Context)
- [Inference-Speed](#Inference-Speed)
- [Quantization-and-Compression](#Quantization-and-Compression)
- [Demos](#Demos)
- [Leaderboards](#Leaderboards)
- [Papers](#Papers)
- [LLM-List](#LLM-List)
- [Pre-trained LLM](#Pre-trained-LLM)
- [Instruction Fine-tuned LLM](#Instruction-finetuned-LLM)
- [Aligned LLM](#Aligned-LLM)
- [Open LLM](#Open-LLM)
- [Popular LLM](#Popular-LLM)
- [LLMOps](#LLMOps)
- [Frameworks for Training](#Frameworks-for-Training)
- [Courses](#Courses)
- [Others](#Others)
- [Other Awesome Lists](#Other-Awesome-Lists)
- [Licenses](#Licenses)
- [Citation](#Citation)

## News
- [2025/08/20] We added the [Anthropomorphic-Taxonomy](#Anthropomorphic-Taxonomy) section.
- [2024/04/26] We added the [Inference-Speed](#Inference-Speed) section.
- [2024/02/26] We added the [Coding-Evaluation](#Coding-Capabilities) section.
- [2024/02/08] We added the [lighteval](https://github.com/huggingface/lighteval) tool from Huggingface.
- [2024/01/15] We added [CRUXEval](https://arxiv.org/abs/2401.03065), [DebugBench](https://github.com/thunlp/DebugBench), [OpenFinData](https://opencompass.org.cn), and [LAiW](https://github.com/Dai-shen/LAiW).
- [2023/12/20] We added the [RAG-Evaluation](#RAG-Evaluation) section.
- [2023/11/15] We added [Instruction-Following-Evaluation](https://github.com/google-research/google-research/tree/master/instruction_following_eval) and [LLMBar](https://github.com/princeton-nlp/LLMBar) for evaluating the instruction following capabilities of LLMs.
- [2023/10/20] We added [SuperCLUE-Agent](https://github.com/CLUEbenchmark/SuperCLUE-Agent) for LLM agent evaluation.
- [2023/09/25] We added [ColossalEval](https://github.com/hpcaitech/ColossalAI/tree/main/applications/ColossalEval) from Colossal-AI.
- [2023/09/22] We added the [LeaderboardFinder](#Leaderboards) chapter.
- [2023/09/20] We added [DeepEval](https://github.com/mr-gpt/deepeval), [FinEval](https://github.com/SUFE-AIFLM-Lab/FinEval), and [SuperCLUE-Safety](https://github.com/CLUEbenchmark/SuperCLUE-Safety) from CLUEbenchmark.
- [2023/09/18] We added [OpenCompass](https://github.com/InternLM/opencompass/tree/main) frExcerpt of 225,025 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:ece75308b445960c, topic:benchmark, topic:evaluation, topic:llm-evaluation
matched fp:ece75308b445960c, topic:llm, topic:llama, topic:qwen