Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
The papers are organized according to our survey: Evaluating Large Language Models: A Comprehensive Survey.
| Date | Stars |
|---|---|
| 2026-07-31 | 804 |
| 2026-08-06 | 804 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome LLMs Evaluation Papers :bookmark_tabs:
The papers are organized according to [our survey](https://arxiv.org/pdf/2310.19736.pdf):
<p align="center"><strong>Evaluating Large Language Models: A Comprehensive Survey</strong></p>
<p align="center">Zishan Guo*, Renren Jin*, Chuang Liu*, Yufei Huang, Dan Shi, Supryadi, </p>
<p align="center">Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, Deyi Xiong†</p>
<p align="center">Tianjin University</p>
<p align="center">(*: Co-first authors, †: Corresponding author)</p>
<div align=center>
<img src="./imgs/Figure_1.png" style="zoom:30%"/>
</div>
If you find our survey useful, please kindly cite our paper:
```bibtex
@article{guo2023evaluating,
title={Evaluating Large Language Models: A Comprehensive Survey},
author={Guo, Zishan and Jin, Renren and Liu, Chuang and Huang, Yufei and Shi, Dan and Yu, Linhao and Liu, Yan and Li, Jiaxuan and Xiong, Bojian and Xiong, Deyi and others},
journal={arXiv preprint arXiv:2310.19736},
year={2023}
}
```
## Contributing to this paper list
Feel free to **open an issue/PR** or e-mail [[email protected]](mailto:[email protected]), [[email protected]](mailto:[email protected]), [[email protected]](mailto:[email protected]) and [[email protected]](mailto:[email protected]) if you find any missing areas, papers, or datasets. We will keep updating this list and survey.
## Updates
* [2023-10-30] Initial Paperlist for LLMs Evaluation from Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Supryadi, Linhao Yu, Jiaxuan Li, Bojian Xiong and Deyi Xiong.
## Survey Introduction
Large language models (LLMs) have demonstrated remarkable capabilities across a broad spectrum of tasks. They have attracted significant attention and been deployed in numerous downstream applications. Nevertheless, akin to a double-edged sword, LLMs also present potential risks. They could suffer from private data leaks or yield inappropriate, harmful, or misleading content. Additionally, the rapid progress of LLMs raises concerns about the potential emergence of superintelligent systems without adequate safeguards. To effectively capitalize on LLM capacities as well as ensure their safe and beneficial development, it is critical to conduct a rigorous and comprehensive evaluation of LLMs.
This survey endeavors to offer a panoramic perspective on the evaluation of LLMs. We categorize the evaluation of LLMs into three major groups: knowledge and capability evaluation, alignment evaluation and safety evaluation. In addition to the comprehensive review on the evaluation methodologies and benchmarks on these three aspects, we collate a compendium of evaluations pertaining to LLMs' performance in specialized domains, and discuss the construction of comprehensive evaluation platforms that covers LLM evaluations on capabilities, alignment, safety, sand applicability.
We hope that this comprehensive overview will stimulate further research interests in the evaluation of LLMs, with the ultimate goal of making evaluation serve as a cornerstone in guiding the responsible development of LLMs. We envision that this will channel their evolution into a direction that maximizes societal benefit while minimizing potential risks.
## Markups
The paper proposes a dataset that can be used for LLMs evaluation.
The paper proposes an evaluation method that can be used for LLMs.
The paper proposes a platform for LLMs evaluation.
The paper examines the performance of LLMs in a particular domain.
## Table of Contents
* [Updates](#updates)
* [Survey Introduction](#survey-introduction)
* [Markups](#markups)
* [Table of Contents](#table-of-contents)
* [Related Surveys for LLMs Evaluation](#related-surveys-for-LLMs-evaluation)
* [Papers](#papersExcerpt of 95,475 characters
Read on GitHub28
20
Zhmin Zhao · Software Analysis and Intelligence Lab (SAIL) & Lab on Maintenance, Construction and Intelligence of Software (MCIS) · Canada
12
Iker García-Ferrero · Multiverse Computing
2
Xingyao Wang · OpenHands / All Hands AI
1
Liang
1
Ikko Eltociear Ashimine · Japan
1
1
Terry Yue Zhuo
1
John Yang
1
Bin Wang · Apodex · Singapore
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:3139563fe5baca4c, llm:Repository is a curated list of papers from a survey 'Evaluating Large Language Models: A Comprehensive Survey' — organizes papers on LLM evaluation.