Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
The official GitHub repository of the paper "Recent advances in large language model benchmarks against data contamination: From static to dynamic evaluation"
| Date | Stars |
|---|---|
| 2026-07-24 | 500 |
| 2026-07-25 | 500 |
| 2026-07-28 | 500 |
| 2026-07-30 | 500 |
| 2026-07-31 | 500 |
| 2026-08-06 | 500 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Recent Advances in Large Language Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation <div align="center"> <img src="img/logo.jpeg" alt="LLM survey" width="1200"><br> </div> <br> <p align="center"> Simin Chen<sup>1</sup>   Yiming Chen<sup>2</sup>   Zexin Li<sup>3</sup>   Yifan Jiang<sup>4</sup>   Zhongwei Wan<sup>5</sup>   Yixin He<sup>4</sup>   Dezhi Ran<sup>6</sup>   Tianle Gu<sup>7</sup>   Haizhou Li<sup>2,8</sup>   Tao Xie<sup>6</sup>   Baishakhi Ray<sup>1</sup>   </p> <p align="center"> <sup>1</sup> Columbia University, <sup>2</sup> National University of Singapore, <sup>3</sup> University of California, Riverside, <sup>4</sup> University of Southern California, <sup>5</sup> The Ohio State University, <sup>6</sup> Peking University, <sup>7</sup> Tsinghua University, <sup>8</sup> The Chinese University of Hong Kong, Shenzhen<br> </p> ## ❤️ Community Support NOTE: As we cannot update the **EMNLP main camera ready** in real time, please refer to this repo for the latest updates, and the paper may be updated later. We also welcome any pull requests or issues to help us make this survey perfect. Your contributions will be acknowledged in the acknowledgements. We will actively maintain this repository by incorporating new research as it emerges. If you have any suggestions regarding our taxonomy, find any missed papers, or update any preprint arXiv paper that has been accepted to some venue, feel free to send us an email or submit a **pull request** using the following markdown format. ```markdown Paper Title, <ins>Conference/Journal/Preprint, Year</ins> [[pdf](link)] [[other resources](link)]. ``` <!-- [^1]: This table was updated Dec 2023. This table will require updates as cool new frameworks are being released frequently and current frameworks continue to mature at an accelerated rate. So please feel free to suggest any important distinguishing features or popular new frameworks--> ## 📌 What is This Survey About? Data contamination has received increasing attention in the era of large language models (LLMs) due to their reliance on vast Internet-derived training corpora. To mitigate the risk of potential data contamination, LLM benchmarking has undergone a transformation from *static* to *dynamic* benchmarking. In this work, we conduct an in-depth analysis of existing *static* to *dynamic* benchmarking methods aimed at reducing data contamination risks. We first examine methods that enhance *static* benchmarks and identify their inherent limitations. We then highlight a critical gap—the lack of standardized criteria for evaluating *dynamic* benchmarks. Based on this observation, we propose a series of optimal design principles for *dynamic* benchmarking and analyze the limitations of existing *dynamic* benchmarks. This survey provides a concise yet comprehensive overview of recent advancements in data contamination research, offering valuable insights and a clear guide for future research efforts. ## 🤔 What is data contamination? Data contamination occurs when benchmark data is inadvertently included in the training phase of language models, leading to an inflated and misleading assessment of their performance. While this issue has been recognized for some time—stemming from the fundamental machine learning principle of separating training and test sets—it has become even more critical with the advent of LLMs. These models often scrape vast amounts of publicly available data from the Internet, significantly increasing the likelihood of contamination. Furthermore, due to privacy and commercial concerns, tracing the exact training data for these models is challenging, if not impossible, complicating efforts to detect and mitigate potential contamination. ## ❓ Why do we need this survey?  This survey is necessary to address the
Excerpt of 27,000 characters
Read on GitHub19
15
12
1
1
1
1
Yue Huang · University of Notre Dame · United States
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:284cc657b2d1b10f, topic:benchmark, topic:evaluation, topic:testing