hkust-nlp/AgentBoard
quality grade F, 34 out of 100An Analytical Evaluation Board of Multi-turn LLM Agents [NeurIPS 2024 Oral]
- stars
- 432
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
953 results
An Analytical Evaluation Board of Multi-turn LLM Agents [NeurIPS 2024 Oral]
Test LLMs on real tasks. Compare models side-by-side.
[TACL'26] VoiceBench: Benchmarking LLM-Based Voice Assistants
Build, evaluate, understand, and fix LLM-based apps
Run evaluation on LLMs using human-eval benchmark
A comprehensive repository of reasoning tasks for LLMs (and beyond)
Testing baseline LLMs performance across various models
[ICLR 2025 Spotlight] An open-sourced LLM judge for evaluating LLM-generated answers.
A comprehensive set of LLM benchmark scores and provider prices. (deprecated, read more in README)
No description
EvaLearn is a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks.
[ACL'24 Outstanding] Data and code for L-Eval, a comprehensive long context language models evaluation benchmark
[ACL2024] T-Eval: Evaluating Tool Utilization Capability of Large Language Models Step by Step
Codes and Data for Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Repo for paper *Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges*
Code for the paper "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples"
Code and data for "Measuring and Narrowing the Compositionality Gap in Language Models"
Code and data for "Lost in the Middle: How Language Models Use Long Contexts"
Code and results accompanying the paper "Refusal in Language Models Is Mediated by a Single Direction".
A benchmark for emotional intelligence in large language models
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
An extensible benchmark for evaluating large language models on planning
Understand and test language model architectures on synthetic tasks.
Red-Teaming Language Models with DSPy
24,523 repositories in the index in total.