aws/fmeval
quality grade D, 41 out of 100Foundation Model Evaluations Library
- stars
- 292
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
951 results
Foundation Model Evaluations Library
The agent responsible for conducting the agent evaluation
Used for adaptive human in the loop evaluation of language and embedding models.
Enterprise-grade deep research skill for Claude Code with 8-phase pipeline, source credibility scoring, and automated validation. Outperforms OpenAI, Gemini, and Claude Desktop in quality and verification.
ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.
GPT-3 found hundreds of security vulnerabilities in this repo - (this was the first real LLM cybersecurity eval!)
Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.
KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)
Evaluate the accuracy of LLM generated outputs
No description
Evaluation tool for LLM QA chains
The official repo for paper, LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.
Arena-Hard-Auto: An automatic LLM benchmark.
Automatic evals for LLMs
LLM Arena by KCORES team
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.
Automatically evaluate your LLMs in Google Colab
PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai
A benchmark for evaluating LLMs on Chinese traditional fortune telling — Bazi (八字) and Ziwei Doushu (紫微斗数).
Doing simple retrieval from LLM models at various context lengths to measure accuracy
Evaluate your LLM-powered apps with TypeScript
HDDM is a python module that implements Hierarchical Bayesian parameter estimation of Drift Diffusion Models (via PyMC).
No description
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
24,537 repositories in the index in total.