NVlabs/verilog-eval
quality grade F, 31 out of 100Verilog evaluation benchmark for large language model
- stars
- 457
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
953 results
Verilog evaluation benchmark for large language model
[ACL 2024 Demo] Official GitHub repo for UltraEval: An open source framework for evaluating foundation models.
FlagEval is an evaluation toolkit for AI large foundation models.
Towards Robust Evaluation for Geospatial Foundation Models
Foundation Model Evaluations Library
The agent responsible for conducting the agent evaluation
A benchmark built to evaluate and improve agent capabilities for supporting legal work.
Used for adaptive human in the loop evaluation of language and embedding models.
"Unit tests" for your agent skills
Measuring frontier coding agents on original, long-horizon engineering tasks
Enterprise-grade deep research skill for Claude Code with 8-phase pipeline, source credibility scoring, and automated validation. Outperforms OpenAI, Gemini, and Claude Desktop in quality and verification.
GEO-first SEO skill for Claude Code. Comprehensive AI search optimization for any website — citability scoring, AI crawler analysis, brand authority, schema markup, platform-specific optimization, and PDF reports. If you want learn how to sell this to real businesses, check out the skool community
ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.
GPT-3 found hundreds of security vulnerabilities in this repo - (this was the first real LLM cybersecurity eval!)
Web Codegen Scorer is a tool for evaluating the quality of web code generated by LLMs.
Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.
KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)
Evaluate the accuracy of LLM generated outputs
No description
Evaluation tool for LLM QA chains
The official repo for paper, LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.
Arena-Hard-Auto: An automatic LLM benchmark.
Automatic evals for LLMs
LLM Arena by KCORES team
24,523 repositories in the index in total.