scaleapi/SWE-bench_Pro-os
quality grade C, 53 out of 100SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- stars
- 491
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
953 results
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
VisualWebArena is a benchmark for multimodal agents.
Repo-level benchmark for real-world Code Agents: from repo understanding → env setup → incremental dev/bug-fixing → task delivery, with cost-aware α metric.
A multi-player tournament benchmark that tests LLMs in social reasoning, strategy, and deception. Players engage in public and private conversations, form alliances, and vote to eliminate each other
Official github repo for SafetyBench, a comprehensive benchmark to evaluate LLMs' safety. [ACL 2024]
Web-Bench is a benchmark designed to evaluate the performance of LLMs in actual Web development.
A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.
BABILong is a benchmark for LLM evaluation using the needle-in-a-haystack approach.
A simple toolkit for benchmarking LLMs on mathematical reasoning tasks. 🧮✨
Python SDK for running evaluations on LLM generated responses
Enterprise RAG Challenge to test accuracy of different LLM-driven assistants
Benchmarking LLMs with Challenging Tasks from Real Users
The lastest paper about detection of LLM-generated text and code
A multi-programming language benchmark for LLMs
Benchmarking LLMs via Uncertainty Quantification
Simple benchmark to test the most popular open source and commercial LLMs with automated OpenCode
A joint community effort to create one central leaderboard for LLMs.
PromtFuzz is an automated tool that generates high-quality fuzz drivers for libraries via a fuzz loop constructed on mutating LLMs' prompts.
Hallucination Detector is a free and open-source tool that helps you verify the accuracy of your LLM generated content instantly.
GPT-Fathom is an open-source and reproducible LLM evaluation suite, benchmarking 10+ leading open-source and closed-source LLMs as well as OpenAI's earlier models on 20+ curated benchmarks under aligned settings.
Codes for our paper "ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate"
Aidan Bench attempts to measure <big_model_smell> in LLMs.
(NeurIPS D&B 2024) STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases
Moonshot - A simple and modular tool to evaluate and red-team any LLM application.
24,523 repositories in the index in total.