lunarmodules/busted
quality grade B, 78 out of 100Elegant Lua unit testing.
- stars
- 1.6k
- stars gained this week
- +1this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
951 results
Elegant Lua unit testing.
Android ViewServer and ADB client
Concurrent browser tests for your Elixir web apps.
AI-native software testing platform — one YAML test language covering API, Web, mobile, messaging, data, and LLM scenarios. Self-hostable.
Various recipes for testing common scenarios with Cypress
HTTP(S) benchmark tools, testing/debugging, & restAPI (RESTful)
Simple JavaScript testing framework for browsers and node.js
An evaluation framework for machine learning models simulating high-throughput materials discovery.
The net:cal calibration framework is a Python 3 library for measuring and mitigating miscalibration of uncertainty estimates, e.g., by a neural network.
Code for the Proceedings of the National Academy of Sciences 2020 article, "Understanding the Role of Individual Units in a Deep Neural Network"
MLCommons Algorithmic Efficiency is a benchmark and competition measuring neural network training speedups due to algorithmic improvements in both training algorithms and models.
alpha-beta-CROWN: An Efficient, Scalable and GPU Accelerated Neural Network Verifier (winner of VNN-COMP 2021, 2022, 2023, 2024, 2025)
Code for "Uncertainty Estimation Using a Single Deep Deterministic Neural Network"
Learning Confidence for Out-of-Distribution Detection in Neural Networks
Codebase for testing whether hidden states of neural networks encode discrete structures.
Reproduce CKA: Similarity of Neural Network Representations Revisited
ETH Robustness Analyzer for Deep Neural Networks
A Collection of Competitive Text-Based Games for Language Model Evaluation and Reinforcement Learning
ExploitBench measures how far AI agents climb, from reaching vulnerable code, to triggering the bug, to building exploit primitives, to arbitrary code execution.
CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities
[ BOF-LAUNCHER ] -> an API for loading, executing and in-memory masking BOFs on Windows and Linux for use in C/Zig/Go/Rust agents/implants. [ Z-BEAC0N ] -> a custom-written stage-1 (aka pre-C2) solution engineered with a small footprint, stealth and modularity in mind. [ DEVELOPED BOFS ] -> cross-platform (12), Linux-only (10), Win-only (2)
[ICML 2025 Oral] Official repo of EmbodiedBench, a comprehensive benchmark designed to evaluate MLLMs as embodied agents.
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)
Agent trace and tool-use safety evaluation lab.
24,537 repositories in the index in total.