yaojingang/yao-meta-skill
quality grade B, 69 out of 100YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
- stars
- 2.1k
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
953 results
YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
The net:cal calibration framework is a Python 3 library for measuring and mitigating miscalibration of uncertainty estimates, e.g., by a neural network.
Code for the Proceedings of the National Academy of Sciences 2020 article, "Understanding the Role of Individual Units in a Deep Neural Network"
MLCommons Algorithmic Efficiency is a benchmark and competition measuring neural network training speedups due to algorithmic improvements in both training algorithms and models.
auto_LiRPA: An Automatic Linear Relaxation based Perturbation Analysis Library for Neural Networks and General Computational Graphs
Code for "Uncertainty Estimation Using a Single Deep Deterministic Neural Network"
Approximating neural network loss landscapes in low-dimensional parameter subspaces for PyTorch
Learning Confidence for Out-of-Distribution Detection in Neural Networks
Codebase for testing whether hidden states of neural networks encode discrete structures.
Reproduce CKA: Similarity of Neural Network Representations Revisited
ETH Robustness Analyzer for Deep Neural Networks
LLM Benchmark for Throughput via Ollama (Local LLMs)
CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities
[ BOF-LAUNCHER ] -> an API for loading, executing and in-memory masking BOFs on Windows and Linux for use in C/Zig/Go/Rust agents/implants. [ Z-BEAC0N ] -> a custom-written stage-1 (aka pre-C2) solution engineered with a small footprint, stealth and modularity in mind. [ DEVELOPED BOFS ] -> cross-platform (12), Linux-only (10), Win-only (2)
Deterministic UI audits for shadcn apps, built for your terminal, your CI, and your AI agent.
[ICML 2025 Oral] Official repo of EmbodiedBench, a comprehensive benchmark designed to evaluate MLLMs as embodied agents.
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)
Agent trace and tool-use safety evaluation lab.
Open source AI Agent evaluation framework for web tasks 🐒🍌
AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent Trajectories
A generative AI-powered framework for testing virtual agents.
Benchmark for automated failure attributions in agentic systems (🏆 ICML 2025 Spotlight)
MCPMark is a comprehensive, stress-testing MCP benchmark designed to evaluate model and agent capabilities in real-world MCP use.
Official repo of VLABench, a large scale benchmark designed for fairly evaluating VLA, Embodied Agent, and VLMs.
24,523 repositories in the index in total.