yaojingang/yao-meta-skill
quality grade C, 64 out of 100YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
- stars
- 2.6k
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
951 results
YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
Evaluation harness for Apodex-1.0 on public deep-research benchmarks.
Embedded systems control library for development, testing and installation
🔭 Powerful tool for testing WebHooks and more
Cargo subcommand to provide various options useful for testing and continuous integration.
A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.
TorchBench is a collection of open source benchmarks used to evaluate PyTorch performance.
AI QA Agent for mobile apps
PySpark test helper methods with beautiful error messages
Symfony extension for PHPStan
Manage & spin up mongodb server binaries with zero(or slight) configuration for tests.
Provider-agnostic, open-source evaluation infrastructure for language models
Playwright Test Visual Studio Code integration
Pa11y CI is a CI-centric accessibility test runner, built using Pa11y
🦞 Official plugin for OpenClaw that exports agent traces to Opik. See and monitor agent behaviour, cost, tokens, errors and more.
GitHub Action to publish unit test results on GitHub
🌏 A tiny 0-dependency thread-safe Java™ lib for setting/viewing dns programmatically without touching host file, make unit/integration testing portable; and a tiny tool for setting/viewing dns of running JVM process.
🧪 single header unit testing framework for C and C++
Automated testing to find logic and performance bugs in database systems
Laravel Dusk provides simple end-to-end testing and browser automation.
Cutleries to help you cook better apps.
Coverage-guided, in-process fuzzing for Node.js
Arquillian provides a component model for integration tests, which includes dependency injection and container life cycle management. Instead of managing a runtime in your test, Arquillian brings your test to the runtime.
Local Azure development. One binary. No account needed. 25+ emulated services for testing, CI and local dev.
24,535 repositories in the index in total.