joke2k/faker
quality grade A, 93 out of 100Faker is a Python package that generates fake data for you.
- stars
- 19k
- stars gained this week
- —
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
953 results
Faker is a Python package that generates fake data for you.
A modern, C++-native, test framework for unit-tests, TDD and BDD - using C++14, C++17 and later (C++11 support is in v2.x branch, and C++03 on the Catch1.x branch)
Opensource IDE For Exploring and Testing API's (lightweight alternative to Postman/Insomnia)
Open-Source API Development Ecosystem • https://hoppscotch.io • Offline, On-Prem & Cloud • Web, Desktop & CLI • Open-Source Alternative to Postman, Insomnia
Storybook is the industry standard workshop for building, documenting, and testing UI components in isolation
🦞 Official plugin for OpenClaw that exports agent traces to Opik. See and monitor agent behaviour, cost, tokens, errors and more.
Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty is an asset class and format for LLM prompts designed to enhance observability, understandability, and portability for developers.
Travel through time in your tests.
TestNG testing framework
Embedded systems control library for development, testing and installation
Terraform & OpenTofu Skill for AI Agents - testing, modules, CI/CD, and production patterns
A comprehensive time-series benchmark evaluating state-of-the-art deep learning architectures (PatchTST, TFT, N-HiTS) against traditional gradient boosting (CatBoost) for accurate 24-hour load prediction.
LiveBench: A Challenging, Contamination-Free LLM Benchmark
A Living Benchmark for Machine Learning on Tabular Data
Mutation testing system
[ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
RocketSim — 30+ tools for Xcode's iOS Simulator. Testing, debugging, network monitoring, captures, accessibility, app actions, and AI agent automation via the RocketSim CLI. Used by 80k+ developers.
SkillsBench evaluates how well skills work and how effective agents are at using them.
An LLM-as-a-judge HTTP proxy to secure agents in production
DoubleML - Double Machine Learning in Python
A development and test oriented OAuth2 mock server
Tool for static code analysis and formatting of Robot Framework language
Cutleries to help you cook better apps.
Release with confidence, state-of-the-art property testing for Scala.
24,523 repositories in the index in total.