sgl-project/genai-bench
quality grade C, 63 out of 100Genai-bench is a powerful benchmark tool designed for comprehensive token-level performance evaluation of large language model (LLM) serving systems.
- stars
- 331
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
951 results
Genai-bench is a powerful benchmark tool designed for comprehensive token-level performance evaluation of large language model (LLM) serving systems.
The open source visual testing platform for teams and AI agents. Review the product, not just the code.
Travel through time in your tests.
Source code for privacytests.org. Includes browser testing code and site rendering.
A faster, simpler way to drive browsers supporting the Chrome DevTools Protocol.
A comprehensive time-series benchmark evaluating state-of-the-art deep learning architectures (PatchTST, TFT, N-HiTS) against traditional gradient boosting (CatBoost) for accurate 24-hour load prediction.
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Interface mocking tool for go generate
Benchmarking physical understanding in generative video models
AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent Trajectories
[TACL'26] VoiceBench: Benchmarking LLM-Based Voice Assistants
TopoBench is a Python library designed to standardize benchmarking and accelerate research in Topological Deep Learning
Collab with OpenAI. A benchmark and harness for finding and exploiting smart contract bugs
Provides automated YAML management and a streamlit workbench. Designed to optimize dev workflows.
MockBukkit is a mocking framework for Bukkit/PaperMC to allow the easy unit testing of Bukkit plugins.
Powerful friendly HTTP mock server & proxy library
PostgreSQL Unit Testing Suite
Pairwise Independent Combinatorial Tool
BDD library for the pytest runner
Ruby Tests Profiling Toolbox
TestNG testing framework
Unified World Model Inference & Evaluation Infrastructure
Approximating neural network loss landscapes in low-dimensional parameter subspaces for PyTorch
A Game of Life simulator Android app and watchface built with Jetpack Compose
24,535 repositories in the index in total.