interpretml/DiCE
quality grade D, 48 out of 100Generate Diverse Counterfactual Explanations for any machine learning model.
- stars
- 1.5k
- stars gained this week
- +1this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
954 results
Generate Diverse Counterfactual Explanations for any machine learning model.
A machine learning toolkit for log parsing [ICSE'19, DSN'16]
Lime: Explaining the predictions of any machine learning classifier
TopoBench is a Python library designed to standardize benchmarking and accelerate research in Topological Deep Learning
Public repository associated with "Deep Learning for ECG Analysis: Benchmarks and Insights from PTB-XL"
Experiments used in "Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning"
Benchmark Suite for Deep Learning
THE Deep Learning Benchmarks
Deep learning-based Video Quality Assessment
Bayesian Deep Learning Benchmarks
Evaluation of Deep Learning Frameworks
The WeightWatcher tool for predicting the accuracy of Deep Neural Networks
A toolbox to iNNvestigate neural networks' predictions!
No description
Generate and evaluate agent skills for code agents like Claude Code, Open Code, OpenAI Codex
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
a high performance library for building cache simulators
List of papers on hallucination detection in LLMs.
UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.
An open-source visual programming environment for battle-testing prompts to LLMs.
AI Red Teaming playground labs to run AI Red Teaming trainings including infrastructure.
TorchBench is a collection of open source benchmarks used to evaluate PyTorch performance.
C2concealer is a command line tool that generates randomized C2 malleable profiles for use in Cobalt Strike.
24,524 repositories in the index in total.