gersteinlab/ML-Bench
quality grade C, 57 out of 100ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code (https://arxiv.org/abs/2311.09835)
- stars
- 316
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
951 results
ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code (https://arxiv.org/abs/2311.09835)
Uncertainty Quantification 360 (UQ360) is an extensible open-source toolkit that can help you estimate, communicate and use uncertainty in machine learning model predictions.
ParaMonte: Parallel Monte Carlo and Machine Learning Library for Python, MATLAB, Fortran, C++, C.
Machine Learning Experiment Manage Platform
Assessing the quality of metagenome-derived genome bins using machine learning
A large-scale benchmark for machine learning methods in fluid dynamics
Applied Machine Learning Explainability Techniques, published by Packt
detect demographic differences in the output of machine learning models or other assessments
A library for debugging/inspecting machine learning classifiers and explaining their predictions
Robustness Gym is an evaluation toolkit for machine learning.
PyRCA: A Python Machine Learning Library for Root Cause Analysis
COVID-19 Projections Using Machine Learning
Machine learning evaluation metrics, implemented in Python, R, Haskell, and MATLAB / Octave
Generate Diverse Counterfactual Explanations for any machine learning model.
A machine learning toolkit for log parsing [ICSE'19, DSN'16]
Lime: Explaining the predictions of any machine learning classifier
Public repository associated with "Deep Learning for ECG Analysis: Benchmarks and Insights from PTB-XL"
Experiments used in "Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning"
Benchmark Suite for Deep Learning
THE Deep Learning Benchmarks
Deep learning-based Video Quality Assessment
A comprehensive time-series benchmark evaluating state-of-the-art deep learning architectures (PatchTST, TFT, N-HiTS) against traditional gradient boosting (CatBoost) for accurate 24-hour load prediction.
Bayesian Deep Learning Benchmarks
Evaluation of Deep Learning Frameworks
24,537 repositories in the index in total.