FranxYao/chain-of-thought-hub
quality grade C, 56 out of 100Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
- stars
- 2.8k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
953 results
Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
Provider-agnostic, open-source evaluation infrastructure for language models
Benchmarking long-form factuality in large language models. Original code for our paper "Long-form factuality in large language models".
A framework for the evaluation of autoregressive code generation language models.
Code for the paper "Evaluating Large Language Models Trained on Code"
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Official github repo for C-Eval, a Chinese evaluation suite for foundation models [NeurIPS 2023]
An open science effort to benchmark legal reasoning in foundation models
③[ICML2024] [IQA, IAA, VQA] All-in-one Foundation Model for visual scoring. Can efficiently fine-tune to downstream datasets.
拼好RAG:手搓并融合了GraphRAG、LightRAG、Neo4j-llm-graph-builder进行知识图谱构建以及搜索;整合DeepSearch技术实现私域RAG的推理;自制针对GraphRAG的评估框架| Integrate GraphRAG, LightRAG, and Neo4j-llm-graph-builder for knowledge graph construction and search. Combine DeepSearch for private RAG reasoning. Create a custom evaluation framework for GraphRAG.
ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code (https://arxiv.org/abs/2311.09835)
Uncertainty Quantification 360 (UQ360) is an extensible open-source toolkit that can help you estimate, communicate and use uncertainty in machine learning model predictions.
ParaMonte: Parallel Monte Carlo and Machine Learning Library for Python, MATLAB, Fortran, C++, C.
Machine Learning Experiment Manage Platform
Assessing the quality of metagenome-derived genome bins using machine learning
A large-scale benchmark for machine learning methods in fluid dynamics
Applied Machine Learning Explainability Techniques, published by Packt
detect demographic differences in the output of machine learning models or other assessments
A library for debugging/inspecting machine learning classifiers and explaining their predictions
Robustness Gym is an evaluation toolkit for machine learning.
PyRCA: A Python Machine Learning Library for Root Cause Analysis
COVID-19 Projections Using Machine Learning
Machine learning evaluation metrics, implemented in Python, R, Haskell, and MATLAB / Octave
Generate Diverse Counterfactual Explanations for any machine learning model.
24,523 repositories in the index in total.