VILA-Lab/ATLAS
quality grade D, 49 out of 100A principled instruction benchmark on formulating effective queries and prompts for large language models (LLMs). Our paper: https://arxiv.org/abs/2312.16171
- stars
- 990
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
951 results
A principled instruction benchmark on formulating effective queries and prompts for large language models (LLMs). Our paper: https://arxiv.org/abs/2312.16171
RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
A prize for finding tasks that cause large language models to show inverse scaling
GAOKAO-Bench is an evaluation framework that utilizes GAOKAO questions as a dataset to evaluate large language models.
This is the repository of HaluEval, a large-scale hallucination evaluation benchmark for Large Language Models.
Can large language models provide useful feedback on research papers? A large-scale empirical analysis.
LDB: A Large Language Model Debugger via Verifying Runtime Execution Step by Step (ACL'24)
This repo contains the source code for RULER: What’s the Real Context Size of Your Long-Context Language Models?
The papers are organized according to our survey: Evaluating Large Language Models: A Comprehensive Survey.
A benchmark to evaluate language models on questions I've previously asked them to solve.
[ICLR 2025 Oral] Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
Official repo for GPTFUZZER : Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
⚡LLM Zoo is a project that provides data, models, and evaluation benchmark for large language models.⚡
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
An MBTI Exploration of Large Language Models
Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
Benchmarking long-form factuality in large language models. Original code for our paper "Long-form factuality in large language models".
A framework for the evaluation of autoregressive code generation language models.
Code for the paper "Evaluating Large Language Models Trained on Code"
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Official github repo for C-Eval, a Chinese evaluation suite for foundation models [NeurIPS 2023]
An open science effort to benchmark legal reasoning in foundation models
③[ICML2024] [IQA, IAA, VQA] All-in-one Foundation Model for visual scoring. Can efficiently fine-tune to downstream datasets.
拼好RAG:手搓并融合了GraphRAG、LightRAG、Neo4j-llm-graph-builder进行知识图谱构建以及搜索;整合DeepSearch技术实现私域RAG的推理;自制针对GraphRAG的评估框架| Integrate GraphRAG, LightRAG, and Neo4j-llm-graph-builder for knowledge graph construction and search. Combine DeepSearch for private RAG reasoning. Create a custom evaluation framework for GraphRAG.
24,537 repositories in the index in total.