Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents
| Date | Stars |
|---|---|
| 2026-07-31 | 3299 |
| 2026-08-03 | 3299 |
| 2026-08-06 | 3300 |
Today
+1 stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Hallucination Leaderboard
Public LLM leaderboard computed using Vectara's Hallucination Evaluation Model, also known as HHEM. This evaluates how often an LLM introduces hallucinations when summarizing a document. We plan to update this regularly as our model and the LLMs get updated over time.
Feel free to check out the [interactive hallucination leaderboard](https://huggingface.co/spaces/vectara/leaderboard) on Hugging Face.
If you are interested in previous versions os this leaderboard:
1. First version based on HHEM-1.0, it is available [here](https://github.com/vectara/hallucination-leaderboard/tree/hhem-1.0-final)
2. Most recent version, based on the previous dataset is available [here](https://github.com/vectara/hallucination-leaderboard/tree/hhem-2.3-old-dataset)
<table style="border-collapse: collapse;">
<tr>
<td style="text-align: center; vertical-align: middle; border: none;">
<img src="img/candle.png" width="50" height="50">
</td>
<td style="text-align: left; vertical-align: middle; border: none;">
In loving memory of <a href="https://www.ivinsfuneralhome.com/obituaries/Simon-Mark-Hughes?obId=30000023">Simon Mark Hughes</a>...
</td>
</tr>
</table>
<!-- LEADERBOARD_START -->
Last updated on May 11, 2026

|Model|Hallucination Rate|Factual Consistency Rate|Answer Rate|Average Summary Length (Words)|
|----|----:|----:|----:|----:|
|antgroup/finix_s1_32b|1.8 %|98.2 %|99.5 %|172.4|
|openai/gpt-5.4-nano-2026-03-17|3.1 %|96.9 %|100.0 %|144.4|
|google/gemini-2.5-flash-lite|3.3 %|96.7 %|99.5 %|95.7|
|microsoft/Phi-4|3.7 %|96.3 %|80.7 %|120.9|
|meta-llama/Llama-3.3-70B-Instruct-Turbo|4.1 %|95.9 %|99.5 %|64.6|
|snowflake/snowflake-arctic-instruct|4.3 %|95.7 %|62.7 %|81.4|
|google/gemma-3-12b-it|4.4 %|95.6 %|97.4 %|89.7|
|mistralai/mistral-large-2411|4.5 %|95.5 %|99.9 %|85.0|
|qwen/qwen3-8b|4.8 %|95.2 %|99.9 %|83.6|
|amazon/nova-pro-v1:0|5.1 %|94.9 %|99.3 %|66.2|
|amazon/nova-2-lite-v1:0|5.1 %|94.9 %|99.6 %|94.1|
|mistralai/mistral-small-2501|5.1 %|94.9 %|97.9 %|98.8|
|ibm-granite/granite-4.0-h-small|5.2 %|94.8 %|100.0 %|107.4|
|google/gemma-4-26b-a4b-it|5.2 %|94.8 %|99.8 %|67.1|
|ai21labs/jamba-mini-2|5.3 %|94.7 %|99.6 %|109.4|
|deepseek-ai/DeepSeek-V3.2-Exp|5.3 %|94.7 %|96.6 %|64.6|
|qwen/qwen3-14b|5.4 %|94.6 %|99.9 %|111.1|
|amazon/nova-micro-v1:0|5.5 %|94.5 %|100.0 %|100.0|
|deepseek-ai/DeepSeek-V3.1|5.5 %|94.5 %|94.5 %|63.7|
|openai/gpt-5.4-mini-2026-03-17|5.5 %|94.5 %|100.0 %|54.7|
|openai/gpt-4.1-2025-04-14|5.6 %|94.4 %|99.9 %|91.7|
|qwen/qwen3-4b|5.7 %|94.3 %|99.9 %|104.7|
|xai-org/grok-3|5.8 %|94.2 %|93.0 %|95.9|
|qwen/qwen3-32b|5.9 %|94.1 %|99.9 %|115.8|
|amazon/nova-lite-v1:0|6.1 %|93.9 %|99.9 %|91.8|
|deepseek-ai/DeepSeek-V3|6.1 %|93.9 %|97.5 %|81.7|
|deepseek-ai/DeepSeek-V3.2|6.3 %|93.7 %|92.6 %|62.0|
|google/gemma-3-4b-it|6.4 %|93.6 %|67.3 %|77.4|
|CohereLabs/command-r-plus-08-2024|6.9 %|93.1 %|95.0 %|91.5|
|arcee-ai/trinity-large-preview|6.9 %|93.1 %|99.0 %|117.3|
|openai/gpt-5.4-2026-03-05|7.0 %|93.0 %|99.9 %|81.7|
|google/gemini-2.5-pro|7.0 %|93.0 %|99.1 %|106.4|
|mistralai/ministral-3b-2410|7.3 %|92.7 %|99.9 %|167.9|
|google/gemma-3-27b-it|7.4 %|92.6 %|98.8 %|96.4|
|google/gemma-4-31b-it|7.4 %|92.6 %|100.0 %|75.8|
|mistralai/ministral-8b-2410|7.4 %|92.6 %|99.9 %|196.0|
|meta-llama/Llama-4-Scout-17B-16E-Instruct|7.7 %|92.3 %|99.0 %|137.3|
|google/gemini-2.5-flash|7.8 %|92.2 %|99.0 %|101.5|
|meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8|8.2 %|91.8 %|100.0 %|106.0|
|google/gemini-3.1-flash-lite-preview|8.2 %|91.8 %|99.6 %|62.6|
|openai/gpt-5.4-pro-2026-03-05|8.3 %|91.7 %|100.0 %|148.5|
|openai/gpt-5.2-low-2025-12-11|8.4 %|91.6 %|100.0 %|126.5|
|deepseek-ai/DeepSeek-V4-Pro|8.6 %|91.4 %|97.2 %|153.8|
|MiniMaxAI/minimax-m2p5|9.1 %|90.9 %|98.2 %|137.2|
|CohereLabs/command-a-03-2025|9.3 %|90.7 %|97.6 %|101.7|
|openai/gpt-5.5|9.3 %|90.7 %|100.0 %|129.6|
|zai-org/GLM-4.5-AIR-FP8|9.3 %|Excerpt of 23,893 characters
Read on GitHubOfer Mendelevitch · Blue Tensor LLC
136
70
50
35
32
9
Ikko Eltociear Ashimine · Japan
3
2
2
1
1
1
Pietro Monticone · Harmonic · United States
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:73a86a0bc03a6a87, topic:llm
matched fp:73a86a0bc03a6a87, name:leaderboard, desc:leaderboard