Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
This benchmark tests how well LLMs incorporate a set of 10 mandatory story elements (characters, objects, core concepts, attributes, motivations, etc.) in a short creative story
| Date | Stars |
|---|---|
| 2026-07-31 | 412 |
| 2026-08-06 | 419 |
Today
+7 stars today
This week
— stars this week
This month
— stars this month
Momentum
28.0
growth rate 0.00%/day
# LLM Creative Story-Writing Benchmark This benchmark compares short stories written to the same constrained creative briefs. Separate evaluator models read matched story pairs and choose which one is better. Those choices are combined into a relative comparison score. Higher scores mean stronger performance against the other models tested. Scores are relative, not grades: zero is near the middle of this comparison set, and overlapping uncertainty ranges can indicate similarly rated models. --- ## Current Results  ### Leaderboard Current comparison set: - 39 rated models - 445 direct model pairings - 53,081 evaluator judgments - the rating combines compatible evaluator-v2 and evaluator-v3 evidence after bridge validation - the chart focuses on current models; the table retains all rated models for historical comparison - striped bars and markers identify models that completed fewer than 400 stories Labels such as `high`, `xhigh`, `max`, and `adaptive` identify the reasoning setting used for that model. Estimated win chance is the model's average expected chance against another model in the full comparison set. | Rank | Model | Comparison score | Estimated win chance | Uncertainty range | |-----:|:------|-----------------:|---------------------:|:------------------| | 1 | Claude Fable 5 (high)§ | 3.3 | 91% | 3.2 to 3.4 | | 2 | GPT-5.5 (xhigh) | 3.0 | 88% | 2.9 to 3.1 | | 3 | Kimi K3 | 2.9 | 87% | 2.8 to 3.0 | | 4 | GPT-5.6 Sol (xhigh) | 2.9 | 87% | 2.8 to 3.0 | | 5 | GPT-5.4 (xhigh) | 2.7 | 86% | 2.6 to 2.9 | | 6 | GPT-5.6 Sol (high) | 2.7 | 85% | 2.6 to 2.8 | | 7 | GPT-5.4 (medium) | 2.7 | 85% | 2.5 to 2.9 | | 8 | Claude Opus 4.7 (adaptive)† | 2.4 | 82% | 2.3 to 2.5 | | 9 | Claude Sonnet 4.6 (16K) | 2.2 | 80% | 2.1 to 2.4 | | 10 | Claude Opus 4.6 (16K) | 1.7 | 74% | 1.5 to 2.0 | | 11 | Muse Spark 1.1 (high) | 1.3 | 69% | 1.2 to 1.5 | | 12 | Claude Opus 4.8 (xhigh) | 1.3 | 69% | 1.2 to 1.4 | | 13 | GPT-5.2 (medium) | 1.0 | 64% | 0.8 to 1.2 | | 14 | GLM-5.2 (max) | 0.9 | 63% | 0.8 to 1.0 | | 15 | Claude Opus 4.8 (high)‡ | 0.8 | 62% | 0.7 to 0.9 | | 16 | Kimi K2.6 | 0.7 | 59% | 0.6 to 0.8 | | 17 | MiniMax-M3 | 0.6 | 58% | 0.4 to 0.7 | | 18 | Mistral Medium 3.1 | 0.2 | 52% | 0.0 to 0.3 | | 19 | DeepSeek V4 Pro | 0.1 | 51% | -0.1 to 0.2 | | 20 | Xiaomi MiMo V2.5 Pro | -0.1 | 48% | -0.2 to 0.1 | | 21 | Qwen 3 Max Preview | -0.1 | 48% | -0.3 to 0.1 | | 22 | Qwen 3.6 Max Preview | -0.4 | 44% | -0.5 to -0.2 | | 23 | GLM-5.1 | -0.5 | 42% | -0.7 to -0.3 | | 24 | Kimi K2.5 | -0.6 | 41% | -0.8 to -0.3 | | 25 | Baidu Ernie 5.1 | -0.7 | 39% | -0.9 to -0.5 | | 26 | Xiaomi MiMo V2 Pro | -0.7 | 39% | -1.0 to -0.5 | | 27 | Mistral Large 3 | -1.3 | 30% | -1.5 to -1.2 | | 28 | Gemma 4 31B Reasoning | -1.4 | 29% | -1.5 to -1.3 | | 29 | Gemini 3.5 Flash | -1.5 | 28% | -1.6 to -1.4 | | 30 | ByteDance Seed 2.0 Pro | -1.5 | 28% | -1.6 to -1.4 | | 31 | Gemini 3.1 Pro Preview | -1.8 | 24% | -1.9 to -1.7 | | 32 | Qwen 3.6 Plus | -1.8 | 24% | -2.0 to -1.6 | | 33 | Mistral Medium 3.5 | -2.0 | 22% | -2.2 to -1.9 | | 34 | Qwen 3.7 Max | -2.1 | 20% | -2.2 to -2.0 | | 35 | DeepSeek V3.2 | -2.4 | 17% | -2.7 to -2.1 | | 36 | GPT-OSS-120B | -2.7 | 15% | -2.8 to -2.6 | | 37 | MiniMax-M2.7 | -3.3 | 9% | -3.5 to -3.2 | | 38 | Grok 4.3 | -3.8 | 6% | -4.0 to -3.6 | | 39 | Grok 4.5 (high) | -4.6 | 2% | -4.7 to -4.5 | ### Coverage Note - † Claude Opus 4.7 completed 347 of 400 stories. Only completed stories were compared. - ‡ Claude Opus 4.8 high completed 399 of 400 stories. Only completed stories were compared. - § Claude Fable 5 high completed 395 of 400 stories. Only completed stories were compared. --- ## Head-to-Head Comparisons  Read each cell by row. Red means the row model performed better, blue means the column model performed better, and grey means the models were not directly compared. Near-whit
Excerpt of 8,618 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:f157a5d7e577c9ab, topic:llm, topic:llama