Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Fully automatic censorship removal for language models
| Date | Stars |
|---|---|
| 2026-07-24 | 26695 |
| 2026-07-25 | 26728 |
| 2026-07-28 | 26728 |
| 2026-07-30 | 26728 |
| 2026-07-31 | 26951 |
| 2026-08-07 | 27177 |
| 2026-08-16 | 27647 |
| 2026-08-18 | 27796 |
| 2026-08-19 | 27842 |
| 2026-08-20 | 27890 |
| 2026-08-21 | 27950 |
| 2026-08-22 | 27991 |
| 2026-08-23 | 28038 |
| 2026-08-24 | 28068 |
| 2026-08-25 | 28099 |
| 2026-08-26 | 28134 |
| 2026-08-27 | 28171 |
| 2026-08-28 | 28383 |
| 2026-08-29 | 28567 |
| 2026-08-30 | 28909 |
| 2026-08-31 | 29429 |
| 2026-09-01 | 29892 |
| 2026-09-02 | 30101 |
| 2026-09-03 | 30248 |
| 2026-09-04 | 30384 |
| 2026-09-05 | 30513 |
| 2026-09-06 | 30670 |
| 2026-09-07 | 30787 |
| 2026-09-08 | 30883 |
| 2026-09-09 | 30988 |
| 2026-09-10 | 31075 |
| 2026-09-11 | 31145 |
| 2026-09-12 | 31211 |
| 2026-09-13 | 31263 |
| 2026-09-14 | 31351 |
| 2026-09-15 | 31449 |
| 2026-09-16 | 31544 |
| 2026-09-17 | 31654 |
| 2026-09-18 | 31764 |
| 2026-09-19 | 31859 |
| 2026-09-20 | 31927 |
Today
+68 stars today
This week
+664 stars this week
This month
+4.0k stars this month
Momentum
941.0
growth rate 2.12%/day
<img width="128" align="right" alt="Logo" src="https://github.com/user-attachments/assets/df5f2840-2f92-4991-aa57-252747d7182e" /> # Heretic: Fully automatic censorship removal for language models<br><br>[](https://discord.gg/gdXc48gSyT) [](https://matrix.to/#/#heretic:matrix.org) [](https://huggingface.co/heretic-org) [](https://codeberg.org/p-e-w/heretic) [](https://trendshift.io/repositories/20538) Heretic is a tool that removes censorship (aka "safety alignment") from transformer-based language models without expensive post-training. It combines an advanced implementation of directional ablation, also known as "abliteration" ([Arditi et al. 2024](https://arxiv.org/abs/2406.11717), Lai 2025 ([1](https://huggingface.co/blog/grimjim/projected-abliteration), [2](https://huggingface.co/blog/grimjim/norm-preserving-biprojected-abliteration))), with a TPE-based parameter optimizer powered by [Optuna](https://optuna.org/). This approach enables Heretic to work **completely automatically.** Heretic finds high-quality abliteration parameters by co-minimizing the number of refusals and the KL divergence from the original model. This results in a decensored model that retains as much of the original model's intelligence as possible. Using Heretic does not require an understanding of transformer internals. In fact, anyone who knows how to run a command-line program can use Heretic to decensor language models. Heretic supports most dense models, including many multimodal models, several different MoE architectures, and even some hybrid models like Qwen3.5. Pure state-space models and certain other research architectures are not yet supported out of the box. <img width="650" height="715" alt="Screenshot" src="https://github.com/user-attachments/assets/d71a5efa-d6be-4705-a817-63332afb2d15" /> Running unsupervised with the default configuration, Heretic can produce decensored models that rival the quality of abliterations created manually by human experts: | Model | Refusals for "harmful" prompts | KL divergence from original model for "harmless" prompts | | :--- | ---: | ---: | | [google/gemma-3-12b-it](https://huggingface.co/google/gemma-3-12b-it) (original) | 97/100 | 0 *(by definition)* | | [mlabonne/gemma-3-12b-it-abliterated-v2](https://huggingface.co/mlabonne/gemma-3-12b-it-abliterated-v2) | 3/100 | 1.04 | | [huihui-ai/gemma-3-12b-it-abliterated](https://huggingface.co/huihui-ai/gemma-3-12b-it-abliterated) | 3/100 | 0.45 | | **[p-e-w/gemma-3-12b-it-heretic](https://huggingface.co/p-e-w/gemma-3-12b-it-heretic) (ours)** | **3/100** | **0.16** | The Heretic version, generated without any human effort, achieves the same level of refusal suppression as other abliterations, but at a much lower KL divergence, indicating less damage to the original model's capabilities. *(You can reproduce those numbers using Heretic's built-in evaluation functionality, e.g. `heretic --model google/gemma-3-12b-it --evaluate-model p-e-w/gemma-3-12b-it-heretic`. Note that the exact values might be platform- and hardware-dependent. The table above was compiled using PyTorch 2.8 on an RTX 5090.)* Of course, mathematical metrics and automated benchmarks never tell the whole story, and are no substitute for human evaluation. Models generated with Heretic have been well-received by users (links and emphasis added): > "I was skeptical before, but I just downloaded > [**GPT-OSS 20B Heretic**](https://huggingface.
Excerpt of 18,207 characters
Read on GitHubPhilipp Emanuel Weidmann
107
20
8
Spiky Moth
6
5
UmranPros
5
Rocker Zhang
3
Nikolai Kolodziej · Germany
3
Vinay Umrethe
3
michaelh
2
Salman Chishti · GitHub, ex-Microsoft · United Kingdom
2
2
Poland
2
Magic · Sayou
1
George
1
Darshan · India
1
Ashar · DigitalOcean
1
Arthur Wuhrmann · @EPFL
1
Brazil
1
zaakir · United Kingdom
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:9842909587f05944, topic:llm, topic:transformer