Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Awesome-Jailbreak-on-LLMs is a collection of state-of-the-art, novel, exciting jailbreak methods on LLMs. It contains papers, codes, datasets, evaluations, and analyses.
| Date | Stars |
|---|---|
| 2026-07-24 | 1541 |
| 2026-07-25 | 1542 |
| 2026-07-28 | 1549 |
| 2026-07-30 | 1549 |
| 2026-07-31 | 1550 |
| 2026-08-06 | 1550 |
Today
— stars today
This week
+1 stars this week
This month
— stars this month
Momentum
1.0
growth rate 0.07%/day
# Awesome-Jailbreak-on-LLMs Awesome-Jailbreak-on-LLMs is a collection of state-of-the-art, novel, exciting jailbreak methods on LLMs. It contains papers, codes, datasets, evaluations, and analyses. Any additional things regarding jailbreak, PRs, issues are welcome and we are glad to add you to the contributor list [here](#contributors). Any problems, please contact [email protected]. If you find this repository useful to your research or work, it is really appreciated to star this repository and cite our papers [here](#Reference). :sparkles: ## Reference If you find this repository helpful for your research, we would greatly appreciate it if you could cite our papers. :sparkles: ``` @article{zhuzhenhao_GuardReasoner_Omni, title={GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, and Video}, author={Zhu, Zhenhao and Liu, Yue and Guo, Yanpei and Qu, Wenjie and Chen, Cancan and He, Yufei and Li, Yibo and Chen, Yulin and Wu, Tianyi and Xu, Huiying and others}, journal={arXiv preprint arXiv:2602.03328}, year={2026} } @article{liuyue_GuardReasoner_VL, title={GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning}, author={Liu, Yue and Zhai, Shengfang and Du, Mingzhe and Chen, Yulin and Cao, Tri and Gao, Hongcheng and Wang, Cheng and Li, Xinfeng and Wang, Kun and Fang, Junfeng and Zhang, Jiaheng and Hooi, Bryan}, journal={arXiv preprint arXiv:2505.11049}, year={2025} } @article{liuyue_GuardReasoner, title={GuardReasoner: Towards Reasoning-based LLM Safeguards}, author={Liu, Yue and Gao, Hongcheng and Zhai, Shengfang and Jun, Xia and Wu, Tianyi and Xue, Zhiwei and Chen, Yulin and Kawaguchi, Kenji and Zhang, Jiaheng and Hooi, Bryan}, journal={arXiv preprint arXiv:2501.18492}, year={2025} } @article{liuyue_FlipAttack, title={FlipAttack: Jailbreak LLMs via Flipping}, author={Liu, Yue and He, Xiaoxin and Xiong, Miao and Fu, Jinlan and Deng, Shumin and Hooi, Bryan}, journal={arXiv preprint arXiv:2410.02832}, year={2024} } @article{wang2025safety, title={Safety in Large Reasoning Models: A Survey}, author={Wang, Cheng and Liu, Yue and Li, Baolong and Zhang, Duzhen and Li, Zhongzhi and Fang, Junfeng}, journal={arXiv preprint arXiv:2504.17704}, year={2025} } ``` ## Bookmarks - [Jailbreak Attack](#jailbreak-attack) - [Attack on LRMs](#attack-on-lrms) - [Black-box Attack](#black-box-attack) - [White-box Attack](#white-box-attack) - [Multi-turn Attack](#multi-turn-attack) - [Attack on RAG-based LLM](#attack-on-rag-based-llm) - [Multi-modal Attack](#multi-modal-attack) - [Jailbreak Defense](#jailbreak-defense) - [Learning-based Defense](#learning-based-defense) - [Strategy-based Defense](#strategy-based-defense) - [Guard Model](#Guard-model) - [Moderation API](#Moderation-API) - [Evaluation & Analysis](#evaluation--analysis) - [Application](#application) ## Papers ### Jailbreak Attack #### Attack on LRMs | Time | Title | Venue | Paper | Code | | ------- | ------------------------------------------------------------ | :---: | :--------------------------------------: | :----------------------------------------------------------: | | 2026.05 | **Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs** | arXiv | [link](https://arxiv.org/abs/2605.24497) | - | | 2025.11 | **BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models** | AAAI'26 | [link](https://arxiv.org/abs/2511.10714) | - | | 2025.10 | **When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models** | arXiv | [link](https://arxiv.org/abs/2510.21285) | - | | 2025.10 | **AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling** | arXiv | [link](https://arxiv.org/abs/2510.05379) | [link](
Excerpt of 114,058 characters
Read on GitHubyueliu1999 · National University of Singapore · Singapore
207
15
6
6
5
4
Xiangyan Liu · Singapore
4
4
3
2
2
2
2
2
2
1
Oscar Wu · Australia
1
DongGeon Lee · @K-intelligence-Midm
1
Jiahao Zhao · Postdoc at CASIA · China
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:138b0f41da6d3227, topic:jailbreak, topic:privacy, name:jailbreak