Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Persuasive Jailbreaker: we can persuade LLMs to jailbreak them!
| Date | Stars |
|---|---|
| 2026-07-31 | 363 |
| 2026-08-06 | 363 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<h1 align='center' style="text-align:center; font-weight:bold; font-size:2.0em;letter-spacing:2.0px;"> How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs </h1>
<p align='center' style="text-align:center;font-size:1.25em;">
<a href="https://www.yi-zeng.com/" target="_blank" style="text-decoration: none;">Yi Zeng<sup>1,*</sup></a> ,
<a href="https://hopelin99.github.io/" target="_blank" style="text-decoration: none;">Hongpeng Lin<sup>2,*</sup></a> ,
<a href="https://communication.ucdavis.edu/people/jingwen-zhang" target="_blank" style="text-decoration: none;">Jingwen Zhang<sup>3</sup></a><br>
<a href="https://cs.stanford.edu/~diyiy/" target="_blank" style="text-decoration: none;">Diyi Yang<sup>4</sup></a> ,
<a href="https://ruoxijia.info/" target="_blank" style="text-decoration: none;">Ruoxi Jia<sup>1,†</sup></a> ,
<a href="https://wyshi.github.io/" target="_blank" style="text-decoration: none;">Weiyan Shi<sup>4,†</sup></a>
<br/>
<sup>1</sup>Virginia Tech <sup>2</sup>Renmin University of China <sup>3</sup>UC, Davis <sup>4</sup>Stanford University<br>
<sup>*</sup>Lead Authors <sup>†</sup>Equal Advising<br/>
</p>
<p align='center';>
<b>
<em>arXiv-Preprint, 2024</em> <br>
</b>
</p>
<p align='center' style="text-align:center;font-size:2.5 em;">
<b>
<a href="https://arxiv.org/abs/2401.06373" target="_blank" style="text-decoration: none;">[arXiv]</a> <a href="https://chats-lab.github.io/persuasive_jailbreaker/" target="_blank" style="text-decoration: none;">[Project Page]</a>
</b>
</p>
## Important update [Oct 9th, 2024] 🚀
Since we disclosed our results earlier, the attack is less effective now after several months, so we decided to release the data in a more open format.
🔍 **What's New?**
The data is now on [huggingface](https://huggingface.co/datasets/CHATS-Lab/Persuasive-Jailbreaker-Data)!
------------
## Important update [April 2nd, 2024] 🚀
We share an alternative method for generating PAPs that eliminates the need to access harmful PAP examples and relies on fine-tuned GPT-3.5.
🔍 **What's New?**
The core of our update lies in the new directory: ```/PAP_Better_Incontext_Sample```.
📚 **How to Use?**
Dive into the ```/PAP_Better_Incontext_Sample``` folder and explore ```test.ipynb``` to begin. This example will walk you through the process of sampling high quality PAPs of the Top-5 persuasive techniques.
## Reproducibility and Codes
For safety concerns, in this repository we only release the persuasion taxonomy and the code for in-context sampling described in our paper. `persuasion_taxonomy.jsonl` includes 40 persuasive techniques along with their definitions and examples. `incontext_sampling_example.ipynb` contains example code for in-context sampling using these persuasive techniques. These techniques and codes can be used to generate Persuasive Adversarial Prompts(PAPs) or for other persuasion tasks.
To train a persuasive paraphraser, researchers can generate questions or use existing ones, employ `incontext_sampling_example.ipynb` for persuasion/attack. Subsequently, the results of these samplings can be evaluated either through manual annotation or by using [GPT-4 Judge](https://llm-tuning-safety.github.io/index.html), thereby generating data suitable for training.
Responsibly, we choose not to publicly release the complete attack code. However, **for safety studies,** researchers can access the data [huggingface](https://huggingface.co/datasets/CHATS-Lab/Persuasive-Jailbreaker-Data)! After signing the release form, you will get the jailbreak data on the [advbench](https://llm-attacks.org/) sub-dataset(refined by [Chao et al.](https://github.com/patrickrchao/JailbreakingLLMs)) to the applicants. Access to the Software is granted on a provisional Excerpt of 13,809 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:8d4a1982f134a2c9, llm:Repository name and description: 'Persuasive Jailbreaker: we can persuade LLMs to jailbreak them!' — indicates techniques for jailbreaking large language models (LLMs).
matched fp:8d4a1982f134a2c9, llm:Repository name and description: 'Persuasive Jailbreaker: we can persuade LLMs to jailbreak them!' — indicates techniques for jailbreaking large language models (LLMs).