CHATS-lab/persuasive_jailbreaker
quality grade C, 61 out of 100Persuasive Jailbreaker: we can persuade LLMs to jailbreak them!
- stars
- 366
- stars gained this week
- —
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Guardrails, PII redaction, prompt-injection defense, interpretability and offensive/defensive AI security.
Signals: ai-safety, ai-security, guardrails, prompt-injection, jailbreak, llm-security, interpretability, explainable-ai
374 results
Persuasive Jailbreaker: we can persuade LLMs to jailbreak them!
sharing NEW strong AI jailbreaks of multiple vendors (LLMs)
[ICML 2024] Binoculars: Zero-Shot Detection of LLM-Generated Text
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks [ICLR 2025]
No description
A template for building LLM-based AI text adventure games, with LLM prompt injection theme as an example story.
No description
Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models know themselves through automated interpretability.
[ICLR 2024] The official implementation of our ICLR2024 paper "AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models".
Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University
RuLES: a benchmark for evaluating rule-following in language models
⚡ Vigil ⚡ Detect prompt injections, jailbreaks, and other potentially risky Large Language Model (LLM) inputs
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
Repository for the Paper (AAAI 2024, Oral) --- Visual Adversarial Examples Jailbreak Large Language Models
This repository collects all relevant resources about interpretability in LLMs
This is The most comprehensive prompt hacking course available, which record our progress on a prompt engineering and prompt hacking course.
Security toolkit for AI agents. Scan your machine for dangerous skills and MCP configs, monitor for supply chain attacks, test prompt injection resistance, and audit live MCP servers for tool poisoning.
Collection of agent skills to find vulnerabilities inside your web/mobile apps.
Hack Claude Code to get any buddy you want
The free build of Claude Code. All telemetry removed, security-prompt guardrails stripped, all experimental features enabled.
2026.3.31 claude code 意外把包含源码的文件上传到 npm 仓库,版本号是 2.1.88,其中 cli.js.map 文件有 57MB 的体积,claude code 的源码在该文件的 sourcesContent 字段里面,解压还原后有 70w 行代码
Hacky repo to see what the Copilot extension sends to the server
Microsoft Security Copilot is a generative AI-powered security solution that helps increase the efficiency and capabilities of defenders to improve security outcomes at machine speed and scale, while remaining compliant to responsible AI principles
A comprehensive security checklist for MCP-based AI tools. Built by SlowMist to safeguard LLM plugin ecosystems.
24,535 repositories in the index in total.