Shiva108/ai-llm-red-team-handbook
quality grade D, 47 out of 100AI / LLM Red Team Field Manual & Consultant’s Handbook
- stars
- 320
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Guardrails, PII redaction, prompt-injection defense, interpretability and offensive/defensive AI security.
Signals: ai-safety, ai-security, guardrails, prompt-injection, jailbreak, llm-security, interpretability, explainable-ai
381 results
AI / LLM Red Team Field Manual & Consultant’s Handbook
AI-native automated software risk analysis skill. LLM-driven, Code-First approach for comprehensive security risk assessment, threat modeling, security testing, penetration testing, and compliance checking.
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours
Open-source LLM red-teaming technique toolkit (162 transforms, 36 mutators, 25 tool surfaces). MIT.
Lightweight LLM firewall: masks PII, routes calls, leaves zero trace.
Persuasive Jailbreaker: we can persuade LLMs to jailbreak them!
This repository provides a benchmark for prompt injection attacks and defenses in LLMs
sharing NEW strong AI jailbreaks of multiple vendors (LLMs)
[ICML 2024] Binoculars: Zero-Shot Detection of LLM-Generated Text
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks [ICLR 2025]
No description
A template for building LLM-based AI text adventure games, with LLM prompt injection theme as an example story.
No description
[ICLR 2024] The official implementation of our ICLR2024 paper "AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models".
Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University
RuLES: a benchmark for evaluating rule-following in language models
[NeurIPS 2025] BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
⚡ Vigil ⚡ Detect prompt injections, jailbreaks, and other potentially risky Large Language Model (LLM) inputs
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
Repository for the Paper (AAAI 2024, Oral) --- Visual Adversarial Examples Jailbreak Large Language Models
This repository collects all relevant resources about interpretability in LLMs
This is The most comprehensive prompt hacking course available, which record our progress on a prompt engineering and prompt hacking course.
A field guide to the visual & copy tics of AI-generated products — and an Agent Skill that scans your project and strips them out. https://killaislop.com
Security toolkit for AI agents. Scan your machine for dangerous skills and MCP configs, monitor for supply chain attacks, test prompt injection resistance, and audit live MCP servers for tool poisoning.
24,523 repositories in the index in total.