Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
| Date | Stars |
|---|---|
| 2026-07-31 | 275 |
| 2026-08-06 | 275 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<div align="center">
<img src="assets/logo.png" alt="Safety Logo" />
<a href="https://www.nowpublishers.com/article/Details/SEC-051" target="_blank"><img src="https://img.shields.io/badge/arXiv-b5212f.svg?logo=arxiv" alt="arXiv"></a>
</div>
## 🚀 About the Survey
This survey provides a systematic review of current safety research on large models, covering **Vision Foundation Models** (VFMs), **Large Language Models** (LLMs), **Vision-Language Pre-training** (VLP) models, **Vision-Language Models** (VLMs), **Diffusion Models** (DMs), and large-model-based **Agents**.
- We present a **comprehensive taxonomy of safety threats** to the above models.
- We review **defense strategies** (if available) proposed for each type of attacks as well as commonly used **datasets and benchmarks**.
- We identify and discuss the **open challenges** in large model safety.
---
## 🌴 Missing Papers
We will keep updating this survey. Please complete the following form to submit your paper for citation.
👉 [Submit Your Paper](https://forms.gle/SkFbVvZYw8r8cJJ17)
We appreciate your contributions and look forward to keeping this survey comprehensive and up to date!
⚠️ Note: We completed a major revision of the paper in August 2025. This version reflects our final planned update, and further substantial revisions may not be possible.
---
The survey have been structured with the following considerations for clarity and readability:
- **Models**. We focus on 6 widely studied model categories, including **VFMs**, **LLMs**, **VLPs**, **VLMs**, **DMs**, and **Agents**, and review the attack and defense methods for each separately. These models represent the most popular large models across various domains.
- **Organization**. For each model category, we classify the reviewed works into attacks and defenses, and identify **10** attack types: **adversarial**, **backdoor**, **poisoning**, **jailbreak**, **prompt injection**, **energy-latency**, **membership inference**, **model extraction**, **data extraction**, and **agent** attacks. When both backdoor and poisoning attacks are present for a model category, we combine them into a single **backdoor & poisoning** category due to their similarities. We review the corresponding defense strategies for each attack type immediately after the attacks.
- **Taxonomy**. For each type of attack or defense, we use a two-level taxonomy: **Category → Subcategory**. The **Category** differentiates attacks and defenses based on the threat model (e.g., white-box, gray-box, black-box) or specific subtasks (e.g., detection, purification, robust training/tuning, and robust inference). The **Subcategory** offers a more detailed classification based on their techniques.
- **Granularity**. To ensure clarity, we simplify the introduction of each reviewed paper, highlighting only its key ideas, objectives, and approaches, while omitting technical details and experimental analyses.
<div align="center">
<figure>
<img src="/assets/stats.png" alt="Paper stats" />
<figcaption>Figure 1: Left: The number of safety research papers published over the past four years. Middle: The distribution of research across
different models. Right: The distribution of research across different types of attacks and defenses. </figcaption>
</figure>
</div>
---
## 🏈 Survey Methodology
First, we conducted a keyword-based search targeting specific model types and threat types to identify relevant papers. Next, we manually filtered out non-safety-related and non-technical papers. For each remaining paper, we categorized its proposed method or framework by analyzing its settings and attack/defense types, assigning them to appropriate categories and subcategories. We collected **390** technical papers, with their distribution across years, model types, and attack/defense strategies illustrated in the following Figures. As shown, safety research on large models has surged significantly since 2023, following the release of Excerpt of 133,834 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:01fc42b6b222dc4e, llm:Repository title and description: 'Awesome-Large-Model-Safety' / 'Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety' — an 'Awesome' list/survey focused on safety of large models and agents.