Awesome Deep Research Agent
We maintain a curated collection of papers exploring the path towards Deep Research (DR) Agents , focusing on formulating core concepts and mapping the research landscape.
⌛️ We’re continuously compiling and updating cutting‑edge insights. Feel free to suggest any related work you find valuable!
Build a digital assistant on your screen. Generated by DALL-E-3 .
🔥 WELCOME CONTRIBUTE!
🔥 This project is actively maintained, and we welcome your contributions. If you have any suggestions, such as missing papers or information, please feel free to open an issue or submit a pull request.
[2025.09.03] 📄 Updated version of our paper "Deep Research Agents: A Systematic Examination And Roadmap" is now available on ArXiv with expanded analysis and refined research directions!
[2025.06.18] 🚀 Our comprehensive survey "Deep Research Agents: A Systematic Examination And Roadmap" is officially released on ArXiv - providing systematic insights into the current state and future of deepresearch agents!
🏗️ Our Works Towards DR Agents
✨✨✨ Deep Research Agents: A Systematic Examination And Roadmap
Table of Contents
Search Engine Integration
Tool Use
Architecture & Workflow
Tuning Methods
Industrial Applications
Benchmarks for DR Agents
Search Engine Integration
📊 Search Engine · API vs Browser Comparison
Legend
✔️ Primary focus
🟫 Secondary/minor focus
— Not present
DR Agent
API
Browser
GAIA
HLE
QA
Base Model
Avatar
🟫
—
—
—
Stark
Claude-3-Opus, GPT-4
CoSearch-Agent
✔️
—
—
—
—
GPT-3.5-turbo
MMAC-Copilot
✔️
—
✔️
—
—
GPT-3.5, GPT-4
Storm
🟫
—
—
—
FreshWiki
GPT-3.5-turbo
OpenResearcher
✔️
—
—
—
Private QA
DeepSeek-V2-Chat
The AI Scientist
✔️
—
—
—
MLE-Bench
GPT-4o, o1-mini, o1-preview
Gemini DR
✔️
✔️
—
✔️
GPQA
Gemini-2.0-Flash
Agent Laboratory
✔️
—
—
—
MLE-Bench
GPT-4o, o1-preview
Search-o1
✔️
—
—
—
GPQA·NQ·TriviaQA
QwQ-32B-preview
WebWalker
—
—
—
—
WebWalkerQA
GPT-4o, Qwen-2.5
Agentic Reasoning
✔️
—
—
—
GPQA
DeepSeek-R1, Qwen2.5
AutoAgent
—
✔️
✔️
—
—
Claude-Sonnet-3.5
Grok DeepSearch
✔️
✔️
—
—
GPQA
Grok 3
OpenAI DR
—
✔️
✔️
✔️
✔️
GPT-o3
Perplexity DR
✔️
🟫
—
✔️
SimoleQA
Flexible
Towards an AI Co-Scientist
✔️
—
—
—
GPQA
Gemini 2.0
Nouswise
—
—
—
—
—
—
AgentRxiv
✔️
—
—
—
GPQA·MedQA
GPT-4o-mini
Agent-R1
✔️
—
—
—
HotpotQA
Qwen2.5-1.5B-Inst
AutoGLM Rumination
—
✔️
—
—
GPQA
GLM-Z1-Air
Copilot Researcher
—
✔️
—
—
—
o3-mini
H2O.ai DR
✔️
✔️
✔️
—
—
h2ogpt-oasst1-512-12b
Manus
✔️
✔️
—
—
—
Claude3.5, GPT-4o
OpenManus
✔️
✔️
—
—
—
Claude3.5, GPT-4o
OWL
✔️
✔️
✔️
—
—
DeepSeek-R1, Gemini-2.5-Pro, GPT-4o
R1-Searcher
🟫
—
—
—
2WikiMultiHopQA, HotpotQA
Llama3.1-8B-Inst, Qwen2.5-7B
ReSearch
🟫
—
—
—
2WikiMultiHopQA, HotpotQA
Qwen2.5-7B, Qwen2.5-7B-Inst
Search-R1
🟫
—
—
—
2WikiMultiHopQA, HotpotQA, NQ, TriviaQA
Llama3.2-3B, Qwen2.5-3B/7B
DeepResearcher
—
✔️
—
—
HotpotQA, NQ, TriviaQA
Qwen2.5-7B-Inst
Genspark Super Agent
✔️
✔️
✔️
—
—
Mixture of 9 LLMs
WebThinker
✔️
—
✔️
✔️
GPQA, WebWalkerQA
QwQ-32B
SWIRL
—
✔️
—
—
HotQA, BeerQA
Gemma 2-27B
SimpleDeepSearcher
—
✔️
✔️
—
2WikiMultiHopQA
Qwen-2.5-7B/32B-In, DeepSeek-D-Qwen-2.5-32B, QwQ-32B
Suna AI
✔️
✔️
—
—
—
GPT-4o, Claude
AgenticSeek
—
✔️
—
—
—
GPT-4o, DeepSeek-R1, Claude
Alita
✔️
✔️
✔️
—
PathVQA
GPT-4o, Claude-Sonnet-4
DeerFlow
✔️
—
—
—
—
Doubao-1.5-Pro-32k, DeepSeek-R1, GPT-4o, Qwen
PANGU DEEPDIVER
✔️
—
—
—
C-SimpleQA, HotpotQA, ProxyQA
Pangu-7B-Reasoner
WebDancer
✔️
—
✔️
—
GAIA, WebWalkerQA
Qwen-2.5, QwQ-32B, DeepSeek-R1, GPT-4o
O-agents
✔️
—
✔️
—
—
GPT-4o, GPT-4.1, Claude-3.7-Sonnet, DeepSeek-R1, Gemini-2.5
Kimi-Researcher
✔️
✔️
—
✔️
SimpleQA
Kimi k1.5/k2
WebSailor
✔️
—
✔️
—
SimpleQA
Qwen-2.5
Agent-KB
✔️
—
✔️
—
SWE-bench
GPT-4o, GPT-4.1, Claude-3.7-Sonnet, o3-mini, Qwen-3, DeepSeek-R1
WebShaper
✔️
—
✔️
—
WebWalkerQA
Qwen-2.5, QwQ-32B
Deep Researcher with Test-Time Diffusion
✔️
—
✔️
✔️
—
Gemini-2.5-Pro
ChatGPT-Agent
—
—
—
—
—
—
AWorld
✔️
✔️
✔️
—
HotpotQA
Gemini-2.5-Pro, GPT-4o
Cognitive Kernel-Pro
✔️
✔️
✔️
—
AgentWebQA, WebWalkerQA, Multi-hop URLQA, DocBench, TableBench
Claude-3.7-Sonnet, CK-Pro-8B
WebWatcher
✔️
—
—
✔️
Browsercom-VL, LiveVQA, MMSearch
Qwen-2.5-VL-32B
WideSearch
✔️
—
—
—
WideSearch
DeepSeek-R1, Doubao-Seed-1.6, Claude Sonnet 4, Gemini-2.5-Pro
MiroRL
✔️
—
✔️
—
—
Qwen3-14B
📊 Tool Use Capabilities Comparison
Legend
✔️ Involved
🟫 Non Disclosure
— Not present
DR Agent
Code Interp.
Data Analytics
Multimodal
Release
CoSearchAgent
—
✔️
—
Feb-2024
Storm
✔️
—
—
Jul-2024
The AI Scientist
✔️
—
—
Aug-2024
Agent Laboratory
✔️
—
—
Jan-2025
Agentic Reasoning
✔️
—
—
Feb-2025
AutoAgent
✔️
—
✔️
Feb-2025
Genspark DR
✔️
✔️
✔️
Feb-2025
Grok DeepSearch
✔️
✔️
✔️
Feb-2025
OpenAI DR
✔️
✔️
✔️
Feb-2025
Perplexity DR
✔️
✔️
✔️
Feb-2025
Towards an AI co-scientist
—
✔️
✔️
Feb-2025
Agent-R1
✔️
—
—
Mar-2025
AutoGLM Romination
✔️
—
✔️
Mar-2025
Copilot Researcher
✔️
✔️
🟫
Mar-2025
Manus
✔️
✔️
✔️
Mar-2025
OpenManus
✔️
✔️
—
Mar-2025
OWL
✔️
✔️
✔️
Mar-2025
H2O.ai DR
✔️
✔️
✔️
Mar-2025
Genspark Super Agent
✔️
✔️
✔️
Apr-2025
WebThinker
✔️
—
—
Apr-2025
Suna Ai
✔️
✔️
—
Apr-2025
Tool-Star
✔️
✔️
—
May-2025
AgenticSeek
✔️
✔️
—
May-2025
Alita
✔️
🟫
🟫
May-2025
DeerFlow
✔️
✔️
—
May-2025
O-agents
✔️
✔️
✔️
Jun-2025
Kimi-Researcher
✔️
✔️
—
Jun-2025
Agent-KB
✔️
✔️
✔️
Jul-2025
AWorld
✔️
✔️
✔️
Jul-2025
Cognitive Kernel-Pro
✔️
✔️
✔️
Aug-2025
WebWatcher
✔️
✔️
✔️
Aug-2025
MiroRL
✔️
✔️
—
Aug-2025
Dynamic Single‑Agent Workflow
Dynamic Multi‑Agent Workflow
📊 Tuning Methods Comparison
Legend
✔️ Implemented
🟫 Details Unknown
— Not present
DR Agent
SFT
RL
Base Model
Data
Reward Design
Gemini DR
🟫
🟫
Gemini-2.0-Flash
—
🟫
WebWalker
—
—
GPT-4o, Qwen-2.5 (7–72B)
WebWalkerQA
—
Grok DeepSearch
—
🟫
Grok 3
—
🟫
OpenAI DR
—
🟫
GPT-o3
—
🟫
Agentic Reasoning
✔️
—
DeepSeek-R1, Qwen2.5
GPQA
Rule-Outcome
AutoAgent
—
✔️
Claude-Sonnet-3.5
—
—
Towards an AI co-scientist
—
—
Gemini 2.0
—
—
Agent-R1
—
PPO · Reinforce++ · GRPO
Qwen2.5-1.5B-Inst
HotpotQA
Rule-Outcome
AutoGLM Rumination
🟫
🟫
GLM-Z1-Air
—
🟫
H2O.ai DR
✔️
🟫
h2ogpt-oasst1-512-12b
—
🟫
Copilot Researcher
🟫
🟫
o3-mini
—
—
ReSearch
—
GRPO
Qwen2.5-7B-Inst · Qwen2.5-32B-Inst
2WikiMultiHopQA
Rule-Outcome
R1-Searcher
✔️
Reinforce++ · GRPO
Qwen2.5-7B-Inst / LLaMA-3.1-8B-Inst
2WikiMultiHopQA · HotpotQA
Rule-Outcome
Search-R1
✔️
PPO · GRPO
Qwen2.5-3B/7B · LLaMA3.2-3B-Inst
NQ · HotpotQA
Rule-Outcome
Nouswise
🟫
🟫
Nouswise
—
🟫
DeepResearcher
—
GRPO
Qwen2.5-7B-Inst
NQ · HotpotQA
Rule-Outcome
Genspark Super Agent
—
🟫
Mixture of Agents
—
🟫
WebThinker
✔️
Iterative Online DPO
QwQ-32B
Expert Dataset
Rule-Outcome
SWIRL
—
Offline-RL
Gemma 2-27B
HotPotQA
—
SimpleDeepSearcher
✔️
PPO
Qwen-2.5-7B-In · Qwen-2.5-32B-In · Deepseek-Distilled-Qwen-32B · QwQ-32B
NQ · HotpotQA · 2WikiMultiHopQA · Musique · SimpleQA · MultiHop-RAG
Process-based reward
PANGU DEEPDIVER
✔️
GRPO
Pangu-7B-Reasoner
WebPuzzle
Rule-Outcome
Tool-Star
✔️
GRPO
Qwen-2.5
NuminaMath · HotpotQA · 2WikiMultiHopQA
Rule-Outcome
WebDancer
✔️
DAPO
Qwen-2.5-7B/32B · QwQ-32B · DeepSeek-R1 · GPT-4o
CRAWLQA · E2HQA
Rule-Outcome
O-agents
—
—
GPT-4o · GPT-4.1 · Claude-3.7-Sonnet · DeepSeek-R1 · Gemini-2.5
—
—
Kimi-Researcher
—
REINFORCE
Kimi k1.5/k2
—
Rule-Outcome
WebSailor
✔️
DUPO
Qwen-2.5-3B/7B/32B/72B
SailorFog-QA
Rule-Outcome
Agent-KB
—
—
GPT-4o · GPT-4.1 · Claude-3.7-Sonnet · o3-mini · Qwen-3 32B · DeepSeek-R1
—
—
WebShaper
✔️
GRPO
Qwen-2.5-3B/7B/32B/72B · QwQ-32B
WebShaper
Rule-Outcome
Cognitive Kernel-Pro
✔️
—
Claude-3.7-Sonnet · CK-Pro-8B
OpenWebVoyager · Multi-hop URLQA · AgentWebQA · WebWalkerQA · DocBench · TableBench
—
WebWatcher
—
GRPO
Qwen-2.5-VL-32B
BrowseComp-VL · Long-tail VQA · Hard VQA
Rule-Outcome
MiroRL
✔️
GRPO
Qwen3-14B
MiroRL-GenQA
Rule-Outcome
Atom-Searcher
✔️
GRPO
Qwen2.5-7B-Inst
2WikiMultiHopQA · HotpotQA
Atomic Thought Reward (ATR)
📊 QA Benchmarks (Hotpot / 2Wiki / NQ / TQ / GPQA)
DR Agent
Base Model
Hotpot
2Wiki
NQ
TQ
GPQA
Release
Search-o1
QwQ-32B-preview
57.3
71.4
49.7
74.1
57.9
Jan-2025
Agentic Reasoning
DeepSeek-R1, Qwen2.5
—
—
—
—
67.0
Feb-2025
Grok DeepSearch
Grok3
—
—
—
—
84.6
Feb-2025
AgentRxiv
GPT-4o-mini
—
—
—
—
41.0
Mar-2025
R1-Searcher
Qwen2.5-7B-Base
71.9
63.8
—
—
—
Mar-2025
ReSearch
Qwen2.5-32B-Inst
67.7
50.0
—
—
—
Mar-2025
Search-R1
Qwen2.5-7B-Inst
34.5
36.9
40.9
55.2
—
Mar-2025
DeepResearcher
Qwen2.5-7B-Inst
64.3
66.6
61.9
85.0
—
Apr-2025
WebThinker
QwQ-32B
—
—
—
—
68.7
Apr-2025
SimpleDeepSearch
QwQ-32B
73.5
—
—
—
—
Apr-2025
SWIRL
Gemma 2-27B
72.0
—
—
—
—
Apr-2025
Tool-Star
Qwen2.5-3B
51.9
40.0
—
—
—
May-2025
📊 GAIA (Test and Val) Benchmarks
DR Agent
Base Model
GAIA L-1
L-2
L-3
Ave.
Release
Split
MMAC-Copilot
GPT-3.5, GPT-4
45.16
20.75
6.12
25.91
Mar-2024
Test
H2O.ai DR
Claude-3.7-Sonnet
89.25
79.87
61.22
79.73
Mar-2025
Test
Alita
Claude-Sonnet-4, GPT-4o
92.47
71.70
55.10
75.42
May-2025
Test
Agent-KB
GPT-4.1, Claude-3.7
84.91
74.42
57.69
75.15
Jul-2025
Test
O-agents
Claude-3.7
83.02
74.42
53.85
73.93
Jun-2025
Test
WebDancer
QwQ-32B
61.5
50.0
25.0
51.5
May-2025
Test
WebShaper
Qwen-2.5-72B
69.2
63.4
16.6
60.1
Jul-2025
Test
Deep Researcher w/ Test-Time Diffusion
Gemini-2.5-Pro
—
—
—
69.1
Jul-2025
Test
Cognitive Kernel-Pro
Claude-3.7-Sonnet
83.02
68.60
53.85
70.91
Aug-2025
Test
AutoAgent
Claude-Sonnet-3.5
71.7
53.5
26.9
55.2
Feb-2025
Dev
OpenAI DR
GPT-o3-customized
78.7
73.2
58.0
67.4
Feb-2025
Dev
Manus
Claude 3.5, GPT-4o
86.5
70.1
57.7
71.4
Mar-2025
Dev
OWL
Claude-3.7-Sonnet
84.9
68.6
42.3
69.7
Mar-2025
Dev
H2O.ai DR
h2ogpt-oasst1-512-12b
67.92
67.44
42.31
63.64
Mar-2025
Dev
Genspark Super Agent
Claude 3 Opus
87.8
72.7
58.8
73.1
Apr-2025
Dev
WebThinker
QwQ-32B
53.8
44.2
16.7
44.7
Apr-2025
Dev
SimpleDeepSearch
QwQ-32B
50.5
45.8
13.8
43.9
Apr-2025
Dev
Alita
Claude-Sonnet-4, GPT-4o
75.15
—
87.27
—
May-2025
Dev
If you find this work helpful, please cite our paper:
@article {huang2025deep ,
title ={ Deep Research Agents: A Systematic Examination And Roadmap} ,
author ={ Huang, Yuxuan and Chen, Yihang and Zhang, Haozheng and Li, Kang and Fang, Meng and Yang, Linyi and Li, Xiaoguang and Shang, Lifeng and Xu, Songcen and Hao, Jianye and others} ,
journal ={ arXiv preprint arXiv:2506.18096} ,
year ={ 2025}
}