This repository curates the latest research papers on the applications and architectural technologies of AI agents. We perform weekly Arxiv searches using specific keywords and pick only those that are particularly interesting. Rather than striving for comprehensiveness, we add papers when they introduce a distinctively new approach or novel concept that stands out from existing methods.
An AI Agent is an autonomous system powered by large language models that can perceive its environment, reason through complex tasks, and use tools to take actions in pursuit of specific goals. It combines reasoning, planning, memory, and tool-use capabilities to operate independently or as part of a multi-agent system.
AI Agent WorkflowsPapers are filed in four layers — capabilities (what an agent can do), architecture (how it's built), operations (how it's run), and applications (where it's used). Each entry links to a curated, date-ordered reading list.
🔥: Recommended papers
📖: Survey papers
⚖️: Benchmark papers
🔄 Badges show papers added in the last 2 months (Aug–Sep 2026); cluster headings show the sum:
(+N)recent additions, 🔥 = high activity. Regenerate:python scripts/update_readme_badges.py.📂 See TAXONOMY.md for the full directory map and the rules for where each paper is filed.
- Agent Capabilities — what an agent can do: cognition, knowledge, action, and how it learns & improves
- Core Cognition — reason, plan, ideate, perceive (+3)
- Reasoning (+1)
- Planning
- Ideation (+2)
- Perception
- Knowledge & Context — what the agent knows and carries between steps (+16) 🔥
- Memory (+6)
- Context Engineering (+5)
- Knowledge Graphs & Ontology (+5)
- Action — acting through tools and reusable skills (+19) 🔥
- Adaptation & Self-Improvement — getting better from experience (+59) 🔥
- Exploration & Discovery (+5)
- Experience & Trajectory Learning (+4)
- Failure Attribution & Error Localization (+10) 🔥
- Self-Correction (+1)
- Verification (+2)
- Self-Evolution (+27) 🔥
- Agent Tuning (+10) 🔥
- Trust & Measurement — is it safe, and how well does it work? (+42) 🔥
- Safety (+17) 🔥
- Agent Evaluation (+25) 🔥
- Other — world models, user profiles, forecasting (+11) 🔥
- Environment (World Models & Simulations) (+5)
- Profile (+2)
- Prediction (+4)
- Core Cognition — reason, plan, ideate, perceive (+3)
- AI Agents Architecture — how an agent is built: single-agent design, multi-agent systems, and the runtime harness (+34) 🔥
- Agent Design & Frameworks (+1)
- Multi-Agent Systems (+3)
- Harness (+30) 🔥
- Operations & Interaction — how an agent is run and works with people: observability, UX, and governance (+22) 🔥
- AI Agents Applications — where agents are deployed, grouped by interface, domain, and task pattern
- By Embodiment / Interface — where the agent acts (+1)
- By Domain / Vertical — the industry it serves (+33) 🔥
- Financial Agents (+6)
- Enterprise Agents (+5)
- AI Scientist (Research Automation) (+13) 🔥
- Vertical / Domain Agents (+9)
- By System Pattern / Task Form — the shape of the task or system (+26) 🔥
- Coding Agents (+10) 🔥
- Data Agents (+8)
- Deep Research Agents (+1)
- World Simulation (+7)
- GenAI Agents Presentations
カテゴリ別の月次トレンド深掘り。2026-06 以降は各論文の arXiv HTML 本文を精読し、図を引用、複数論文で裏付けたファクトを中心にまとめています(作成手順は .claude/skills/newsletter)。
2026-08
- Harness · Safety · Agent Evaluation · Self-Evolution · Skills · Failure Attribution · Agent Tuning · Governance
2026-07
2026-06
2026-05
2026-04
2026-02
2026-01
Monthly curated picks (a handful of standout papers per month) are archived under highlights/:
