Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[EMNLP2025] From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
| Date | Stars |
|---|---|
| 2026-07-31 | 424 |
| 2026-08-06 | 426 |
Today
+2 stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome LLM Scientific Discovery [](https://awesome.re)
A curated list of pioneering research papers, tools, and resources at the intersection of Large Language Models (LLMs) and Scientific Discovery.
Survey: ***From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery.*** ([https://arxiv.org/abs/2505.13259])
The survey delineates the evolving role of LLMs in science through a three-level autonomy framework:
* **Level 1: LLM as Tool:** LLMs augmenting human researchers for specific, well-defined tasks.
* **Level 2: LLM as Analyst:** LLMs exhibiting greater autonomy in processing complex information and offering insights.
* **Level 3: LLM as Scientist:** LLM-based systems autonomously conducting major research stages.
Below is a visual representation of this taxonomy:

We aim to provide a comprehensive overview for researchers, developers, and enthusiasts interested in this rapidly advancing field.
> **Last major update: 2026.07.** This refresh adds a large batch of 2025–2026 papers and a dedicated section on frontier industry-lab systems (Google DeepMind, OpenAI, Microsoft Research, Meta FAIR, FutureHouse, Sakana AI, and others). Contributions and PRs are very welcome — see [Contributing](#contributing).
## Contents
* [Level 1: LLM as Tool](#level-1-llm-as-tool)
* [Literature Review and Information Gathering](#literature-review-and-information-gathering)
* [Idea Generation and Hypothesis Formulation](#idea-generation-and-hypothesis-formulation)
* [Experiment Planning and Execution](#experiment-planning-and-execution)
* [Data Analysis and Organization](#data-analysis-and-organization)
* [Conclusion and Hypothesis Validation](#conclusion-and-hypothesis-validation)
* [Iteration and Refinement](#iteration-and-refinement)
* [Level 2: LLM as Analyst](#level-2-llm-as-analyst)
* [Machine Learning Research](#machine-learning-research)
* [Data Modeling and Analysis](#data-modeling-and-analysis)
* [Function Discovery](#function-discovery)
* [Natural Science Research](#natural-science-research)
* [General Research](#general-research)
* [Survey Generation](#survey-generation)
* [Level 3: LLM as Scientist](#level-3-llm-as-scientist)
* [General-Purpose Autonomous Research Agents](#general-purpose-autonomous-research-agents)
* [Discovery-Oriented Scientific Systems](#discovery-oriented-scientific-systems)
* [Autonomous Research Ecosystems and Infrastructure](#autonomous-research-ecosystems-and-infrastructure)
* [Frontier Labs and Foundation Models for Science](#frontier-labs-and-foundation-models-for-science)
* [Other Related Works](#other-related-works)
* [Contributing](#contributing)
---
## Level 1: LLM as Tool
At this foundational level, LLMs function as tailored tools under direct human supervision, designed to execute specific, well-defined tasks within a single stage of the scientific method. Their primary goal is to enhance researcher efficiency.
### Literature Review and Information Gathering
Automating literature search, retrieval, synthesis, structuring, and organization.
* **SCIMON : Scientific Inspiration Machines Optimized for Novelty** [](https://arxiv.org/pdf/2305.14259) - *Wang et al. (2023.05)*
* **ResearchAgent: Iterative research idea generation over scientific literature with Large Language Models** [](https://arxiv.org/pdf/2404.07738) - *Baek et al. (2024.04)*
* **Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction** [](https://arxiv.org/pdf/2404.14215) - *Deng et al. (2024.04)*
* **TKGT: Redefinition and A New Way of teExcerpt of 53,430 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:8d07a7571aa49a26, llm:Repository title and description: '[EMNLP2025] From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery' — a survey/awesome list about LLMs applied to scientific discovery.
matched fp:8d07a7571aa49a26, llm:Repository title and description: '[EMNLP2025] From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery' — a survey/awesome list about LLMs applied to scientific discovery.
matched fp:8d07a7571aa49a26, llm:Repository title and description: '[EMNLP2025] From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery' — a survey/awesome list about LLMs applied to scientific discovery.