This repository is a collection of world modeling progress based on Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond, covering 1,000+ papers and benchmarks grouped by section in reverse chronological order and available in a public Notion database. Released under the MIT License. Check out our poster.
Note
📚 If you find this resource useful, please cite and the repo:
@article{chu2026agenticworldmodelingfoundations,
title = {Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond},
author = {Meng Chu and Xuan Billy Zhang and Kevin Qinghong Lin and Lingdong Kong and Jize Zhang and Teng Tu and Weijian Ma and Ziqi Huang and Senqiao Yang and Wei Huang and Yeying Jin and Zhefan Rao and Jinhui Ye and Xinyu Lin and Xichen Zhang and Qisheng Hu and Shuai Yang and Leyang Shen and Wei Chow and Yifei Dong and Fengyi Wu and Quanyu Long and Bin Xia and Shaozuo Yu and Mingkang Zhu and Wenhu Zhang and Jiehui Huang and Haokun Gui and Runyi Li and Chenyu Tang and Dong Huang and Xuhang Chen and Rui Liu and Chengzu Li and Shiyi Du and Xu Huang and Haoxuan Che and Long Chen and Qifeng Chen and Wenya Wang and Wenxuan Zhang and Xiaojuan Qi and Yang Deng and Yanwei Li and Mike Zheng Shou and Zhi-Qi Cheng and See-Kiong Ng and Ziwei Liu and Philip Torr and Jiaya Jia},
year = {2026},
eprint = {2604.22748},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2604.22748}
}Tip
👋 Welcome to join the discussion on or
, share your work, collaborate on distillation, and help us grow the agentic world modeling community together.
- 🚧 More updates coming soon, stay tuned!
- Taxonomy Overview
- L1: Predictor
- L2: Simulator
- L3: Evolver
- Benchmarks & Evaluation
- Related Surveys
- Welcome to Contribute
| Level | Definition | Key Capability | Physical | Digital | Social | Scientific |
|---|---|---|---|---|---|---|
| L1 Predictor | One-step local transition | Prediction accuracy, robustness, identifiability | RSSM, V-JEPA, TD-MPC2 | LLM pred., Othello-WM | ToMnet, BToM | GNN, FNO |
| L2 Simulator | Multi-step rollout respecting governing laws | Long-horizon coherence, intervention sensitivity, constraint consistency | DreamerV3, Sora, Cosmos | WebDreamer, Code2World | Generative Agents, CICERO | GraphCast, NeuralGCM |
| L3 Evolver | Design → Execute → Observe → Reflect with model revision | Active information expansion, autonomous execution, belief revision | AdaptSim, Self-Modeling | AlphaEvolve, FunSearch | Evolving Constitutions, AgentSociety | A-Lab, AI Scientist |
Methods learning local one-step operators: state inference, forward dynamics, observation decoding, and inverse dynamics.
- A Compositional Object-Based Approach to Learning Physical Dynamics (ICLR, 2016)
- A Game Theoretic Framework for Model Based Reinforcement Learning (ICML, 2020)
- A Generalist Agent (Trans. Mach. Learn. Res., 2022)
- A Generalist Dynamics Model for Control (arXiv, 2023)
- A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures (arXiv, 2026)
- A Path Towards Autonomous Machine Intelligence (OpenReview preprint, 2022)
- A-JEPA: Joint-Embedding Predictive Architecture Can Listen (arXiv, 2023)
- Accelerating Model-Based Reinforcement Learning with State-Space World Models (arXiv, 2025)
- ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning (arXiv, 2025)
- Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning (arXiv, 2026)
- AD3: Implicit Action is the Key for World Models to Distinguish the Diverse Visual Distractors (ICML, 2024)
- Adaptive World Models: Learning Behaviors by Latent Imagination Under Non-Stationarity (arXiv, 2024)
- Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning (RSS, 2024)
- Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers (CVPR, 2025)
- Affordances Enable Partial World Modeling with LLMs (arXiv, 2026)
- AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation (arXiv, 2025)
- AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models (arXiv, 2026)
- Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations (arXiv, 2026)
- Beyond Euclidean Proximity: Repairing Latent World Models with Horizon-Matched Trajectory Reachability Metrics (arXiv, 2026)
- BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models (arXiv, 2025)
- Bootstrap Off-policy with World Model (arXiv, 2025)
- Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models (arXiv, 2025)
- Bounded Exploration with World Model Uncertainty in Soft Actor-Critic Reinforcement Learning Algorithm (arXiv, 2024)
- Bounding Distributional Shifts in World Modeling through Novelty Detection (IEEE/RJS International Conference on Intelligent RObots and Systems, 2025)
- BRICKS-WM: Building Reusability via Interface Composition Kinetics for Structured World Models (arXiv, 2026)
- Building spatial world models from sparse transitional episodic memories (arXiv, 2025)
- Can World Models Benefit VLMs for World Dynamics? (arXiv, 2025)
- Cardiac Copilot: Automatic Probe Guidance for Echocardiography with World Model (arXiv, 2024)
- CarFormer: Self-Driving with Learned Object-Centric Representations (ECCV, 2024)
- Causal-JEPA: Learning World Models through Object-Level Latent Interventions (arXiv, 2026)
- CausalVAE as a Plug-in for World Models: Towards Reliable Counterfactual Dynamics (arXiv, 2026)
- Chain of World: World Model Thinking in Latent Motion (arXiv, 2026)
- CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization (arXiv, 2026)
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models (arXiv, 2025)
- Closing the Train-Test Gap in World Models for Gradient-Based Planning (arXiv, 2025)
- Cognitively Inspired Energy-Based World Models (arXiv, 2024)
- CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving (arXiv, 2025)
- Context and Diversity Matter: The Emergence of In-Context Learning in World Models (arXiv, 2025)
- Contextual Latent World Models for Offline Meta Reinforcement Learning (arXiv, 2026)
- CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models (CVPR, 2025)
- CWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent Pipelines (arXiv, 2026)
- DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions (arXiv, 2025)
- DayDreamer: World Models for Physical Robot Learning (CoRL, 2022)
- Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models (NeurIPS, 2018)
- Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation (arXiv, 2026)
- Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models (arXiv, 2026)
- Deterministic World Models for Verification of Closed-loop Vision-based Systems (arXiv, 2025)
- DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA (arXiv, 2026)
- Differentiable Raycasting for Self-supervised Occupancy Forecasting (ECCV, 2022)
- Diffusion Transformer World-Action Model for AV Scene Prediction (arXiv, 2026)
- DiLA: Disentangled Latent Action World Models (arXiv, 2026)
- DINO-Foresight: Looking into the Future with DINO (arXiv, 2024)
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning (ICML, 2024)
- Discrete Codebook World Models for Continuous Control (ICLR, 2025)
- Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning (ICCV, 2025)
- DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control (arXiv, 2026)
- DiWA: Diffusion Policy Adaptation with World Models (arXiv, 2025)
- DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving (arXiv, 2026)
- Do Transformer World Models Give Better Policy Gradients? (ICML, 2024)
- Dream to Control: Learning Behaviors by Latent Imagination (ICLR, 2019)
- Dream to Drive With Predictive Individual World Model (IEEE Transactions on Intelligent Vehicles, 2024)
- DREAM-Chunk: Reactive Action Chunking with Latent World Model (arXiv, 2026)
- DREAMer-VXS: A Latent World Model for Sample-Efficient AGV Exploration in Stochastic, Unobserved Environments (arXiv, 2025)
- DreamerV3-XP: Optimizing exploration through uncertainty estimation (arXiv, 2025)
- Dreaming of Many Worlds: Learning Contextual World Models Aids Zero-Shot Generalization (RLJ, 2024)
- Dreaming the Unseen: World Model-regularized Diffusion Policy for Out-of-Distribution Robustness (arXiv, 2026)
- Dreaming when Necessary: Advancing World Action Models with Adaptive Multi-Modal Reasoning (arXiv, 2026)
- Dreaming: Model-based Reinforcement Learning by Latent Imagination without Reconstruction (ICRA, 2020)
- DreamSAC: Learning Hamiltonian World Models via Symmetry Exploration (arXiv, 2026)
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge (arXiv, 2025)
- Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving (arXiv, 2026)
- Driving Beyond Privilege: Distilling Dense-Reward Knowledge into Sparse-Reward Policies (arXiv, 2025)
- Dynamic Sparsity: Challenging Common Sparsity Assumptions for Learning World Models in Robotic Reinforcement Learning Benchmarks (AAAI, 2025)
- Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot Generalization (arXiv, 2025)
- DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control (NeurIPS, 2024)
- DynaWM: Dynamics-Aware Distillation with World Model and Momentum Targets for Smooth Locomotion over Continuous Stairs (arXiv, 2026)
- DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving (arXiv, 2026)
- DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving (arXiv, 2026)
- Efficient Image-Goal Navigation with Representative Latent World Model (arXiv, 2025)
- EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds (ICCV, 2025)
- Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images (NeurIPS, 2015)
- Embodied World Models Emerge from Navigational Task in Open-Ended Environments (arXiv, 2025)
- Enhance Sample Efficiency and Robustness of End-to-End Urban Autonomous Driving via Semantic Masked World Model (IEEE transactions on intelligent transportation systems (Print), 2022) — End-to-end autonomous driving provides a feasible way to automatically maximize overall driving system performance by directly mapping the raw pixels from a front-facing camera to control signals.
- Enhancing End-to-End Autonomous Driving with Latent World Model (ICLR, 2024)
- Enhancing Policy Learning with World-Action Model (arXiv, 2026)
- Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction (arXiv, 2026)
- EWAM: An Enhanced World Action Model for Closed-Loop Online Adaptation in Embodied Intelligence (arXiv, 2026)
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions (arXiv, 2025)
- Factored Latent Action World Models (arXiv, 2026)
- Factored World Models for Zero-Shot Generalization in Robotic Manipulation (arXiv, 2022)
- Factorized Latent Dynamics for Video JEPA: An Empirical Study of Auxiliary Objectives (arXiv, 2026)
- FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models (arXiv, 2025)
- FAWAM: Force-Aware World Action Models for Closed-Loop Contact-Rich Manipulation (arXiv, 2026)
- FF-JEPA: Long-Horizon Planning in World Models with Latent Planners (arXiv, 2026)
- Finetuning Offline World Models in the Real World (CoRL, 2023)
- FLARE: Learning Future-Aware Latent Representations from Vision-Language Models for Autonomous Driving (arXiv, 2026)
- FLARE: Robot Learning with Implicit World Modeling (arXiv, 2025)
- Flash-WAM: Modality-Aware Distillation for World Action Models (arXiv, 2026)
- Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments (arXiv, 2026)
- FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models (arXiv, 2025)
- FOCUS: Object-Centric World Models for Robotics Manipulation (arXiv, 2023)
- Foresight Diffusion: Improving Sampling Consistency in Predictive Diffusion Models (arXiv, 2025)
- Foundational World Models Accurately Detect Bimanual Manipulator Failures (arXiv, 2026)
- FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment (arXiv, 2026)
- FSF-Net: Enhance 4D Occupancy Forecasting with Coarse BEV Scene Flow for Autonomous Driving (arXiv, 2024)
- GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction (CVPR, 2024)
- GenAD: Generative End-to-End Autonomous Driving (ECCV, 2024)
- General agents contain world models (arXiv, 2025)
- GeoSem-WAM: Geometry- and Semantic-Aware World Action Models (arXiv, 2026)
- GeoWorld-VLM: Geometry from World Models for Vision-Language Models (arXiv, 2026)
- Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation (arXiv, 2025)
- Gradient-based Planning with World Models (arXiv, 2023)
- GraphWorld: Long-Horizon Planning with World Models for End-to-End Autonomous Driving (arXiv, 2026)
- Grounding Large Language Models In Embodied Environment With Imperfect World Models (arXiv, 2024)
- Grounding Video Models to Actions through Goal Conditioned Exploration (ICLR, 2024)
- HarmonyDream: Task Harmonization Inside World Models (ICML, 2023)
- HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models (arXiv, 2026)
- Hierarchical JEPA Meets Predictive Remote Control in Beyond 5G Networks (arXiv, 2026)
- Hindsight for Foresight: Unsupervised Structured Dynamics Models from Physical Interaction (IEEE/RJS International Conference on Intelligent RObots and Systems, 2020)
- How Hard is it to Confuse a World Model? (arXiv, 2025)
- IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI (arXiv, 2024)
- In-Context Reinforcement Learning via Communicative World Models (arXiv, 2025)
- In-Context World Modeling for Robotic Control (arXiv, 2026)
- InDRiVE: Intrinsic Disagreement based Reinforcement for Vehicle Exploration through Curiosity Driven Generalized World Model (arXiv, 2025)
- InDRiVE: Reward-Free World-Model Pretraining for Autonomous Driving via Latent Disagreement (arXiv, 2025)
- Inference-time Physics Alignment of Video Generative Models with Latent World Models (arXiv, 2026)
- Inter-environmental world modeling for continuous and compositional dynamics (arXiv, 2025)
- Intuitive physics understanding emerges from self-supervised pretraining on natural videos (arXiv, 2025)
- IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model (arXiv, 2025)
- Is Conditional Generative Modeling all you need for Decision-Making? (ICLR, 2022)
- Iso-Dream: Isolating and Leveraging Noncontrollable Visual Dynamics in World Models (NeurIPS, 2022)
- JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning (arXiv, 2026)
- JEPA for RL: Investigating Joint-Embedding Predictive Architectures for Reinforcement Learning (arXiv, 2025)
- JEPA-VLA: Video Predictive Embedding is Needed for VLA Models (arXiv, 2026)
- KAN-Dreamer: Benchmarking Kolmogorov-Arnold Networks as Function Approximators in World Models (arXiv, 2025)
- Kinodynamic Motion Planning for Mobile Robot Navigation across Inconsistent World Models (arXiv, 2025)
- Knowledge Graphs as World Models for Semantic Material-Aware Obstacle Handling in Autonomous Vehicles (arXiv, 2025)
- Label-Efficient Grasp Joint Prediction with Point-JEPA (arXiv, 2025)
- Language-conditioned world model improves policy generalization by reading environmental descriptions (arXiv, 2025)
- LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving (arXiv, 2026)
- Latent Action Pretraining from Videos (arXiv, 2024)
- Latent Action Pretraining Through World Modeling (arXiv, 2025)
- Latent Action World Models for Control with Unlabeled Trajectories (arXiv, 2025)
- Latent Policy Steering with Embodiment-Agnostic Pretrained World Models (arXiv, 2025)
- Latent Video Prediction Learns Better World Models (arXiv, 2026)
- LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies (arXiv, 2026)
- LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion (arXiv, 2026)
- Learned Perceptive Forward Dynamics Model for Safe and Platform-aware Robotic Navigation (Robotics, 2025)
- Learning Abstract World Models with a Group-Structured Latent Space (arXiv, 2025)
- Learning and Leveraging World Models in Visual Representation Learning (arXiv, 2024)
- Learning Humanoid Locomotion with World Model Reconstruction (arXiv, 2025)
- Learning Invariant Visual Representations for Planning with Joint-Embedding Predictive World Models (arXiv, 2026)
- Learning Latent Action World Models In The Wild (arXiv, 2026)
- Learning Latent Dynamic Robust Representations for World Models (ICML, 2024)
- Learning Latent Dynamics for Planning from Pixels (ICML, 2018)
- Learning Robot Manipulation from Audio World Models (arXiv, 2025)
- Learning to Act from Actionless Videos through Dense Correspondences (ICLR, 2023)
- Learning to Poke by Poking: Experiential Learning of Intuitive Physics (Advances in Neural Information Processing Systems, 2016)
- Learning to Reach Goals via Iterated Supervised Learning (ICLR, 2021)
- Learning to Sample: Reinforcement Learning-Guided Sampling for Autonomous Vehicle Motion Planning (arXiv, 2025)
- Learning to unfold cloth: Scaling up world models to deformable object manipulation (arXiv, 2026)
- Learning World Models for Unconstrained Goal Navigation (NeurIPS, 2024)
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels (arXiv, 2026)
- Linear Spatial World Models Emerge in Large Language Models (arXiv, 2025)
- LSRE: Latent Semantic Rule Encoding for Real-Time Semantic Risk Detection in Autonomous Driving (arXiv, 2025)
- LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model (arXiv, 2026)
- Making Foresight Actionable: Repurposing Representation Alignment in World Action Models (arXiv, 2026)
- MAMBA: an Effective World Model Approach for Meta-Reinforcement Learning (ICLR, 2024)
- ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model (IEEE/RJS International Conference on Intelligent RObots and Systems, 2025)
- ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation (ECCV, 2024)
- Mask World Model: Predicting What Matters for Robust Robot Policy Learning (arXiv, 2026)
- Masked World Models for Visual Control (CoRL, 2022)
- Mastering diverse control tasks through world models (Nature, 2025)
- Mastering Diverse Domains through World Models (arXiv, 2023)
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning (ICLR, 2022)
- MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features (arXiv, 2023)
- Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation (arXiv, 2026)
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs (arXiv, 2025)
- MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation (arXiv, 2026)
- Mobile Manipulation with Active Inference for Long-Horizon Rearrangement Tasks (arXiv, 2025)
- MoDem-V2: Visuo-Motor World Models for Real-World Robot Manipulation (ICRA, 2023)
- MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios (AAAI, 2025)
- Multi-Camera Unified Pre-Training via 3D Scene Reconstruction (IEEE Robotics and Automation Letters, 2023)
- Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning (ICML, 2025)
- Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model (arXiv, 2026)
- Multimodal Dreaming: A Global Workspace Approach to World Model-Based Reinforcement Learning (arXiv, 2025)
- NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning (IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024)
- NavigateDiff: Visual Predictors are Zero-Shot Navigation Assistants (arXiv, 2025)
- Neural Fields as World Models (arXiv, 2026)
- Neural Motion Simulator Pushing the Limit of World Models in Reinforcement Learning (CVPR, 2025)
- Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning (ICRA, 2017)
- Next Embedding Prediction Makes World Models Stronger (arXiv, 2026)
- Next Forcing: Causal World Modeling with Multi-Chunk Prediction (arXiv, 2026)
- NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards (arXiv, 2025)
- NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models (IEEE Transactions on Image Processing, 2025)
- Object-Centric World Model for Language-Guided Manipulation (arXiv, 2025)
- Object-Centric World Models for Causality-Aware Reinforcement Learning (AAAI, 2025)
- OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision (arXiv, 2025)
- On Memory: A comparison of memory mechanisms in world models (arXiv, 2025)
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning (arXiv, 2025)
- One-shot World Models Using a Transformer Trained on a Synthetic Prior (arXiv, 2024)
- PACT: Perception-Action Causal Transformer for Autoregressive Robotics Pre-Training (IEEE/RJS International Conference on Intelligent RObots and Systems, 2022)
- PH-Dreamer: A Physics-Driven World Model via Port-Hamiltonian Generative Dynamics (arXiv, 2026)
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining (arXiv, 2025)
- Physically Interpretable World Models via Weakly Supervised Representation Learning (arXiv, 2024)
- PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement Learning (ICML, 2025)
- PILCO: A Model-Based and Data-Efficient Approach to Policy Search (ICML, 2011)
- PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation (NeurIPS, 2024)
- PLAN-S: Bridging Planning with Latent Style Dynamics for Autonomous Driving World Models (arXiv, 2026)
- Planning from Pixels using Inverse Dynamics Models (ICLR, 2020)
- PLUME: Probabilistic Latent Unified World Modeling and Parameter Estimation for Multi-Finger Manipulation (arXiv, 2026)
- Point Cloud Forecasting as a Proxy for 4D Occupancy Forecasting (CVPR, 2023)
- Point Tracking Improves World Action Models (arXiv, 2026)
- Pre-training Contextualized World Models with In-the-wild Videos for Reinforcement Learning (NeurIPS, 2023)
- Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts (arXiv, 2026)
- Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation (arXiv, 2024)
- Prismatic World Model: Learning Compositional Dynamics for Planning in Hybrid Systems (arXiv, 2025)
- Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models (arXiv, 2026)
- Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling (arXiv, 2025)
- ProTerrain: Probabilistic Physics-Informed Rough Terrain World Modeling (arXiv, 2025)
- R2-Dreamer: Redundancy-Reduced World Models without Decoders or Augmentation (arXiv, 2026)
- Real-World Robot Control by Deep Active Inference With a Temporally Hierarchical World Model (IEEE Robotics and Automation Letters, 2025)
- RECON: Rapid Exploration for Open-World Navigation with Latent Goal Models (CoRL, 2021)
- Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models (arXiv, 2026)
- Reinforcement Learning with Inverse Rewards for World Model Post-training (arXiv, 2025)
- Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement (arXiv, 2026)
- Representing Positional Information in Generative World Models for Object Manipulation (European Conference on Artificial Intelligence, 2024)
- RESBev: Making BEV Perception More Robust (arXiv, 2026)
- Reset-free Reinforcement Learning with World Models (Trans. Mach. Learn. Res., 2024)
- ResWM: Residual-Action World Model for Visual RL (arXiv, 2026)
- Revisiting Feature Prediction for Learning Visual Representations from Video (Trans. Mach. Learn. Res., 2024)
- Reward-free World Models for Online Imitation Learning (ICML, 2024)
- RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation (arXiv, 2026)
- Safety Certification in the Latent space using Control Barrier Functions and World Models (International Conference on Intelligent Cloud Computing, 2025)
- SANTS: A State-Adaptive Scheduler for World Action Models (arXiv, 2026)
- Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling (arXiv, 2025)
- Self-Consistent Model-based Adaptation for Visual Reinforcement Learning (IJCAI, 2025)
- Self-Supervised JEPA-based World Models for LiDAR Occupancy Completion and Forecasting (arXiv, 2026)
- Self-supervised Multi-future Occupancy Forecasting for Autonomous Driving (Robotics, 2024)
- Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object Detection (AAAI, 2025)
- Semantic Belief-State World Model for 3D Human Motion Prediction (arXiv, 2026)
- SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models (ICML, 2025)
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models (arXiv, 2025)
- SimuDICE: Offline Policy Optimization Through World Model Updates and DICE Estimation (arXiv, 2024)
- SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors (arXiv, 2026)
- SPARTAN: A Sparse Transformer World Model Attending to What Matters (arXiv, 2024)
- Spatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecasting (CVPR, 2024)
- Spiking World Model with Multi-Compartment Neurons for Model-based Reinforcement Learning (Proceedings of the National Academy of Sciences of the United States of America, 2025)
- SR-AIF: Solving Sparse-Reward Robotic Tasks From Pixels with Active Inference and World Models (ICRA, 2024)
- State Representation Learning for Control: An Overview (Neural Networks, 2018)
- Statler: State-Maintaining Language Models for Embodied Reasoning (ICRA, 2023)
- Structure Abstraction and Generalization in a Hippocampal-Entorhinal Inspired World Model (arXiv, 2026)
- Structure-aware World Model for Probe Guidance via Large-scale Self-supervised Pre-train (arXiv, 2024)
- Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models (arXiv, 2026)
- SWAP: Symmetric Equivariant World-Model for Agile Robot Parkour (arXiv, 2026)
- SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution (arXiv, 2026)
- TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation (arXiv, 2026)
- Taming generative video models for zero-shot optical flow extraction (arXiv, 2025)
- TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement Learning (arXiv, 2025)
- TD-MPC2: Scalable, Robust World Models for Continuous Control (ICLR, 2023)
- Temporal Difference Learning for Model Predictive Control (ICML, 2022)
- Temporal Straightening for Latent Planning (arXiv, 2026)
- Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles (arXiv, 2025)
- ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model (arXiv, 2026)
- Time-Aware World Model for Adaptive Prediction and Control (ICML, 2025)
- Time-Series JEPA for Predictive Remote Control under Capacity-Limited Networks (arXiv, 2024)
- Towards Unraveling and Improving Generalization in World Models (arXiv, 2024)
- TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving (arXiv, 2026)
- TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos (arXiv, 2025)
- TransDreamer: Reinforcement Learning with Transformer World Models (arXiv, 2022)
- Transformers and Slot Encoding for Sample Efficient Physical World Modelling (arXiv, 2024)
- UMAD: Unsupervised Mask-Level Anomaly Detection for Autonomous Driving (BMVC Workshops, 2024)
- UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model (AAAI, 2025)
- Unified Vision-Language-Action Model (arXiv, 2025)
- Unifying (Machine) Vision via Counterfactual World Modeling (arXiv, 2023)
- Unifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic Tasks (arXiv, 2026)
- UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning (arXiv, 2025)
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions (Robotics, 2025)
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation (ICLR, 2023)
- UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent (ICML, 2025)
- UWM-JEPA: Predictive World Models That Imagine in Belief Space (arXiv, 2026)
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning (arXiv, 2026)
- VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis (arXiv, 2026)
- Value-guided action planning with JEPA world models (arXiv, 2025)
- VDRive: Leveraging Reinforced VLA and Diffusion Policy for End-to-end Autonomous Driving (arXiv, 2025)
- VEGA-3D: Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding (arXiv, 2026)
- VFMF: World Modeling by Forecasting Vision Foundation Model Features (arXiv, 2025)
- VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward (arXiv, 2026)
- VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model (arXiv, 2026)
- Video generation models as world simulators (OpenAI, 2024)
- Video Prediction Models as Rewards for Reinforcement Learning (NeurIPS, 2023)
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations (ICML, 2024)
- Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling (arXiv, 2025)
- VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation (NeurIPS, 2024)
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models (arXiv, 2025)
- ViPRA: Video Prediction for Robot Actions (arXiv, 2025)
- Visual Point Cloud Forecasting Enables Scalable Autonomous Driving (CVPR, 2023)
- Visuomotor Grasping with World Models for Surgical Robots (arXiv, 2025)
- ViVa: A Video-Generative Value Model for Robot Reinforcement Learning (arXiv, 2026)
- VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model (arXiv, 2026)
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models (arXiv, 2023)
- WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation (arXiv, 2026)
- WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems (arXiv, 2026)
- What Drives Success in Physical Planning with Joint-Embedding Predictive World Models? (arXiv, 2025)
- What You Don't Know Can Hurt You: How Well do Latent Safety Filters Understand Partially Observable Safety Constraints? (arXiv, 2025)
- When Does LeJEPA Learn a World Model? (arXiv, 2026)
- When Object-Centric World Models Meet Policy Learning: From Pixels to Policies, and Where It Breaks (arXiv, 2025)
- When to Trust Your Model: Model-Based Policy Optimization (NeurIPS, 2019)
- Why and How Auxiliary Tasks Improve JEPA Representations (arXiv, 2025)
- WIMLE: Uncertainty-Aware World Models with IMLE for Sample-Efficient Continuous Control (arXiv, 2026)
- World Model Failure Classification and Anomaly Detection for Autonomous Inspection (arXiv, 2026)
- World Model for AI Autonomous Navigation in Mechanical Thrombectomy (International Conference on Medical Image Computing and Computer-Assisted Intervention, 2025)
- World Model Implanting for Test-time Adaptation of Embodied Agents (ICML, 2025)
- World Model Robustness via Surprise Recognition (arXiv, 2025)
- World Model-Based Perception for Visual Legged Locomotion (ICRA, 2024)
- World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning (ACL, 2025)
- World Models as Group Actions (arXiv, 2026)
- World Models as Reference Trajectories for Rapid Motor Adaptation (arXiv, 2025)
- World Models for Anomaly Detection during Model-Based Reinforcement Learning Inference (IEEE/RJS International Conference on Intelligent RObots and Systems, 2025)
- World Models for Autonomous Navigation of Terrestrial Robots from LIDAR Observations (Journal of Intelligent & Fuzzy Systems: Applications in Engineering and Technology, 2025)
- World Models That Know When They Don't Know: Controllable Video Generation with Calibrated Uncertainty (arXiv, 2025)
- World Pilot: Steering Vision-Language-Action Models with World-Action Priors (arXiv, 2026)
- World Value Models for Robotic Manipulation (arXiv, 2026)
- World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis (arXiv, 2026)
- World-Task Factorization for Robot Learning (arXiv, 2026)
- World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning (arXiv, 2026)
- WorldVLA: Towards Autoregressive Action World Model (arXiv, 2025)
- Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation (arXiv, 2026)
- Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models (ICLR, 2023)
- Zero-shot World Models Are Developmentally Efficient Learners (arXiv, 2026)
- A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment (ICML, 2024)
- A Generalist Agent (Trans. Mach. Learn. Res., 2022)
- A Simple Framework for Contrastive Learning of Visual Representations (ICML, 2020)
- A0C: Alpha Zero in Continuous Action Space (arXiv, 2018)
- Accurate and Efficient World Modeling with Masked Latent Transformers (ICML, 2025)
- Adapting a World Model for Trajectory Following in a 3D Game (arXiv, 2025)
- ARROW: Augmented Replay for RObust World models (arXiv, 2026)
- Auto-Encoding Variational Bayes (ICLR, 2013)
- $beta$-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework (ICLR, 2017)
- Better World Models Can Lead to Better Post-Training Performance (arXiv, 2025)
- BRo-JEPA: Learning Modular Arithmetic in Latent Space (arXiv, 2026)
- BWArea Model: Learning World Model, Inverse Dynamics, and Policy for Controllable Language Generation (arXiv, 2024)
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models (arXiv, 2025)
- CoDreamer: Communication-Based Decentralised World Models (arXiv, 2024)
- Compete and Compose: Learning Independent Mechanisms for Modular World Models (arXiv, 2024)
- Cross-View World Models (arXiv, 2026)
- CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models (arXiv, 2026)
- Curiosity-Driven Exploration by Self-Supervised Prediction (ICML, 2017)
- CURL: Contrastive Unsupervised Representations for Reinforcement Learning (ICML, 2020)
- Data-Efficient Reinforcement Learning with Self-Predictive Representations (ICLR, 2020)
- Deep SPI: Safe Policy Improvement via World Models (arXiv, 2025)
- DeepMDP: Learning Continuous Latent Space Models for Representation Learning (ICML, 2019)
- Design and Optimization of Reinforcement Learning-Based Agents in Text-Based Games (International Journal of Computer Science & Information Technology (IJCSIT), 2025) — As AI technology advances, research in playing text-based games with agents has become progressively popular.
- Diffusion for World Modeling: Visual Details Matter in Atari (NeurIPS, 2024)
- DINOv2: Learning Robust Visual Features without Supervision (Trans. Mach. Learn. Res., 2023)
- Discrete JEPA: Learning Discrete Token Representations without Reconstruction (arXiv, 2025)
- DreamerV3-XP: Optimizing exploration through uncertainty estimation (arXiv, 2025)
- DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing (ICLR, 2023)
- DyMoDreamer: World Modeling with Dynamic Modulation (arXiv, 2025)
- Dyn-O: Building Structured World Models with Object-Centric Representations (arXiv, 2025)
- Dyna, an integrated architecture for learning, planning, and reacting (ACM SIGART Bulletin, 1991)
- EDELINE: Enhancing Memory in Diffusion-based World Models via Linear-Time Sequence Modeling (arXiv, 2025)
- Efficient Exploration and Discriminative World Model Learning with an Object-Centric Abstraction (ICLR, 2024)
- Efficient World Models with Context-Aware Tokenization (ICML, 2024)
- Finite Automata Extraction: Low-data World Model Learning as Programs from Gameplay Video (arXiv, 2025)
- From Curiosity to Competence: How World Models Interact with the Dynamics of Exploration (Annual Meeting of the Cognitive Science Society, 2025)
- From Observations to Events: Event-Aware World Model for Reinforcement Learning (arXiv, 2026)
- GLAM: Global-Local Variation Awareness in Mamba-based World Model (AAAI, 2025)
- Hieros: Hierarchical Imagination on Structured State Space Sequence World Models (ICML, 2023)
- High-Resolution Image Synthesis with Latent Diffusion Models (CVPR, 2021)
- Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task (arXiv, 2025)
- Identifiable Token Correspondence for World Models (arXiv, 2026)
- Improving Token-Based World Models with Parallel Observation Prediction (ICML, 2024)
- Improving Transformer World Models for Data-Efficient RL (ICML, 2025)
- Improving World Models using Deep Supervision with Linear Probes (arXiv, 2025)
- InCoder-32B-Thinking: Industrial Code World Model for Thinking (arXiv, 2026)
- Joint Learning of Hierarchical Neural Options and Abstract World Model (arXiv, 2026)
- Large Language Models as Commonsense Knowledge for Large-Scale Task Planning (NeurIPS, 2023)
- Learning Local Causal World Models with State Space Models and Attention (arXiv, 2025)
- Learning To Explore With Predictive World Model Via Self-Supervised Learning (arXiv, 2025)
- Learning Transformer-based World Models with Contrastive Predictive Coding (ICLR, 2025)
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures (arXiv, 2025)
- Making Large Language Models into World Models with Precondition and Effect Knowledge (International Conference on Computational Linguistics, 2024)
- Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning (NeurIPS, 2023)
- Markov Decision Processes: Discrete Stochastic Dynamic Programming (John Wiley & Sons, 1994)
- Mastering Atari Games with Limited Data (NeurIPS, 2021)
- Mastering Atari with Discrete World Models (ICLR, 2020)
- Mastering Atari, Go, chess and shogi by planning with a learned model (Nature, 2019)
- Mastering Diverse Domains through World Models (arXiv, 2023)
- Mastering Memory Tasks with World Models (ICLR, 2024)
- Meta-Reinforcement Learning with Discrete World Models for Adaptive Load Balancing (ACM Southeast Regional Conference, 2025)
- MetaOthello: A Controlled Study of Multiple World Models in Transformers (arXiv, 2026)
- Model-Based Reinforcement Learning for Atari (ICLR, 2019)
- Momentum Contrast for Unsupervised Visual Representation Learning (CVPR, 2019)
- Neural Discrete Representation Learning (NeurIPS, 2017)
- Next-Latent Prediction Transformers Learn Compact World Models (arXiv, 2025)
- Novelty Detection in Reinforcement Learning with World Models (ICML, 2023)
- Object-Centric World Models from Few-Shot Annotations for Sample-Efficient Reinforcement Learning (arXiv, 2025)
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning (arXiv, 2025)
- Performance Asymmetry in Model-Based Reinforcement Learning (arXiv, 2025)
- Planning and Acting in Partially Observable Stochastic Domains (Artificial Intelligence, 1998) — In this paper, we bring techniques from operations research to bear on the problem of choosing optimal actions in partially observable stochastic domains.
- Policy and World Modeling Co-Training for Language Agents (arXiv, 2026)
- Recurrent World Models Facilitate Policy Evolution (Advances in Neural Information Processing Systems, 2018)
- Representation Learning with Contrastive Predictive Coding (arXiv, 2018)
- Revisiting the Othello World Model Hypothesis (arXiv, 2025)
- SafeDream: Safety World Model for Proactive Early Jailbreak Detection (arXiv, 2026)
- Scalable Diffusion Models with Transformers (IEEE/CVF International Conference on Computer Vision, 2023)
- Self-supervised Hierarchical Visual Reasoning with World Model (arXiv, 2026)
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture (CVPR, 2023)
- Simple, Good, Fast: Self-Supervised World Models Free of Baggage (ICLR, 2025)
- Simulus: Combining Improvements in Sample-Efficient World Model Agents (arXiv, 2025)
- Slots, Transitions, Loops: Learning Composable World Models for ARC (arXiv, 2026)
- Stochastic Backpropagation and Approximate Inference in Deep Generative Models (ICML, 2014)
- STORM: Efficient Stochastic Transformer based World Models for Reinforcement Learning (NeurIPS, 2023)
- Structured Latent Dynamics in Wireless CSI via Homomorphic World Models (arXiv, 2026)
- Task Aware Dreamer for Task Generalization in Reinforcement Learning (arXiv, 2023)
- Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction (arXiv, 2025)
- Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation (arXiv, 2026)
- Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning (bioRxiv, 2026)
- Toward Accurate Image Generation via Dynamic Generative Image Transformer (IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026)
- TransDreamerV3: Implanting Transformer In DreamerV3 (arXiv, 2025)
- Transformer-based World Models Are Happy With 100k Interactions (ICLR, 2023)
- Transformers are Sample Efficient World Models (ICLR, 2022)
- Transformers Linearly Represent Highly Structured World Models (arXiv, 2026)
- Transformers Use Causal World Models in Maze-Solving Tasks (arXiv, 2024)
- Unsupervised State Representation Learning in Atari (Advances in Neural Information Processing Systems, 2019)
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos (Advances in Neural Information Processing Systems, 2022)
- World Modelling Improves Language Model Agents (arXiv, 2025)
- World Models (arXiv, 2018)
- Zero-shot World Models via Search in Memory (arXiv, 2025)
- Active Confusion Expression in Large Language Models: Leveraging World Models toward Better Social Reasoning (arXiv, 2025)
- AffectVerse: Emotional World Models for Multimodal Affective Computing (arXiv, 2026)
- GAWM: Global-Aware World Model for Multi-Agent Reinforcement Learning (arXiv, 2025)
- Large Emotional World Model (arXiv, 2025)
- LSRE: Latent Semantic Rule Encoding for Real-Time Semantic Risk Detection in Autonomous Driving (arXiv, 2025)
- Machine theory of mind (ICML, 2018)
- Neural Sabermetrics with World Model: Play-by-play Predictive Modeling with Large Language Model (arXiv, 2026)
- Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning (Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems, 2026)
- Social-JEPA: Emergent Geometric Isomorphism (arXiv, 2026)
- Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech (arXiv, 2025)
- The Mouth is Not the Brain: Bridging Energy-Based World Models and Language Generation (arXiv, 2026)
- A multimodal and temporal foundation model for virtual patient representations at healthcare system scale (arXiv, 2026)
- A World Model of Radiologist Reading for Medical Image Representation Learning (arXiv, 2026)
- Accurate Prediction of Protein Structures and Interactions Using a Three-Track Neural Network (Science, 2021)
- Accurate Structure Prediction of Biomolecular Interactions with AlphaFold 3 (Nature, 2024)
- Cardiac Copilot: Automatic Probe Guidance for Echocardiography with World Model (arXiv, 2024)
- CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning (CVPR, 2025)
- Chreode: A Cell World Model for One-Step Temporal Dynamics and Perturbation Prediction (arXiv, 2026)
- DyNeMoC: A semi-supervised architecture for classifying time series brain data (International Conference on Learning Representations Workshop, 2023)
- Evolutionary-Scale Prediction of Atomic-Level Protein Structure with a Language Model (Science, 2023)
- Fast transient networks in spontaneous human brain activity (eLife, 2014)
- From Kepler to Newton: Inductive Biases Guide Learned World Models in Transformers (arXiv, 2026)
- Highly Accurate Protein Structure Prediction with AlphaFold (Nature, 2021)
- Improved Protein Structure Prediction Using Potentials from Deep Learning (Nature, 2020)
- Large-scale cortical functional networks are organized in structured cycles (Nature Neuroscience, 2025)
- Mixtures of large-scale dynamic functional brain network modes (NeuroImage, 2022)
- Predictive Coding in the Visual Cortex: a Functional Interpretation of Some Extra-Classical Receptive-Field Effects (Nature Neuroscience, 1999)
- Sonata: A Hybrid World Model for Inertial Kinematics under Clinical Data Scarcity (arXiv, 2026)
- Spontaneous cortical activity transiently organises into frequency specific phase-coupling networks (Nature Communications, 2018)
- Structure-aware World Model for Probe Guidance via Large-scale Self-supervised Pre-train (arXiv, 2024)
- The Free-Energy Principle: A Unified Brain Theory? (Nature Reviews Neuroscience, 2010)
- The Patient is not a Moving Document: A World Model Training Paradigm for Longitudinal EHR (arXiv, 2026)
- Vector Quantization in the Brain: Grid-like Codes in World Models (arXiv, 2025)
- When do World Models Successfully Learn Dynamical Systems? (arXiv, 2025)
- World Model for AI Autonomous Navigation in Mechanical Thrombectomy (International Conference on Medical Image Computing and Computer-Assisted Intervention, 2025)
- World Models Unlock Optimal Foraging Strategies in Reinforcement Learning Agents (arXiv, 2025)
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing (arXiv, 2025)
- Xray2Xray: World Model from Chest X-rays with Volumetric Context (arXiv, 2025)
Methods rolling out multi-step futures that respect governing laws: long-horizon coherence, intervention sensitivity, and constraint consistency.
- τ_0-WM: A Unified Video-Action World Model for Robotic Manipulation (arXiv, 2026)
- 3D-Belief: Embodied Belief Inference via Generative 3D World Modeling (arXiv, 2026)
- 3D-VLA: A 3D Vision-Language-Action Generative World Model (ICML, 2024)
- 3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation (AAAI, 2025)
- 3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model (arXiv, 2025)
- 4D Driving Scene Generation With Stereo Forcing (arXiv, 2025)
- A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens (arXiv, 2026)
- A Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World Models (arXiv, 2025)
- ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment (arXiv, 2026)
- AcceRL: A Distributed Asynchronous Reinforcement Learning and World Model Framework for Vision-Language-Action Models (arXiv, 2026)
- Act2Goal: From World Model To General Goal-conditioned Policy (arXiv, 2025)
- Action Images: End-to-End Policy Learning via Multiview Video Generation (arXiv, 2026)
- Active World-Model with 4D-informed Retrieval for Exploration and Awareness (arXiv, 2026)
- ActWorld: From Explorable to Interactive World Model via Action-Aware Memory (arXiv, 2026)
- AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models (arXiv, 2025)
- AdaPower: Specializing World Foundation Models for Predictive Manipulation (arXiv, 2025)
- Adapting World Models with Latent-State Dynamics Residuals (arXiv, 2025)
- AdaWM: Adaptive World Model based Planning for Autonomous Driving (ICLR, 2025)
- AdaWorld: Learning Adaptable World Models with Latent Actions (ICML, 2025)
- AdaWorldPolicy: World-Model-Driven Diffusion Policy with Online Adaptive Learning for Robotic Manipulation (arXiv, 2026)
- Addressing Corner Cases in Autonomous Driving: A World Model-based Approach with Mixture of Experts and LLMs (arXiv, 2025)
- ADriver-I: A General World Model for Autonomous Driving (arXiv, 2023)
- Advancing Open-source World Models (arXiv, 2026)
- Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space (arXiv, 2025)
- AETHER: Geometric-Aware Unified World Modeling (ICCV, 2025)
- AHA-WAM: Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing (arXiv, 2026)
- AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps (arXiv, 2026)
- AirScape: An Aerial Generative World Model with Motion Controllability (ACM Multimedia, 2025)
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning (ICLR, 2020)
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail (arXiv, 2025)
- An Efficient and Multi-Modal Navigation System with One-Step World Model (arXiv, 2026)
- An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training (arXiv, 2024)
- AnchorDrive: LLM Scenario Rollout with Anchor-Guided Diffusion Regeneration for Safety-Critical Scenario Generation (arXiv, 2026)
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories (arXiv, 2026)
- AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization (arXiv, 2026)
- Astra: General Interactive World Model with Autoregressive Denoising (arXiv, 2025)
- AstraNav-World: World Model for Foresight Control and Consistency (arXiv, 2025)
- Audio-Visual World Models: Towards Multisensory Imagination in Sight and Sound (arXiv, 2025)
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation (arXiv, 2025)
- AutoWorld: Scaling Multi-Agent Traffic Simulation with Self-Supervised World Models (arXiv, 2026)
- AVID: Adapting Video Diffusion Models to World Models (arXiv, 2024)
- Back to the Features: DINO as a Foundation for Video World Models (arXiv, 2025)
- BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation (arXiv, 2026)
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models (arXiv, 2025)
- Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement Learning (arXiv, 2026)
- Being-H0.7: A Latent World-Action Model from Egocentric Videos (arXiv, 2026)
- BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents (arXiv, 2024)
- Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation (arXiv, 2026)
- Beyond Pixel Histories: World Models with Persistent 3D State (arXiv, 2026)
- Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation (Advances in Neural Information Processing Systems, 2021)
- BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks (arXiv, 2026)
- Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation (arXiv, 2026)
- Bridging the Gap Between Multimodal Foundation Models and World Models (arXiv, 2025)
- BuilderBench: The Building Blocks of Intelligent Agents (arXiv, 2025)
- Building Explicit World Model for Zero-Shot Open-World Object Manipulation (arXiv, 2026)
- Can Test-Time Scaling Improve World Foundation Model? (arXiv, 2025)
- Captain Safari: A World Engine (arXiv, 2025)
- CARLA: An Open Urban Driving Simulator (CoRL, 2017)
- Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation (arXiv, 2026)
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation (arXiv, 2026)
- Causal World Modeling for Robot Control (arXiv, 2026)
- Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models (arXiv, 2026)
- Challenger: Affordable Adversarial Driving Video Generation (arXiv, 2025)
- Choreographing a World of Dynamic Objects (arXiv, 2026)
- ChronoDreamer: Action-Conditioned World Model as an Online Simulator for Robotic Planning (arXiv, 2025)
- ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation (arXiv, 2025)
- ChronosObserver: Taming 4D World with Hyperspace Diffusion Sampling (arXiv, 2025)
- CKT-WAM: Parameter-Efficient Context Knowledge Transfer Between World Action Models (arXiv, 2026)
- Co-Evolving Latent Action World Models (arXiv, 2025)
- Coding Agent Is Good As World Simulator (arXiv, 2026)
- CoGen: 3D Consistent Video Generation via Adaptive Conditioning for Autonomous Driving (arXiv, 2025)
- COMBO: Compositional World Models for Embodied Multi-Agent Cooperation (ICLR, 2024)
- COME: Adding Scene-Centric Forecasting Control to Occupancy World Model (arXiv, 2025)
- Composition of Memory Experts for Diffusion World Models (arXiv, 2026)
- Compositional Foundation Models for Hierarchical Planning (NeurIPS, 2023)
- Compression and Retrieval: Implicit Memory Retrieval for Video World Models (arXiv, 2026)
- ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask (arXiv, 2026)
- Constructing the Umwelt: Cognitive Planning through Belief-Intent Co-Evolution (arXiv, 2025)
- ContactGaussian-WM: Learning Physics-Grounded World Model from Videos (arXiv, 2026)
- Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval (ACM SIGGRAPH Conference and Exhibition on Computer Graphics and Interactive Techniques in Asia, 2025)
- Continual Reinforcement Learning by Planning with Online World Models (ICML, 2025)
- Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion (ICLR, 2023)
- CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving (Proceedings of the AAAI Conference on Artificial Intelligence, 2025)
- CorrectionPlanner: Self-Correction Planner with Reinforcement Learning in Autonomous Driving (arXiv, 2026)
- Cortex 2.0: Grounding World Models in Real-World Industrial Deployment (arXiv, 2026)
- Cosmos 3: Omnimodal World Models for Physical AI (arXiv, 2026)
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning (arXiv, 2026)
- Cosmos World Foundation Model Platform for Physical AI (arXiv, 2025)
- Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models (arXiv, 2025)
- Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling (arXiv, 2025)
- Cosmos-Surg-DVRK: World Foundation Model-Based Automated Online Evaluation of Surgical Robot Policy Learning (IEEE Robotics and Automation Letters, 2025)
- Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control (arXiv, 2025)
- Counterfactual World Models via Digital Twin-conditioned Video Diffusion (arXiv, 2025)
- CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation (arXiv, 2026)
- Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning (arXiv, 2025)
- Coupled Local and Global World Models for Efficient First Order RL (arXiv, 2026)
- CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving (arXiv, 2026)
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation (arXiv, 2025)
- CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving (arXiv, 2025)
- DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation (arXiv, 2026)
- DDP-WM: Disentangled Dynamics Prediction for Efficient World Models (arXiv, 2026)
- DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory (arXiv, 2026)
- Decoupled Diffusion Sparks Adaptive Scene Generation (ICCV, 2025)
- Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation (arXiv, 2025)
- Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression (arXiv, 2025)
- DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving (arXiv, 2026)
- DeepVerse: 4D Autoregressive Video Generation as a World Model (arXiv, 2025)
- DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos (arXiv, 2026)
- Denoising Hamiltonian Network for Physical Reasoning (Trans. Mach. Learn. Res., 2025)
- DeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied Tracking (arXiv, 2026)
- DexSim2Real$^{\mathbf{2}}$: Building Explicit World Model for Precise Articulated Object Dexterous Manipulation (IEEE Transactions on robotics, 2024)
- Dexterous World Models (arXiv, 2025)
- DexWorldModel: Causal Latent World Modeling towards Automated Learning of Embodied Tasks (arXiv, 2026)
- Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models (arXiv, 2025)
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion (NeurIPS, 2024)
- Diffusion Models for Video Prediction and Infilling (arXiv, 2022)
- Diffusion World Model (arXiv, 2024)
- Diffusion-Based Imaginative Coordination for Bimanual Manipulation (arXiv, 2025)
- Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning (arXiv, 2026)
- DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation (ICCV, 2025)
- Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving (arXiv, 2026)
- DiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer (arXiv, 2025)
- dm_control: Software and Tasks for Continuous Control (Software Impacts, 2020)
- DMWM: Dual-Mind World Model with Long-Term Imagination (arXiv, 2025)
- Doe-1: Closed-Loop Autonomous Driving with Large World Model (arXiv, 2024)
- DOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model (arXiv, 2024)
- DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration (arXiv, 2025)
- Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination (ICLR, 2024)
- Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation (arXiv, 2026)
- Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow (arXiv, 2025)
- DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos (arXiv, 2026)
- DreamDrive: Generative 4D Scene Modeling from Street View Images (ICRA, 2024)
- DreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving (arXiv, 2026)
- DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical Representations (ICML, 2021)
- DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes (arXiv, 2024)
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models (arXiv, 2025)
- DreamingV2: Reinforcement Learning with Discrete World Models without Reconstruction (IEEE/RJS International Conference on Intelligent RObots and Systems, 2022)
- Dreamitate: Real-World Visuomotor Policy Learning via Video Generation (CoRL, 2024)
- Dreamland: Controllable World Creation with Simulator and Generative Models (arXiv, 2025)
- DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models (arXiv, 2026)
- DreamWorld: Unified World Modeling in Video Generation (arXiv, 2026)
- DreamX-World 1.0: A General-Purpose Interactive World Model (arXiv, 2026)
- Drift-Resistant Navigation World Model with Anchored Epipolar Guidance (arXiv, 2026)
- DriveArena: A Closed-Loop Generative Simulation Platform for Autonomous Driving (ICCV, 2024)
- DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation (arXiv, 2024)
- DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning (arXiv, 2026)
- DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation (CVPR, 2024)
- DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving (arXiv, 2023)
- DriveFuture: Future-Aware Latent World Models for Autonomous Driving (arXiv, 2026)
- DriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous Driving (2024 IEEE International Automated Vehicle Validation Conference (IAVVC), 2024) — The advancement of autonomous driving technologies necessitates increasingly sophisticated methods for understanding and predicting real-world scenarios.
- DriveLaW:Unifying Planning and Video Generation in a Latent Driving World (arXiv, 2025)
- Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout (arXiv, 2026)
- DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment (ACM Multimedia, 2025)
- DriveVA: Video Action Models are Zero-Shot Drivers (arXiv, 2026)
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving (arXiv, 2025)
- DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving (arXiv, 2026)
- DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving (arXiv, 2026)
- DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous Driving (CVPR, 2024)
- DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving (ICCV, 2025)
- Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving (AAAI, 2024)
- Driving Into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving (CVPR, 2023)
- DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model (arXiv, 2023)
- DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes (IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025)
- DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive Transformers (ICCV, 2024)
- DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation (CVPR, 2024)
- DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT (arXiv, 2024)
- DSG-World: Learning a 3D Gaussian World Model from Dual State Videos (arXiv, 2025)
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model (arXiv, 2025)
- dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model (arXiv, 2026)
- DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes (ICLR, 2024)
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling (arXiv, 2025)
- DyWA: Dynamics-Adaptive World Action Model for Generalizable Non-Prehensile Manipulation (ICCV, 2025)
- E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control (arXiv, 2026)
- EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields (arXiv, 2026)
- EA-WM: Event-Aware World Models with Task-Specification Grounding for Long-Horizon Manipulation (arXiv, 2026)
- EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance (CVPR, 2025)
- Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents (CVPR, 2024)
- Efficient Reinforcement Learning by Guiding Generalist World Models with Non-Curated Data (arXiv, 2025)
- Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination (arXiv, 2026)
- Ego-centric Learning of Communicative World Models for Autonomous Driving (arXiv, 2025)
- Ego-Vision World Model for Humanoid Contact Planning (arXiv, 2025)
- Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis (arXiv, 2026)
- EgoExo-WM: Unlocking Exo Video for Ego World Models (arXiv, 2026)
- EgoForge: Goal-Directed Egocentric World Simulator (arXiv, 2026)
- EgoSim: Egocentric World Simulator for Embodied Interaction Generation (arXiv, 2026)
- Einstein World Models (arXiv, 2026)
- Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model (arXiv, 2025)
- EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling (arXiv, 2025)
- Embody4D: A Generalist 4D World Model for Embodied AI (arXiv, 2026)
- Empowering Multi-Robot Cooperation via Sequential World Models (arXiv, 2025)
- Emu3.5: Native Multimodal Models are World Learners (arXiv, 2025)
- Enactor: From Traffic Simulators to Surrogate World Models (Proceedings of the 12th International Conference on Vehicle Technology and Intelligent Transport Systems, 2026)
- End-to-End Driving with Online Trajectory Evaluation via BEV World Model (ICCV, 2025)
- End-to-End Training for Autoregressive Video Diffusion via Self-Resampling (arXiv, 2025)
- EnerVerse-AC: Envisioning Embodied Environments with Action Condition (arXiv, 2025)
- EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation (arXiv, 2025)
- Enhancing Physical Consistency in Lightweight World Models (arXiv, 2025)
- EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting (arXiv, 2026)
- Epona: Autoregressive Diffusion World Model for Autonomous Driving (ICCV, 2025)
- EponaV2: Driving World Model with Comprehensive Future Reasoning (arXiv, 2026)
- EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards (arXiv, 2026)
- EVA: An Embodied World Model for Future Video Anticipation (arXiv, 2024)
- ω-EVA: Envision, Verify, and Act with Latent Interactive World Models (arXiv, 2026)
- Evaluating Gemini Robotics Policies in a Veo World Simulator (arXiv, 2025)
- EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory (arXiv, 2025)
- Explicit World Models for Reliable Human-Robot Collaboration (arXiv, 2026)
- ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving (arXiv, 2026)
- FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction (arXiv, 2025)
- FAR-Drive: Frame-AutoRegressive Video Generation in Closed-Loop Autonomous Driving (arXiv, 2026)
- Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention (arXiv, 2026)
- Fast-WAM: Do World Action Models Need Test-time Future Imagination? (arXiv, 2026)
- Feedback World Model Enables Precise Guidance of Diffusion Policy (arXiv, 2026)
- FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model (arXiv, 2024)
- FlowDreamer: A RGB-D World Model With Flow-Based Motion Representations for Robot Manipulation (IEEE Robotics and Automation Letters, 2025)
- FlyMirage: A Fully Automated Generation Pipeline for Diverse and Scalable UAV Flight Data via Generative World Model (arXiv, 2026)
- FOLIAGE: Towards Physical Intelligence World Models Via Unbounded Surface Evolution (arXiv, 2025)
- Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals (arXiv, 2025)
- FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making (ICML, 2025)
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models (arXiv, 2025)
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction (arXiv, 2025)
- From Imitation to Exploration: End-to-end Autonomous Driving based on World Model (arXiv, 2024)
- From Pixels to Cooperation Multi Agent Reinforcement Learning based on Multimodal World Models (arXiv, 2025)
- From Virtual Games to Real-World Play (arXiv, 2025)
- Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion (ICML, 2026)
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving (arXiv, 2025)
- FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model (arXiv, 2026)
- FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model (arXiv, 2025)
- GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation (arXiv, 2025)
- GAIA-1: A Generative World Model for Autonomous Driving (arXiv, 2023)
- GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving (arXiv, 2025)
- Game-Theoretic Risk-Shaped Reinforcement Learning for Safe Autonomous Driving (arXiv, 2025)
- Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players (arXiv, 2026)
- GaussianAD: Gaussian-Centric End-to-End Autonomous Driving (arXiv, 2024)
- GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation (arXiv, 2026)
- GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation (arXiv, 2025)
- GaussTwin: Unified Simulation and Correction with Gaussian Splatting for Robotic Digital Twins (arXiv, 2026)
- GE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulation (arXiv, 2026)
- GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation (arXiv, 2026)
- GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control (CVPR, 2024)
- GEM: Gaussian Evolution Model for Occupancy Forecasting and Motion Planning (arXiv, 2026)
- GEM: Generating LiDAR World Model via Deformable Mamba (arXiv, 2026)
- Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation (arXiv, 2024)
- Generalization of World Models under Environmental Variability for Vision-based Quadrotor Navigation (arXiv, 2026)
- Generalized Predictive Model for Autonomous Driving (CVPR, 2024)
- Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control (arXiv, 2026)
- Generating Multimodal Driving Scenes via Next-Scene Prediction (CVPR, 2025)
- Generating Out-of-Distribution Scenarios Using Language Models (ICRA, 2024)
- Generating Traffic Scenarios via In-Context Learning to Learn Better Motion Planner (AAAI, 2024)
- Generative Scenario Rollouts for End-to-End Autonomous Driving (arXiv, 2026)
- Generative World Explorer (arXiv, 2024)
- Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report (arXiv, 2025)
- Generative World Models of Tasks: LLM-Driven Hierarchical Scaffolding for Embodied Agents (arXiv, 2025)
- Genesis: A Generative and Universal Physics Engine for Robotics and Beyond (arXiv, 2024)
- GenEx: Generating an Explorable World (arXiv, 2024)
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation (arXiv, 2025)
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation (arXiv, 2025)
- GenRL: Multimodal-foundation world models for generalization in embodied agents (NeurIPS, 2024)
- GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control (arXiv, 2025)
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling (arXiv, 2025)
- Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context (arXiv, 2026)
- Geometry-aware 4D Video Generation for Robot Manipulation (arXiv, 2025)
- Geometry-Aware Implicit Memory for Video World Models (arXiv, 2026)
- Geometry-Aware Rotary Position Embedding for Consistent Video World Model (arXiv, 2026)
- GeoWorld: Geometric World Models (arXiv, 2026)
- GHIL-Glue: Hierarchical Control with Filtered Subgoal Images (ICRA, 2024)
- GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning (arXiv, 2026)
- GigaBrain-0: A World Model-Powered Vision-Language-Action Model (arXiv, 2025)
- GigaWorld-0: World Models as Data Engine to Empower Embodied AI (arXiv, 2025)
- GigaWorld-Policy: An Efficient Action-Centered World--Action Model (arXiv, 2026)
- Glad: A Streaming Scene Generator for Autonomous Driving (ICLR, 2025)
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation (arXiv, 2024)
- GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment (arXiv, 2025)
- Grounded World Model for Semantically Generalizable Planning (arXiv, 2026)
- Grounding World Simulation Models in a Real-World Metropolis (arXiv, 2026)
- GRUtopia: Dream General Robots in a City at Scale (arXiv, 2024)
- GSDrive: Reinforcing Driving Policies by Multi-mode Future Trajectory Probing with 3D Gaussian Splatting Environment (arXiv, 2026)
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation (ICCV, 2025)
- H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model (arXiv, 2026)
- HAIC: Humanoid Agile Object Interaction Control via Dynamics-Aware World Model (arXiv, 2026)
- HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning (arXiv, 2026)
- HaM-World: Soft-Hamiltonian World Models with Selective Memory for Planning (arXiv, 2026)
- Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures (arXiv, 2026)
- HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models (arXiv, 2026)
- HARP: Autoregressive Latent Video Prediction with High-Fidelity Image Generator (arXiv, 2022)
- HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation (arXiv, 2026)
- HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation (ICCV, 2025)
- HERO: Hierarchical Extrapolation and Refresh for Efficient World Models (arXiv, 2025)
- Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training (arXiv, 2026)
- Hierarchical Model-Based Imitation Learning for Planning in Autonomous Driving (IEEE/RJS International Conference on Intelligent RObots and Systems, 2022)
- Hierarchical Planning with Latent World Models (arXiv, 2026)
- Hierarchical World Models as Visual Whole-Body Humanoid Controllers (ICLR, 2024)
- HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation (arXiv, 2026)
- History-Guided Video Diffusion (ICML, 2025)
- Holo-World: Unified Camera, Object and Weather Control for Video World Model (ICLR, 2026)
- HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving (arXiv, 2024)
- HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation (arXiv, 2026)
- Humanoid World Models: Open World Foundation Models for Humanoid Robotics (arXiv, 2025)
- HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels (arXiv, 2025)
- HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds (arXiv, 2026)
- $I^{\mathbf{2}}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting (ICCV, 2025)
- IC-World: In-Context Generation for Shared World Modeling (arXiv, 2025)
- IGen: Scalable Data Generation for Robot Learning from Open-World Images (arXiv, 2025)
- iMaC: Translating Actions into Motion and Contact Images for Embodied World Models (arXiv, 2026)
- ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving (arXiv, 2025)
- Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation (arXiv, 2025)
- Imagine-2-Drive: Leveraging High-Fidelity World Models via Multi-Modal Diffusion Policies (IEEE/RJS International Conference on Intelligent RObots and Systems, 2024)
- iMoWM: Taming Interactive Multi-Modal World Model for Robotic Manipulation (arXiv, 2025)
- IMWM: Intuition Models Complement World Models for Latent Planning (arXiv, 2026)
- Inference Time Policy Optimization for Offline RL with Differentiable World Models (arXiv, 2026)
- Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling (IEEE Robotics and Automation Letters, 2025)
- Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation (arXiv, 2025)
- InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models (ICCV, 2024)
- Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory (arXiv, 2026)
- Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout (arXiv, 2025)
- InfinityDrive: Breaking Time Limits in Driving World Models (arXiv, 2024)
- INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling (arXiv, 2026)
- InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model for Spatial Intelligence (arXiv, 2026)
- InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation (ICCV, 2025)
- Interactive World Simulator for Robot Policy Training and Evaluation (arXiv, 2026)
- Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation (arXiv, 2026)
- InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation (arXiv, 2026)
- IOI: Decoupling Kinematics and Physics for Interactive World Models (arXiv, 2026)
- IRASim: A Fine-Grained World Model for Robot Manipulation (ICCV, 2024)
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning (NeurIPS Datasets and Benchmarks, 2021)
- Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning (arXiv, 2025)
- iVideoGPT: Interactive VideoGPTs are Scalable World Models (NeurIPS, 2024)
- "Just in Time" World Modeling Supports Human Planning and Reasoning (arXiv, 2026)
- KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models (arXiv, 2025)
- Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation (arXiv, 2026)
- Kinematics-Aware Latent World Models for Data-Efficient Autonomous Driving (arXiv, 2026)
- Kling-Omni Technical Report (arXiv, 2025)
- LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation (arXiv, 2025)
- LaGen: Towards Autoregressive LiDAR Scene Generation (arXiv, 2025)
- Language Models Meet World Models: Embodied Experiences Enhance Language Models (NeurIPS, 2023)
- Language-Guided Traffic Simulation via Scene-Level Diffusion (CoRL, 2023)
- Large Video Planner Enables Generalizable Robot Control (arXiv, 2025)
- Latent Chain-of-Thought World Modeling for End-to-End Driving (arXiv, 2025)
- Latent Geometry Beyond Search: Amortizing Planning in World Models (arXiv, 2026)
- Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics Modeling (arXiv, 2026)
- Latent Spatial Memory for Video World Models (arXiv, 2026)
- Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving (arXiv, 2026)
- LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation (arXiv, 2025)
- LaWM: Least Action World Models for Long-Horizon Physical Consistency from Visual Observations (arXiv, 2026)
- Learn2Fold: Structured Origami Generation with World Model Planning (arXiv, 2026)
- Learning 3D Persistent Embodied World Models (arXiv, 2025)
- Learning Action-Conditional and Object-Centric Gaussian Splatting World Models for Rigid Objects (arXiv, 2026)
- Learning Bilevel Policies over Symbolic World Models for Long-Horizon Planning (arXiv, 2026)
- Learning Interactive Real-World Simulators (ICLR, 2023)
- Learning Interactive World Model for Object-Centric Reinforcement Learning (arXiv, 2025)
- Learning Massively Multitask World Models for Continuous Control (arXiv, 2025)
- Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving (ICRA, 2024)
- Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation (arXiv, 2026)
- Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression (arXiv, 2025)
- Learning Rollout from Sampling: An R1-Style Tokenized Traffic Simulation Model (IEEE Robotics and Automation Letters, 2026)
- Learning to Drive from a World Model (2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2025) — Most self-driving systems rely on hand-coded perception outputs and engineered driving rules.
- Learning to Generate 4D LiDAR Sequences (arXiv, 2025)
- Learning to Model the World with Language (ICML, 2023)
- Learning to Simulate Complex Physics with Graph Networks (ICML, 2020)
- Learning Universal Policies via Text-Guided Video Generation (NeurIPS, 2023)
- Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control (arXiv, 2025)
- Learning Vision-Language-Action World Models for Autonomous Driving (arXiv, 2026)
- Learning Visual Feature-Based World Models via Residual Latent Action (arXiv, 2026)
- Learning World Models for Interactive Video Generation (arXiv, 2025)
- LEGS: Fine-Tuning Teleop-Free VLAs for Humanoid Loco-manipulation in an Embodied Gaussian Splatting World (arXiv, 2026)
- LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences (AAAI, 2025)
- LidarDM: Generative LiDAR Simulation in a Generated World (ICRA, 2024)
- Lifting Embodied World Models for Planning and Control (arXiv, 2026)
- LinguaSim: Interactive Multi-Vehicle Testing Scenario Generation via Natural Language Instruction Based on Large Language Models (2025 IEEE 28th International Conference on Intelligent Transportation Systems (ITSC), 2025) — The generation of testing and training scenarios for autonomous vehicles has drawn significant attention.
- LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving (arXiv, 2025)
- LIVE: Long-horizon Interactive Video World Modeling (arXiv, 2026)
- LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models (arXiv, 2026)
- LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines (arXiv, 2024)
- LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving (arXiv, 2026)
- Logic-Guided Socially-aware Robot Navigation World Model (arXiv, 2025)
- Logic-in-frames: Dynamic keyframe search via visual semantic-logical verification for long video understanding (arXiv, 2025)
- LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model (arXiv, 2026)
- Long-Context Autoregressive Video Modeling with Next-Frame Prediction (arXiv, 2025)
- LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model (arXiv, 2025)
- LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE (arXiv, 2025)
- LongVie 2: Multimodal Controllable Ultra-Long Video World Model (arXiv, 2025)
- LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding (Proceedings of the AAAI Conference on Artificial Intelligence, 2025)
- Lumiere: A Space-Time Diffusion Model for Video Generation (SIGGRAPH Asia, 2024)
- LUMOS: Language-Conditioned Imitation Learning with World Models (ICRA, 2025)
- Lyra 2.0: Explorable Generative 3D Worlds (arXiv, 2026)
- MAD: Motion Appearance Decoupling for efficient Driving World Models (arXiv, 2026)
- MAGI-1: Autoregressive Video Generation at Scale (arXiv, 2025)
- MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control (ICCV, 2024)
- MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes (IEEE Workshop/Winter Conference on Applications of Computer Vision, 2024)
- MagicDrive: Street View Generation with Diverse 3D Geometry Control (ICLR, 2023)
- MagicTime: Time-Lapse Video Generation Models as Metamorphic Simulators (IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024)
- MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration (arXiv, 2025)
- ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance (IEEE International Conference on Acoustics, Speech, and Signal Processing, 2025)
- ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI (arXiv, 2024)
- Map-World: Masked Action planning and Path-Integral World Model for Autonomous Driving (arXiv, 2025)
- MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving (arXiv, 2026)
- Marble: A Multimodal World Model (arXiv, 2025)
- Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions (arXiv, 2025)
- MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction (CVPR, 2025)
- MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models (arXiv, 2026)
- Matrix-3D: Omnidirectional Explorable 3D World Generation (arXiv, 2025)
- Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation (arXiv, 2026)
- MemoryWAM: Efficient World Action Modeling with Persistent Memory (arXiv, 2026)
- MetaUrban: An Embodied AI Simulation Platform for Urban Micromobility (ICLR, 2024)
- MetaWorld-X: Hierarchical World Modeling via VLM-Orchestrated Experts for Humanoid Loco-Manipulation (arXiv, 2026)
- MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data (arXiv, 2026)
- MetaWorld: Skill Transfer and Composition in a Hierarchical World Model for Grounding High-Level Instructions (arXiv, 2026)
- MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving (arXiv, 2025)
- MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis (arXiv, 2025)
- MindCube: Spatial Mental Modeling from Limited Views (arXiv, 2025)
- MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving (arXiv, 2025)
- MindJourney: Test-Time Scaling with World Models for Spatial Reasoning (arXiv, 2025)
- Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models (arXiv, 2024)
- Model-Based Imitation Learning for Urban Driving (NeurIPS, 2022)
- MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator (arXiv, 2025)
- MosaicMem: Hybrid Spatial Memory for Controllable Video World Models (arXiv, 2026)
- Motion Prompting: Controlling Video Generation with Motion Trajectories (CVPR, 2024)
- MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation (arXiv, 2026)
- MotuBrain: An Advanced World Action Model for Robot Control (arXiv, 2026)
- Motus: A Unified Latent Action World Model (arXiv, 2025)
- MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer (arXiv, 2025)
- MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation (arXiv, 2025)
- MuJoCo: A Physics Engine for Model-Based Control (IROS, 2012)
- Multi-Task Interactive Robot Fleet Learning with Visual World Models (CoRL, 2024)
- MultiWorld: Scalable Multi-Agent Multi-View Video World Models (arXiv, 2026)
- MUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations (2025 IEEE Intelligent Vehicles Symposium (IV), 2023) — World models for autonomous driving have the potential to dramatically improve the reasoning capabilities of today's systems.
- MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation (arXiv, 2026)
- MWM: Mobile World Models for Action-Conditioned Consistent Prediction (arXiv, 2026)
- Nano World Models: A Minimalist Implementation of Future Video Prediction (arXiv, 2026)
- NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction (arXiv, 2025)
- Navigation World Models (CVPR, 2024)
- NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation (arXiv, 2026)
- NavWM: A Unified Navigation World Model for Foresight-Driven Planning (arXiv, 2026)
- NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos (arXiv, 2026)
- Neural World Models for Computer Vision (arXiv, 2023)
- NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models (arXiv, 2026)
- Not All Actions Are Equal: Rethinking Conditioning for Dexterous World Model (arXiv, 2026)
- NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos (arXiv, 2025)
- NuiWorld: Exploring a Scalable Framework for End-to-End Controllable World Generation (arXiv, 2026)
- NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation (arXiv, 2026)
- OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation (arXiv, 2026)
- Oasis: A Universe in a Transformer (arXiv, 2024)
- Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models (ICRA, 2025)
- OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving (arXiv, 2024)
- OccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner Framework (arXiv, 2025)
- OccSim: Multi-kilometer Simulation with Long-horizon Occupancy World Models (arXiv, 2026)
- OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving (IEEE Transactions on Image Processing, 2024)
- OccTENS: 3D Occupancy World Model via Temporal Next-Scale Prediction (IEEE Robotics and Automation Letters, 2025)
- Occupancy World Model for Robots (arXiv, 2025)
- OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving (ECCV, 2023)
- Olaf-World: Orienting Latent Actions for Video World Modeling (arXiv, 2026)
- OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation (arXiv, 2026)
- OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving (ACM Multimedia, 2025)
- OmniNWM: Omniscient Driving Navigation World Models (arXiv, 2025)
- OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation (arXiv, 2026)
- OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation (arXiv, 2026)
- One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy (arXiv, 2026)
- OptiWorld: Optimal Control for Video World Generation under Physical Constraints (arXiv, 2026)
- Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models (arXiv, 2025)
- OrbiSim: World Models as Differentiable Physics Engines for Embodied Intelligence (arXiv, 2026)
- ORV: 4D Occupancy-centric Robot Video Generation (arXiv, 2025)
- OSCAR: Omni-Embodiment Skeleton-Conditioned World Action Model for Robotics (arXiv, 2026)
- OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation (arXiv, 2025)
- Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space (AAAI, 2025)
- Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models (arXiv, 2026)
- Owl-1: Omni World Model for Consistent Long Video Generation (arXiv, 2024)
- PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation (arXiv, 2026)
- PAN: A World Model for General, Interactable, and Long-Horizon World Simulation (arXiv, 2025)
- Panacea: Panoramic and Controllable Video Generation for Autonomous Driving (CVPR, 2023)
- Pandora: Towards General World Model with Natural Language Actions and Video States (arXiv, 2024)
- PanoWorld: Geometry-Consistent Panoramic Video World Modeling (arXiv, 2026)
- Parallel Stochastic Gradient-Based Planning for World Models (arXiv, 2026)
- Particle-Grid Neural Dynamics for Learning Deformable Object Models from RGB-D Videos (Robotics, 2025)
- ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation (arXiv, 2025)
- Pathdreamer: A World Model for Indoor Navigation (ALVR, 2021)
- PCASim: Promptable Closed-loop Adversarial Simulation for Urban Traffic Environment (arXiv, 2026)
- PerlAD: Towards Enhanced Closed-Loop End-to-End Autonomous Driving With Pseudo-Simulation-Based Reinforcement Learning (IEEE Robotics and Automation Letters, 2026)
- PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation (arXiv, 2026)
- Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning (arXiv, 2026)
- Perspective-Shifted Neuro-Symbolic World Models: A Framework for Socially-Aware Robot Navigation (IEEE International Symposium on Robot and Human Interactive Communication, 2025)
- PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation (ECCV, 2024)
- Physical Informed Driving World Model (arXiv, 2024)
- PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models (arXiv, 2025)
- Physically Embodied Gaussian Splatting: A Realtime Correctable World Model for Robotics (arXiv, 2024)
- PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos (ICCV, 2025)
- PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis (arXiv, 2025)
- PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image (arXiv, 2025)
- PhyWorld: Physics-Faithful World Model for Video Generation (arXiv, 2026)
- PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation (arXiv, 2026)
- PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation (Robotics, 2025)
- PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop (ICML, 2025)
- Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model (arXiv, 2026)
- Planning to Explore via Self-Supervised World Models (ICML, 2020)
- Planning with Adaptive World Models for Autonomous Driving (ICRA, 2024)
- Planning with Diffusion for Flexible Behavior Synthesis (ICML, 2022)
- Planning with Reasoning using Vision Language World Model (arXiv, 2025)
- Playable Environments: Video Manipulation in Space and Time (CVPR, 2022)
- PlayerOne: Egocentric World Simulator (arXiv, 2025)
- PlayWorld: Learning Robot World Models from Autonomous Play (arXiv, 2026)
- PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation (arXiv, 2026)
- Policy-Driven World Model Adaptation for Robust Offline Model-based Reinforcement Learning (arXiv, 2025)
- Policy-Guided World Model Planning for Language-Conditioned Visual Navigation (arXiv, 2026)
- PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth (IEEE/RJS International Conference on Intelligent RObots and Systems, 2025)
- Pre-Trained Video Generative Models as World Simulators (Proceedings of the AAAI Conference on Artificial Intelligence, 2025)
- Prediction with Action: Visual Policy Learning via Joint Denoising Process (arXiv, 2024)
- Predictive but Not Plannable: RC-aux for Latent World Models (arXiv, 2026)
- Prisma-World: Camera-Controllable Multi-Agent Video World Model (arXiv, 2026)
- ProDrive: Proactive Planning for Autonomous Driving via Ego-Environment Co-Evolution (arXiv, 2026)
- Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins (Robotics, 2025)
- ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos (arXiv, 2025)
- ProPhy: Progressive Physical Alignment for Dynamic World Simulation (arXiv, 2025)
- PWM: Policy Learning with Multi-Task World Models (ICLR, 2024)
- Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation (arXiv, 2026)
- RadarGen: Automotive Radar Point Cloud Generation from Cameras (arXiv, 2025)
- RAE-NWM: Navigation World Model in Dense Visual Representation Space (arXiv, 2026)
- RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO (arXiv, 2026)
- Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2) (arXiv, 2025)
- RAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray Space (arXiv, 2026)
- RealEngine: Simulating Autonomous Driving in Realistic Context (arXiv, 2025)
- RealWonder: Real-Time Physical Action-Conditioned Video Generation (arXiv, 2026)
- Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving (arXiv, 2026)
- ReconDreamer++: Harmonizing Generative and Reconstructive Models for Driving Scene Representation (ICCV, 2025)
- ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration (CVPR, 2024)
- Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control (arXiv, 2025)
- Reinforcing Action Policies by Prophesying (arXiv, 2025)
- Reinforcing VLAs in Task-Agnostic World Models (arXiv, 2026)
- Remote Sensing-Oriented World Model (arXiv, 2025)
- Renderworld: World Model with Self-Supervised 3D Label (ICRA, 2024)
- RepWAM: World Action Modeling with Representation Visual-Action Tokenizers (arXiv, 2026)
- ReSim: Reliable World Simulation for Autonomous Driving (arXiv, 2025)
- ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving (arXiv, 2026)
- Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks (arXiv, 2025)
- Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation (arXiv, 2025)
- ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models (arXiv, 2026)
- RISE: Self-Improving Robot Policy with Compositional World Model (arXiv, 2026)
- Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving (arXiv, 2026)
- Risk-Controllable Multi-View Diffusion for Driving Scenario Generation (arXiv, 2026)
- RLVR-World: Training World Models with Reinforcement Learning (arXiv, 2025)
- RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models (arXiv, 2026)
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots (arXiv, 2024)
- RoboDream: Compositional World Models for Scalable Robot Data Synthesis (arXiv, 2026)
- RoboDreamer: Learning Compositional World Models for Robot Imagination (ICML, 2024)
- RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation (IEEE/RJS International Conference on Intelligent RObots and Systems, 2025)
- RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation (arXiv, 2025)
- RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL (arXiv, 2025)
- RoboScape: Physics-informed Embodied World Model (arXiv, 2025)
- Robot Learning from a Physical World Model (arXiv, 2025)
- Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics (arXiv, 2025)
- RTFM: A Real-Time Frame Model (arXiv, 2025)
- RynnVLA-002: A Unified Vision-Language-Action and World Model (arXiv, 2025)
- S-NeRF++: Autonomous Driving Simulation via Neural Reconstruction and Generation (IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024)
- Safe Planning and Policy Optimization via World Model Learning (European Conference on Artificial Intelligence, 2025)
- Safe Urban Traffic Control via Uncertainty-Aware Conformal Prediction and World-Model Reinforcement Learning (arXiv, 2026)
- SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models (arXiv, 2025)
- SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer (arXiv, 2026)
- SAPIEN: A SimulAted Part-based Interactive ENvironment (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020)
- SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation (arXiv, 2026)
- Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation (arXiv, 2026)
- Scalable Policy Evaluation with Video World Models (arXiv, 2025)
- Scaling Cross-Embodiment World Models for Dexterous Manipulation (arXiv, 2025)
- Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds (arXiv, 2026)
- Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method (IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025)
- Scaling World Model for Hierarchical Manipulation Policies (arXiv, 2026)
- Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization (arXiv, 2026)
- SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model (CVPR, 2025)
- SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and Rollout (NeurIPS, 2024)
- See Tomorrow, Act Today: Foresight-Driven Autonomous Driving (arXiv, 2026)
- Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention (arXiv, 2024)
- Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation (arXiv, 2025)
- SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation (arXiv, 2025)
- Self-Correcting VLA: Online Action Refinement via Sparse World Imagination (arXiv, 2026)
- Self-Forcing++: Towards Minute-Scale High-Quality Video Generation (arXiv, 2025)
- Semantic Communications with World Models (arXiv, 2025)
- Semantic World Models (arXiv, 2025)
- Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving (ICLR, 2025)
- Sensorimotor World Models: Perception for Action via Inverse Dynamics (arXiv, 2026)
- ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling (arXiv, 2026)
- SimGen: Simulator-conditioned Driving Scene Generation (NeurIPS, 2024)
- Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation (arXiv, 2026)
- SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds (arXiv, 2025)
- SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models (arXiv, 2026)
- SlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric Models (ICLR, 2022)
- Solving Motion Planning Tasks with a Scalable Generative Model (ECCV, 2024)
- Sora Generates Videos with Stunning Geometrical Consistency (arXiv, 2024)
- SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Time (arXiv, 2025)
- Sparse Imagination for Efficient Visual World Model Planning (arXiv, 2025)
- SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning (arXiv, 2026)
- SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model (arXiv, 2025)
- SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries (Proceedings of the AAAI Conference on Artificial Intelligence, 2025)
- SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation (arXiv, 2026)
- Spatia: Video Generation with Updatable Spatial Memory (arXiv, 2025)
- ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation (arXiv, 2026)
- StableWorld: Towards Stable and Consistent Long Interactive Video Generation (arXiv, 2026)
- Stag-1: Towards Realistic 4D Driving Simulation with Video Generation Model (arXiv, 2024)
- STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation (IEEE/RJS International Conference on Intelligent RObots and Systems, 2025)
- STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation (arXiv, 2026)
- StateSpaceDiffuser: Bringing Long Context to Diffusion World Models (arXiv, 2025)
- Stereo World Model: Camera-Guided Stereo Video Generation (arXiv, 2026)
- Stochastic Variational Video Prediction (ICLR, 2018)
- STORM: Search-Guided Generative World Models for Robotic Manipulation (arXiv, 2025)
- Street Gaussians: Modeling Dynamic Urban Scenes with Gaussian Splatting (arXiv, 2024)
- Structured World Models from Human Videos (RSS, 2023)
- SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation (arXiv, 2026)
- Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training (arXiv, 2026)
- Symphony: Learning Realistic and Diverse Agents for Autonomous Driving Simulation (ICRA, 2022)
- Synthetic Video Enhances Physical Fidelity in Video Synthesis (arXiv, 2025)
- T3Former: Delta-Triplane Transformers as Occupancy World Models (arXiv, 2025)
- Tactile-WAM: Touch-Aware World Action Model with Tactile Asymmetric Attention (arXiv, 2026)
- TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion (arXiv, 2026)
- Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution (arXiv, 2026)
- TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model (arXiv, 2025)
- Temporal Difference Flows (ICML, 2025)
- TeraSim-World: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous Driving (arXiv, 2025)
- Terra: Explorable Native 3D World Model with Point Latents (arXiv, 2025)
- TesserAct: Learning 4D Embodied World Models (arXiv, 2025)
- Test-Time Mixture of World Models for Embodied Agents in Dynamic Environments (arXiv, 2026)
- The DAWN of World-Action Interactive Models (arXiv, 2026)
- The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control (arXiv, 2024)
- The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text (arXiv, 2025)
- Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration? (ICLR, 2026)
- Think2Drive: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving (in CARLA-v2) (arXiv, 2024)
- Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators (arXiv, 2026)
- Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning (arXiv, 2026)
- Toward Physically Consistent Driving Video World Models under Challenging Trajectories (arXiv, 2026)
- Toward Safe Autonomous Robotic Endovascular Interventions using World Models (arXiv, 2026)
- Towards Accurate Generative Models of Video: A New Metric & Challenges (arXiv, 2018)
- Towards foundational LiDAR world models with efficient latent flow matching (arXiv, 2025)
- Towards High-Consistency Embodied World Model with Multi-View Trajectory Videos (arXiv, 2025)
- Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models (arXiv, 2026)
- TrafficBots: Towards World Models for Autonomous Driving Simulation and Motion Prediction (ICRA, 2023)
- Trajectory World Models for Heterogeneous Environments (ICML, 2025)
- TRAP: Tail-aware Ranking Attack for World-Model Planning (arXiv, 2026)
- TriVLA: A Triple-System-Based Unified Vision-Language-Action Model with Episodic World Modeling for General Robot Control (arXiv, 2025)
- U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences (arXiv, 2025)
- UCM: Unifying Camera Control and Memory with Time-aware Positional Encoding Warping for World Models (arXiv, 2026)
- Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning Work on Real Robots (arXiv, 2025)
- Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving (arXiv, 2026)
- UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving (arXiv, 2026)
- UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving (arXiv, 2026)
- UniDWM: Towards a Unified Driving World Model via Multifaceted Representation Learning (arXiv, 2026)
- Unified 3D Scene Understanding Through Physical World Modeling (arXiv, 2026)
- Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising (arXiv, 2026)
- Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process (arXiv, 2025)
- Unified Driving Tokens: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and Planning (arXiv, 2026)
- Unified Video Action Model (Robotics, 2025)
- Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets (Robotics, 2025)
- UniFuture: A 4D Driving World Model for Future Generation and Perception (arXiv, 2025)
- UniMLVG: Unified Framework for Multi-View Long Video Generation with Comprehensive Control Capabilities for Autonomous Driving (ICCV, 2024)
- UniScene: Unified Occupancy-centric Driving Scene Generation (CVPR, 2024)
- UniSim: A Neural Closed-Loop Sensor Simulator (CVPR, 2023)
- UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling (arXiv, 2026)
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation (arXiv, 2025)
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving (arXiv, 2025)
- UniWorld: Autonomous Driving Pre-training via World Models (arXiv, 2023)
- Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation (arXiv, 2024)
- Unlocking Efficient Vehicle Dynamics Modeling via Analytic World Models (AAAI, 2025)
- UnO: Unsupervised Occupancy Fields for Perception and Forecasting (CVPR, 2024)
- Unraveling the Effects of Synthetic Data on End-to-End Autonomous Driving (ICCV, 2025)
- UrbanWorld: An Urban World Model for 3D City Generation (arXiv, 2024)
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv, 2025)
- VaViM and VaVAM: Autonomous Driving through Video Generative Modeling (arXiv, 2025)
- VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models (IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025)
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness (arXiv, 2025)
- VChain: Chain-of-Visual-Thought for Reasoning in Video Generation (arXiv, 2025)
- VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation (arXiv, 2025)
- VectorWorld: Efficient Streaming World Model via Diffusion Flow on Vector Graphs (arXiv, 2026)
- Vega: Learning to Drive with Natural Language Instructions (arXiv, 2026)
- Vehicle Dynamics Embedded World Models for Autonomous Driving (IEEE transactions on intelligent transportation systems (Print), 2025) — World models have gained significant attention as a promising approach for autonomous driving.
- VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control (arXiv, 2026)
- Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation (CVPR, 2025)
- Vid2World: Crafting Video Diffusion Models to Interactive World Models (arXiv, 2025)
- Vidar: Embodied Video Diffusion Model for Generalist Manipulation (arXiv, 2025)
- Vidarc: Embodied Video Diffusion Model for Closed-loop Control (arXiv, 2025)
- Video Generators are Robot Policies (arXiv, 2025)
- Video Language Planning (ICLR, 2023)
- Video models are zero-shot learners and reasoners (arXiv, 2025)
- Video World Models with Long-term Spatial Memory (arXiv, 2025)
- Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024)
- VideoPoet: A Large Language Model for Zero-Shot Video Generation (ICML, 2023)
- VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory (arXiv, 2025)
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators (arXiv, 2025)
- Vision-Centric 4D Occupancy Forecasting and Planning via Implicit Residual World Models (arXiv, 2025)
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future (arXiv, 2025)
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability (NeurIPS, 2024)
- VISTAv2: World Imagination for Indoor Vision-and-Language Navigation (arXiv, 2025)
- Visual Foresight for Robotic Stow: A Diffusion-Based World Model from Sparse Snapshots (arXiv, 2026)
- VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning (ICLR, 2024)
- Visuo-Tactile World Models (arXiv, 2026)
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search (arXiv, 2025)
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators (arXiv, 2025)
- VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model (arXiv, 2026)
- VLM-DEWM: Dynamic External World Model for Verifiable and Resilient Vision-Language Planning in Manufacturing (arXiv, 2026)
- VLM-SAFE: Vision-Language Model-Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving (arXiv, 2025)
- VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory (ICCV, 2025)
- Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation (ACM Transactions on Graphics, 2025)
- Walk through Paintings: Egocentric World Models from Internet Priors (arXiv, 2026)
- WALL-WM: Carving World Action Modeling at the Event Joints (arXiv, 2026)
- WAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT (arXiv, 2026)
- WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation (arXiv, 2026)
- Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data (arXiv, 2026)
- WHALE: Towards Generalizable and Scalable World Models for Embodied Decision-making (arXiv, 2024)
- When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning (arXiv, 2026)
- When Digital Twins Meet Large Language Models: Realistic, Interactive, and Editable Simulation for Autonomous Driving (IEEE Robotics & Automation Magazine, 2025)
- When to Trust Imagination: Adaptive Action Execution for World Action Models (arXiv, 2026)
- When World Models Dream Wrong: Physical-Conditioned Adversarial Attacks against World Models (arXiv, 2026)
- Whole-Body Conditioned Egocentric Video Prediction (arXiv, 2025)
- WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation (arXiv, 2025)
- WM-DAgger: Enabling Efficient Data Aggregation for Imitation Learning with World Models (arXiv, 2026)
- WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation (IEEE/RJS International Conference on Intelligent RObots and Systems, 2025)
- WMPO: World Model-based Policy Optimization for Vision-Language-Action Models (arXiv, 2025)
- WoMAP: World Models For Embodied Open-Vocabulary Object Localization (arXiv, 2025)
- WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration (arXiv, 2025)
- WonderTurbo: Generating Interactive 3D World in 0.72 Seconds (arXiv, 2025)
- WonderZoom: Multi-Scale 3D World Generation (arXiv, 2025)
- World Action Models are Zero-shot Policies (arXiv, 2026)
- World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays (arXiv, 2026)
- World Guidance: World Modeling in Condition Space for Action Generation (arXiv, 2026)
- World Model Self-Distillation: Training World Models to Solve General Tasks (arXiv, 2026)
- World model-based end-to-end scene generation for accident anticipation in autonomous driving (Communications Engineer, 2025)
- World Model-Based Learning for Long-Term Age of Information Minimization in Vehicular Networks (Global Communications Conference, 2025)
- World Model-Enabled Causal Digital Twins for Semantic Communications in Physical AI Systems (arXiv, 2026)
- World Modeling with Probabilistic Structure Integration (arXiv, 2025)
- World Models for General Surgical Grasping (arXiv, 2024)
- World Models for Learning Dexterous Hand-Object Interactions from Human Videos (arXiv, 2025)
- World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning (arXiv, 2026)
- World Models via Policy-Guided Trajectory Diffusion (Trans. Mach. Learn. Res., 2023)
- World Models with Hints of Large Language Models for Goal Achieving (NAACL, 2024)
- World Simulation with Video Foundation Models for Physical AI (arXiv, 2025)
- World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks (arXiv, 2026)
- World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training (arXiv, 2025)
- World-R1: Reinforcing 3D Constraints for Text-to-Video Generation (arXiv, 2026)
- World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems (arXiv, 2026)
- World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy (arXiv, 2026)
- World2Act: Latent Action Post-Training via Skill-Compositional World Models (arXiv, 2026)
- World4Drive: End-to-End Autonomous Driving via Intention-Aware Physical Latent World Model (ICCV, 2025)
- World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation (arXiv, 2025)
- WorldAgents: Can Foundation Image Models be Agents for 3D World Models? (arXiv, 2026)
- WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching (arXiv, 2026)
- WorldCache: Content-Aware Caching for Accelerated Video World Models (arXiv, 2026)
- WorldCompass: Reinforcement Learning for Long-Horizon World Models (arXiv, 2026)
- WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models (arXiv, 2026)
- WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens (arXiv, 2024)
- WorldEval: World Model as Real-World Robot Policies Evaluator (arXiv, 2025)
- WorldExplorer: Towards Generating Fully Navigable 3D Scenes (ACM SIGGRAPH Conference and Exhibition on Computer Graphics and Interactive Techniques in Asia, 2025)
- WorldFlow3D: Flowing Through 3D Distributions for Unbounded World Generation (arXiv, 2026)
- WorldGrow: Generating Infinite 3D World (Proceedings of the AAAI Conference on Artificial Intelligence, 2025)
- WorldGym: World Model as An Environment for Policy Evaluation (arXiv, 2025)
- WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World (arXiv, 2025)
- WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models (arXiv, 2026)
- WorldPlanner: Monte Carlo Tree Search and MPC with Action-Conditioned Visual World Models (arXiv, 2025)
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling (arXiv, 2025)
- WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving (AAAI, 2025)
- WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning (arXiv, 2026)
- WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation (arXiv, 2026)
- WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation (ECCV, 2023)
- WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL (arXiv, 2026)
- WoW: Towards a World omniscient World model Through Embodied Interaction (arXiv, 2025)
- WPT: World-to-Policy Transfer via Online World Model Distillation (arXiv, 2025)
- WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation (arXiv, 2025)
- X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference (arXiv, 2026)
- X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling (arXiv, 2026)
- X-MOBILITY: End-to-End Generalizable Navigation via World Modeling (ICRA, 2024)
- X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability (arXiv, 2025)
- X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving (arXiv, 2026)
- Xiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving (arXiv, 2026)
- Yume-1.5: A Text-Controlled Interactive World Generation Model (arXiv, 2025)
- Yume: An Interactive World Generation Model (arXiv, 2025)
- ActionParty: Multi-Subject Action Binding in Generative Video Games (arXiv, 2026)
- ActWorld: From Explorable to Interactive World Model via Action-Aware Memory (arXiv, 2026)
- Advancing Open-source World Models (arXiv, 2026)
- Agent Learning via Early Experience (arXiv, 2025)
- Agent Planning with World Knowledge Model (NeurIPS, 2024)
- Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning (arXiv, 2026)
- Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback (arXiv, 2025)
- Agentic World Modeling for 6G: Near-Real-Time Generative State-Space Reasoning (arXiv, 2025)
- Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction (ICML, 2024)
- AGWM: Affordance-Grounded World Models for Environments with Compositional Prerequisites (arXiv, 2026)
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning (ICLR, 2020)
- AlignUSER: Human-Aligned LLM Agents via World Models for Recommender System Evaluation (arXiv, 2026)
- AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction (ICCV, 2025)
- Better Decisions through the Right Causal World Model (arXiv, 2025)
- Beyond Static Forecasting: Unleashing the Power of World Models for Mobile Traffic Extrapolation (arXiv, 2026)
- Bringing Stories Alive: Generating Interactive Fiction Worlds (Artificial Intelligence and Interactive Digital Entertainment Conference, 2020)
- Code World Models for General Game Playing (arXiv, 2025)
- Code World Models for Parameter Control in Evolutionary Algorithms (arXiv, 2026)
- Code2World: A GUI World Model via Renderable Code Generation (arXiv, 2026)
- COMBAT: Conditional World Models for Behavioral Agent Training (arXiv, 2026)
- Computer-Using World Model (arXiv, 2026)
- Computing Machinery and Intelligence (Mind, 1950)
- CWM: An Open-Weights LLM for Research on Code Generation with World Models (arXiv, 2025)
- Debugging code world models (arXiv, 2026)
- Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation (arXiv, 2026)
- Diffusion Models Are Real-Time Game Engines (ICLR, 2024)
- Distilling Game Code World Model Generation into Lightweight Large Language Models (arXiv, 2026)
- DreamX-World 1.0: A General-Purpose Interactive World Model (arXiv, 2026)
- Dual-Mind World Models: A General Framework for Learning in Dynamic Wireless Networks (arXiv, 2025)
- Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI Agents (arXiv, 2025)
- DynaWeb: Model-Based Reinforcement Learning of Web Agents (arXiv, 2026)
- ECHO: Terminal Agents Learn World Models for Free (arXiv, 2026)
- Emergence of Social Norms in Large Language Model-based Agent Societies (arXiv, 2024)
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task (ICLR, 2023)
- Explainable Reinforcement Learning Agents Using World Models (arXiv, 2025)
- From Virtual Games to Real-World Play (arXiv, 2025)
- From Word to World: Can Large Language Models be Implicit Text-based World Models? (arXiv, 2025)
- From Zero to Hero: Training-Free Custom Concept Spawning in World Models (arXiv, 2026)
- Game-tars: Pretrained foundation models for scalable generalist multimodal game agents (arXiv, 2025)
- GameFactorly: Creating New Games with Generative Interactive Videos (ICCV, 2025)
- GameGen-X: Interactive Open-world Game Video Generation (ICLR, 2024)
- Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players (arXiv, 2026)
- General Agentic Planning Through Simulative Reasoning with World Models (arXiv, 2025)
- Generating Code World Models with Large Language Models Guided by Monte Carlo Tree Search (NeurIPS, 2024)
- Generating Symbolic World Models via Test-time Scaling of Large Language Models (Trans. Mach. Learn. Res., 2025)
- Generative Agents: Interactive Simulacra of Human Behavior (ACM Symposium on User Interface Software and Technology, 2023)
- Generative Visual Code Mobile World Models (arXiv, 2026)
- Genie: Generative Interactive Environments (ICML, 2024)
- Geometry-Aware Rotary Position Embedding for Consistent Video World Model (arXiv, 2026)
- Graph World Model (ICML, 2025)
- Grounded Answers for Multi-agent Decision-making Problem through Generative World Model (NeurIPS, 2024)
- GTM: Simulating the World of Tools for AI Agents (arXiv, 2025)
- Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models (arXiv, 2026)
- How Mobile World Model Guides GUI Agents? (arXiv, 2026)
- Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model (arXiv, 2025)
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition (arXiv, 2025)
- HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels (arXiv, 2025)
- Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models (arXiv, 2026)
- Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models (arXiv, 2026)
- Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation (arXiv, 2025)
- IPR-1: Interactive Physical Reasoner (arXiv, 2025)
- Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents (Trans. Mach. Learn. Res., 2024)
- Language Agents Meet Causality - Bridging LLMs and Causal World Models (arXiv, 2024)
- Learning Ad Hoc Network Dynamics via Graph-Structured World Models (arXiv, 2026)
- Learning Interactive Real-World Simulators (ICLR, 2023)
- Learning POMDP World Models from Observations with Language-Model Priors (arXiv, 2026)
- Learning to Simulate Dynamic Environments With GameGAN (CVPR, 2020)
- Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task Planning (NeurIPS, 2023)
- LIVE: Long-horizon Interactive Video World Modeling (arXiv, 2026)
- LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training (arXiv, 2025)
- Long-Context State-Space Video World Models (ICCV, 2025)
- LongLive: Real-time Interactive Long Video Generation (arXiv, 2025)
- macOSWorld: A Multilingual Interactive Benchmark for GUI Agents (arXiv, 2025)
- MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model (arXiv, 2026)
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory (arXiv, 2026)
- Matrix-Game: Interactive World Foundation Model (arXiv, 2025)
- MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments (arXiv, 2026)
- Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft (arXiv, 2025)
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework (ICLR, 2024)
- MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World Control (arXiv, 2024)
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft (arXiv, 2025)
- MobileDreamer: Generative Sketch World Model for GUI Agent (arXiv, 2026)
- MobiWorld: World Models for Mobile Wireless Network (arXiv, 2025)
- Model as a Game: On Numerical and Spatial Consistency for Generative Games (2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025) — Recent advances in generative models have significantly impacted game generation.
- MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines (arXiv, 2026)
- MultiWorld: Scalable Multi-Agent Multi-View Video World Models (arXiv, 2026)
- NeuralOS: Towards Simulating Operating Systems via Neural Generative Models (arXiv, 2025)
- Neuro-Symbolic Synergy for Interactive World Modeling (arXiv, 2026)
- NitroGen: An Open Foundation Model for Generalist Gaming Agents (arXiv, 2026)
- NuiWorld: Exploring a Scalable Framework for End-to-End Controllable World Generation (arXiv, 2026)
- Object-Centric World Models Meet Monte Carlo Tree Search (arXiv, 2026)
- OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation (arXiv, 2026)
- One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration (arXiv, 2025)
- PAN: A World Model for General, Interactable, and Long-Horizon World Simulation (arXiv, 2025)
- PathWise: Planning through World Model for Automated Heuristic Design via Self-Evolving LLMs (arXiv, 2026)
- Playable Video Generation (CVPR, 2021)
- PoE-World: Compositional World Modeling with Products of Programmatic Experts (arXiv, 2025)
- Probabilistic Dreaming for World Models (arXiv, 2026)
- Project Sid: Many-agent simulations toward AI civilization (arXiv, 2024)
- Promptable Game Models: Text-guided Game Simulation via Masked Diffusion Models (ACM Transactions on Graphics, 2023)
- PROWL: Prioritized Regret-Driven Optimization for World Model Learning (arXiv, 2026)
- R-WoM: Retrieval-augmented World Model For Computer-use Agents (arXiv, 2025)
- RAGEN-2: Reasoning Collapse in Agentic RL (arXiv, 2026)
- Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning (arXiv, 2025)
- ReactiveGWM: Steering NPC in Reactive Game World Models (arXiv, 2026)
- Reasoning with Language Model is Planning with World Model (EMNLP, 2023)
- Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention (arXiv, 2025)
- Reinforcement World Model Learning for LLM-based Agents (arXiv, 2026)
- RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy (arXiv, 2025)
- RLVR-World: Training World Models with Reinforcement Learning (arXiv, 2025)
- Rolling Forcing: Autoregressive Long Video Diffusion in Real Time (arXiv, 2025)
- SafePred: A Predictive Guardrail for Computer-Using Agents via World Models (arXiv, 2026)
- Scaling Agent Learning via Experience Synthesis (arXiv, 2025)
- Scaling Instructable Agents Across Many Simulated Worlds (arXiv, 2024)
- SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models (arXiv, 2026)
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion (arXiv, 2025)
- ShowUI: One vision-language-action model for GUI visual agent (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025)
- Simulating Environments with Reasoning Models for Agent Training (arXiv, 2025)
- Solaris: Building a Multiplayer Video World Model in Minecraft (arXiv, 2026)
- Solving the Frame Problem: A Mathematical Investigation of the Common Sense Law of Inertia (MIT Press, 1997)
- Some Philosophical Problems from the Standpoint of Artificial Intelligence (Machine Intelligence, 1969)
- StableWorld: Towards Stable and Consistent Long Interactive Video Generation (arXiv, 2026)
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (NeurIPS, 2024)
- SWE-World: Building Software Engineering Agents in Docker-Free Environments (arXiv, 2026)
- Synthesizing world models for bilevel planning (Trans. Mach. Learn. Res., 2025)
- Telecom World Models: Unifying Digital Twins, Foundation Models, and Predictive Planning for 6G (arXiv, 2026)
- The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control (arXiv, 2024)
- Training Agents Inside of Scalable World Models (arXiv, 2025)
- Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning (arXiv, 2025)
- UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026)
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents (arXiv, 2025)
- UISim: An Interactive Image-Based UI Simulator for Dynamic Mobile Environments (arXiv, 2025)
- Unbounded: A Generative Infinite Game of Character Life Simulation (ICLR, 2024)
- Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach (EMNLP, 2025)
- VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents (arXiv, 2025)
- Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024)
- ViMo: A Generative Visual GUI World Model for App Agent (arXiv, 2025)
- VJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World Models (arXiv, 2026)
- Voyager: An Open-Ended Embodied Agent with Large Language Models (TMLR, 2024)
- Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation (ICLR, 2024)
- Web World Models (arXiv, 2025)
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents (Advances in Neural Information Processing Systems, 2022)
- WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis (arXiv, 2025)
- WebWorld: A Large-Scale World Model for Web Agent Training (arXiv, 2026)
- World Craft: Agentic Framework to Create Visualizable Worlds via Text (arXiv, 2026)
- World Model on Million-Length Video And Language With Blockwise RingAttention (ICLR, 2024)
- World Models for Policy Refinement in StarCraft II (arXiv, 2026)
- World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning (arXiv, 2026)
- World of Bits: An Open-Domain Platform for Web-Based Agents (ICML, 2017)
- World-Model-Augmented Web Agents with Action Correction (arXiv, 2026)
- WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation (arXiv, 2026)
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment (NeurIPS, 2024)
- WorldGen: From Text to Traversable and Interactive 3D Worlds (arXiv, 2025)
- WorldKV: Efficient World Memory with World Retrieval and Compression (arXiv, 2026)
- WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling (arXiv, 2025)
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling (arXiv, 2025)
- Yan: Foundational Interactive Video Generation (arXiv, 2025)
- A game-theoretic approach to normative multi-agent systems (Normative Multi-agent Systems, 2007)
- ActionParty: Multi-Subject Action Binding in Generative Video Games (arXiv, 2026)
- Agentifying Agentic AI (arXiv, 2025)
- AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles (arXiv, 2026)
- AlignUSER: Human-Aligned LLM Agents via World Models for Recommender System Evaluation (arXiv, 2026)
- Bayesian theory of mind: Modeling joint belief-desire attribution (Annual Meeting of the Cognitive Science Society, 2011)
- BDI agents: From theory to practice (International Conference on Multi-Agent Systems, 1995)
- Boosting LLM agents with recursive contemplation for effective deception handling (ACL, 2024)
- CAMS: A CityGPT-Powered Agentic Framework for Urban Human Mobility Simulation (arXiv, 2025)
- Causal Cartographer: From Mapping to Reasoning Over Counterfactual Worlds (arXiv, 2025)
- Character-LLM: A trainable agent for role-playing (EMNLP, 2023)
- ChatDev: Communicative Agents for Software Development (ACL, 2024)
- ChatHaruhi: Reviving Anime Character in Reality via Large Language Model (arXiv, 2023)
- Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents (Advances in Neural Information Processing Systems, 2024)
- Cultural Evolution of Cooperation among LLM Agents (International Conference on Autonomous Agents and Multiagent Systems, 2024)
- Deal or No Deal? End-to-End Learning of Negotiation Dialogues (EMNLP, 2017)
- Decentralized Collective World Model for Emergent Communication and Coordination (International Conference on Development and Learning, 2025)
- Decoupling strategy and generation in negotiation dialogues (EMNLP, 2018)
- DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration (arXiv, 2025)
- Efficient Generation of Diverse Cooperative Agents with World Models (arXiv, 2025)
- Ego-centric Learning of Communicative World Models for Autonomous Driving (arXiv, 2025)
- Emergence of Social Norms in Large Language Model-based Agent Societies (arXiv, 2024)
- Emergent social conventions and collective bias in LLM populations (Science Advances, 2025)
- Explicit World Models for Reliable Human-Robot Collaboration (arXiv, 2026)
- From Word Models to World Models: Translating from Natural Language to the Probabilistic Language of Thought (arXiv, 2023)
- Game-theoretic LLM: Agent Workflow for Negotiation Games (arXiv, 2024)
- Generative Agents: Interactive Simulacra of Human Behavior (ACM Symposium on User Interface Software and Technology, 2023)
- Generative Social Choice (ACM Conference on Economics and Computation, 2024)
- Generative World Models of Tasks: LLM-Driven Hierarchical Scaffolding for Embodied Agents (arXiv, 2025)
- Grounded Answers for Multi-agent Decision-making Problem through Generative World Model (NeurIPS, 2024)
- GRUtopia: Dream General Robots in a City at Scale (arXiv, 2024)
- How well can LLMs negotiate? negotiationarena platform and analysis (arXiv, 2024)
- Human-level play in the game of diplomacy by combining language models with strategic reasoning (Science, 2022)
- Hypothesis-driven theory-of-mind reasoning for large language models (arXiv, 2025)
- K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic Reasoning (Findings of the Association for Computational Linguistics: NAACL, 2025)
- LaMP: When Large Language Models Meet Personalization (ACL, 2024)
- Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game (ICML, 2023)
- LLM-based agent society investigation: Collaboration and confrontation in avalon gameplay (EMNLP, 2024)
- LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment (EMNLP, 2025)
- Logic-Guided Socially-aware Robot Navigation World Model (arXiv, 2025)
- MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model (arXiv, 2026)
- Maintenance of Social Commitments in Multiagent Systems (AAAI, 2021)
- MASim: Multilingual Agent-Based Simulation for Social Science (arXiv, 2025)
- Minding language models (lack of) theory of mind: A plug-and-play multi-character belief tracker (ACL, 2023)
- Network formation and dynamics among multi-LLMs (PNAS Nexus, 2024)
- OASIS: Open Agent Social Interaction Simulations with One Million Agents (arXiv, 2024)
- On the formal specification of electronic institutions (Agent Mediated Electronic Commerce: The European AgentLink Perspective, 2001)
- Out of One, Many: Using Language Models to Simulate Human Samples (Political Analysis, 2023)
- PersonaGym: Evaluating Persona Agents and LLMs (EMNLP, 2025)
- Perspective-Shifted Neuro-Symbolic World Models: A Framework for Socially-Aware Robot Navigation (IEEE International Symposium on Robot and Human Interactive Communication, 2025)
- Persuasion for good: Towards a personalized persuasive dialogue system for social good (ACL, 2019)
- Policy4OOD: A Knowledge-Guided World Model for Policy Intervention Simulation against the Opioid Overdose Crisis (arXiv, 2026)
- PolicySim: An LLM-Based Agent Social Simulation Sandbox for Proactive Policy Optimization (Proceedings of the ACM Web Conference 2026, 2026)
- Pragmatic language interpretation as probabilistic inference (Trends in cognitive sciences, 2016)
- Project Sid: Many-agent simulations toward AI civilization (arXiv, 2024)
- Quantifying the Impact of Large Language Models on Collective Opinion Dynamics (arXiv, 2023)
- Role play with large language models (Nature, 2023)
- RoleLLM: Benchmarking, eliciting, and enhancing role-playing abilities of large language models (ACL, 2024)
- S3: Social-network Simulation System with Large Language Model-Empowered Agents (Social Science Research Network, 2023)
- Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting Pot (ICML, 2021)
- Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy Planning (Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025)
- Simulating Human-like Daily Activities with Desire-driven Autonomy (ICLR, 2024)
- SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds (arXiv, 2025)
- Social Simulacra: Creating Populated Prototyping Communities for Social Computing Research (ACM Symposium on User Interface Software and Technology, 2022)
- Social World Model-Augmented Mechanism Design Policy Learning (arXiv, 2025)
- SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World Users (arXiv, 2025)
- Sotopia-$pi$: Interactive learning of socially intelligent language agents (ACL, 2024)
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents (ICLR, 2023)
- The Consensus Game: Language Model Generation via Equilibrium Search (ICLR, 2024)
- The rational speech act framework (Annual Review of Linguistics, 2023)
- The traitors: Deception and trust in multi-agent language model simulations (arXiv, 2025)
- Theory of mind in large language models: Assessment and enhancement (ACL, 2025)
- Think Twice: Perspective-Taking Improves Large Language Models' Theory-of-Mind Capabilities (ACL, 2023)
- Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning (arXiv, 2025)
- Understanding epistemic language with a language-augmented bayesian theory of mind (Transactions of the Association for Computational Linguistics, 2025)
- 4D VQ-GAN: Synthesising Medical Scans at Any Time Point for Personalised Disease Progression Modelling of Idiopathic Pulmonary Fibrosis (arXiv, 2025)
- A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature (arXiv, 2026)
- A foundation model for the Earth system (Nature, 2025)
- Accurate Medium-Range Global Weather Forecasting with 3D Neural Networks (Nature, 2023)
- Agentifying Patient Dynamics within LLMs through Interacting with Clinical World Model (arXiv, 2026)
- An autonomous laboratory for the accelerated synthesis of inorganic materials (Nature, 2023)
- Auditory closed-loop stimulation of the sleep slow oscillation enhances memory (Neuron, 2013)
- Beyond Patient Invariance: Learning Cardiac Dynamics via Action-Conditioned JEPAs (arXiv, 2026)
- BLINK: Behavioral Latent Modeling of NK Cell Cytotoxicity (arXiv, 2026)
- Boltzmann Generators: Sampling Equilibrium States of Many-Body Systems with Deep Learning (Science, 2019)
- Brain-WM: Brain Glioblastoma World Model (arXiv, 2026)
- CellFlow: Simulating Cellular Morphology Changes via Flow Matching (ICML, 2025)
- ChemBO: Bayesian Optimization of Small Organic Molecules with Synthesizable Recommendations (International Conference on Artificial Intelligence and Statistics, 2020)
- ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area (AAAI, 2025)
- ChronoMedicalWorld: A Medical World Model for Learning Patient Trajectories from Longitudinal Care Data (arXiv, 2026)
- CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space (arXiv, 2025)
- ClimaX: A foundation model for weather and climate (ICML, 2023)
- Contextual Invertible World Models: A Neuro-Symbolic Agentic Framework for Colorectal Cancer Drug Response (arXiv, 2026)
- Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling (arXiv, 2025)
- Cosmos-Surg-DVRK: World Foundation Model-Based Automated Online Evaluation of Surgical Robot Policy Learning (IEEE Robotics and Automation Letters, 2025)
- DeepEarth: Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding (arXiv, 2026)
- DreamReg: Belief-Driven World Model for 2D-3D Ultrasound Registration (arXiv, 2026)
- Earth-o1: A Grid-free Observation-native Atmospheric World Model (arXiv, 2026)
- ECG-WM: A Physiology-Informed ECG World Model for Clinical Intervention Simulation (arXiv, 2026)
- EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance (CVPR, 2025)
- EHRWorld: A Patient-Centric Medical World Model for Long-Horizon Clinical Trajectories (arXiv, 2026)
- Entrainment of brain oscillations by transcranial alternating current stimulation (Current biology, 2014)
- EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting (arXiv, 2026)
- Evolving Interactive Diagnostic Agents in a Virtual Clinical Environment (arXiv, 2025)
- EyeWorld: A Generative World Model of Ocular State and Dynamics (arXiv, 2026)
- FieldSeer I: Physics-Guided World Models for Long-Horizon Electromagnetic Dynamics under Partial Observability (arXiv, 2025)
- Fourier Neural Operator for Parametric Partial Differential Equations (ICLR, 2020)
- GenCast: Diffusion-based ensemble forecasting for medium-range weather (arXiv, 2023)
- Generalized Neural-Network Representation of High-Dimensional Potential-Energy Surfaces (Physical Review Letters, 2007)
- Generating Counterfactual Patient Timelines from Real-World Data (arXiv, 2026)
- Generative Medical Event Models Improve with Scale (arXiv, 2025)
- Genie 3: A new frontier for world models (Google DeepMind, 2025)
- GraphCast: Learning skillful medium-range global weather forecasting (arXiv, 2022)
- Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells (arXiv, 2026)
- LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment (EMNLP, 2025)
- Machine Learning for Molecular Simulation (Annual Review of Physical Chemistry, 2020)
- medDreamer: Model-Based Reinforcement Learning with Latent Imagination on Complex EHRs for Clinical Decision Support (arXiv, 2025)
- Medical World Model (ICCV, 2025)
- MolWorld: Molecule World Models for Actionable Molecular Optimization (arXiv, 2026)
- MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses (ICLR, 2024)
- MRI Contrast Enhancement Kinetics World Model (arXiv, 2026)
- Multi Time Scale World Models (Advances in Neural Information Processing Systems, 2023)
- Neural general circulation models for weather and climate (Nature, 2023)
- Neural Operator: Learning Maps Between Function Spaces with Applications to PDEs (JMLR, 2023)
- ODesign: A World Model for Biomolecular Interaction Design (arXiv, 2025)
- On-the-fly closed-loop materials discovery via Bayesian active learning (Nature Communications, 2020)
- Pangu-Weather: A 3D High-Resolution Model for Fast and Accurate Global Weather Forecast (arXiv, 2022)
- PhysFire-WM: A Physics-Informed World Model for Emulating Fire Spread Dynamics (arXiv, 2025)
- Physics-Informed Deep Neural Operator Networks (Machine Learning in Modeling and Simulation: Methods and Applications, 2023)
- Physics-Informed Neural Operator for Learning Partial Differential Equations (ACM/IMS Journal of Data Science, 2024)
- PI-JEPA: Label-Free Surrogate Pretraining for Coupled Multiphysics Simulation via Operator-Split Latent Prediction (arXiv, 2026)
- Policy4OOD: A Knowledge-Guided World Model for Policy Intervention Simulation against the Opioid Overdose Crisis (arXiv, 2026)
- Population-based black-box optimization for biological sequence design (ICML, 2020)
- Probabilistic Weather Forecasting with Machine Learning (Nature, 2024)
- SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation (arXiv, 2026)
- Scaling Recurrence-aware Foundation Models for Clinical Records via Next-Visit Prediction (arXiv, 2026)
- SchNet: A Continuous-Filter Convolutional Neural Network for Modeling Quantum Interactions (Advances in Neural Information Processing Systems, 2017)
- Simulating clinical interventions with a generative multimodal model of human physiology (arXiv, 2026)
- Spatiotemporal Forecasting as Planning: A Model-Based Reinforcement Learning Approach with Generative World Models (arXiv, 2025)
- Surfing Uncertainty: Prediction, Action, and the Embodied Mind (Oxford University Press, 2015)
- Surgical Vision World Model (DEMI@MICCAI, 2025)
- SurgVista: Long-Horizon Surgical World Modeling with Plausible Instrument-Tissue Dynamics (arXiv, 2026)
- SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation (arXiv, 2026)
- Targeted materials discovery using Bayesian algorithm execution (NPJ Computational Materials, 2024)
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery (arXiv, 2024)
- Toward Safe Autonomous Robotic Endovascular Interventions using World Models (arXiv, 2026)
- Toward World Modeling of Physiological Signals with Chaos-Theoretic Balancing and Latent Dynamics (arXiv, 2026)
- Towards Suturing World Models: Learning Predictive Models for Robotic Surgical Tasks (arXiv, 2025)
- Validated Synthetic Patient Generation for Small Longitudinal Cohorts: Coagulation Dynamics Across Pregnancy (arXiv, 2026)
- VCWorld: A Biological World Model for Virtual Cell Simulation (arXiv, 2025)
Methods that design, execute, observe, and reflect — revising the model itself: active information expansion, autonomous execution, and belief revision.
- AdaptSim: Task-Driven Simulation Adaptation for Sim-to-Real Transfer (CoRL, 2023)
- Aligning Agentic World Models via Knowledgeable Experience Learning (arXiv, 2026)
- Egocentric Visual Self-Modeling for Autonomous Robot Dynamics Prediction and Adaptation (NPJ Robotics, 2025)
- Foundation World Models for Agents that Learn, Verify, and Adapt Reliably Beyond Static Environments (Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems, 2026)
- Kairos: A Native World Model Stack for Physical AI (arXiv, 2026)
- NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments (ICCV, 2025)
- Orchestrate, Generate, Reflect: A VLM-Based Multi-Agent Collaboration Framework for Automated Driving Policy Learning (arXiv, 2025)
- RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization (arXiv, 2026)
- Self-Improving Embodied Foundation Models (arXiv, 2025)
- Self-Improving World Modelling with Latent Actions (arXiv, 2026)
- SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents (arXiv, 2026)
- VideoAgent: Self-Improving Video Generation (arXiv, 2024)
- World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry (arXiv, 2026)
- World-Gymnast: Training Robots with Reinforcement Learning in a World Model (arXiv, 2026)
- Active Intelligence in Video Avatars via Closed-loop World Modeling (arXiv, 2025)
- AgentGym: Evolving Large Language Model-based Agents across Diverse Environments (arXiv, 2024)
- AlphaEvolve: A coding agent for scientific and algorithmic discovery (arXiv, 2025)
- Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models (arXiv, 2026)
- CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay (ICML, 2024)
- CoEx - Co-evolving World-model and Exploration (EMNLP, 2025)
- COMAP: Co-Evolving World Models and Agent Policies for LLM Agents (arXiv, 2026)
- Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025)
- Do Embodied Agents Dream of Pixelated Sheep?: Embodied Decision Making using Language Guided World Modelling (ICML, 2023)
- EvolvingAgent: Curriculum Self-evolving Agent with Continual World Model for Long-Horizon Tasks (arXiv, 2025)
- Learning to learn by gradient descent by gradient descent (Advances in Neural Information Processing Systems, 2016)
- Mathematical discoveries from program search with large language models (Nature, 2024)
- MindForge: Empowering Embodied Agents with Theory of Mind for Lifelong Cultural Learning (arXiv, 2024)
- Neuro-Symbolic World Models for Adapting to Open World Novelty (Adaptive Agents and Multi-Agent Systems, 2023)
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds (arXiv, 2025)
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning (arXiv, 2025)
- WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents (arXiv, 2025)
- WALL-E: World Alignment by Rule Learning Improves World Model-based LLM Agents (arXiv, 2024)
- WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model (EMNLP, 2025)
- WorldLLM: Improving LLMs' world modeling using curiosity-driven theory-making (arXiv, 2025)
- AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society (arXiv, 2025)
- Evolving Interpretable Constitutions for Multi-Agent Coordination (arXiv, 2026)
- MindForge: Empowering Embodied Agents with Theory of Mind for Lifelong Cultural Learning (arXiv, 2024)
- Systematic Biases in LLM Simulations of Debates (EMNLP, 2024)
- YuLan-OneSim: Towards the Next Generation of Social Simulator with Large Language Models (arXiv, 2025)
- AI-Supervisor: Autonomous AI Research Supervision via a Persistent Research World Model (arXiv, 2026)
- AlphaEvolve: A coding agent for scientific and algorithmic discovery (arXiv, 2025)
- Autonomous Chemical Research with Large Language Models (Nature, 2023)
- BacterAI Maps Microbial Metabolism Without Prior Knowledge (Nature Microbiology, 2023)
- BioLab: End-to-End Autonomous Life Sciences Research with Multi-Agents System Integrating Biological Foundation Models (bioRxiv, 2025)
- Biomni: A General-Purpose Biomedical AI Agent (bioRxiv, 2025)
- Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery (arXiv, 2026)
- Closed-Loop Cycles of Experiment Design, Execution, and Learning Accelerate Systems Biology Model Development in Yeast (Proceedings of the National Academy of Sciences, 2019)
- Delocalized, asynchronous, closed-loop discovery of organic laser emitters (Science, 2024)
- Large Language Models for Automated Open-domain Scientific Hypotheses Discovery (ACL, 2024)
- Learning to Generate Research Idea with Dynamic Control (Advances in Neural Information Processing Systems Workshop, 2024)
- LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography (arXiv, 2026)
- MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search (arXiv, 2025)
- On-the-fly Closed-loop Autonomous Materials Discovery via Bayesian Active Learning (arXiv, 2020)
- OriGene: A Self-Evolving Virtual Disease Biologist Automating Therapeutic Target Discovery (bioRxiv, 2025)
- ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition (arXiv, 2025)
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search (arXiv, 2025)
- The Structure of Scientific Revolutions (University of Chicago Press, 1962)
- Towards an AI co-scientist (arXiv, 2025)
- Towards Robot Scientists for Autonomous Scientific Discovery (Automated Experimentation, 2010)
Benchmarks and evaluation suites for world models, grouped by world.
- 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models (arXiv, 2025)
- A Control-Centric Benchmark for Video Prediction (arXiv, 2023)
- ACT-Bench: Towards Action Controllable World Models for Autonomous Driving (arXiv, 2024)
- ACWM-Phys: Investigating Generalized Physical Interaction in Action-Conditioned Video World Models (arXiv, 2026)
- Adapting Vision-Language Models for Evaluating World Models (arXiv, 2025)
- Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks (arXiv, 2025)
- AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models (arXiv, 2024)
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems (arXiv, 2025)
- ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control (arXiv, 2026)
- Are AI-Generated Driving Videos Ready for Autonomous Driving? A Diagnostic Evaluation Framework (arXiv, 2025)
- BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation (arXiv, 2024)
- Bench2Drive-VL: Benchmarks for Closed-Loop Autonomous Driving with Vision-Language Models (arXiv, 2026)
- Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving (arXiv, 2024)
- Beyond Simulation: Benchmarking World Models for Planning and Causality in Autonomous Driving (ICRA, 2025)
- Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA (arXiv, 2026)
- CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks (IEEE Robotics and Automation Letters, 2021)
- Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications (CVPR, 2023)
- Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency (arXiv, 2026)
- Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark (arXiv, 2025)
- CarDreamer: Open-Source Learning Platform for World-Model-Based Autonomous Driving (IEEE Internet of Things Journal, 2024)
- ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation (arXiv, 2024)
- CoPhy: Counterfactual Learning of Physical Dynamics (ICLR, 2020)
- CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models (arXiv, 2026)
- Current Agents Fail to Leverage World Model as Tool for Foresight (arXiv, 2026)
- Current World Models Lack a Persistent State Core (arXiv, 2026)
- DeepMind Control Suite (arXiv, 2018)
- Do generative video models understand physical principles? (IEEE Workshop/Winter Conference on Applications of Computer Vision, 2025)
- Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning (arXiv, 2026)
- Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation (ACL, 2025)
- Do World Action Models Generalize Better than VLAs? A Robustness Study (arXiv, 2026)
- Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models (IEEE/RJS International Conference on Intelligent RObots and Systems, 2025)
- DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving (arXiv, 2026)
- Driving in Corner Case: A Real-World Adversarial Closed-Loop Evaluation Platform for End-to-End Autonomous Driving (arXiv, 2025)
- DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model (NeurIPS, 2024)
- DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving (arXiv, 2026)
- Echo-Memory: A Controlled Study of Memory in Action World Models (arXiv, 2026)
- Elements of World Knowledge (EWOK): A cognition-inspired framework for evaluating basic world knowledge in language models (Transactions of the Association for Computational Linguistics, 2024)
- EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment (arXiv, 2024)
- ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction (arXiv, 2025)
- EvalCrafter: Benchmarking and Evaluating Large Video Generation Models (CVPR, 2024)
- EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models (arXiv, 2025)
- FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation (NeurIPS, 2023)
- From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving (arXiv, 2025)
- Habitat 2.0: Training Home Assistants to Rearrange their Habitat (NeurIPS, 2021)
- Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots (ICLR, 2024)
- Habitat: A Platform for Embodied AI Research (IEEE/CVF International Conference on Computer Vision, 2019)
- Hallucination in World Models is Predictable and Preventable (arXiv, 2026)
- HandoverSim: A Simulation Framework and Benchmark for Human-to-Robot Object Handovers (ICRA, 2022)
- HM3D-OVON: A Dataset and Benchmark for Open-Vocabulary Object Goal Navigation (IROS, 2024)
- How Far is Video Generation from World Model: A Physical Law Perspective (ICML, 2024)
- HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation (RSS, 2024)
- ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models (arXiv, 2026)
- iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks (CoRL, 2021)
- Imagine the Unseen World: A Benchmark for Systematic Generalization in Visual World Models (NeurIPS, 2023)
- IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments (arXiv, 2025)
- Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models (arXiv, 2026)
- Is Your Driving World Model an All-Around Player? (arXiv, 2026)
- iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework (arXiv, 2026)
- JailWAM: Jailbreaking World Action Models in Robot Control (arXiv, 2026)
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning (NeurIPS, 2023)
- LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference (arXiv, 2025)
- LLM world models are mental: Output layer evidence of brittle world model use in LLM mechanical reasoning (arXiv, 2025)
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills (arXiv, 2023)
- MBench: A Comprehensive Benchmark on Memory Capability for Video World Models (arXiv, 2026)
- Meta-World+: An Improved, Standardized, RL Benchmark (NeurIPS, 2025)
- Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning (CoRL, 2019)
- MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models (arXiv, 2026)
- MMGR: Multi-Modal Generative Reasoning (arXiv, 2025)
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos (arXiv, 2024)
- MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models (arXiv, 2026)
- NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking (NeurIPS, 2024)
- NeoRL: A Near Real-World Benchmark for Offline Reinforcement Learning (NeurIPS, 2021)
- nuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles (arXiv, 2021)
- nuScenes: A Multimodal Dataset for Autonomous Driving (CVPR, 2019)
- OccSTeP: Benchmarking 4D Occupancy Spatio-Temporal Persistence (arXiv, 2025)
- Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models (arXiv, 2026)
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling (arXiv, 2025)
- Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics (arXiv, 2026)
- Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models (arXiv, 2026)
- PAI-Bench: A Comprehensive Benchmark For Physical AI (arXiv, 2025)
- PhyGround: Benchmarking Physical Reasoning in Generative World Models (arXiv, 2026)
- PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models (arXiv, 2026)
- PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models (arXiv, 2026)
- Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical Properties (NeurIPS, 2023)
- PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models (ICLR, 2026)
- Probing Multimodal LLMs as World Models for Driving (IEEE Robotics and Automation Letters, 2024)
- Quantitative Video World Model Evaluation for Geometric-Consistency (arXiv, 2026)
- Reference-Free Assessment of Physical Consistency in World Model-based Video Generation (arXiv, 2026)
- Rethinking Video Generation Model for the Embodied World (arXiv, 2026)
- RLBench: The Robot Learning Benchmark & Learning Environment (IEEE Robotics and Automation Letters, 2020)
- RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots (arXiv, 2026)
- RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation (arXiv, 2026)
- RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation (arXiv, 2026)
- SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation (arXiv, 2026)
- Sekai: A Video Dataset towards World Exploration (arXiv, 2025)
- SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model (IEEE/RJS International Conference on Intelligent RObots and Systems, 2025)
- stable-worldmodel: A Platform for Reproducible World Modeling Research and Evaluation (arXiv, 2026)
- T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation (arXiv, 2025)
- Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets? (arXiv, 2025)
- Thinking Ahead: Foresight Intelligence in MLLMs and World Models (arXiv, 2025)
- Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments (Pattern Recognition, 2025)
- Towards Autonomous Micromobility through Scalable Urban Simulation (CVPR, 2025)
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation (arXiv, 2024)
- Uniocc: a Unified Benchmark for Occupancy Forecasting and Prediction in Autonomous Driving (ICCV, 2025)
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation (arXiv, 2025)
- VideoPhy: Evaluating Physical Commonsense for Video Generation (arXiv, 2024)
- VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos? (arXiv, 2025)
- WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation (arXiv, 2026)
- What if? Emulative Simulation with World Models for Situated Reasoning (arXiv, 2026)
- What-If World: A Causal Benchmark for General World Models in Embodied Scenarios (arXiv, 2026)
- World Reasoning Arena (arXiv, 2026)
- World-in-World: World Models in a Closed-Loop World (arXiv, 2025)
- WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform (arXiv, 2026)
- WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models (arXiv, 2026)
- WorldBench: Disambiguating Physics for Diagnostic Evaluation of World Models (arXiv, 2026)
- WorldMark: A Unified Benchmark Suite for Interactive Video World Models (arXiv, 2026)
- WorldModelBench: Judging Video Generation Models As World Models (arXiv, 2025)
- WorldOlympiad: Can Your World Model Survive a Triathlon? (arXiv, 2026)
- WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning (arXiv, 2025)
- WorldScore: A Unified Evaluation Benchmark for World Generation (ICCV, 2025)
- WorldSimBench: Towards Video Generation Models as World Simulators (arXiv, 2024)
- Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test (arXiv, 2026)
- 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models (arXiv, 2025)
- AgentBench: Evaluating LLMs as Agents (ICLR, 2024)
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents (ICLR, 2025)
- AppAgent: Multimodal Agents as Smartphone Users (CHI Conference on Human Factors in Computing Systems, 2025)
- Assessing Adaptive World Models in Machines with Novel Games (arXiv, 2025)
- Benchmarking the Spectrum of Agent Capabilities (arXiv, 2021)
- Benchmarking World-Model Learning with Environment-Level Queries (arXiv, 2025)
- Can Language Models Serve as Text-Based World Simulators? (ACL, 2024)
- CausalARC: Abstract Reasoning with Causal World Models (arXiv, 2025)
- Computer-Use Agents as Judges for Generative User Interface (arXiv, 2025)
- Current Agents Fail to Leverage World Model as Tool for Foresight (arXiv, 2026)
- Deep Reinforcement Learning at the Edge of the Statistical Precipice (Advances in Neural Information Processing Systems, 2021)
- Deep Reinforcement Learning that Matters (AAAI, 2018)
- Do LLMs Build Spatial World Models? Evidence from Grid-World Maze Tasks (arXiv, 2026)
- EgoCS-400K: An Egocentric Gameplay Dataset for World Models (arXiv, 2026)
- Evaluating the World Model Implicit in a Generative Model (NeurIPS, 2024)
- GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents (arXiv, 2026)
- GEBench: Benchmarking Image Generation Models as GUI Environments (arXiv, 2026)
- Hallucination in World Models is Predictable and Preventable (arXiv, 2026)
- HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models (arXiv, 2026)
- Interactive Fiction Games: A Colossal Adventure (AAAI, 2019)
- Katakomba: Tools and Benchmarks for Data-Driven NetHack (Advances in Neural Information Processing Systems, 2023)
- Learning to Generalize with Object-Centric Agents in the Open World Survival Game Crafter (IEEE Transactions on Games, 2024)
- LLM-Based World Models Can Make Decisions Solely, But Rigorous Evaluations are Needed (arXiv, 2024)
- LoopNav: Benchmarking Spatial Consistency in World Models (arXiv, 2025)
- MCU: An Evaluation Framework for Open-Ended Game Agents (ICML, 2023)
- Melting Pot 2.0 (arXiv, 2022)
- Mind2Web: Towards a Generalist Agent for the Web (Advances in Neural Information Processing Systems, 2023)
- MIND: Benchmarking Memory Consistency and Action Control in World Models (arXiv, 2026)
- MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge (NeurIPS, 2022)
- MineRL: A Large-Scale Dataset of Minecraft Demonstrations (arXiv, 2019)
- Minigrid & Miniworld: Modular & Customizable Reinforcement Learning Environments for Goal-Oriented Tasks (arXiv, 2023)
- MiniHack the Planet: A Sandbox for Open-Ended Reinforcement Learning Research (arXiv, 2021)
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos (arXiv, 2024)
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents (arXiv, 2025)
- OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web (arXiv, 2024)
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments (NeurIPS, 2024)
- ScienceWorld: Is your Agent Smarter than a 5th Grader? (EMNLP, 2022)
- SmallWorlds: Assessing Dynamics Understanding of World Models in Isolated Environments (arXiv, 2025)
- stable-worldmodel: A Platform for Reproducible World Modeling Research and Evaluation (arXiv, 2026)
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (ICLR, 2023)
- Text2World: Benchmarking Large Language Models for Symbolic World Model Generation (ACL, 2025)
- TextWorld: A Learning Environment for Text-based Games (CGW@IJCAI, 2018)
- The Arcade Learning Environment: An Evaluation Platform for General Agents (Journal of Artificial Intelligence Research, 2012)
- UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI (ICCV, 2025)
- VBench: Comprehensive Benchmark Suite for Video Generative Models (CVPR, 2023)
- Video-Bench: Human-Aligned Video Generation Benchmark (CVPR, 2025)
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks (arXiv, 2024)
- WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation (arXiv, 2026)
- WebArena: A Realistic Web Environment for Building Autonomous Agents (ICLR, 2023)
- WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG (arXiv, 2026)
- WMAttack: Automated Attack Search for Adversarial Evaluation of World-Model Agents (arXiv, 2026)
- World of Workflows: a Benchmark for Bringing World Models to Enterprise Systems (arXiv, 2026)
- WorldMark: A Unified Benchmark Suite for Interactive Video World Models (arXiv, 2026)
- WorldSimBench: Towards Video Generation Models as World Simulators (arXiv, 2024)
- Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning (ICLR, 2025)
- FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions (EMNLP, 2023)
- From Text to Tactic: Evaluating LLMs Playing the Game of Avalon (arXiv, 2023)
- FutureSim: Replaying World Events to Evaluate Adaptive Agents (arXiv, 2026)
- FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction (arXiv, 2025)
- Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots (ICLR, 2024)
- HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models (EMNLP, 2023)
- HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models (arXiv, 2025)
- Melting Pot 2.0 (arXiv, 2022)
- OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models (Association for Computational Linguistics, 2024)
- PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics (arXiv, 2025)
- Revisiting the Evaluation of Theory of Mind through Question Answering (EMNLP, 2019)
- Understanding Social Reasoning in Language Models with Language Models (Advances in Neural Information Processing Systems, 2023)
- Augmenting large language models with chemistry tools (Nature Machine Intelligence, 2024)
- ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation (arXiv, 2024)
- DiscoveryBench: Towards Data-Driven Discovery with Large Language Models (ICLR, 2024)
- Expert evaluation of LLM world models: A high-Tc superconductivity case study (Proceedings of the National Academy of Sciences of the United States of America, 2025)
- How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment (arXiv, 2025)
- MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation (arXiv, 2026)
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos (arXiv, 2024)
- Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics (arXiv, 2026)
- ScienceWorld: Is your Agent Smarter than a 5th Grader? (EMNLP, 2022)
- What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models (ICML, 2025)
Surveys, reviews, and roadmaps of world models and closely related areas.
- A Generalization Theory for JEPA-Based World Models (arXiv, 2026)
- A Mechanistic View on Video Generation as World Models: State and Dynamics (arXiv, 2026)
- A Tutorial on World Models and Physical AI (ACM Computing Surveys, 2026)
- Against the Monolithic Wireless World Model: Why NextG Needs Composable and Agentic Intelligence (arXiv, 2026)
- Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond (arXiv, 2026)
- Autonomous Video Generation with Counterfactual Controllability for Self-Evolving World Models (arXiv, 2026)
- Bridging the Agent-World Gap: Text World Models for LLM-based Agents (arXiv, 2026)
- Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models (arXiv, 2026)
- From Digital Twins to World Models:Opportunities, Challenges, and Applications for Mobile Edge General Intelligence (arXiv, 2026)
- From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models (arXiv, 2026)
- From Static Risk to Dynamic Trajectories: Toward World-Model-Inspired Clinical Prediction (arXiv, 2026)
- Graph World Models: Concepts, Taxonomy, and Future Directions (arXiv, 2026)
- Grounding Clinical AI Competency in Human Cognition Through the Clinical World Model and Skill-Mix Framework (arXiv, 2026)
- Human Cognition in Machines: A Unified Perspective of World Models (arXiv, 2026)
- Imperfect World Models are Exploitable (arXiv, 2026)
- Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception (arXiv, 2026)
- Interpreting Physics in Video World Models (arXiv, 2026)
- Latent State Design for World Models under Sufficiency Constraints (arXiv, 2026)
- Latent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open Challenges (arXiv, 2026)
- Medical world models: representing medical states, modelling clinical dynamics and guiding intervention policies (arXiv, 2026)
- OpenWorldLib: A Unified Codebase and Definition of Advanced World Models (arXiv, 2026)
- Physically Native World Models: A Hamiltonian Perspective on Generative World Modeling (arXiv, 2026)
- Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses (arXiv, 2026)
- The Trinity of Consistency as a Defining Principle for General World Models (arXiv, 2026)
- Toward World Models for Epidemiology (arXiv, 2026)
- Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends (arXiv, 2026)
- Towards World Models in Biomedical Research (arXiv, 2026)
- Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms (arXiv, 2026)
- Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling (arXiv, 2026)
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models (arXiv, 2026)
- Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform (arXiv, 2026)
- World Action Models: The Next Frontier in Embodied AI (arXiv, 2026)
- World Model for Robot Learning: A Comprehensive Survey (arXiv, 2026)
- World Models for Robotic Manipulation: A Survey (arXiv, 2026)
- 3D and 4D World Modeling: A Survey (arXiv, 2025)
- A "Good" Regulator May Provide a World Model for Intelligent Systems (Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 2025)
- A Comprehensive Survey on World Models for Embodied AI (arXiv, 2025)
- A Step Toward World Models: A Survey on Robotic Manipulation (arXiv, 2025)
- A Survey of Interactive Generative Video (arXiv, 2025)
- A Survey of World Models for Autonomous Driving (arXiv, 2025)
- A Survey on World Models Grounded in Acoustic Physical Information (arXiv, 2025)
- A Survey: Learning Embodied Intelligence from Physical Simulators and World Models (arXiv, 2025)
- A Unified Definition of Hallucination: It's The World Model, Stupid! (arXiv, 2025)
- Beyond Generative AI: World Models for Clinical Prediction, Counterfactuals, and Planning (arXiv, 2025)
- Beyond World Models: Rethinking Understanding in AI Models (AAAI, 2025)
- Critiques of World Models (arXiv, 2025)
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges (arXiv, 2025)
- Embodied AI Agents: Modeling the World (arXiv, 2025)
- Embodied AI: From LLMs to World Models [Feature] (IEEE Circuits and Systems Magazine, 2025)
- Exploring the Evolution of Physics Cognition in Video Generation: A Survey (arXiv, 2025)
- Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds (arXiv, 2025)
- Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis (IEEE Open Journal of Intelligent Transportation Systems, 2025)
- Four Principles for Physically Interpretable World Models (NeuS, 2025)
- From 2D to 3D Cognition: A Brief Survey of General World Models (arXiv, 2025)
- From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery (arXiv, 2025)
- From Masks to Worlds: A Hitchhiker's Guide to World Models (arXiv, 2025)
- Generative Physical AI in Vision: A Survey (arXiv, 2025)
- Humanoid Locomotion and Manipulation: Current Progress and Challenges in Control, Planning, and Learning (IEEE/ASME transactions on mechatronics, 2025)
- Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review (arXiv, 2025)
- Modeling the Mental World for Embodied AI: A Comprehensive Review (arXiv, 2025)
- Natural Building Blocks for Structured World Models: Theory, Evidence, and Scaling (arXiv, 2025)
- Position: Interactive Generative Video as Next-Generation Game Engine (arXiv, 2025)
- Rethinking the Simulation vs. Rendering Dichotomy: No Free Lunch in Spatial World Modelling (arXiv, 2025)
- Simulating the Real World: A Unified Survey of Multimodal Generative Models (IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025)
- Simulating the Visual World with Artificial Intelligence: A Roadmap (arXiv, 2025)
- The brain-AI convergence: Predictive and generative world models for general-purpose computation (arXiv, 2025)
- The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey (arXiv, 2025)
- The Safety Challenge of World Models for Embodied AI Agents: A Review (arXiv, 2025)
- What Does it Mean for a Neural Network to Learn a "World Model"? (arXiv, 2025)
- When do neural networks learn world models? (ICML, 2025)
- World Models for Cognitive Agents: Transforming Edge Intelligence in Future Networks (arXiv, 2025)
- World Models in Artificial Intelligence: Sensing, Learning, and Reasoning Like a Child (arXiv, 2025)
- World Models Should Prioritize the Unification of Physical and Social Dynamics (arXiv, 2025)
- Aligning Cyber Space With Physical World: A Comprehensive Survey on Embodied AI (IEEE/ASME transactions on mechatronics, 2024)
- Cognitive Architectures for Language Agents (TMLR, 2024)
- Emergence of Implicit World Models from Mortal Agents (arXiv, 2024)
- Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey (arXiv, 2024)
- From Efficient Multimodal Models to World Models: A Survey (arXiv, 2024)
- Generative Emergent Communication: Large Language Model is a Collective World Model (Advanced Robotics, 2024)
- Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond (arXiv, 2024)
- Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling (arXiv, 2024)
- Scaling Laws for Pre-training Agents and World Models (arXiv, 2024)
- Sora and V-JEPA Have Not Learned The Complete Real World Model -- A Philosophical Analysis of Video AIs Through the Theory of Productive Imagination (arXiv, 2024)
- Sora as a World Model? A Complete Survey on Text-to-Video Generation (arXiv, 2024)
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models (arXiv, 2024)
- Understanding World or Predicting Future? A Comprehensive Survey of World Models (ACM CSUR, 2024)
- Video as the New Language for Real-World Decision Making (ICML, 2024)
- World Models for Autonomous Driving: An Initial Survey (IEEE Transactions on Intelligent Vehicles, 2024)
- Facing Off World Model Backbones: RNNs, Transformers, and S4 (arXiv, 2023)
- Language Models, Agent Models, and World Models: The LAW for Machine Reasoning and Planning (arXiv, 2023)
- A Survey of Embodied AI: From Simulators to Research Tasks (IEEE Transactions on Emerging Topics in Computational Intelligence, 2021)
- Model-based Reinforcement Learning: A Survey (Found. Trends Mach. Learn., 2020)
- Building Machines That Learn and Think Like People (Behavioral and Brain Sciences, 2017)
- World Models: The Safety Perspective (2024 IEEE 35th International Symposium on Software Reliability Engineering Workshops (ISSREW), 2024) — With the proliferation of the Large Language Model (LLM), the concept of World Models (WM) has recently attracted a great deal of attention in the AI research c.
We welcome contributions! Open the 📄 Add a paper / benchmark issue and give just three things:
- arXiv URL
- Category —
L1 Predictor·L2 Simulator·L3 Evolver·Survey·Benchmark - Law —
Physical·Digital·Social·Scientific
Everything else — title, authors, venue, year, abstract, code link — is fetched automatically from arXiv. Leave the optional fields blank unless you want to override an auto-filled value.
