| 🌟 Classic | ✅ Open Source | ❌️ Not Open Source |
Badge Guide. 🌟 marks representative or category-defining work, with the entire paper entry bolded. ✅ indicates that public resources such as code, data, model weights, demos, or an official project page are available. ❌️ indicates that no public implementation or resource has been found yet. The [YYYY-MM] prefix indicates the arXiv submission month when parsed from an arXiv identifier, or a known release/update month for non-arXiv resources.
Major updates and announcements are shown below. Scroll for full timeline.
🔥 [2026-5] Repository Launch — Awesome Interactive World Models is now live! We're building a curated collection of Interactive World Model papers, systematically categorized by key research problems. See CONTRIBUTING.md for how to contribute.
💡 [Ongoing] Community Contributions Welcome — Help us maintain the most up-to-date world models resource! Submit papers via PR or contact us at email.
⭐ [Ongoing] Support This Project — If you find this useful, please cite our work and give us a star. Share with your research community!
[2026-08-03] ✅WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
[2026-08-02] ✅MiniWorld: Democratizing the Training of Video World Models from Scratch
[2026-07-31] ✅WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation
[2026-07-31] ❌️BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning
[2026-07-30] ✅PhiZero: A World Model Built Around Physical Language
[2026-07-30] ✅ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
[2026-07-29] ✅StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
[2026-07-29] ✅ContactFlow: A video action conditioning that transfers across embodiments
[2026-07-29] ❌️CG-World: A Large-Scale World-State Dataset and Protocol for World Models
[2026-07-28] ✅Wonder: Video World Model Done Better
[2026-07-23] ✅Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
[2026-07-21] ✅Generative World Renderer at the Speed of Play
[2026-07-15] ❌️From Pixels to States: Rethinking Interactive World Models as Game Engines
[2026-07-15] ❌️M⁴World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming
[2026-07-15] ❌️Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Chassis Adaptation
[2026-07-13] ❌️Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency
[2026-07-13] ✅ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space
[2026-07-13] ✅Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
[2026-07-12] ❌️LIDAR-AD: A Decoder-Free Latent-Interaction Dreamer with Action-Residual Chains for Autonomous Driving
[2026-07-12] ❌️Is Energy Guidance All You Need? Training-Free Norm Injection for Driving World Models
[2026-07-12] ❌️World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning
[2026-07-11] ❌️Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models
[2026-07-10] ✅PanoWorld: Real-World Panoramic Generation
[2026-07-10] ❌️Causally Debiased Latent Action Model for Embodied Action Conditioned World Models
[2026-07-08] ✅Infinite Worlds with Versatile Interactions
[2026-07-07] ✅RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation
[2026-07-07] ✅AlayaWorld: Long-Horizon and Playable Video World Generation
[2026-07-07] ✅MoWorld: A Flash World Model
[2026-07-06] ✅Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
[2026-07-06] ✅Multiplayer Interactive World Models with Representation Autoencoders
[2026-07-05] ✅Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
[2026-07-04] ❌️Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
[2026-07-03] ✅Vidu S1: A Real-Time Interactive Video Generation Model
[2026-07-02] ✅WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory
[2026-07-02] ✅PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
[2026-07-01] ❌️RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation
- 🔥 Recent Paper Updates
- 🌐 General World Model
- 🤝 Multi-Agent World Model
- 📊 Benchmark
- 💾 Dataset
- 🎬 Long Video Generation
- 🧠 Memory Consistency
- 🎮 Action Control
- 🖱️ Interactive
- 🔧 Post-Training
- ⚙️ Physics
- 📏 Evaluation
- 📚 Survey
- 📝 Citation
[2026-06] ✅🌟(Omnimodal world model for understanding, generation, simulation, and action)Cosmos 3: Omnimodal World Models for Physical AI
[2026-05] ✅🌟(Full-stack open-source real-time interactive video world model framework)minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
[2026-01] ✅🌟(Minute-level Memory)Lingbot-World:Advancing Open-source World Models
[2025-08] ✅🌟(First open-source real-time interactive World Model)Matrix-game 2.0: An open-source real-time and streaming interactive world model
[2026-08] ✅MiniWorld: Democratizing the Training of Video World Models from Scratch
[2026-07] ❌️BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning
[2026-07] ✅PhiZero: A World Model Built Around Physical Language
[2026-07] ✅StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
[2026-07] ✅Wonder: Video World Model Done Better
[2026-07] ✅Generative World Renderer at the Speed of Play
[2026-07] ❌️Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Chassis Adaptation
[2026-07] ✅ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space
[2026-07] ✅Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
[2026-07] ❌️LIDAR-AD: A Decoder-Free Latent-Interaction Dreamer with Action-Residual Chains for Autonomous Driving
[2026-07] ✅PanoWorld: Real-World Panoramic Generation
[2026-07] ✅Infinite Worlds with Versatile Interactions
[2026-07] ❌️AlayaWorld: Long-Horizon and Playable Video World Generation
[2026-07] ✅MoWorld: A Flash World Model
[2026-06] ❌️DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
[2026-06] ✅DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model
[2026-06] ❌️MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold
[2026-06] ✅Next Forcing: Causal World Modeling with Multi-Chunk Prediction
[2026-06] ✅DreamX-World 1.0: A General-Purpose Interactive World Model
[2026-06] ❌️GeoStream: Toward Precise Camera Controlled Streaming Video Generation
[2026-05] ✅Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models
[2026-05] ✅SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
[2026-05] ✅Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video
[2026-04] ✅Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
[2026-03] ✅InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model
[2025-12] ✅Yume-1.5: A Text-Controlled Interactive World Generation Model
[2025-12] ❌️RELIC: Interactive Video World Model with Long-Horizon Memory
[2025-12] ✅Astra: General Interactive World Model with Autoregressive Denoising
[2025-11] ❌️Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
[2025-11] ❌️PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
[2025-11] ✅MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
[2025-08] ❌️Yan: Foundational Interactive Video Generation
[2025-07] ✅Yume: An Interactive World Generation Model
[2025-06] ✅Matrix-Game: Interactive World Foundation Model
[2025-06] ✅Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
[2024-02] ❌️Genie: Generative Interactive Environments
[2026-07] ✅Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
[2026-07] ✅Multiplayer Interactive World Models with Representation Autoencoders
[2026-06] ✅Prisma-World: Camera-Controllable Multi-Agent Video World Model
[2026-06] ❌️MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data
[2026-05] ✅Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
[2026-04] ✅MultiWorld: Scalable Multi-Agent Multi-View Video World Models
[2026-05] ✅🌟(Multi-turn interactive world model evaluation)WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
[2026-02] ✅🌟(First Open Domain Memory & Action Benchmark)MIND: Benchmarking Memory Consistency and Action Control in World Models
[2026-08] ✅WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
[2026-06] ❌️WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
[2026-06] ✅WorldOlympiad: Can Your World Model Survive a Triathlon?
[2025-05] ✅Toward Memory-Aided World Models: Benchmarking via Spatial Consistency
[2025-04] ✅WorldScore: A Unified Evaluation Benchmark for World Generation
[2025-02] ✅WorldModelBench: Judging Video Generation Models As World Models
[2024-10] ✅WorldSimBench: Towards Video Generation Models as World Simulators
[2026-07] ❌️ CG-World: A Large-Scale World-State Dataset and Protocol for World Models
[2026-07] ✅ Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
[2026-06] ✅ PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models
[2026-06] ✅ EgoCS-400K: An Egocentric Gameplay Dataset for World Models
[2025-09] ✅ SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
[2025-09] ✅ OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
[2025-06] ✅ Sekai: A Video Dataset towards World Exploration
[2024-11] ✅ GameGen-X: Open-World Video Game Dataset
[2026-05] ✅ 🌟(First architecture supporting arbitrary step OPD) AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
[2026-03] ✅🌟(Scale up Real-Time model to 14B)Helios: Real Real-Time Long Video Generation Model
[2026-02] ✅🌟(ODE perspective)Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation
[2025-12] ❌️🌟(Compression recovery training)Pretraining Frame Preservation in Autoregressive Video Memory Compression
[2025-10] ❌️🌟(Solves Self-Forcing issues)Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
[2025-08] ❌️🌟(Learnable sparse attention routing)Mixture of Contexts for Long Video Generation
[2025-06] ✅🌟(DMD perspective)Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
[2025-04] ✅🌟(Context compression)Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
[2024-12] ✅🌟(Teacher->Student acceleration)From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
[2024-07] ✅🌟(Training techniques)Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
[2026-06] ✅Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation
[2026-05] ✅DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
[2026-05] ✅VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
[2026-05] ✅AdaState: Self-Evolving Anchors for Streaming Video Generation
[2026-05] ❌️Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
[2026-05] ✅Q-ARVD: Quantizing Autoregressive Video Diffusion Models
[2026-05] ❌️DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
[2026-05] ✅FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching
[2026-05] ✅LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
[2026-05] ✅(One Step)One-Forcing: Towards Stable One-Step Autoregressive Video Generation
[2026-03] ✅Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
[2026-02] ✅Context Forcing: Consistent Autoregressive Video Generation with Long Context
[2026-02] ✅Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
[2026-02] ❌️Mode Seeking meets Mean Seeking for Fast Long Video Generation
[2026-02] ❌️LIVE: Long-horizon Interactive Video World Modeling
[2026-02] ❌️LIGHT FORCING: Accelerating Autoregressive Video Diffusion via Sparse Attention
[2026-02] ✅Pathwise Test-Time Correction for Autoregressive Long Video Diffusion Models
[2025-12] ✅VideoSSM: Autoregressive Long Video Generation with Hybrid State Space Memory
[2025-09] ✅Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
[2025-09] ✅LongLive: Real-time Interactive Long Video Generation
[2025-03] ✅FAR: Frame Autoregressive Model for Both Short- and Long-Context Video Modeling
[2025-03] ✅NOVA: Autoregressive Video Generation without Vector Quantization
[2024-10] ✅Progressive Autoregressive Video Diffusion Models
[2026-02] ✅🌟(w/o Pose Memory)Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory
[2025-06] ❌️🌟(w/ Pose FOV overlap retrieval Memory)Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
[2026-06] ✅MemLearner: Learning to Query Context memory for Video World Models
[2026-06] ❌️Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation
[2026-06] ❌️TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment
[2026-06] ✅Latent Spatial Memory for Video World Models
[2026-06] ✅Echo-Memory: A Controlled Study of Memory in Action World Models
[2026-06] ✅Geometry-Aware Implicit Memory for Video World Models
[2026-05] ✅OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation
[2026-05] ✅E3C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control
[2026-05] ✅WorldKV: Efficient World Memory with World Retrieval and Compression
[2026-03] ❌️MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
[2026-02] ✅AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
[2026-01] ✅StableWorld: Towards Stable and Consistent Long Interactive Video Generation
[2025-10] ❌️Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
[2025-06] ✅VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
[2025-06] ❌️Video World Models with Long-term Spatial Memory
[2025-06] ✅DeepVerse: 4D Autoregressive Video Generation as a World Model
[2025-05] ✅Learning World Models for Interactive Video Generation
[2025-04] ✅WorldMem: Long-term Consistent World Simulation with Memory
[2026-03] ✅🌟(First out-of-view Event Dynamic Memory)LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
[2026-02] ✅🌟(First multiplayer Minecraft World Model)Solaris: Building a Multiplayer Video World Model in Minecraft
[2026-01] ✅🌟(First out-of-view Dynamic Memory)Flow Equivariant World Models: Structured Dynamics Outside the Field of View
[2026-07] ✅WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory
[2026-05] ✅Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
[2026-03] ✅Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
[2026-02] ✅🌟(Transferable skills)Olaf-World: Orienting Latent Actions for Video World Modeling
[2025-03] ✅🌟(No longer limited to fixed controls)AdaWorld: Learning Adaptable World Models with Latent Actions
[2025-01] ❌️🌟(Action control from games)GameFactory: Creating New Games with Generative Interactive Videos
[2026-07] ✅ContactFlow: A video action conditioning that transfers across embodiments
[2026-07] ❌️Is Energy Guidance All You Need? Training-Free Norm Injection for Driving World Models
[2026-07] ❌️Causally Debiased Latent Action Model for Embodied Action Conditioned World Models
[2026-07] ✅RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation
[2026-07] ✅Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
[2026-07] ❌️Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
[2026-06] ✅Streaming Video Generation with Streaming Force Control
[2026-05] ✅E3C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control
[2026-05] ✅SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
[2026-05] ✅ReactiveGWM: Steering NPC in Reactive Game World Models
[2025-12] ✅🌟(Edit objects: add/remove/recolor)Spatia: Video Generation with Updatable Spatial Memory
[2025-11] ❌️🌟(Text interaction: spawn weapons, affect environment)Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
[2026-07] ✅Vidu S1: A Real-Time Interactive Video Generation Model
[2026-07] ❌️World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning
[2026-04] ✅World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
[2026-02] ✅WorldCompass: Reinforcement Learning for Long-Horizon World Models
[2025-05] ✅RLVR-World: Training World Models with Reinforcement Learning
[2026-07] ✅PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
[2026-06] ❌️IOI: Decoupling Kinematics and Physics for Interactive World Models
[2026-03] ✅RealWonder: Real-Time Physical Action-Conditioned Video Generation
[2026-08] ✅WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation
[2026-07] ❌️RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation
[2026-06] ✅BadWorld: Adversarial Attacks on World Models
[2026-03] ✅Interactive World Simulator for Robot Policy Training and Evaluation
[2026-04] ✅🌟(Defines advanced world models around perception, interaction, and long-term memory, and introduces OpenWorldLib as a unified inference framework.) OpenWorldLib: A Unified Codebase and Definition of Advanced World Models
[2026-07] ❌️From Pixels to States: Rethinking Interactive World Models as Game Engines
[2026-05] ✅Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends
[2026-03] ❌️Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms
[2026-07] ❌️Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models
[2025-05] ✅Vid2World: Crafting Video Diffusion Models to Interactive World Models
If you find this repository useful, please consider citing this list:
@misc{ye2026awesomeiwm,
title = {Awesome-Interactive-World-Model},
author = {Yixuan Ye and Ruiqi Wu},
journal = {GitHub repository},
url = {https://github.com/EasonTuT/Awesome-Interactive-World-Model},
year = {2026},
}