| 🌟 Classic | ✅ Open Source | ❌️ Not Open Source |
Badge Guide. 🌟 marks representative or category-defining work, with the entire paper entry bolded. ✅ indicates that public resources such as code, data, model weights, demos, or an official project page are available. ❌️ indicates that no public implementation or resource has been found yet. The [YYYY-MM] prefix indicates the arXiv submission month when parsed from an arXiv identifier, or a known release/update month for non-arXiv resources.
Major updates and announcements are shown below. Scroll for full timeline.
🔥 [2026-5] Repository Launch — Awesome Interactive World Models is now live! We're building a curated collection of Interactive World Model papers, systematically categorized by key research problems. See CONTRIBUTING.md for how to contribute.
💡 [Ongoing] Community Contributions Welcome — Help us maintain the most up-to-date world models resource! Submit papers via PR or contact us at email.
⭐ [Ongoing] Support This Project — If you find this useful, please cite our work and give us a star. Share with your research community!
[2026-09-11] ✅CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
[2026-09-10] ✅Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence
[2026-09-10] ✅World in World: Explore the World with World Models
[2026-09-09] ✅Programmable World Model
[2026-09-08] ✅ActionSplice: In-Flight Action Editing for Interactive World Models
[2026-09-07] ❌️TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image
[2026-09-05] ✅PhysWeep: Does a Video Generator Realize the Physics You Ask For?
[2026-09-03] ✅Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
[2026-09-03] ✅WorldReward: Reward Modeling for Camera-Conditioned World Models
[2026-09-03] ✅OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mapping
[2026-09-03] ❌️Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation
[2026-09-03] ✅PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
[2026-09-02] ✅SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
[2026-09-01] ✅H3-World: Turning Language Understanding into World Control
[2026-09-01] ✅MeRoPE: Metric Rotary Position Embedding for Camera-Controlled Video Generation
[2026-09-01] ✅AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization
[2026-08-31] ✅CAER: Causal Action Effect Reweighting for World Model Training
[2026-08-31] ✅Can Video World Models Track Unobserved World States?
[2026-08-30] ✅Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
[2026-08-28] ❌️RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction
[2026-08-28] ✅Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting
[2026-08-27] ✅R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models
[2026-08-27] ✅4DStreamCtrl: Interactive Video Generation with Online 4D Control
[2026-08-27] ✅NVIDIA Cosmos-H-Dreams: Real-Time Generative Physics Simulation for Surgical Robotics
[2026-08-26] ❌️StreamAV-Bench: A Comprehensive Benchmark for Streaming Audio-Video Generation
[2026-08-26] ❌️WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression
[2026-08-25] ✅Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning
[2026-08-25] ✅Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
[2026-08-24] ✅ReWorld: An Interactive World Model with Long-Horizon Memory
[2026-08-24] ❌️EchoWM: Open and Enterable Omnimodal World Models
[2026-08-24] ✅Population-Scalable Multi-Agent World Modeling
[2026-08-23] ❌️Where World Models Break: Natural-Input Failure Discovery
[2026-08-19] ✅WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
[2026-08-19] ❌️CamWorldQA: Perceptual Quality Assessment of Camera-Controlled World Video Generation
[2026-08-18] ✅WorldMind: Decoupled Game World Model for State-Aware NPC Behavior
[2026-08-18] ✅Hydra-0: Action Flow for Generalist World Modeling and Control
[2026-08-18] ❌️Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
[2026-08-17] ❌️GaussianDWM++: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation
[2026-08-15] ✅SCOPE: Score-Isolated Agentic Optimization for Video World Models
[2026-08-14] ✅Marionette: Predicting World States, Rendering Geometry, Painting Appearance
[2026-08-14] ✅ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
[2026-08-14] ✅PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
[2026-08-13] ✅DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
[2026-08-13] ✅H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
[2026-08-12] ❌️How Can Driving World Models Do Counterfactual Prediction?
[2026-08-11] ✅Sekai2: From World Exploration to Interactive World Modeling
[2026-08-10] ✅RealWeather: Realistic and Scene-Faithful Weather Translation with Driving World Models
[2026-08-10] ✅WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation
[2026-08-10] ❌️Twin Rollouts: Noise-Coupled Counterfactual Branching in Interactive Video World Models
[2026-08-10] ✅MASS: Multiplayer World Models with Authoritative Shared State
[2026-08-08] ✅Distilling Physical Priors into Streaming World Models
[2026-08-07] ✅Addressable Memory for Video World Models
[2026-08-07] ❌️Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts
[2026-08-06] ✅GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
[2026-08-05] ✅HelloWorld: Enabling Socially Interactive Characters in Video World Models
[2026-08-05] ❌️muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards
[2026-08-05] ✅WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
[2026-08-05] ❌️Overcoming Statistical Bias in Action-Controllable World Models
[2026-08-04] ✅EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
[2026-08-04] ❌️Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
[2026-08-03] ✅WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
[2026-08-02] ✅MiniWorld: Democratizing the Training of Video World Models from Scratch
[2026-07-31] ✅WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation
[2026-07-31] ❌️BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning
[2026-07-30] ✅PhiZero: A World Model Built Around Physical Language
[2026-07-30] ✅ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
[2026-07-29] ✅StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
[2026-07-29] ✅ContactFlow: A video action conditioning that transfers across embodiments
[2026-07-29] ❌️CG-World: A Large-Scale World-State Dataset and Protocol for World Models
[2026-07-28] ✅Wonder: Video World Model Done Better
[2026-07-23] ✅Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
[2026-07-21] ✅Generative World Renderer at the Speed of Play
[2026-07-15] ❌️From Pixels to States: Rethinking Interactive World Models as Game Engines
[2026-07-15] ❌️M⁴World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming
[2026-07-15] ❌️Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Chassis Adaptation
[2026-07-13] ❌️Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency
[2026-07-13] ✅ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space
[2026-07-13] ✅Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
[2026-07-12] ❌️LIDAR-AD: A Decoder-Free Latent-Interaction Dreamer with Action-Residual Chains for Autonomous Driving
[2026-07-12] ❌️Is Energy Guidance All You Need? Training-Free Norm Injection for Driving World Models
[2026-07-12] ❌️World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning
[2026-07-11] ❌️Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models
[2026-07-10] ✅PanoWorld: Real-World Panoramic Generation
[2026-07-10] ❌️Causally Debiased Latent Action Model for Embodied Action Conditioned World Models
[2026-07-08] ✅Infinite Worlds with Versatile Interactions
[2026-07-07] ✅RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation
[2026-07-07] ✅AlayaWorld: Long-Horizon and Playable Video World Generation
[2026-07-07] ✅MoWorld: A Flash World Model
[2026-07-06] ✅Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
[2026-07-06] ✅Multiplayer Interactive World Models with Representation Autoencoders
[2026-07-05] ✅Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
[2026-07-04] ❌️Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
[2026-07-03] ✅Vidu S1: A Real-Time Interactive Video Generation Model
[2026-07-02] ✅WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory
[2026-07-02] ✅PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
[2026-07-01] ❌️RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation
- 🔥 Recent Paper Updates
- 🌐 General World Model
- 🤝 Multi-Agent World Model
- 📊 Benchmark
- 💾 Dataset
- 🎬 Long Video Generation
- 🧠 Memory Consistency
- 🎮 Action Control
- 🖱️ Interactive
- 🔧 Post-Training
- ⚙️ Physics
- 📏 Evaluation
- 📚 Survey
- 📝 Citation
[2026-06] ✅🌟(Omnimodal world model for understanding, generation, simulation, and action)Cosmos 3: Omnimodal World Models for Physical AI
[2026-05] ✅🌟(Full-stack open-source real-time interactive video world model framework)minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
[2026-01] ✅🌟(Minute-level Memory)Lingbot-World:Advancing Open-source World Models
[2025-08] ✅🌟(First open-source real-time interactive World Model)Matrix-game 2.0: An open-source real-time and streaming interactive world model
[2026-09] ✅Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence
[2026-09] ✅Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
[2026-09] ✅SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
[2026-08] ✅NVIDIA Cosmos-H-Dreams: Real-Time Generative Physics Simulation for Surgical Robotics
[2026-08] ❌️EchoWM: Open and Enterable Omnimodal World Models
[2026-08] ✅Marionette: Predicting World States, Rendering Geometry, Painting Appearance
[2026-08] ✅ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
[2026-08] ✅GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
[2026-08] ✅MiniWorld: Democratizing the Training of Video World Models from Scratch
[2026-07] ❌️BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning
[2026-07] ✅PhiZero: A World Model Built Around Physical Language
[2026-07] ✅StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
[2026-07] ✅Wonder: Video World Model Done Better
[2026-07] ✅Generative World Renderer at the Speed of Play
[2026-07] ❌️Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Chassis Adaptation
[2026-07] ✅ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space
[2026-07] ✅Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
[2026-07] ❌️LIDAR-AD: A Decoder-Free Latent-Interaction Dreamer with Action-Residual Chains for Autonomous Driving
[2026-07] ✅PanoWorld: Real-World Panoramic Generation
[2026-07] ✅Infinite Worlds with Versatile Interactions
[2026-07] ❌️AlayaWorld: Long-Horizon and Playable Video World Generation
[2026-07] ✅MoWorld: A Flash World Model
[2026-06] ❌️DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
[2026-06] ✅DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model
[2026-06] ❌️MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold
[2026-06] ✅Next Forcing: Causal World Modeling with Multi-Chunk Prediction
[2026-06] ✅DreamX-World 1.0: A General-Purpose Interactive World Model
[2026-06] ❌️GeoStream: Toward Precise Camera Controlled Streaming Video Generation
[2026-05] ✅Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models
[2026-05] ✅SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
[2026-05] ✅Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video
[2026-04] ✅Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
[2026-03] ✅InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model
[2025-12] ✅Yume-1.5: A Text-Controlled Interactive World Generation Model
[2025-12] ❌️RELIC: Interactive Video World Model with Long-Horizon Memory
[2025-12] ✅Astra: General Interactive World Model with Autoregressive Denoising
[2025-11] ❌️Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
[2025-11] ❌️PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
[2025-11] ✅MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
[2025-08] ❌️Yan: Foundational Interactive Video Generation
[2025-07] ✅Yume: An Interactive World Generation Model
[2025-06] ✅Matrix-Game: Interactive World Foundation Model
[2025-06] ✅Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
[2024-02] ❌️Genie: Generative Interactive Environments
[2026-08] ✅Population-Scalable Multi-Agent World Modeling
[2026-08] ✅MASS: Multiplayer World Models with Authoritative Shared State
[2026-07] ✅Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
[2026-07] ✅Multiplayer Interactive World Models with Representation Autoencoders
[2026-06] ✅Prisma-World: Camera-Controllable Multi-Agent Video World Model
[2026-06] ❌️MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data
[2026-05] ✅Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
[2026-04] ✅MultiWorld: Scalable Multi-Agent Multi-View Video World Models
[2026-05] ✅🌟(Multi-turn interactive world model evaluation)WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
[2026-02] ✅🌟(First Open Domain Memory & Action Benchmark)MIND: Benchmarking Memory Consistency and Action Control in World Models
[2026-09] ✅PhysWeep: Does a Video Generator Realize the Physics You Ask For?
[2026-08] ❌️RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction
[2026-08] ✅PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
[2026-08] ✅R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models
[2026-08] ❌️StreamAV-Bench: A Comprehensive Benchmark for Streaming Audio-Video Generation
[2026-08] ❌️CamWorldQA: Perceptual Quality Assessment of Camera-Controlled World Video Generation
[2026-08] ✅PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
[2026-08] ✅H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
[2026-08] ✅WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
[2026-06] ❌️WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
[2026-06] ✅WorldOlympiad: Can Your World Model Survive a Triathlon?
[2025-05] ✅Toward Memory-Aided World Models: Benchmarking via Spatial Consistency
[2025-04] ✅WorldScore: A Unified Evaluation Benchmark for World Generation
[2025-02] ✅WorldModelBench: Judging Video Generation Models As World Models
[2024-10] ✅WorldSimBench: Towards Video Generation Models as World Simulators
[2026-09] ❌️ Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation
[2026-08] ✅ WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
[2026-08] ✅ Sekai2: From World Exploration to Interactive World Modeling
[2026-07] ❌️ CG-World: A Large-Scale World-State Dataset and Protocol for World Models
[2026-07] ✅ Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
[2026-06] ✅ PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models
[2026-06] ✅ EgoCS-400K: An Egocentric Gameplay Dataset for World Models
[2025-09] ✅ SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
[2025-09] ✅ OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
[2025-06] ✅ Sekai: A Video Dataset towards World Exploration
[2024-11] ✅ GameGen-X: Open-World Video Game Dataset
[2026-05] ✅ 🌟(First architecture supporting arbitrary step OPD) AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
[2026-03] ✅🌟(Scale up Real-Time model to 14B)Helios: Real Real-Time Long Video Generation Model
[2026-02] ✅🌟(ODE perspective)Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation
[2025-12] ❌️🌟(Compression recovery training)Pretraining Frame Preservation in Autoregressive Video Memory Compression
[2025-10] ❌️🌟(Solves Self-Forcing issues)Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
[2025-08] ❌️🌟(Learnable sparse attention routing)Mixture of Contexts for Long Video Generation
[2025-06] ✅🌟(DMD perspective)Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
[2025-04] ✅🌟(Context compression)Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
[2024-12] ✅🌟(Teacher->Student acceleration)From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
[2024-07] ✅🌟(Training techniques)Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
[2026-08] ❌️WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression
[2026-08] ✅Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
[2026-06] ✅Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation
[2026-05] ✅DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
[2026-05] ✅VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
[2026-05] ✅AdaState: Self-Evolving Anchors for Streaming Video Generation
[2026-05] ❌️Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
[2026-05] ✅Q-ARVD: Quantizing Autoregressive Video Diffusion Models
[2026-05] ❌️DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
[2026-05] ✅FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching
[2026-05] ✅LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
[2026-05] ✅(One Step)One-Forcing: Towards Stable One-Step Autoregressive Video Generation
[2026-03] ✅Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
[2026-02] ✅Context Forcing: Consistent Autoregressive Video Generation with Long Context
[2026-02] ✅Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
[2026-02] ❌️Mode Seeking meets Mean Seeking for Fast Long Video Generation
[2026-02] ❌️LIVE: Long-horizon Interactive Video World Modeling
[2026-02] ❌️LIGHT FORCING: Accelerating Autoregressive Video Diffusion via Sparse Attention
[2026-02] ✅Pathwise Test-Time Correction for Autoregressive Long Video Diffusion Models
[2025-12] ✅VideoSSM: Autoregressive Long Video Generation with Hybrid State Space Memory
[2025-09] ✅Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
[2025-09] ✅LongLive: Real-time Interactive Long Video Generation
[2025-03] ✅FAR: Frame Autoregressive Model for Both Short- and Long-Context Video Modeling
[2025-03] ✅NOVA: Autoregressive Video Generation without Vector Quantization
[2024-10] ✅Progressive Autoregressive Video Diffusion Models
[2026-02] ✅🌟(w/o Pose Memory)Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory
[2025-06] ❌️🌟(w/ Pose FOV overlap retrieval Memory)Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
[2026-09] ✅OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mapping
[2026-08] ✅Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
[2026-08] ✅ReWorld: An Interactive World Model with Long-Horizon Memory
[2026-08] ❌️Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
[2026-08] ✅Addressable Memory for Video World Models
[2026-06] ✅MemLearner: Learning to Query Context memory for Video World Models
[2026-06] ❌️Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation
[2026-06] ❌️TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment
[2026-06] ✅Latent Spatial Memory for Video World Models
[2026-06] ✅Echo-Memory: A Controlled Study of Memory in Action World Models
[2026-06] ✅Geometry-Aware Implicit Memory for Video World Models
[2026-05] ✅OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation
[2026-05] ✅E3C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control
[2026-05] ✅WorldKV: Efficient World Memory with World Retrieval and Compression
[2026-03] ❌️MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
[2026-02] ✅AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
[2026-01] ✅StableWorld: Towards Stable and Consistent Long Interactive Video Generation
[2025-10] ❌️Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
[2025-06] ✅VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
[2025-06] ❌️Video World Models with Long-term Spatial Memory
[2025-06] ✅DeepVerse: 4D Autoregressive Video Generation as a World Model
[2025-05] ✅Learning World Models for Interactive Video Generation
[2025-04] ✅WorldMem: Long-term Consistent World Simulation with Memory
[2026-03] ✅🌟(First out-of-view Event Dynamic Memory)LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
[2026-02] ✅🌟(First multiplayer Minecraft World Model)Solaris: Building a Multiplayer Video World Model in Minecraft
[2026-01] ✅🌟(First out-of-view Dynamic Memory)Flow Equivariant World Models: Structured Dynamics Outside the Field of View
[2026-07] ✅WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory
[2026-05] ✅Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
[2026-03] ✅Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
[2026-02] ✅🌟(Transferable skills)Olaf-World: Orienting Latent Actions for Video World Modeling
[2025-03] ✅🌟(No longer limited to fixed controls)AdaWorld: Learning Adaptable World Models with Latent Actions
[2025-01] ❌️🌟(Action control from games)GameFactory: Creating New Games with Generative Interactive Videos
[2026-09] ✅World in World: Explore the World with World Models
[2026-09] ✅ActionSplice: In-Flight Action Editing for Interactive World Models
[2026-09] ✅H3-World: Turning Language Understanding into World Control
[2026-09] ✅MeRoPE: Metric Rotary Position Embedding for Camera-Controlled Video Generation
[2026-08] ✅CAER: Causal Action Effect Reweighting for World Model Training
[2026-08] ✅AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization
[2026-08] ✅Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting
[2026-08] ✅CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
[2026-08] ✅4DStreamCtrl: Interactive Video Generation with Online 4D Control
[2026-08] ✅WorldMind: Decoupled Game World Model for State-Aware NPC Behavior
[2026-08] ✅Hydra-0: Action Flow for Generalist World Modeling and Control
[2026-08] ✅DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
[2026-08] ❌️Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts
[2026-08] ❌️Overcoming Statistical Bias in Action-Controllable World Models
[2026-08] ✅EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
[2026-08] ❌️Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
[2026-07] ✅ContactFlow: A video action conditioning that transfers across embodiments
[2026-07] ❌️Is Energy Guidance All You Need? Training-Free Norm Injection for Driving World Models
[2026-07] ❌️Causally Debiased Latent Action Model for Embodied Action Conditioned World Models
[2026-07] ✅RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation
[2026-07] ✅Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
[2026-07] ❌️Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
[2026-06] ✅Streaming Video Generation with Streaming Force Control
[2026-05] ✅E3C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control
[2026-05] ✅SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
[2026-05] ✅ReactiveGWM: Steering NPC in Reactive Game World Models
[2025-12] ✅🌟(Edit objects: add/remove/recolor)Spatia: Video Generation with Updatable Spatial Memory
[2025-11] ❌️🌟(Text interaction: spawn weapons, affect environment)Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
[2026-09] ✅Programmable World Model
[2026-08] ✅HelloWorld: Enabling Socially Interactive Characters in Video World Models
[2026-08] ✅RealWeather: Realistic and Scene-Faithful Weather Translation with Driving World Models
[2026-07] ✅Vidu S1: A Real-Time Interactive Video Generation Model
[2026-08] ✅SCOPE: Score-Isolated Agentic Optimization for Video World Models
[2026-08] ✅WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
[2026-07] ❌️World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning
[2026-04] ✅World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
[2026-02] ✅WorldCompass: Reinforcement Learning for Long-Horizon World Models
[2025-05] ✅RLVR-World: Training World Models with Reinforcement Learning
[2026-09] ❌️TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image
[2026-08] ✅Distilling Physical Priors into Streaming World Models
[2026-08] ❌️muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards
[2026-07] ✅PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
[2026-06] ❌️IOI: Decoupling Kinematics and Physics for Interactive World Models
[2026-03] ✅RealWonder: Real-Time Physical Action-Conditioned Video Generation
[2026-09] ✅WorldReward: Reward Modeling for Camera-Conditioned World Models
[2026-08] ✅Can Video World Models Track Unobserved World States?
[2026-08] ❌️Where World Models Break: Natural-Input Failure Discovery
[2026-08] ❌️How Can Driving World Models Do Counterfactual Prediction?
[2026-08] ❌️Twin Rollouts: Noise-Coupled Counterfactual Branching in Interactive Video World Models
[2026-08] ✅WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation
[2026-07] ❌️RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation
[2026-06] ✅BadWorld: Adversarial Attacks on World Models
[2026-03] ✅Interactive World Simulator for Robot Policy Training and Evaluation
[2026-04] ✅🌟(Defines advanced world models around perception, interaction, and long-term memory, and introduces OpenWorldLib as a unified inference framework.) OpenWorldLib: A Unified Codebase and Definition of Advanced World Models
[2026-07] ❌️From Pixels to States: Rethinking Interactive World Models as Game Engines
[2026-05] ✅Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends
[2026-03] ❌️Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms
[2026-07] ❌️Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models
[2025-05] ✅Vid2World: Crafting Video Diffusion Models to Interactive World Models
If you find this repository useful, please consider citing this list:
@misc{ye2026awesomeiwm,
title = {Awesome-Interactive-World-Model},
author = {Yixuan Ye and Ruiqi Wu},
journal = {GitHub repository},
url = {https://github.com/EasonTuT/Awesome-Interactive-World-Model},
year = {2026},
}