A miniature garden
AN INTERACTIVE WORLD MODEL

WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon

Worlds that evolve with you.

Video Results · Long-Horizon Consistency & Flexible Control
WORLDPLAY2 / EXPLORING POSSIBILITY
TL;DR

WorldPlay2 generalizes across scenes and characters, achieving long-horizon consistency alongside flexible control.

OVERVIEW VIDEO
1 min 58 s · sound on · Watch on YouTube ↗
Overview video preview
ABSTRACT

Interactive world models require responding in real time to versatile controls and maintaining long-horizon consistency. However, modeling heterogeneous controls remains difficult, while explosive contexts and unstable distillation impede achieving both long-horizon consistency and real-time responsiveness. In this paper, we present WorldPlay2, an interactive world model that couples a factorized hybrid control interface with a co-design of compressed memory and stable distillation. 1) Our factorized hybrid control interface integrates frame-aligned action control with structured semantic control that explicitly disentangles scene appearance, character identity, and dynamic semantic events, thereby facilitating effective control learning. 2) To achieve efficient long-horizon modeling, we compress historical contexts into compact memory tokens shared by the autoregressive student and the bidirectional teacher. This design enables clip-wise, memory-conditioned score evaluation instead of jointly processing an entire long rollout, substantially reducing distillation overhead. 3) We further propose Stable Forcing, which initializes the autoregressive student via a few-step strategy and leverages full-rollout replay to preserve the quality of long-horizon rollouts, ensuring robust and stable distillation. Extensive experiments demonstrate the strong generalizability of our model and its superior performance compared to existing methods.

METHOD

Three pieces. One real-time world.

A factorized control interface to support flexible interactions, a compact memory for efficient long-horizon modeling and distillation, and a stable distillation framework for stable distillation.

Figure 2: overview of WorldPlay2 — factorized hybrid control feeds an autoregressive world model whose history is compressed into sink, compressed, and temporal memory tokens
Figure 2Factorized hybrid control splits the input into frame-aligned actions and structured semantic control; a history compressor turns past frames into compact memory tokens for efficient long-horizon inference.
01

Factorized hybrid control interface

Frame-aligned action control. Continuous camera pitch and yaw, discrete longitudinal and lateral movements, the camera perspective, and a jump action are embedded into a unified action representation.

Structured semantic control. A structured caption partitions high-level signals into three decoupled fields: scene for environmental content, character for the identity and appearance of the controlled entity, and event for the current semantic change.

At = {at, dt},  dt = (dscene, dcharacter, devent)
02

Distillation-oriented compressed memory

A learnable history compressor encodes the history into compact memory tokens, consisting of sink tokens, compressed tokens, and adjacent temporal tokens.

Compared with full-resolution contexts, the sequence length is reduced by approximately l·s2 = 32× (l = 2 temporal and s = 4 spatial downsampling).

m<t = [xsink; Cφ(x<t, xlr<t); xtmp]
03

Stable Forcing

Few-step initialization. The autoregressive student is warm-started by extending PDD to the memory-augmented model, so that its few-step rollout distribution remains close to the teacher distribution.

Full-rollout replay. Each chunk undergoes full few-step sampling during rollout, while a randomly selected intermediate step is cached and replayed with gradients.

Efficient score evaluation. The long rollout is partitioned into clips, and real and fake scores are computed per clip, conditioned on compact memory tokens.

sfake/real = v(xσ[iL:(i+1)L], σ, m<iL, A≤(i+1)L)
Figure 4: Stable Forcing — full few-step rollout, memory-augmented bidirectional teachers for real and fake scores, and gradient replay of one cached step per chunk
Figure 4Stable Forcing combines few-step initialization, full-rollout replay, and efficient score evaluation for stable long-horizon distillation.
01 LONG-HORIZON CONSISTENCY

Keep exploring. The world remembers.

02 FLEXIBLE CONTROL

Your world. Your move.

03 COMPARISONS

See the difference. Side by side.

COMPARE WITH

Responding to semantic events

Cloud and balloon events in the same city scene. Compare how each model responds to the requested changes.

WorldPlayBASELINE
WorldPlay2OURS
SYNCHRONIZED PLAYBACK0s / 13s
Same scene · shared playback controlsFigure 5