By jointly generating panoramic observer, World Observer continuously tracks and updates regions beyond the Actor’s current view.
Observers are not tied to the actor. We can choose which regions to observe independently of the actor’s position.
We can place multiple Observers where we want, and track what we cannot see from the single Actor/Observer.
World Observer is trained on synchronized actor and observer videos covering diverse observer configurations. Real panoramic videos provide observers that move with the actor, while CARLA-rendered synthetic data adds observers placed independently of the actor and multiple observers at once.
A single video DiT jointly generates the actor and panoramic observer streams in one token sequence, so objects that leave the actor's view keep evolving in the observer.
A query on a re-entering object attends to the observer region
holding its out-of-view state; the non-joint baseline loses the object.
An object leaves the actor's view and comes back. Existing models often fall into a frame-locked, lost, frozen, or impostor state, while World Observer preserves its identity and state evolution across the out-of-view interval.
“A person pulling two suitcases walks across the lobby, exits the view, keeps moving naturally, and re-enters.”
“A white armored truck driving straight ahead exits the view, keeps moving naturally, and re-enters.”
“A pedestrian walking straight ahead moves out of the view, keeps walking naturally, and re-enters.”
“A dark blue shuttle bus driving straight moves out of the view, keeps moving naturally, and re-enters.”
Decoupling the observer gives it its own prompt, so text can drive events beyond the actor’s view.
“Santa Claus walks in from the hallway, crosses the room, and sits down in an armchair.”
How can a world model continuously observe regions beyond the actor’s current view? Video world models simulate how an environment evolves from an agent’s actions, yet remain actor-centric. Once an object leaves the actor’s view, they lose direct evidence of its evolution, often failing to preserve its state and dynamics upon re-entry. To address this, we introduce World Observer, which decouples observing from acting by jointly generating a perspective actor for the agent-centric view with one or more panoramic observers that watch selected world regions. This allows objects that leave the actor’s view to remain visually evolving in an observer, so their updated states are reflected when they re-enter. We ground the actor and observers by warping from a shared panoramic source for explicit geometric correspondence, and introduce an Observer Sink of high-resolution perspective references to restore fine appearance upon re-entry. Since the observers are decoupled from the actor, they can be placed freely across the scene, extended to multiple locations for broader coverage, and driven by control signals to steer out-of-view evolution. To evaluate out-of-view evolution, we further introduce world-space metrics and a benchmark spanning real and synthetic scenes. World Observer substantially improves out-of-view dynamics while remaining competitive in visual fidelity, camera control, and 3D adherence.
@misc{choi2026worldobserverjointactorobserver,
title={World Observer: Joint Actor-Observer Generation for Persistent World Modeling},
author={Hyunwook Choi and Dahyun Chung and Hyunsung Kim and Siyoon Jin and Jinhyeok Choi and Junyoung Seo and Seungryong Kim},
year={2026},
eprint={2610.02162},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.02162},
}