World Observer:
Joint Actor-Observer Generation
for Persistent World Modeling

KAIST AI
arXiv preprint
0:00
AI-Generated Video by WorldObserver
TL;DR

How can a world model continuously observe regions beyond the actor’s current view? World Observer jointly generates an actor view with panoramic observers that keep unseen regions observable. Placing observers where we want, it evolves and controls the actor’s unseen regions.

Evolve World State Beyond the Actor’s View

By jointly generating panoramic observer, World Observer continuously tracks and updates regions beyond the Actor’s current view.

AI-Generated Video by WorldObserver

Place Observer Freely

Observers are not tied to the actor. We can choose which regions to observe independently of the actor’s position.

AI-Generated Video by WorldObserver

Place Multiple Observers

We can place multiple Observers where we want, and track what we cannot see from the single Actor/Observer.

AI-Generated Video by WorldObserver

Control World State Beyond the Actor’s View

Decoupling the observer from the actor also enables independent observer control, allowing us to control beyond the Actor’s current view.

AI-Generated Video by WorldObserver

Data Overview

World Observer is trained on synchronized actor and observer videos covering diverse observer configurations. Real panoramic videos provide observers that move with the actor, while CARLA-rendered synthetic data adds observers placed independently of the actor and multiple observers at once.

Moving Observer
Actor
Observer
Flexible Observer
Actor
Observer
Multi-Observer
Actor
Observer 1
Observer 2

Model Overview

A single video DiT jointly generates the actor and panoramic observer streams in one token sequence, so objects that leave the actor's view keep evolving in the observer.

World Observer architecture and Observer Sink
  • Joint actor-observer modeling: objects leaving the actor’s view keep evolving in the observer, and both streams share this world state. Separate prompts let the observer act as both memory and control.
  • Panoramic warping & flexible resolution: warping both streams from a shared panorama fixes each stream’s camera condition and aligns actor and observer geometry. Since the observer only maintains the world state rather than being displayed, it is generated at a flexible resolution.
  • Observer Sink: high-resolution perspective views of the surroundings, cropped from the panorama, restore fine appearance when a region re-enters the actor view.
Self-attention visualization of a re-entering object A query on a re-entering object attends to the observer region holding its out-of-view state; the non-joint baseline loses the object.

Comparison with World Models

An object leaves the actor's view and comes back. Existing models often fall into a frame-locked, lost, frozen, or impostor state, while World Observer preserves its identity and state evolution across the out-of-view interval.

“A person pulling two suitcases walks across the lobby, exits the view, keeps moving naturally, and re-enters.”

Ground Truth
Ours (Single)
DreamX-World
FantasyWorld
HY-World 1.5
HyDRA
LingBot-World
Matrix-Game 3
OmniRoam
PanoWorld

“A white armored truck driving straight ahead exits the view, keeps moving naturally, and re-enters.”

Ground Truth
Ours (Single)
DreamX-World
FantasyWorld
HY-World 1.5
HyDRA
LingBot-World
Matrix-Game 3
OmniRoam
PanoWorld

“A pedestrian walking straight ahead moves out of the view, keeps walking naturally, and re-enters.”

Ground Truth
Ours (Single)
Ours (Multi)
DreamX-World
FantasyWorld
HY-World 1.5
HyDRA
LingBot-World
Matrix-Game 3
OmniRoam
PanoWorld

“A dark blue shuttle bus driving straight moves out of the view, keeps moving naturally, and re-enters.”

Ground Truth
Ours (Single)
Ours (Multi)
DreamX-World
FantasyWorld
HY-World 1.5
HyDRA
LingBot-World
Matrix-Game 3
OmniRoam
PanoWorld

Ablation Study

Components

OursFull model
Actor
Observer
Out-of-view objects keep their motion and re-enter with their appearance intact.
w/oObserver Sink
Actor
Observer
Dynamics survive, but fine appearance degrades under large viewpoint changes.
w/oObserver
Actor
No observer
Generated from the actor alone, out-of-view dynamics are not maintained.
w/oActor
Actor (crop of panorama)
Observer
Generating only the panorama keeps out-of-view content, but the distorted panorama loses visual fidelity.

Prompt Placement

Decoupling the observer gives it its own prompt, so text can drive events beyond the actor’s view.

Prompt

“Santa Claus walks in from the hallway, crosses the room, and sits down in an armchair.”

altPrompting Actor
Actor
Observer
With the prompt only on the actor, text control is hallucinated inside the actor view.
OursPrompting Decoupled Observer
Actor
Observer
A separate observer prompt enables control over the entire scene the observer sees.

Abstract

How can a world model continuously observe regions beyond the actor’s current view? Video world models simulate how an environment evolves from an agent’s actions, yet remain actor-centric. Once an object leaves the actor’s view, they lose direct evidence of its evolution, often failing to preserve its state and dynamics upon re-entry. To address this, we introduce World Observer, which decouples observing from acting by jointly generating a perspective actor for the agent-centric view with one or more panoramic observers that watch selected world regions. This allows objects that leave the actor’s view to remain visually evolving in an observer, so their updated states are reflected when they re-enter. We ground the actor and observers by warping from a shared panoramic source for explicit geometric correspondence, and introduce an Observer Sink of high-resolution perspective references to restore fine appearance upon re-entry. Since the observers are decoupled from the actor, they can be placed freely across the scene, extended to multiple locations for broader coverage, and driven by control signals to steer out-of-view evolution. To evaluate out-of-view evolution, we further introduce world-space metrics and a benchmark spanning real and synthetic scenes. World Observer substantially improves out-of-view dynamics while remaining competitive in visual fidelity, camera control, and 3D adherence.

Citation

@misc{choi2026worldobserverjointactorobserver,
  title={World Observer: Joint Actor-Observer Generation for Persistent World Modeling},
  author={Hyunwook Choi and Dahyun Chung and Hyunsung Kim and Siyoon Jin and Jinhyeok Choi and Junyoung Seo and Seungryong Kim},
  year={2026},
  eprint={2610.02162},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2610.02162},
}