ME-World
Multi-Agent Egocentric World Model
with Fine-Grained Embodied Interaction

KAIST AI
†: Co-corresponding authors
arXiv Preprint 2026
TL;DR

ME-World generates synchronized ego streams for multiple interacting agents in a shared world. Joint denoising, shared action conditioning, and shared environment memory keep actions, the environment, and interaction outcomes consistent across views.

Real Interactions: Two People, One Kitchen

Both agents wear head-mounted cameras while cooking together. Given each agent's head trajectory, both agents' body motion and the shared observation history, ME-World generates the two ego streams jointly, so the partner's hands, the counter and the objects being handled agree across the two views.

Agent 1Agent 2
Agent 1Agent 2
Agent 1Agent 2
Agent 1Agent 2
Synchronized ego streams generated jointly on the real benchmark. Each clip shows Agent 1 (left) and Agent 2 (right) at the same instant.

Synthetic Worlds: Distinct Characters, Shared Scenes

On the synthetic benchmark each agent is a different rigged avatar, so the model has to keep the counterpart's identity stable while both viewpoints move. The third-person view (TPV, the square tile at the right of each pair) is a separate camera placed in the scene so the interaction can be read from outside; it is neither generated nor given to the model.

Agent 1Agent 2
TPV (reference only)
Agent 1Agent 2
TPV (reference only)
Agent 1Agent 2
TPV (reference only)
Agent 1Agent 2
TPV (reference only)
Generated ego pairs (left) with the corresponding third-person view, TPV (right).

Beyond Two Agents, Beyond One Chunk

The formulation is written for N agents. Without changing the model, we append a third stream to the joint sequence and generate three synchronized ego views of three interacting avatars.

Agent 1Agent 2Agent 3
TPV (reference only)
Agent 1Agent 2Agent 3
TPV (reference only)
Agent 1Agent 2Agent 3
TPV (reference only)
Agent 1Agent 2Agent 3
TPV (reference only)
Three-agent generation. The third-person view is a visual reference only.

Autoregressive Long-Horizon Generation

Chunks are chained by conditioning each new chunk on two frames of the previous one. The model is trained with teacher forcing: the conditioning frames come from ground truth, lightly perturbed with noise, so it learns to continue from imperfect handoff frames without ever rolling out during training. At inference a two-agent sequence of 221 frames keeps the same agents and the same room throughout.

Agent 1Agent 2
221 frames, autoregressive
TPV (reference only)
Agent 1Agent 2
221 frames, autoregressive
TPV (reference only)
Autoregressive two-agent generation over 221 frames. The third-person view is a visual reference only.

Comparison with Existing Methods

We compare against three families using released checkpoints without fine-tuning: multi-view video generation (GEN3C, AnyView-DVS), which receives one agent's ground-truth stream as the source view; single-ego world models (EgoSim, JointControlVideo); and general world models (DreamX-World, LingBot-World), which generate the two streams independently. Pick an example below.

Ground truth
ME-World (Ours)
GEN3C
AnyView-DVS
EgoSim
JointControlVideo
DreamX-World
LingBot-World
Each tile shows Agent 1 (left) and Agent 2 (right). Multi-view baselines are given Agent 1's ground-truth stream as input.

Data Overview

Training needs synchronized multi-ego recordings that share one world state. We combine real two-person recordings from CoMind with a synthetic set rendered from interacting avatars, and process both into the same conditioning format: scene, pose, camera, reference images and text.

Real Recordings

Two people wear head-mounted cameras while cooking together. Body and hand joints of both agents are rendered into each view with one fixed colour per person, so the same person keeps the same colour in every stream.

Agent 1: ego stream
Agent 1: shared action condition
Agent 2: ego stream
Agent 2: shared action condition

Synthetic Multi-Agent Interactions

Paired motions from Inter-X and InterHuman are retargeted onto rigged VRM avatars, placed at validated anchors in indoor and outdoor Blender scenes, and rendered synchronously from two head-mounted cameras. Rendered depth and instance masks give exact warped-scene, visibility, pose and ray conditions for every frame.

Agent 1: ego stream
Agent 1: shared action condition
Agent 2: ego stream
Agent 2: shared action condition
Third-person render of the same synthetic interaction.

Model Overview

ME-World fine-tunes a video diffusion transformer to denoise all agents' ego streams as one shared token sequence. Each stream is conditioned on its own viewing rays and on every agent's body motion, and the whole sequence is grounded by a shared environment memory built from all agents' observation history.

ME-World architecture: condition construction per agent and joint denoising with cross-stream anchor memory

Joint Multi-Agent Generation

The noisy latents of all agents are concatenated along the token dimension so that self-attention exchanges information across streams at every layer. Streams share one positional frame and are told apart by a learnable agent embedding. Text conditioning stays per stream through separate cross-attention paths.

Shared Action Conditioning

Instead of feeding each stream only its own action, we render the body and hand skeletons of all agents into each target view, with a consistent colour per identity. The target agent's own hands and arms appear as ego-visible motion, and the other agents appear as observed motion. Head motion is expressed as per-pixel rays in a shared canonical frame, so camera trajectories across streams are geometrically comparable.

Shared action conditioning on a synthetic clip. Left: both agents' bodies and head cameras in the shared world. Right: all agents' motion projected into each agent's camera becomes that stream's pose condition, next to the generated frame.

Shared Environment Memory

All agents' past observations are pooled. The stream-aligned geometric memory path picks the history frame with the largest reprojected coverage of each target view, masks out the agents, and warps it into that view together with a visibility mask. The cross-stream anchor memory path greedily selects history frames that cover complementary regions and appends them, unwarped, as clean reference tokens that every stream attends to during denoising.

Shared environment memory on a real clip. Left: per agent, the warped history frame, the shared action condition and the generated frame. Right: the scene lifted to 3D with both agents' camera trajectories; the yellow frustum is the history frame warped into the current view, and the green frustums are the eight shared anchor references.

Components Ablation

Starting from Cosmos-Predict2.5 with image and text conditioning, we add the conditions one at a time. Joint generation couples the streams so appearance and environment stop diverging. Shared action renders every agent's pose into each view so the same interaction appears from both sides. Shared environment memory anchors both streams to the same room.

1Cosmos-Predict2.5, zero-shot
2Cosmos-Predict2.5, fine-tuned
3+ self hand pose + history warp
4(3) + joint generation
5(4) + shared action
6(4) + shared environment
7(3) + shared action + env, no joint
8Full model (Ours)
Variant numbers follow Table 2 of the paper. Each tile shows Agent 1 (left) and Agent 2 (right).

Measuring Shared-World Consistency

Cross-view agreement can come for free from shared inputs. To isolate what the model itself keeps consistent, we compare generated streams only where the input gives no direct answer, using co-visible correspondences from ground-truth depth and cameras.

Alongside these we report camera control, self and other-agent action control, and standard video quality. ME-World improves all of them over multi-view, single-ego and general world models on both benchmarks. See the paper for the tables.

Abstract

Egocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared environment. Existing multi-agent world models generate observations for multiple agents, but rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored. We formulate embodied multi-agent world modeling as synchronized ego-stream generation for multiple agents interacting through fine-grained actions in a shared world. This requires cross-view action consistency, shared-environment consistency, and consistent propagation of interaction-induced state updates. We propose ME-World, which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory. We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency. Experiments show ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.

Citation

@article{chung2026meworld,
  title={{ME-World: Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction}},
  author={Chung, Dahyun and Jin, Siyoon and Choi, Hyunwook and An, Honggyu and Seo, Junyoung and Kim, Hyunsung and Kim, Seung Wook and Kim, Seungryong},
  journal={arXiv preprint arXiv:XXXX.XXXXX},
  year={2026}
}