Can Learning to Answer Spatial Questions Improve
Novel-View Generation?

by Jahyeok Koo, Hyeonseo Yu, Seungho Jang, Jaewoo Jung, Jiwon Kang, Seokju Cho, and Seungryong Kim†

KAIST AI  ·  †Corresponding author

Ask4NV learns novel-view generation jointly with 3D language understanding in one unified multimodal model. Spatial question answering and visual grounding supervise the generation transformer through joint self-attention, improving image fidelity and geometric consistency in and out of domain.

arXiv 2026

Results

Finding 1Ask4NV improves novel-view image quality across multiple benchmarks.

Why does joint training help?

Two baselines from the same Lance checkpoint: Lance-FT (generation only) and Lance-FT + ViT (same ViT conditioning and mask, no understanding training).

Finding 23D language supervision improves cross-view correspondence beyond ViT conditioning.
Ask4NVLance-FTlayer 26
Query view · query token in red
query view
Finding 3Joint training improves similarity in a 3D-aware feature space.

Patch features from Video-3D-LLM's 3D-tuned SigLIP, mean cosine similarity between generated and ground-truth views. Ask4NV is highest on all four datasets; the margin over Lance-FT + ViT is the effect of joint training alone.

BibTeX

TBD