KAIST AI · †Corresponding author
Ask4NV learns novel-view generation jointly with 3D language understanding in one unified multimodal model. Spatial question answering and visual grounding supervise the generation transformer through joint self-attention, improving image fidelity and geometric consistency in and out of domain.
arXiv 2026
Two baselines from the same Lance checkpoint: Lance-FT (generation only) and Lance-FT + ViT (same ViT conditioning and mask, no understanding training).
Patch features from Video-3D-LLM's 3D-tuned SigLIP, mean cosine similarity between generated and ground-truth views. Ask4NV is highest on all four datasets; the margin over Lance-FT + ViT is the effect of joint training alone.
TBD