Papers
arxiv:2610.04722

NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis

Published on Oct 3
ยท Submitted by
Ruslan Rakhimov
on Oct 8
Authors:
,
,

Abstract

Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor this generation process to explicit camera geometry, we propose Multi-scale Projective Pose Encoding, which injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention at every resolution. NAMVIS further combines global conditioning with dense geometry-aware cross-attention, enabling the model to preserve source-view appearance while maintaining target-view consistency. Across Objaverse, GSO, and OmniObject3D, NAMVIS outperforms diffusion-based baselines in PSNR, SSIM, and LPIPS, while running over 3 times faster than the evaluated diffusion baselines under the same evaluation setting. These results suggest that geometry-conditioned next-scale autoregression is a promising and efficient alternative to diffusion for sparse-view multi-view synthesis. Additional qualitative results, videos, and resources are available at https://corl-team.github.io/namvis/

Community

Paper author Paper submitter

NAMVIS replaces diffusion with next-scale autoregression for sparse-view multi-view synthesis: target views are generated coarse-to-fine in 7 scale steps, with all tokens of a scale predicted in parallel and a multi-scale PRoPE injecting camera geometry into attention. It takes 0.6 s per view vs 2.0 s for the fastest diffusion baseline we tested, with better PSNR/SSIM/LPIPS on Objaverse, GSO and OmniObject3D. Code, the 1B checkpoint and ~200K Objaverse-XL renders are open.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.04722
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 2

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.