Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?
Abstract
Native 3D texture generation synthesizes colors directly in 3D space for a given geometry, conditioned on multi-view reference images. It is generally believed that training such models requires large-scale, high-quality real 3D asset data, whose acquisition remains a long-standing and challenging problem. In this work, we propose Tex-Zero, demonstrating that a high-fidelity native 3D texture generation framework can be trained without 3D assets. Our key observation is that only high-quality and fine-grained color information is essential for 3D texture training, while the required geometric information is less critical and can be manually constructed rather than obtained from real 3D assets. This finding makes it possible to transform abundant, high-quality 2D images into effective training samples for 3D texture generation. Specifically, we convert high-quality 2D images into 3D training samples by representing each image as a plane in 3D space and applying patch-wise random rotations and aggregation to construct complex geometric structures. Using these constructed image data, we train the Tex-Zero VAE, which can reconstruct real 3D assets with high quality despite never observing them during training. Building upon the Tex-Zero VAE, we train the Tex-Zero DiT also exclusively on the constructed image data, where the conditioning 2D multi-view images are transformed into planes in 3D space and also encoded by the Tex-Zero VAE, thereby reducing the representation gap and improving generation quality. Extensive experiments show that Tex-Zero generates high-fidelity 3D textures with fine-grained details solely using images as training data, offering a promising perspective on the data paradigm for scaling 3D texture generation.
Community
🚀 We introduce Tex-Zero, a native 3D texture generation framework trained entirely without real textured 3D assets.
Our key finding is that high-quality, fine-grained color information matters more than real 3D geometry for texture learning. Tex-Zero turns abundant 2D images into effective 3D training samples by mapping image patches to planes, randomly rotating them, and aggregating them into complex geometric structures. A VAE rained only on this constructed image data can reconstruct both 2D image and 3D assets with high quality. DiT trained on image data can generate high-fidelity textures for real 3D assets at inference time.
This offers a new data-scaling path for native 3D texture generation: leveraging high-quality 2D imagery instead of expensive large-scale textured 3D datasets.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding (2026)
- TaoTex: Boosting Texture Detail Fidelity for Native 3D Material Generation (2026)
- Fusion-Aware Direct 3D Gaussian Generation with Structured Patch Latent Flows (2026)
- Luce: Relightable Gaussians for 3D Asset Generation (2026)
- UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing (2026)
- DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion (2026)
- GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.34621 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper