MODUS-Hunyuan

MODUS-Hunyuan applies the MODUS any-to-any adaptation recipe to the HunyuanImage-3.0 backbone. It supports independent prediction and chained generation across the same 16-modality interface as MODUS-BAGEL: text, caption, RGB, depth, normals, detection, segmentation, edges, and global or local visual features.

This release contains the EMA weights from 16,750 completed optimizer steps. It is exported as a self-contained BF16 inference checkpoint with its tokenizer, VAE, modality registry, MODUS adapter, and 40 weight shards. It does not contain optimizer, scheduler, data-loader, or training state.

Use

Inference code and examples are provided in the MODUS repository. The checkpoint is approximately 158 GiB and the released inference path was validated on four GPUs.

hf download EPFL-VILAB/MODUS-Hunyuan --local-dir models/modus-hunyuan

python infer.py --backend hunyuan --condition rgb --target depth \
  checkpoint_path=models/modus-hunyuan \
  input_image=test_images/01_basil_cathedral.jpg

Verification

The export was checked bitwise against the source EMA checkpoint: 5,198 model tensors, including 5 MODUS adapter tensors and 280 VAE tensors, passed. The three inference examples documented in the MODUS README were also executed successfully against this export.

License and restrictions

The model is a derivative of HunyuanImage-3.0 and is distributed under the Tencent Hunyuan Community License. That license includes territorial and use restrictions; read it before downloading, using, or redistributing the model. In particular, the licensed Territory excludes the European Union, United Kingdom, and South Korea.

This model is provided by the EPFL Visual Intelligence and Learning Lab. Tencent is not affiliated with, associated with, sponsoring, or endorsing MODUS or this release. See NOTICE for the required attribution.

Citation

@article{ye2026modus,
  title   = {MODUS: Decoder-only Any-to-Any Modeling of Diverse Modalities},
  author  = {Ye, Mingqiao and An, Zhaochong and Gao, Zhitong and Liu, Xian
             and Fleuret, Fran\c{c}ois and Li, Chuan and Zadeh, Amir
             and Belongie, Serge and Dehghan, Afshin and Allardice, Jesse
             and Mizrahi, David and Kar, O\u{g}uzhan Fatih and Bachmann, Roman
             and Zamir, Amir},
  journal = {arXiv preprint arXiv:2607.25948},
  year    = {2026},
}
Downloads last month
-
Safetensors
Model size
83B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EPFL-VILAB/MODUS-Hunyuan

Finetuned
(4)
this model

Paper for EPFL-VILAB/MODUS-Hunyuan