CG-MLLM (v0.1)
CG-MLLM: Captioning and Generating 3D Content via Multi-modal Large Language Models (ICML 2026)
CG-MLLM is a 3D multimodal large language model built on a Qwen3-VL backbone and a Hunyuan3D-2.1 VAE, for 3D captioning and 3D generation.
This repository is the v0.1 multi-task checkpoint, on Qwen3-VL-2B-Instruct. It covers image-to-3D, text-to-3D, image understanding, and 3D understanding.
The 4B image-to-3D checkpoint (Qwen3-VL-4B, specialized on HY3D-Bench) is at JreamH/CGMLLM-4B-i2o.
Links
- Code (how to run): https://github.com/dreaming-huang/CG-MLLM
- Paper (arXiv): https://arxiv.org/abs/2601.21798
- Hugging Face Papers: https://huggingface.co/papers/2601.21798
- Project page: https://cv.jream.top/CG-MLLM-page/
- ICML poster: https://icml.cc/virtual/2026/poster/63909
Model Details
| Item | Value |
|---|---|
| Version | v0.1, multi-task |
| Vision-language backbone | Qwen3-VL-2B-Instruct |
| 3D latent tokenizer | Hunyuan3D-2.1 VAE |
| Tasks | Image-to-3D, text-to-3D, image understanding, 3D understanding |
| Checkpoint file | ema.safetensors |
| Venue | ICML 2026 |
| License | Apache-2.0 |
| Authors | Junming Huang, Chi Wang, Letian Li, Guangkai Xu, Donglin Huang, Hao Chen, Qiang Dai, Weiwei Xu |
| Affiliations | Zhejiang University; LIGHTSPEED |
Usage is documented in the GitHub repository. This card does not duplicate the inference commands.
Citation
If you find this work useful, please cite:
@article{huang2026cg,
title={CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models},
author={Huang, Junming and Wang, Chi and Li, Letian and Xu, Guangkai and Huang, Donglin and Chen, Hao and Dai, Qiang and Xu, Weiwei},
journal={arXiv preprint arXiv:2601.21798},
year={2026}
}
Acknowledgments
This work builds on Qwen3-VL, BAGEL, and the Hunyuan3D-2.1 shape VAE. Please follow their licenses when using those components.