CG-MLLM (v0.1)

CG-MLLM: Captioning and Generating 3D Content via Multi-modal Large Language Models (ICML 2026)

CG-MLLM is a 3D multimodal large language model built on a Qwen3-VL backbone and a Hunyuan3D-2.1 VAE, for 3D captioning and 3D generation.

This repository is the v0.1 multi-task checkpoint, on Qwen3-VL-2B-Instruct. It covers image-to-3D, text-to-3D, image understanding, and 3D understanding.

The 4B image-to-3D checkpoint (Qwen3-VL-4B, specialized on HY3D-Bench) is at JreamH/CGMLLM-4B-i2o.

Links

Model Details

Item Value
Version v0.1, multi-task
Vision-language backbone Qwen3-VL-2B-Instruct
3D latent tokenizer Hunyuan3D-2.1 VAE
Tasks Image-to-3D, text-to-3D, image understanding, 3D understanding
Checkpoint file ema.safetensors
Venue ICML 2026
License Apache-2.0
Authors Junming Huang, Chi Wang, Letian Li, Guangkai Xu, Donglin Huang, Hao Chen, Qiang Dai, Weiwei Xu
Affiliations Zhejiang University; LIGHTSPEED

Usage is documented in the GitHub repository. This card does not duplicate the inference commands.

Citation

If you find this work useful, please cite:

@article{huang2026cg,
  title={CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models},
  author={Huang, Junming and Wang, Chi and Li, Letian and Xu, Guangkai and Huang, Donglin and Chen, Hao and Dai, Qiang and Xu, Weiwei},
  journal={arXiv preprint arXiv:2601.21798},
  year={2026}
}

Acknowledgments

This work builds on Qwen3-VL, BAGEL, and the Hunyuan3D-2.1 shape VAE. Please follow their licenses when using those components.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for JreamH/CGMLLM