Distilling Multi-view Diffusion Models into 3D Generators

Apr 3, 2025·
秦皓
秦皓
,
Luyuan Chen
,
Ming Kong
,
Mengxu Lu
,
Qiang Zhu
· 1 min read
Abstract
DD3G distills a multi-view diffusion model into a 3D Gaussian generator. It aligns teacher and student representation spaces, introduces a pattern extraction and progressive decoding generator, and produces 3D Gaussians from a single image in 0.06 seconds.
Type
Publication
IEEE Transactions on Multimedia
publications

DD3G transfers visual and spatial knowledge from a multi-view diffusion model into an efficient feed-forward 3D Gaussian generator.

秦皓
Authors
Ph.D. Student at Zhejiang University

I am a Ph.D. student in the College of Computer Science and Technology at Zhejiang University. My research focuses on spatial intelligence, 3D-AIGC, multi-agent systems, latent reasoning for VLMs, and contrastive learning, with a broader interest in building intelligent systems that connect perception, reasoning, and controllable creation in the world.

Email: haoqin@zju.edu.cn