
CAST3D is a training-free framework for transforming arbitrary 2D assets into coherent 3D objects or scenes under textual guidance. It targets controllable 3D composition and editing while preserving semantic and geometric consistency.

AGE is a multi-agent framework for continual and flexible 3D Gaussian editing. It decomposes complex user goals into modular editing skills, maintains structured memory for multi-turn interaction, supports backtracking, and grounds planning and reflection in 3D perceptual analysis.

This work proposes a diffusion-based method for 3D-aware image composition. Users specify an object's 3D bounding box, and the method generates high-fidelity composites guided by image, object identity, and depth constraints.

NGS-Marker provides native watermarking for 3D Gaussian Splatting assets. Instead of relying on rendered 2D images, it embeds and detects ownership signals in the 3D domain, improving robustness when protected assets are partially reused or transformed.

VF-Editor enables native editing of 3D Gaussian primitives across scenes and instructions. By modeling editable variation directly in the 3D representation, it aims to improve flexibility, efficiency, and cross-view consistency over indirect 2D-to-3D editing pipelines.

GGCN is a robust gait recognition model that uses a generate network, encoder network, and feature mapping network to reduce covariate interference and learn more discriminative gait representations.

DD3G distills a multi-view diffusion model into a 3D Gaussian generator. It aligns teacher and student representation spaces, introduces a pattern extraction and progressive decoding generator, and produces 3D Gaussians from a single image in 0.06 seconds.

This work studies semantic-aware positive sample generation for contrastive learning through multi-source and multi-modal prompt alignment. It uses large multimodal model capabilities to improve semantic consistency and sample diversity.

ProSL progressively optimizes pseudo-label generation in self-supervised contrastive learning for skeleton-based action recognition. It builds a semantic codebook from clustering and iteratively improves representation learning on multiple downstream tasks.

PRPose improves 3D human pose estimation by fitting the hidden probability distribution of the 2D-to-3D lifting process and using adaptive noise sampling to generate plausible multi-hypothesis 3D poses.