Academic Journal

CoShMDM: Contact and Shape-Aware Latent Motion Diffusion Model for Human Interaction Generation.

Bibliographic Details
Title: CoShMDM: Contact and Shape-Aware Latent Motion Diffusion Model for Human Interaction Generation.
Authors: Manjotho AA, Tewolde TT, Duma RA, Niu Z
Source: IEEE transactions on visualization and computer graphics [IEEE Trans Vis Comput Graph] 2026 Jul; Vol. 32 (7), pp. 5911-5924.
Publication Type: Journal Article
Language: English
Journal Info: Publisher: IEEE Computer Society Country of Publication: United States NLM ID: 9891704 Publication Model: Print Cited Medium: Internet ISSN: 1941-0506 (Electronic) Linking ISSN: 10772626 NLM ISO Abbreviation: IEEE Trans Vis Comput Graph Subsets: MEDLINE
Imprint Name(s): Original Publication: New York, NY : IEEE Computer Society, c1995-
MeSH Terms: Image Processing, Computer-Assisted*/methods , Computer Graphics*, Humans ; Algorithms ; Motion ; Reinforcement Machine Learning
Abstract: Generating realistic two-person interaction motions from text holds immense potential in computer vision and animations. While existing latent motion diffusion models offer compact and efficient representations, they often fail to produce physically plausible contacts and are typically constrained to a single canonical body shape. As a result, the generated motion sequences exhibit substantial mesh penetrations and lack interaction realism. To address these limitations, we propose a contact and shape-aware latent motion representation and diffusion model (CoShMDM) for generating realistic two-person interactions from text. Our framework begins by constructing contact-compatible motion using SMPL-based meshes and a normal alignment-based mesh contact matrix to capture fine-grained mesh-level contacts. To account for shape diversity, we incorporate SMPL shape parameters and iteratively learn contact dynamics across different body shapes. Additionally, a reinforcement learning-based mesh penetration avoidance policy network, guided by signed distance fields, is introduced to minimize mesh penetrations while preserving contact fidelity and shape-aware motion. We further employ a dual-encoder VQ-VAE to learn disentangled latent representations for motion and contacts, which are then utilized in a text- and body-shape-conditioned diffusion model. To ensure spatial, temporal, and semantic coherence, we integrate a novel contact and motion consistency module into the diffusion transformer. Extensive evaluations on the InterHuman and InterX datasets demonstrate that our method outperforms state-of-the-art approaches achieving lowest FID scores (4.801 and 0.013), with 19% and 17.3% reductions in mesh penetrations, and 17.8% and 33.2% gains in contact similarity, respectively.
Entry Date(s): Date Created: 20260319 Date Completed: 20260623 Latest Revision: 20260624
Update Code: 20260624
DOI: 10.1109/TVCG.2026.3675725
PMID: 41855059
Database: MEDLINE
Description
ISSN:1941-0506
DOI:10.1109/TVCG.2026.3675725