MoChat: Joints-Grouped Spatio-Temporal Grounding Multimodal Large Language Model for Multi-Turn Motion Comprehension and Description.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: MoChat: Joints-Grouped Spatio-Temporal Grounding Multimodal Large Language Model for Multi-Turn Motion Comprehension and Description.
Συγγραφείς: Mo J, Chen Y, Lin R, Ni Y, Liang F, Zeng M, Hu X, Li M
Πηγή: IEEE journal of biomedical and health informatics [IEEE J Biomed Health Inform] 2026 Mar; Vol. 30 (3), pp. 1972-1985.
Τύπος έκδοσης: Journal Article
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: Institute of Electrical and Electronics Engineers Country of Publication: United States NLM ID: 101604520 Publication Model: Print Cited Medium: Internet ISSN: 2168-2208 (Electronic) Linking ISSN: 21682194 NLM ISO Abbreviation: IEEE J Biomed Health Inform Subsets: MEDLINE
Imprint Name(s): Original Publication: New York, NY : Institute of Electrical and Electronics Engineers, 2013-
Ιατρικοί όροι (MeSH): Image Processing, Computer-Assisted*/methods , Deep Learning*, Movement/physiology ; Humans ; Large Language Models
Περίληψη: Despite continuous advancements in deep learning for understanding human motion, existing models often struggle to accurately identify action timing and specific body parts, typically supporting only single-round interaction. This limitation is particularly pronounced in home exercise monitoring, neurological disorder assessment, and rehabilitation, where precise motion analysis is crucial for ensuring exercise efficacy, detecting early signs of neurological conditions, and guiding personalized recovery programs. In this paper, we propose MoChat, a multimodal large language model capable of spatio-temporal grounding of human motion and multi-turn dialogue understanding. To achieve this, we first group spatial features in skeleton frames according to human anatomical structures and process them through a Joints-Grouped Skeleton Encoder. The encoder's outputs are fused with large language model embeddings to generate spatio-aware representations. A cross-attention-based Regression Head module is then designed to align hidden-layer embeddings and skeletal sequence embeddings, enabling precise temporal grounding. Furthermore, we develop a pipeline for temporal grounding task to extract timestamps from skeleton-text pairs and construct a multi-turn instruction dialogues for spatial grounding task. Finally, various task instructions are generated for jointly training. Experimental results demonstrate that MoChat achieves state-of-the-art performance across multiple metrics in motion understanding tasks, making it as the first model capable of fine-grained spatio-temporal grounding of human motion.
Entry Date(s): Date Created: 20251110 Date Completed: 20260306 Latest Revision: 20260309
Update Code: 20260309
DOI: 10.1109/JBHI.2025.3631045
PMID: 41212709
Βάση Δεδομένων: MEDLINE
Περιγραφή
ISSN:2168-2208
DOI:10.1109/JBHI.2025.3631045