Academic Journal

EmoPoseFace: Head Pose Aware Speech-Driven 3D Emotional Facial Animation Using Latent Diffusion.

Bibliographic Details
Title: EmoPoseFace: Head Pose Aware Speech-Driven 3D Emotional Facial Animation Using Latent Diffusion.
Authors: Zhao X, Dai J, Zhou F, Wang H, Song Z, Hao A, Qin H, Gao Y
Source: IEEE transactions on visualization and computer graphics [IEEE Trans Vis Comput Graph] 2026 Aug; Vol. 32 (8), pp. 7631-7644.
Publication Type: Journal Article
Language: English
Journal Info: Publisher: IEEE Computer Society Country of Publication: United States NLM ID: 9891704 Publication Model: Print Cited Medium: Internet ISSN: 1941-0506 (Electronic) Linking ISSN: 10772626 NLM ISO Abbreviation: IEEE Trans Vis Comput Graph Subsets: MEDLINE
Imprint Name(s): Original Publication: New York, NY : IEEE Computer Society, c1995-
MeSH Terms: Emotions*/physiology , Emotions*/classification , Imaging, Three-Dimensional*/methods , Speech*/physiology , Facial Expression* , Computer Graphics*, Head/physiology ; Face/diagnostic imaging ; Humans ; Avatar
Abstract: Speech-driven 3D facial animation has notable applications in the VR domain, including virtual anchors and digital avatars, etc. However, producing facial animations that convey complex emotional expressions remains a substantial challenge. Existing methods struggle to simultaneously achieve accurate lip synchronization, natural facial expressions, and realistic emotional representation. Significantly, the impact of head pose on boosting facial emotional expressiveness has not been thoroughly investigated. To address these issues, we propose EmoPoseFace, a novel Diffusion-based network to generate speech-driven 3D emotional facial animations with synchronized head poses. Our method employs a dual-branch conditional generation architecture to separately model facial expressions and head poses, integrating emotion and head-pose conditions for coherent facial expression-pose control. In addition, we design the Global-local Facial Fine-grained Editing Module (GL-FFE), which achieves emotional enhancement of facial expressions and fine-grained facial modification, while maintains the naturalness and authenticity of facial movements. Extensive experiments demonstrate that our approach outperforms existing methods in lip-sync accuracy and emotional detail preservation. The introduction of head pose control and GL-FFE significantly expands the expressiveness of emotional virtual facial animation, and the fine-grained editing is widely approved in perceptual user studies.
Entry Date(s): Date Created: 20260610 Date Completed: 20260702 Latest Revision: 20260702
Update Code: 20260703
DOI: 10.1109/TVCG.2026.3702133
PMID: 42268773
Database: MEDLINE
Description
ISSN:1941-0506
DOI:10.1109/TVCG.2026.3702133