Project Page

SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis

A semantic-aware framework for holistic co-speech gesture generation

1Wuhan University  ·   2Tongyi Lab, Alibaba Group  ·   3Zhejiang University
ICCV 2025
SemTalk teaser
Why SemTalk is needed. Co-speech motion is not only rhythm. Most frames carry ordinary beat-aligned movement, but a few frames carry semantic emphasis that makes a gesture feel intentional. SemTalk separates these two sources of motion and fuses them frame by frame, so the generated speaker can stay rhythmically stable while still producing sparse, meaningful gestures.

Abstract

Co-speech gesture generation must balance frequent rhythm-aligned movements with sparse but essential semantic gestures. We present SemTalk, a holistic co-speech gesture generation framework with frame-level semantic emphasis. Our key insight is to separately learn general motions and sparse motions, and then adaptively fuse them. In particular, rhythmic consistency learning is explored to establish rhythm-related base motion, ensuring a coherent foundation that synchronizes gestures with the speech rhythm. Subsequently, semantic emphasis learning is designed to generate semantic-aware sparse motion, focusing on frame-level semantic cues. Finally, to integrate sparse motion into the base motion and generate semantic-emphasized co-speech gestures, we further leverage a learned semantic score for adaptive synthesis. Qualitative and quantitative comparisons on two public datasets demonstrate that our method outperforms the state-of-the-art, delivering high-quality co-speech motion with enhanced semantic richness over a stable base motion. The code is available at https://github.com/Xiangyue-Zhang/SemTalk.

Demo

Method

SemTalk uses a two-stream design: one stream learns a stable rhythm-aligned base, while another stream activates sparse semantic motion only when speech calls for emphasis.

SemTalk framework
Architecture of SemTalk. (a) Base Motion Generation uses rhythmic consistency learning to produce rhythm-aligned codes \( q^r \), conditioned on rhythmic features \( \gamma_b \), \( \gamma_h \), seed pose \( \tilde{m} \), and \( id \). (b) Sparse Motion Generation employs semantic emphasis learning to generate semantic codes \( q^s \), activated by semantic score \( \psi \). (c) Adaptively Fusion automatically combines \( q^r \) and \( q^s \) based on \( \psi \) to produce mixed codes \( q^m \) at frame level for rhythmically aligned and contextually rich motions.
Two streams Separate rhythmic base motion from sparse semantic gestures.
Frame-level Fuse base and semantic codes according to a learned semantic score.
BEAT2 + SHOW Evaluated on both in-domain and cross-dataset holistic motion generation.

Semantic Emphasis

The central design choice is not to make every frame more expressive. Instead, SemTalk learns when the speech contains a phrase that should trigger stronger, sparse gestures.

Semantic score visualization
Semantic score. The learned score highlights frame ranges where speech semantics require visible emphasis. These frames receive sparse semantic motion codes, while ordinary frames keep the rhythm-aligned base motion. This avoids the common failure mode where generated motion is either too flat everywhere or overly active everywhere.

Qualitative Results

The visual comparisons show what the semantic branch changes: gestures become more intentional at meaningful words while staying stable during ordinary speech.

Comparison on BEAT2
Comparison on BEAT2 Dataset. SemTalk* refers to the model trained solely on the Base Motion Generation stage. In contrast, SemTalk successfully emphasizes sparse yet vivid motions. For example, when the phrase "my opinion" is spoken, SemTalk-driven characters raise both hands and make the gesture of extending their index finger to emphasize the statement.
Comparison on SHOW
Comparison on SHOW Dataset. SemTalk shows more agile gestures than TalkSHOW, EMAGE, and DiffSHEG, when applied to unseen data. Our method captures natural and contextually rich gestures, particularly in moments of emphasis such as "I like to do" and "relaxing."
Facial comparison on BEAT2
Facial Comparison on BEAT2. Our approach synchronizes facial expressions closely with phonetic and semantic cues in speech, generating natural lip movements that enhance clarity and expressiveness.
SemTalk user study
User study. Human preference results support the same conclusion as the qualitative figures: separating rhythm and semantic emphasis improves perceived realism and semantic match rather than only optimizing numeric metrics.

Quantitative Results

Lower is better for FGD, MSE, and LVD. BC and DIV follow the benchmark convention shown in each table. SemTalk now provides results for the original Speaker 2 protocol and a released 25-speaker BEAT2 checkpoint.

BEAT2 1-Speaker Setting (Speaker 2 Paper Protocol)
Method Venue FGD lower BC higher DIV higher MSE lower LVD lower
CaMN ECCV 2022 6.644 6.769 10.86 - -
DSG IJCAI 2023 8.811 7.241 11.49 - -
TalkSHOW CVPR 2023 6.209 6.947 13.47 7.791 7.771
EMAGE CVPR 2024 5.512 7.724 13.06 7.680 7.556
DiffSHEG CVPR 2024 8.986 7.142 11.91 7.665 8.673
SemTalk ICCV 2025 4.278 7.770 12.91 6.153 6.938
EchoMask ACM MM 2025 4.623 7.738 13.37 6.761 7.290
GlobalDiff AAAI 2026 4.780 7.050 13.73 6.330
StreamTalk ECCV 2026 3.830 7.040 13.18
SHOW
Method FGD lower BC higher DIV higher MSE lower LVD lower
CaMN 22.12 7.712 10.37 - -
DSG 24.84 8.027 10.23 - -
Habibie et al. 27.22 8.209 8.541 145.6 47.35
TalkSHOW 24.43 8.249 10.98 139.6 45.17
EMAGE 22.12 8.280 12.46 136.1 42.44
DiffSHEG 24.87 8.061 10.79 139.0 45.77
SemTalk 20.18 8.304 11.36 134.1 39.15
BEAT2 All-Speakers Setting (25 English Speakers)
Method Venue FGD ↓ BeatAlign / BC (near GT) Diversity / DIV (near GT) MSE ↓ (×10-8) LVD ↓ (×10-5)
GT 0.477 7.29
CaMN ECCV 2022 0.512 0.200 5.58
Audio2Photoreal CVPR 2024 0.849 0.326 6.24
ReMoDiffuse ICCV 2023 1.120 0.218 5.06
DSG IJCAI 2023 1.174 0.734 11.12
HoloGest 3DV 2025 0.646 0.803 13.53
EMAGE CVPR 2024 0.692 0.284 6.06 6.908
RAG-GESTURE CVPR 2025 0.487 0.514 9.94
SemTalk ICCV 2025 0.356 0.510 8.409 4.439 1.435
EchoMask ACM MM 2025 0.566 0.495 9.299 4.700 6.090
GlobalDiff AAAI 2026 0.263 0.404 8.24 4.144
StreamTalk ECCV 2026 0.293 0.616 7.27
The four related project rows use a shared raw-value scale across all project pages; unavailable metrics are marked “—”. The SemTalk row is from the released seed-43 Sparse checkpoint at epoch 246. Speaker 2 and All-Speakers use different training protocols and should not be interpreted as a controlled single-speaker versus multi-speaker ablation.

BibTeX

@inproceedings{zhang2025semtalk,
  title={SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis},
  author={Zhang, Xiangyue and Li, Jianfang and Zhang, Jiaxu and Dang, Ziqiang and Ren, Jianqiang and Bo, Liefeng and Tu, Zhigang},
  booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
  pages={13761--13771},
  year={2025}
}