A semantic-aware framework for holistic co-speech gesture generation
Co-speech gesture generation must balance frequent rhythm-aligned movements with sparse but essential semantic gestures. We present SemTalk, a holistic co-speech gesture generation framework with frame-level semantic emphasis. Our key insight is to separately learn general motions and sparse motions, and then adaptively fuse them. In particular, rhythmic consistency learning is explored to establish rhythm-related base motion, ensuring a coherent foundation that synchronizes gestures with the speech rhythm. Subsequently, semantic emphasis learning is designed to generate semantic-aware sparse motion, focusing on frame-level semantic cues. Finally, to integrate sparse motion into the base motion and generate semantic-emphasized co-speech gestures, we further leverage a learned semantic score for adaptive synthesis. Qualitative and quantitative comparisons on two public datasets demonstrate that our method outperforms the state-of-the-art, delivering high-quality co-speech motion with enhanced semantic richness over a stable base motion. The code is available at https://github.com/Xiangyue-Zhang/SemTalk.
SemTalk uses a two-stream design: one stream learns a stable rhythm-aligned base, while another stream activates sparse semantic motion only when speech calls for emphasis.
The central design choice is not to make every frame more expressive. Instead, SemTalk learns when the speech contains a phrase that should trigger stronger, sparse gestures.
The visual comparisons show what the semantic branch changes: gestures become more intentional at meaningful words while staying stable during ordinary speech.
Lower is better for FGD, MSE, and LVD. BC and DIV follow the benchmark convention shown in each table. SemTalk now provides results for the original Speaker 2 protocol and a released 25-speaker BEAT2 checkpoint.
| Method | Venue | FGD lower | BC higher | DIV higher | MSE lower | LVD lower |
|---|---|---|---|---|---|---|
| CaMN | ECCV 2022 | 6.644 | 6.769 | 10.86 | - | - |
| DSG | IJCAI 2023 | 8.811 | 7.241 | 11.49 | - | - |
| TalkSHOW | CVPR 2023 | 6.209 | 6.947 | 13.47 | 7.791 | 7.771 |
| EMAGE | CVPR 2024 | 5.512 | 7.724 | 13.06 | 7.680 | 7.556 |
| DiffSHEG | CVPR 2024 | 8.986 | 7.142 | 11.91 | 7.665 | 8.673 |
| SemTalk ↗ | ICCV 2025 | 4.278 | 7.770 | 12.91 | 6.153 | 6.938 |
| EchoMask ↗ | ACM MM 2025 | 4.623 | 7.738 | 13.37 | 6.761 | 7.290 |
| GlobalDiff ↗ | AAAI 2026 | 4.780 | 7.050 | 13.73 | 6.330 | — |
| StreamTalk ↗ | ECCV 2026 | 3.830 | 7.040 | 13.18 | — | — |
| Method | FGD lower | BC higher | DIV higher | MSE lower | LVD lower |
|---|---|---|---|---|---|
| CaMN | 22.12 | 7.712 | 10.37 | - | - |
| DSG | 24.84 | 8.027 | 10.23 | - | - |
| Habibie et al. | 27.22 | 8.209 | 8.541 | 145.6 | 47.35 |
| TalkSHOW | 24.43 | 8.249 | 10.98 | 139.6 | 45.17 |
| EMAGE | 22.12 | 8.280 | 12.46 | 136.1 | 42.44 |
| DiffSHEG | 24.87 | 8.061 | 10.79 | 139.0 | 45.77 |
| SemTalk | 20.18 | 8.304 | 11.36 | 134.1 | 39.15 |
| Method | Venue | FGD ↓ | BeatAlign / BC (near GT) | Diversity / DIV (near GT) | MSE ↓ (×10-8) | LVD ↓ (×10-5) |
|---|---|---|---|---|---|---|
| GT | — | — | 0.477 | 7.29 | — | — |
| CaMN | ECCV 2022 | 0.512 | 0.200 | 5.58 | — | — |
| Audio2Photoreal | CVPR 2024 | 0.849 | 0.326 | 6.24 | — | — |
| ReMoDiffuse | ICCV 2023 | 1.120 | 0.218 | 5.06 | — | — |
| DSG | IJCAI 2023 | 1.174 | 0.734 | 11.12 | — | — |
| HoloGest | 3DV 2025 | 0.646 | 0.803 | 13.53 | — | — |
| EMAGE | CVPR 2024 | 0.692 | 0.284 | 6.06 | 6.908 | — |
| RAG-GESTURE | CVPR 2025 | 0.487 | 0.514 | 9.94 | — | — |
| SemTalk ↗ | ICCV 2025 | 0.356 | 0.510 | 8.409 | 4.439 | 1.435 |
| EchoMask ↗ | ACM MM 2025 | 0.566 | 0.495 | 9.299 | 4.700 | 6.090 |
| GlobalDiff ↗ | AAAI 2026 | 0.263 | 0.404 | 8.24 | 4.144 | — |
| StreamTalk ↗ | ECCV 2026 | 0.293 | 0.616 | 7.27 | — | — |
@inproceedings{zhang2025semtalk,
title={SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis},
author={Zhang, Xiangyue and Li, Jianfang and Zhang, Jiaxu and Dang, Ziqiang and Ren, Jianqiang and Bo, Liefeng and Tu, Zhigang},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
pages={13761--13771},
year={2025}
}