Speech-guided masked modeling for holistic co-speech gesture generation
Co-speech gesture generation benefits from masked modeling, but existing strategies struggle to identify semantically significant frames for effective motion masking. We propose EchoMask, a speech-queried attention-based mask modeling framework for holistic co-speech gesture generation. Our key insight is to leverage motion-aligned speech features to guide the masked motion modeling process, selectively masking rhythm-related and semantically expressive motion frames. Specifically, we first propose a motion-audio alignment module (MAM) to construct a latent motion-audio joint space. In this space, both low-level and high-level speech features are projected, enabling motion-aligned speech representation using learnable speech queries. Then, a speech-queried attention mechanism (SQA) is introduced to compute frame-level attention scores through interactions between motion keys and speech queries, guiding selective masking toward motion frames with high attention scores. Finally, the motion-aligned speech features are also injected into the generation network to facilitate co-speech motion generation. Qualitative and quantitative evaluations confirm that our method outperforms existing state-of-the-art approaches, successfully producing high-quality co-speech motion. The code is available at https://github.com/Xiangyue-Zhang/EchoMask.
The main question is where masking should happen. Random masking is blind, and loss-based masking can over-focus on hard but uninformative transitions. EchoMask uses speech as the query signal.
EchoMask first aligns audio and motion in a shared latent space, then uses speech-queried attention to decide which motion frames should be masked and reconstructed.
The examples show that speech-guided masking improves both body gestures and facial articulation, especially where semantic cues should trigger clearer motion.
Lower is better for FGD, MSE, and LVD. BC and DIV follow the benchmark convention shown in each table. EchoMask now provides results for the original Speaker 2 protocol and a released 25-speaker BEAT2 checkpoint.
| Setting | Method | Venue | FGD lower | BC higher | DIV higher | MSE lower | LVD lower |
|---|---|---|---|---|---|---|---|
| Facial | FaceFormer | CVPR 2022 | - | - | - | 7.787 | 7.593 |
| Facial | CodeTalker | CVPR 2023 | - | - | - | 8.026 | 7.766 |
| Non-facial | DisCo | ACM MM 2022 | 9.680 | 6.441 | 9.892 | - | - |
| Non-facial | HA2G | CVPR 2022 | 12.14 | 6.711 | 8.916 | - | - |
| Non-facial | CaMN | ECCV 2022 | 6.644 | 6.769 | 10.86 | - | - |
| Non-facial | LivelySpeaker | ICCV 2023 | 11.80 | 6.659 | 11.28 | - | - |
| Non-facial | DSG | IJCAI 2023 | 8.811 | 7.241 | 11.49 | - | - |
| Holistic | Habibie et al. | IVA 2021 | 9.040 | 7.716 | 8.213 | 8.614 | 8.043 |
| Holistic | TalkSHOW | CVPR 2023 | 6.209 | 6.947 | 13.47 | 7.791 | 7.771 |
| Holistic | EMAGE | CVPR 2024 | 5.512 | 7.724 | 13.06 | 7.680 | 7.556 |
| Holistic | DiffSHEG | CVPR 2024 | 8.986 | 7.142 | 11.91 | 7.665 | 8.673 |
| Holistic | EchoMask ↗ | ACM MM 2025 | 4.623 | 7.738 | 13.37 | 6.761 | 7.290 |
| Holistic | SemTalk ↗ | ICCV 2025 | 4.278 | 7.770 | 12.91 | 6.153 | 6.938 |
| Holistic | GlobalDiff ↗ | AAAI 2026 | 4.780 | 7.050 | 13.73 | 6.330 | — |
| Holistic | StreamTalk ↗ | ECCV 2026 | 3.830 | 7.040 | 13.18 | — | — |
| Method | Venue | FGD ↓ | BeatAlign / BC (near GT) | Diversity / DIV (near GT) | MSE ↓ (×10-8) | LVD ↓ (×10-5) |
|---|---|---|---|---|---|---|
| GT | — | — | 0.477 | 7.29 | — | — |
| CaMN | ECCV 2022 | 0.512 | 0.200 | 5.58 | — | — |
| Audio2Photoreal | CVPR 2024 | 0.849 | 0.326 | 6.24 | — | — |
| ReMoDiffuse | ICCV 2023 | 1.120 | 0.218 | 5.06 | — | — |
| DSG | IJCAI 2023 | 1.174 | 0.734 | 11.12 | — | — |
| HoloGest | 3DV 2025 | 0.646 | 0.803 | 13.53 | — | — |
| EMAGE | CVPR 2024 | 0.692 | 0.284 | 6.06 | 6.908 | — |
| RAG-GESTURE | CVPR 2025 | 0.487 | 0.514 | 9.94 | — | — |
| SemTalk ↗ | ICCV 2025 | 0.356 | 0.510 | 8.409 | 4.439 | 1.435 |
| EchoMask ↗ | ACM MM 2025 | 0.566 | 0.495 | 9.299 | 4.700 | 6.090 |
| GlobalDiff ↗ | AAAI 2026 | 0.263 | 0.404 | 8.24 | 4.144 | — |
| StreamTalk ↗ | ECCV 2026 | 0.293 | 0.616 | 7.27 | — | — |
@inproceedings{zhang2025echomask,
title={EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation},
author={Zhang, Xiangyue and Li, Jianfang and Zhang, Jiaxu and Ren, Jianqiang and Bo, Liefeng and Tu, Zhigang},
booktitle={Proceedings of the 33rd ACM International Conference on Multimedia},
pages={10827--10836},
year={2025}
}