Project Page

EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation

Speech-guided masked modeling for holistic co-speech gesture generation

1Wuhan University  ·   2Tongyi Lab, Alibaba Group
ACM MM 2025
EchoMask teaser
Why EchoMask is needed. Masked motion modeling is powerful, but random masking does not know which gesture frames matter. EchoMask lets speech query the motion sequence, so training focuses on frames that carry semantic or rhythm information instead of treating all frames equally.

Abstract

Co-speech gesture generation benefits from masked modeling, but existing strategies struggle to identify semantically significant frames for effective motion masking. We propose EchoMask, a speech-queried attention-based mask modeling framework for holistic co-speech gesture generation. Our key insight is to leverage motion-aligned speech features to guide the masked motion modeling process, selectively masking rhythm-related and semantically expressive motion frames. Specifically, we first propose a motion-audio alignment module (MAM) to construct a latent motion-audio joint space. In this space, both low-level and high-level speech features are projected, enabling motion-aligned speech representation using learnable speech queries. Then, a speech-queried attention mechanism (SQA) is introduced to compute frame-level attention scores through interactions between motion keys and speech queries, guiding selective masking toward motion frames with high attention scores. Finally, the motion-aligned speech features are also injected into the generation network to facilitate co-speech motion generation. Qualitative and quantitative evaluations confirm that our method outperforms existing state-of-the-art approaches, successfully producing high-quality co-speech motion. The code is available at https://github.com/Xiangyue-Zhang/EchoMask.

Demo

Masking Problem

The main question is where masking should happen. Random masking is blind, and loss-based masking can over-focus on hard but uninformative transitions. EchoMask uses speech as the query signal.

Comparison of masking strategies
Existing methods predominantly adopt random (a) or loss-based (b) masking strategies. Random masking often fails to target semantically meaningful regions. Loss-based masking prioritizes frames with high reconstruction error, but high loss may simply reflect abrupt yet uninformative transitions. Our EchoMask (c) uses speech-queried attention to identify semantically important frames.

Method

EchoMask first aligns audio and motion in a shared latent space, then uses speech-queried attention to decide which motion frames should be masked and reconstructed.

EchoMask framework
Architecture of EchoMask. (a) MAM projects motion and audio into a shared latent space. Learnable speech queries \(Q'\) are refined through hierarchical cross-attention with HuBERT features (\( \gamma_l, \gamma_h \)) and jointly processed with quantized latent motion \( \tilde{z}_m \) via a shared transformer, optimized with contrastive loss. (b) Given \( m \), mask transformer teacher computes a cross-attention map \( \mathcal{M} \) between latent poses \(p\) and motion-aligned speech features \( Q \), identifying semantically important frames. These frames are masked via a Soft2Hard strategy to produce \( \tilde{m} \), which the student transformer uses to generate motion tokens.
Speech queries Use speech features to find motion frames worth reconstructing.
MAM Build a shared motion-audio latent space before masking.
Holistic Evaluate body and facial motion in a unified co-speech setting.

Qualitative Results

The examples show that speech-guided masking improves both body gestures and facial articulation, especially where semantic cues should trigger clearer motion.

Body comparison on BEAT2
Comparison on BEAT2 Dataset. Red boxes highlight implausible or uncoordinated motions, while green boxes indicate coherent and semantically appropriate results. Our EchoMask consistently generates co-speech motions that are semantically aligned with ground truth.
Facial comparison on BEAT2
Facial Comparison on BEAT2. Our approach tightly synchronizes facial expressions with both phonetic and semantic cues in speech, producing natural and articulate lip movements.

Quantitative Results

Lower is better for FGD, MSE, and LVD. BC and DIV follow the benchmark convention shown in each table. EchoMask now provides results for the original Speaker 2 protocol and a released 25-speaker BEAT2 checkpoint.

BEAT2 1-Speaker Setting (Speaker 2 Paper Protocol)
Setting Method Venue FGD lower BC higher DIV higher MSE lower LVD lower
Facial FaceFormer CVPR 2022 - - - 7.787 7.593
Facial CodeTalker CVPR 2023 - - - 8.026 7.766
Non-facial DisCo ACM MM 2022 9.680 6.441 9.892 - -
Non-facial HA2G CVPR 2022 12.14 6.711 8.916 - -
Non-facial CaMN ECCV 2022 6.644 6.769 10.86 - -
Non-facial LivelySpeaker ICCV 2023 11.80 6.659 11.28 - -
Non-facial DSG IJCAI 2023 8.811 7.241 11.49 - -
Holistic Habibie et al. IVA 2021 9.040 7.716 8.213 8.614 8.043
Holistic TalkSHOW CVPR 2023 6.209 6.947 13.47 7.791 7.771
Holistic EMAGE CVPR 2024 5.512 7.724 13.06 7.680 7.556
Holistic DiffSHEG CVPR 2024 8.986 7.142 11.91 7.665 8.673
Holistic EchoMask ↗ ACM MM 2025 4.623 7.738 13.37 6.761 7.290
Holistic SemTalk ↗ ICCV 2025 4.278 7.770 12.91 6.153 6.938
Holistic GlobalDiff ↗ AAAI 2026 4.780 7.050 13.73 6.330 —
Holistic StreamTalk ↗ ECCV 2026 3.830 7.040 13.18 — —
BEAT2 All-Speakers Setting (25 English Speakers)
Method Venue FGD ↓ BeatAlign / BC (near GT) Diversity / DIV (near GT) MSE ↓ (×10-8) LVD ↓ (×10-5)
GT — — 0.477 7.29 — —
CaMN ECCV 2022 0.512 0.200 5.58 — —
Audio2Photoreal CVPR 2024 0.849 0.326 6.24 — —
ReMoDiffuse ICCV 2023 1.120 0.218 5.06 — —
DSG IJCAI 2023 1.174 0.734 11.12 — —
HoloGest 3DV 2025 0.646 0.803 13.53 — —
EMAGE CVPR 2024 0.692 0.284 6.06 6.908 —
RAG-GESTURE CVPR 2025 0.487 0.514 9.94 — —
SemTalk ↗ ICCV 2025 0.356 0.510 8.409 4.439 1.435
EchoMask ↗ ACM MM 2025 0.566 0.495 9.299 4.700 6.090
GlobalDiff ↗ AAAI 2026 0.263 0.404 8.24 4.144 —
StreamTalk ↗ ECCV 2026 0.293 0.616 7.27 — —
The four related project rows use a shared raw-value scale across all project pages; unavailable metrics are marked “—”. The EchoMask row is from the released seed-44 checkpoint at epoch 380. Speaker 2 and All-Speakers use different training protocols and should not be interpreted as a controlled single-speaker versus multi-speaker ablation.

BibTeX

@inproceedings{zhang2025echomask,
  title={EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation},
  author={Zhang, Xiangyue and Li, Jianfang and Zhang, Jiaxu and Ren, Jianqiang and Bo, Liefeng and Tu, Zhigang},
  booktitle={Proceedings of the 33rd ACM International Conference on Multimedia},
  pages={10827--10836},
  year={2025}
}