ECHO-GEmbodied Co-speech Humanoid mOtion Generation
Full-body robot motion from speech audio and timed transcripts.
Main video
Real-robot demonstrations
Conditioning comparison
Shared audio
Quantitative results
BEAT2 speaker-held-out evaluation
Architecture
Abstract
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework that jointly conditions full-body robot-motion generation on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) integrates frame-aligned acoustic features into the motion stream and retrieves token-level linguistic context through global and temporally biased cross-attention. Trained with rectified flow matching, SGDiT models the one-to-many relationship between utterances and accompanying gestures directly in robot space. To support training and evaluation, we introduce a BEAT2-derived dataset pairing audio and timed transcripts with robot motion, together with a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio–text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating user study also favors joint conditioning over the compared configurations.