ECHO-GEmbodied Co-speech Humanoid mOtion Generation

Full-body robot motion from speech audio and timed transcripts.

Main video

Real-robot demonstrations

15 trials

Conditioning comparison

Shared audio

0:00 / 0:15

Quantitative results

BEAT2 speaker-held-out evaluation

Bold bestUnderline second-best↓ lower is better   ·   ↑ higher is better

Architecture

Abstract

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework that jointly conditions full-body robot-motion generation on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) integrates frame-aligned acoustic features into the motion stream and retrieves token-level linguistic context through global and temporally biased cross-attention. Trained with rectified flow matching, SGDiT models the one-to-many relationship between utterances and accompanying gestures directly in robot space. To support training and evaluation, we introduce a BEAT2-derived dataset pairing audio and timed transcripts with robot motion, together with a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio–text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating user study also favors joint conditioning over the compared configurations.

Project overview

Open original file