Skip to content

← ArchivePaper2026

Conversational Gesture Model (CGM): Extending Speaker-Centric Audio-Driven Motion Generation to Full Conversation Gestures

Tomer Koren, Adi Rosenthal, Doron Friedman, Ariel Shamir

EurographicsAcademic0 citesMotion Synthesis

A cross attention model fuses interlocutor audio, text, and gesture cues so one character can generate both speaking and listening gestures in dialogue.

How to read this

Category
Academic Eurographics paper on co speech gesture generation, extending single speaker gesture synthesis to full two person conversation
Contributions
  • Introduces the Conversational Gesture Model, CGM, which generates gestures for a character in both speaking and listening roles within the same dialogue, rather than only when that character is talking
  • Uses a cross attention architecture to fuse the interlocutor's audio, text and gesture cues with the character's own gesture encoding, so the model conditions on what the other person is saying and doing
  • Extends prior speaker centric audio driven motion generation, which only modeled the person currently talking, to the bidirectional dynamics of a real conversation
Context
This directly extends the speaker centric audio driven co speech gesture line of research, its own framing broadens that prior work to cover listening behavior and interlocutor cues, not just speaking. The archive lists no builds_on entry, but the paper names the gap it closes plainly, existing models generate good speaking gestures but ignore how a listener reacts to the other person.
Correctness
This is a full Eurographics paper in Computer Graphics Forum, the cross attention design is a concrete architectural claim, but web search did not surface the paper's quantitative evaluation section. Treat the specific quality of the listening gestures as unconfirmed until the full paper or its video results are checked directly.
Clarity
Likely moderately technical, cross attention fusion of multiple modalities, audio, text, gesture, is standard motion synthesis architecture, readable for someone comfortable with recent gesture generation literature, less so for a pure production artist without that background.
How to read it
First pass, read the abstract for the core idea, one model, both speaking and listening gestures, driven by interlocutor cues. Second pass, once available, read the cross attention fusion architecture and how interlocutor features are represented and injected. Third pass is worth a full read for anyone building conversational or dialogue driven virtual characters where listening body language matters as much as speaking gestures.

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →