DyadTalk: A Unified Pure Audio-Driven Framework for Dyadic 3D Talking Head Generation
Abstract
Dyadic face-to-face interaction is a fundamental form of human communication, where facial motions are continuously shaped by mutual engagement between two interlocutors rather than being produced independently. In many real-world applications, however, direct visual access to the conversation partner is often restricted or unavailable due to strict privacy requirements, making privacy-preserving audio-only generation especially desirable. Existing methods confront three main limitations: (1) single-speaker 3D talking head approaches generate isolated facial animations, failing to model mutual interactive dependencies; (2) current dyadic interaction frameworks predominantly rely on the partner's visual input as conditioning, rendering them difficult to apply in privacy-sensitive scenarios; and (3) asymmetric speaker-listener paradigms fall short of handling the dynamic turn-taking behavior in genuine bidirectional conversations. To tackle these challenges, we formally define the pure audio-driven dyadic 3D talking head generation task and propose DyadTalk, a unified framework requiring no visual cues during inference. The core of DyadTalk is the Hybrid Interaction Module (HIM) featuring a "Mix-then-Decouple" design. It effectively mixes dual speech streams to capture shared conversational dynamics while decoupling the representations to preserve speaker-specific motion patterns. Extensive experiments validate that DyadTalk consistently outperforms the audio-adapted baseline, DualTalk*. Specifically, it achieves a 43% reduction in expression Fréchet Distance (FD-EXP) and a 35% decrease in Mean Squared Error (MSE-EXP) on the in-domain test split. Furthermore, DyadTalk demonstrates strong out-of-domain robustness, reducing generalization errors by over 23% across key perceptual metrics. By successfully recovering realistic and coherent bidirectional motions solely from synchronized audio, our work provides a viable solution for privacy-friendly conversational avatar generation.
Qualitative comparison of pure audio-driven dyadic facial motion synthesis. Compared with DualTalk*, our method generates more expressive and interaction-aware facial motions for both interlocutors, with facial behaviors that more closely resemble the ground truth. The highlighted regions illustrate improved mouth articulation, expression realism, and cross-speaker responsiveness in multi-turn conversations.
Additional qualitative comparison among ground truth, our audio-only model, DualTalk*, and the original vision-dependent DualTalk. The selected examples show that our method produces more natural conversational facial motions than the direct audio-only baseline and narrows the quality gap to the vision-dependent setting.