XTalker

Diffusion-Based Talking-Head Avatar Generation in Motion Space

XTalker generates talking-head videos from a single source image and an audio clip. It predicts compact speech-driven motion representations and renders stable, identity-preserving video frames through a LivePortrait-style pipeline.

Paper submission in progress. Content is being continuously organized and optimized.

Code: https://github.com/EngineeringAI-LAB/xtalker