Cross-Lingual Adaptation of Phoneme--Motion Memory for Turkish Talking Head Generation
Hüseyin Temiz, Berk Gökberk, Lale Akarun
2026 34th Signal Processing and Communications Applications Conference (SIU)
Abstract
Speech-driven 3D facial animation is largely developed on English datasets, leaving cross-lingual generalization underexplored. We study the transferability of memory-based motion representations to Turkish using MemoryTalker, a two-stage architecture with a Facial Motion Memory and audio-guided stylization. The model is trained on VOCASET and evaluated zero-shot on TDV, a Turkish 3D facial motion dataset. Results reveal a significant cross-lingual gap, especially in lip articulation, while fine-tuning with limited Turkish data effectively adapts the motion memory, achieving near in-domain performance without sacrificing temporal consistency. These findings show that memory-based motion representations enable efficient cross-lingual transfer and rapid adaptation to language-specific articulation.