RADIO: Reference-Agnostic Dubbing Video Synthesis

Lee, Dongyeun; Kim, Chaewon; Yu, Sangjoon; Yoo, Jaejun; Park, Gyeong-Moon

Computer Science > Computer Vision and Pattern Recognition

arXiv:2309.01950 (cs)

[Submitted on 5 Sep 2023 (v1), last revised 6 Nov 2023 (this version, v2)]

Title:RADIO: Reference-Agnostic Dubbing Video Synthesis

Authors:Dongyeun Lee, Chaewon Kim, Sangjoon Yu, Jaejun Yoo, Gyeong-Moon Park

View PDF

Abstract:One of the most challenging problems in audio-driven talking head generation is achieving high-fidelity detail while ensuring precise synchronization. Given only a single reference image, extracting meaningful identity attributes becomes even more challenging, often causing the network to mirror the facial and lip structures too closely. To address these issues, we introduce RADIO, a framework engineered to yield high-quality dubbed videos regardless of the pose or expression in reference images. The key is to modulate the decoder layers using latent space composed of audio and reference features. Additionally, we incorporate ViT blocks into the decoder to emphasize high-fidelity details, especially in the lip region. Our experimental results demonstrate that RADIO displays high synchronization without the loss of fidelity. Especially in harsh scenarios where the reference frame deviates significantly from the ground truth, our method outperforms state-of-the-art methods, highlighting its robustness.

Comments:	Accepted by WACV 2024
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2309.01950 [cs.CV]
	(or arXiv:2309.01950v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2309.01950

Submission history

From: Dongyeun Lee [view email]
[v1] Tue, 5 Sep 2023 04:56:18 UTC (6,359 KB)
[v2] Mon, 6 Nov 2023 05:06:54 UTC (6,361 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:RADIO: Reference-Agnostic Dubbing Video Synthesis

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:RADIO: Reference-Agnostic Dubbing Video Synthesis

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators