The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance

Zhang, Lin; Wang, Xin; Cooper, Erica; Evans, Nicholas; Yamagishi, Junichi

doi:10.1109/TASLP.2022.3233236

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2204.05177 (eess)

[Submitted on 11 Apr 2022 (v1), last revised 30 Jan 2023 (this version, v3)]

Title:The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance

Authors:Lin Zhang, Xin Wang, Erica Cooper, Nicholas Evans, Junichi Yamagishi

View PDF

Abstract:Automatic speaker verification is susceptible to various manipulations and spoofing, such as text-to-speech synthesis, voice conversion, replay, tampering, adversarial attacks, and so on. We consider a new spoofing scenario called "Partial Spoof" (PS) in which synthesized or transformed speech segments are embedded into a bona fide utterance. While existing countermeasures (CMs) can detect fully spoofed utterances, there is a need for their adaptation or extension to the PS scenario. We propose various improvements to construct a significantly more accurate CM that can detect and locate short-generated spoofed speech segments at finer temporal resolutions. First, we introduce newly developed self-supervised pre-trained models as enhanced feature extractors. Second, we extend our PartialSpoof database by adding segment labels for various temporal resolutions. Since the short spoofed speech segments to be embedded by attackers are of variable length, six different temporal resolutions are considered, ranging from as short as 20 ms to as large as 640 ms. Third, we propose a new CM that enables the simultaneous use of the segment-level labels at different temporal resolutions as well as utterance-level labels to execute utterance- and segment-level detection at the same time. We also show that the proposed CM is capable of detecting spoofing at the utterance level with low error rates in the PS scenario as well as in a related logical access (LA) scenario. The equal error rates of utterance-level detection on the PartialSpoof database and ASVspoof 2019 LA database were 0.77 and 0.90%, respectively.

Comments:	Published in IEEE/ACM Transactions on Audio, Speech, and Language Processing (DOI: https://doi.org/10.1109/TASLP.2022.3233236)
Subjects:	Audio and Speech Processing (eess.AS); Cryptography and Security (cs.CR); Sound (cs.SD)
Cite as:	arXiv:2204.05177 [eess.AS]
	(or arXiv:2204.05177v3 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2204.05177
Journal reference:	IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 813-825, 2023
Related DOI:	https://doi.org/10.1109/TASLP.2022.3233236

Submission history

From: Lin Zhang [view email]
[v1] Mon, 11 Apr 2022 15:09:07 UTC (2,198 KB)
[v2] Thu, 28 Apr 2022 11:40:35 UTC (2,181 KB)
[v3] Mon, 30 Jan 2023 10:39:53 UTC (2,214 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators