Uniphore's submission to Fearless Steps Challenge Phase-2

S, Karthik Pandia D; Spera, Cosimo

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2006.05747 (eess)

[Submitted on 10 Jun 2020]

Title:Uniphore's submission to Fearless Steps Challenge Phase-2

Authors:Karthik Pandia D S, Cosimo Spera

View PDF

Abstract:We propose supervised systems for speech activity detection (SAD) and speaker identification (SID) tasks in Fearless Steps Challenge Phase-2. The proposed systems for both the tasks share a common convolutional neural network (CNN) architecture. Mel spectrogram is used as features. For speech activity detection, the spectrogram is divided into smaller overlap** chunks. The network is trained to recognize the chunks. The network architecture and the training steps used for the SID task are similar to that of the SAD task, except that longer spectrogram chunks are used. We propose a two-level identification method for SID task. First, for each chunk, a set of speakers is hypothesized based on the neural network posterior probabilities. Finally, the speaker identity of the utterance is identified using the chunk-level hypotheses by applying a voting rule. On SAD task, a detection cost function score of 5.96%, and 5.33% are obtained on dev and eval sets, respectively. A top 5 retrieval accuracy of 82.07% and 82.42% are obtained on the dev and eval sets for SID task. A brief analysis is made on the results to provide insights into the miss-classified cases in both the tasks.

Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD)
Cite as:	arXiv:2006.05747 [eess.AS]
	(or arXiv:2006.05747v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2006.05747

Submission history

From: Karthik Pandia Durai [view email]
[v1] Wed, 10 Jun 2020 09:34:46 UTC (1,504 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Uniphore's submission to Fearless Steps Challenge Phase-2

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Uniphore's submission to Fearless Steps Challenge Phase-2

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators