Showing 1–2 of 2 results for author: Shekofteh, Y

Search v0.5.6 released 2020-02-24

arXiv:2211.09956 [pdf]

eess.AS cs.AI cs.SD

A Persian ASR-based SER: Modification of Sharif Emotional Speech Database and Investigation of Persian Text Corpora

Authors: Ali Yazdani, Yasser Shekofteh

Abstract: Speech Emotion Recognition (SER) is one of the essential perceptual methods of humans in understanding the situation and how to interact with others, therefore, in recent years, it has been tried to add the ability to recognize emotions to human-machine communication systems. Since the SER process relies on labeled data, databases are essential for it. Incomplete, low-quality or defective data may… ▽ More Speech Emotion Recognition (SER) is one of the essential perceptual methods of humans in understanding the situation and how to interact with others, therefore, in recent years, it has been tried to add the ability to recognize emotions to human-machine communication systems. Since the SER process relies on labeled data, databases are essential for it. Incomplete, low-quality or defective data may lead to inaccurate predictions. In this paper, we fixed the inconsistencies in Sharif Emotional Speech Database (ShEMO), as a Persian database, by using an Automatic Speech Recognition (ASR) system and investigating the effect of Farsi language models obtained from accessible Persian text corpora. We also introduced a Persian/Farsi ASR-based SER system that uses linguistic features of the ASR outputs and Deep Learning-based models. △ Less

Submitted 18 November, 2022; originally announced November 2022.

Comments: 7 pages, 4 figures, 8 tables

MSC Class: 68T10 (Primary) 68T50; 68T07 (Secondary) ACM Class: I.2
arXiv:2204.13601 [pdf]

cs.SD cs.AI eess.AS

doi 10.1109/ICCKE54056.2021.9721504

Emotion Recognition In Persian Speech Using Deep Neural Networks

Authors: Ali Yazdani, Hossein Simchi, Yasser Shekofteh

Abstract: Speech Emotion Recognition (SER) is of great importance in Human-Computer Interaction (HCI), as it provides a deeper understanding of the situation and results in better interaction. In recent years, various machine learning and Deep Learning (DL) algorithms have been developed to improve SER techniques. Recognition of the spoken emotions depends on the type of expression that varies between diffe… ▽ More Speech Emotion Recognition (SER) is of great importance in Human-Computer Interaction (HCI), as it provides a deeper understanding of the situation and results in better interaction. In recent years, various machine learning and Deep Learning (DL) algorithms have been developed to improve SER techniques. Recognition of the spoken emotions depends on the type of expression that varies between different languages. In this paper, to further study important factors in the Farsi language, we examine various DL techniques on a Farsi/Persian dataset, Sharif Emotional Speech Database (ShEMO), which was released in 2018. Using signal features in low- and high-level descriptions and different deep neural networks and machine learning techniques, Unweighted Accuracy (UA) of 65.20% and Weighted Accuracy (WA) of 78.29% are achieved. △ Less

Submitted 12 November, 2022; v1 submitted 28 April, 2022; originally announced April 2022.

Comments: 5 pages, 1 figure, 3 tables

MSC Class: 68T10 (primary) 68T07 (secondary) ACM Class: I.2

Journal ref: 11th International Conference on Computer and Knowledge Engineering (ICCKE 2021)

Search v0.5.6 released 2020-02-24