BLEU, METEOR, BERTScore: Evaluation of Metrics Performance in Assessing Critical Translation Errors in Sentiment-oriented Text

Saadany, Hadeel; Orasan, Constantin

doi:10.26615/978-954-452-071-7_006

Computer Science > Computation and Language

arXiv:2109.14250 (cs)

[Submitted on 29 Sep 2021]

Title:BLEU, METEOR, BERTScore: Evaluation of Metrics Performance in Assessing Critical Translation Errors in Sentiment-oriented Text

Authors:Hadeel Saadany, Constantin Orasan

View PDF

Abstract:Social media companies as well as authorities make extensive use of artificial intelligence (AI) tools to monitor postings of hate speech, celebrations of violence or profanity. Since AI software requires massive volumes of data to train computers, Machine Translation (MT) of the online content is commonly used to process posts written in several languages and hence augment the data needed for training. However, MT mistakes are a regular occurrence when translating sentiment-oriented user-generated content (UGC), especially when a low-resource language is involved. The adequacy of the whole process relies on the assumption that the evaluation metrics used give a reliable indication of the quality of the translation. In this paper, we assess the ability of automatic quality metrics to detect critical machine translation errors which can cause serious misunderstanding of the affect message. We compare the performance of three canonical metrics on meaningless translations where the semantic content is seriously impaired as compared to meaningful translations with a critical error which exclusively distorts the sentiment of the source text. We conclude that there is a need for fine-tuning of automatic metrics to make them more robust in detecting sentiment critical errors.

Comments:	Accepted for TRITON (TRanslation and Interpreting Technology ONline) 2021
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2109.14250 [cs.CL]
	(or arXiv:2109.14250v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2109.14250
Journal reference:	TRITON (2021) 48-56
Related DOI:	https://doi.org/10.26615/978-954-452-071-7_006

Submission history

From: Hadeel Saadany [view email]
[v1] Wed, 29 Sep 2021 07:51:17 UTC (1,011 KB)

Computer Science > Computation and Language

Title:BLEU, METEOR, BERTScore: Evaluation of Metrics Performance in Assessing Critical Translation Errors in Sentiment-oriented Text

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:BLEU, METEOR, BERTScore: Evaluation of Metrics Performance in Assessing Critical Translation Errors in Sentiment-oriented Text

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators