Search | arXiv e-print repository

Sentiment Analysis Using Aligned Word Embeddings for Uralic Languages

Authors: Khalid Alnajjar, Mika Hämäläinen, Jack Rueter

Abstract: In this paper, we present an approach for translating word embeddings from a majority language into 4 minority languages: Erzya, Moksha, Udmurt and Komi-Zyrian. Furthermore, we align these word embeddings and present a novel neural network model that is trained on English data to conduct sentiment analysis and then applied on endangered language data through the aligned word embeddings. To test ou… ▽ More In this paper, we present an approach for translating word embeddings from a majority language into 4 minority languages: Erzya, Moksha, Udmurt and Komi-Zyrian. Furthermore, we align these word embeddings and present a novel neural network model that is trained on English data to conduct sentiment analysis and then applied on endangered language data through the aligned word embeddings. To test our model, we annotated a small sentiment analysis corpus for the 4 endangered languages and Finnish. Our method reached at least 56\% accuracy for each endangered language. The models and the sentiment corpus will be released together with this paper. Our research shows that state-of-the-art neural models can be used with endangered languages with the only requirement being a dictionary between the endangered language and a majority language. △ Less

Submitted 24 May, 2023; originally announced May 2023.

Comments: Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023)

arXiv:2301.01134 [pdf, other]

Ring That Bell: A Corpus and Method for Multimodal Metaphor Detection in Videos

Authors: Khalid Alnajjar, Mika Hämäläinen, Shuo Zhang

Abstract: We present the first openly available multimodal metaphor annotated corpus. The corpus consists of videos including audio and subtitles that have been annotated by experts. Furthermore, we present a method for detecting metaphors in the new dataset based on the textual content of the videos. The method achieves a high F1-score (62\%) for metaphorical labels. We also experiment with other modalitie… ▽ More We present the first openly available multimodal metaphor annotated corpus. The corpus consists of videos including audio and subtitles that have been annotated by experts. Furthermore, we present a method for detecting metaphors in the new dataset based on the textual content of the videos. The method achieves a high F1-score (62\%) for metaphorical labels. We also experiment with other modalities and multimodal methods; however, these methods did not out-perform the text-based model. In our error analysis, we do identify that there are cases where video could help in disambiguating metaphors, however, the visual cues are too subtle for our model to capture. The data is available on Zenodo. △ Less

Submitted 15 December, 2022; originally announced January 2023.

Comments: Figlang 2022

arXiv:2212.02911 [pdf, ps, other]

Modern French Poetry Generation with RoBERTa and GPT-2

Authors: Mika Hämäläinen, Khalid Alnajjar, Thierry Poibeau

Abstract: We present a novel neural model for modern poetry generation in French. The model consists of two pretrained neural models that are fine-tuned for the poem generation task. The encoder of the model is a RoBERTa based one while the decoder is based on GPT-2. This way the model can benefit from the superior natural language understanding performance of RoBERTa and the good natural language generatio… ▽ More We present a novel neural model for modern poetry generation in French. The model consists of two pretrained neural models that are fine-tuned for the poem generation task. The encoder of the model is a RoBERTa based one while the decoder is based on GPT-2. This way the model can benefit from the superior natural language understanding performance of RoBERTa and the good natural language generation performance of GPT-2. Our evaluation shows that the model can create French poetry successfully. On a 5 point scale, the lowest score of 3.57 was given by human judges to typicality and emotionality of the output poetry while the best score of 3.79 was given to understandability. △ Less

Submitted 6 December, 2022; originally announced December 2022.

Comments: ICCC 2022

arXiv:2212.02907 [pdf, other]

Emotion Conditioned Creative Dialog Generation

Authors: Khalid Alnajjar, Mika Hämäläinen

Abstract: We present a DialGPT based model for generating creative dialog responses that are conditioned based on one of the following emotions: anger, disgust, fear, happiness, pain, sadness and surprise. Our model is capable of producing a contextually apt response given an input sentence and a desired emotion label. Our model is capable of expressing the desired emotion with an accuracy of 0.6. The best… ▽ More We present a DialGPT based model for generating creative dialog responses that are conditioned based on one of the following emotions: anger, disgust, fear, happiness, pain, sadness and surprise. Our model is capable of producing a contextually apt response given an input sentence and a desired emotion label. Our model is capable of expressing the desired emotion with an accuracy of 0.6. The best performing emotions are neutral, fear and disgust. When measuring the strength of the expressed emotion, we find that anger, fear and disgust are expressed in the most strong fashion by the model. △ Less

Submitted 6 December, 2022; originally announced December 2022.

Comments: NLP4DH 2022

arXiv:2212.02170 [pdf, other]

Automatic Generation of Factual News Headlines in Finnish

Authors: Maximilian Koppatz, Khalid Alnajjar, Mika Hämäläinen, Thierry Poibeau

Abstract: We present a novel approach to generating news headlines in Finnish for a given news story. We model this as a summarization task where a model is given a news article, and its task is to produce a concise headline describing the main topic of the article. Because there are no openly available GPT-2 models for Finnish, we will first build such a model using several corpora. The model is then fine-… ▽ More We present a novel approach to generating news headlines in Finnish for a given news story. We model this as a summarization task where a model is given a news article, and its task is to produce a concise headline describing the main topic of the article. Because there are no openly available GPT-2 models for Finnish, we will first build such a model using several corpora. The model is then fine-tuned for the headline generation task using a massive news corpus. The system is evaluated by 3 expert journalists working in a Finnish media house. The results showcase the usability of the presented approach as a headline suggestion tool to facilitate the news production process. △ Less

Submitted 5 December, 2022; originally announced December 2022.

Comments: INLG 2022

arXiv:2212.02168 [pdf, ps, other]

Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog

Authors: Mika Hämäläinen, Khalid Alnajjar, Thierry Poibeau

Abstract: We present a method for extracting a multilingual sentiment annotated dialog data set from Fallout New Vegas. The game developers have preannotated every line of dialog in the game in one of the 8 different sentiments: \textit{anger, disgust, fear, happy, neutral, pained, sad } and \textit{surprised}. The game has been translated into English, Spanish, German, French and Italian. We conduct experi… ▽ More We present a method for extracting a multilingual sentiment annotated dialog data set from Fallout New Vegas. The game developers have preannotated every line of dialog in the game in one of the 8 different sentiments: \textit{anger, disgust, fear, happy, neutral, pained, sad } and \textit{surprised}. The game has been translated into English, Spanish, German, French and Italian. We conduct experiments on multilingual, multilabel sentiment analysis on the extracted data set using multilingual BERT, XLMRoBERTa and language specific BERT models. In our experiments, multilingual BERT outperformed XLMRoBERTa for most of the languages, also language specific models were slightly better than multilingual BERT for most of the languages. The best overall accuracy was 54\% and it was achieved by using multilingual BERT on Spanish data. The extracted data set presents a challenging task for sentiment analysis. We have released the data, including the testing and training splits, openly on Zenodo. The data set has been shuffled for copyright reasons. △ Less

Submitted 5 December, 2022; originally announced December 2022.

Comments: FDG 2022

arXiv:2211.01889 [pdf, other]

When to Laugh and How Hard? A Multimodal Approach to Detecting Humor and its Intensity

Authors: Khalid Alnajjar, Mika Hämäläinen, Jörg Tiedemann, Jorma Laaksonen, Mikko Kurimo

Abstract: Prerecorded laughter accompanying dialog in comedy TV shows encourages the audience to laugh by clearly marking humorous moments in the show. We present an approach for automatically detecting humor in the Friends TV show using multimodal data. Our model is capable of recognizing whether an utterance is humorous or not and assess the intensity of it. We use the prerecorded laughter in the show as… ▽ More Prerecorded laughter accompanying dialog in comedy TV shows encourages the audience to laugh by clearly marking humorous moments in the show. We present an approach for automatically detecting humor in the Friends TV show using multimodal data. Our model is capable of recognizing whether an utterance is humorous or not and assess the intensity of it. We use the prerecorded laughter in the show as annotation as it marks humor and the length of the audience's laughter tells us how funny a given joke is. We evaluate the model on episodes the model has not been exposed to during the training phase. Our results show that the model is capable of correctly detecting whether an utterance is humorous 78% of the time and how long the audience's laughter reaction should last with a mean absolute error of 600 milliseconds. △ Less

Submitted 3 November, 2022; originally announced November 2022.

Comments: Outstanding paper award in COLING 2022

arXiv:2207.04453 [pdf]

Multilingual Persuasion Detection: Video Games as an Invaluable Data Source for NLP

Authors: Teemu Pöyhönen, Mika Hämäläinen, Khalid Alnajjar

Abstract: Role-playing games (RPGs) have a considerable amount of text in video game dialogues. Quite often this text is semi-annotated by the game developers. In this paper, we extract a multilingual dataset of persuasive dialogue from several RPGs. We show the viability of this data in building a persuasion detection system using a natural language processing (NLP) model called BERT. We believe that video… ▽ More Role-playing games (RPGs) have a considerable amount of text in video game dialogues. Quite often this text is semi-annotated by the game developers. In this paper, we extract a multilingual dataset of persuasive dialogue from several RPGs. We show the viability of this data in building a persuasion detection system using a natural language processing (NLP) model called BERT. We believe that video games have a lot of unused potential as a datasource for a variety of NLP tasks. The code and data described in this paper are available on Zenodo. △ Less

Submitted 10 July, 2022; originally announced July 2022.

Comments: DiGRA 2022

arXiv:2205.08024 [pdf, ps, other]

Harnessing Multilingual Resources to Question Answering in Arabic

Authors: Khalid Alnajjar, Mika Hämäläinen

Abstract: The goal of the paper is to predict answers to questions given a passage of Qur'an. The answers are always found in the passage, so the task of the model is to predict where an answer starts and where it ends. As the initial data set is rather small for training, we make use of multilingual BERT so that we can augment the training data by using data available for languages other than Arabic. Furth… ▽ More The goal of the paper is to predict answers to questions given a passage of Qur'an. The answers are always found in the passage, so the task of the model is to predict where an answer starts and where it ends. As the initial data set is rather small for training, we make use of multilingual BERT so that we can augment the training data by using data available for languages other than Arabic. Furthermore, we crawl a large Arabic corpus that is domain specific to religious discourse. Our approach consists of two steps, first we train a BERT model to predict a set of possible answers in a passage. Finally, we use another BERT based model to rank the candidate answers produced by the first BERT model. △ Less

Submitted 16 May, 2022; originally announced May 2022.

arXiv:2112.14153 [pdf, other]

Processing M.A. Castrén's Materials: Multilingual Typed and Handwritten Manuscripts

Authors: Niko Partanen, Jack Rueter, Mika Hämäläinen, Khalid Alnajjar

Abstract: The study forms a technical report of various tasks that have been performed on the materials collected and published by Finnish ethnographer and linguist, Matthias Alexander Castrén (1813-1852). The Finno-Ugrian Society is publishing Castrén's manuscripts as new critical and digital editions, and at the same time different research groups have also paid attention to these materials. We discuss th… ▽ More The study forms a technical report of various tasks that have been performed on the materials collected and published by Finnish ethnographer and linguist, Matthias Alexander Castrén (1813-1852). The Finno-Ugrian Society is publishing Castrén's manuscripts as new critical and digital editions, and at the same time different research groups have also paid attention to these materials. We discuss the workflows and technical infrastructure used, and consider how datasets that benefit different computational tasks could be created to further improve the usability of these materials, and also to aid the further processing of similar archived collections. We specifically focus on the parts of the collections that are processed in a way that improves their usability in more technical applications, complementing the earlier work on the cultural and linguistic aspects of these materials. Most of these datasets are openly available in Zenodo. The study points to specific areas where further research is needed, and provides benchmarks for text recognition tasks. △ Less

Submitted 28 December, 2021; originally announced December 2021.

Comments: Proceedings of the Workshop on Natural Language Processing for Digital Humanities

arXiv:2112.12489 [pdf, other]

TFW2V: An Enhanced Document Similarity Method for the Morphologically Rich Finnish Language

Authors: Quan Duong, Mika Hämäläinen, Khalid Alnajjar

Abstract: Measuring the semantic similarity of different texts has many important applications in Digital Humanities research such as information retrieval, document clustering and text summarization. The performance of different methods depends on the length of the text, the domain and the language. This study focuses on experimenting with some of the current approaches to Finnish, which is a morphological… ▽ More Measuring the semantic similarity of different texts has many important applications in Digital Humanities research such as information retrieval, document clustering and text summarization. The performance of different methods depends on the length of the text, the domain and the language. This study focuses on experimenting with some of the current approaches to Finnish, which is a morphologically rich language. At the same time, we propose a simple method, TFW2V, which shows high efficiency in handling both long text documents and limited amounts of data. Furthermore, we design an objective evaluation method which can be used as a framework for benchmarking text similarity approaches. △ Less

Submitted 23 December, 2021; originally announced December 2021.

Comments: Workshop on Natural Language Processing for Digital Humanities (NLP4DH)

arXiv:2111.04574 [pdf]

Detecting Depression in Thai Blog Posts: a Dataset and a Baseline

Authors: Mika Hämäläinen, Pattama Patpong, Khalid Alnajjar, Niko Partanen, Jack Rueter

Abstract: We present the first openly available corpus for detecting depression in Thai. Our corpus is compiled by expert verified cases of depression in several online blogs. We experiment with two different LSTM based models and two different BERT based models. We achieve a 77.53\% accuracy with a Thai BERT model in detecting depression. This establishes a good baseline for future researcher on the same c… ▽ More We present the first openly available corpus for detecting depression in Thai. Our corpus is compiled by expert verified cases of depression in several online blogs. We experiment with two different LSTM based models and two different BERT based models. We achieve a 77.53\% accuracy with a Thai BERT model in detecting depression. This establishes a good baseline for future researcher on the same corpus. Furthermore, we identify a need for Thai embeddings that have been trained on a more varied corpus than Wikipedia. Our corpus, code and trained models have been released openly on Zenodo. △ Less

Submitted 8 November, 2021; originally announced November 2021.

Comments: Workshop on Noisy User-generated Text (at EMNLP)

arXiv:2111.03800 [pdf, other]

Finnish Dialect Identification: The Effect of Audio and Text

Authors: Mika Hämäläinen, Khalid Alnajjar, Niko Partanen, Jack Rueter

Abstract: Finnish is a language with multiple dialects that not only differ from each other in terms of accent (pronunciation) but also in terms of morphological forms and lexical choice. We present the first approach to automatically detect the dialect of a speaker based on a dialect transcript and transcript with audio recording in a dataset consisting of 23 different dialects. Our results show that the b… ▽ More Finnish is a language with multiple dialects that not only differ from each other in terms of accent (pronunciation) but also in terms of morphological forms and lexical choice. We present the first approach to automatically detect the dialect of a speaker based on a dialect transcript and transcript with audio recording in a dataset consisting of 23 different dialects. Our results show that the best accuracy is received by combining both of the modalities, as text only reaches to an overall accuracy of 57\%, where as text and audio reach to 85\%. Our code, models and data have been released openly on Github and Zenodo. △ Less

Submitted 6 November, 2021; originally announced November 2021.

Comments: EMNLP 2021

arXiv:2109.11326 [pdf, ps, other]

The Current State of Finnish NLP

Authors: Mika Hämäläinen, Khalid Alnajjar

Abstract: There are a lot of tools and resources available for processing Finnish. In this paper, we survey recent papers focusing on Finnish NLP related to many different subcategories of NLP such as parsing, generation, semantics and speech. NLP research is conducted in many different research groups in Finland, and it is frequently the case that NLP tools and models resulting from academic research are m… ▽ More There are a lot of tools and resources available for processing Finnish. In this paper, we survey recent papers focusing on Finnish NLP related to many different subcategories of NLP such as parsing, generation, semantics and speech. NLP research is conducted in many different research groups in Finland, and it is frequently the case that NLP tools and models resulting from academic research are made available for others to use on platforms such as Github. △ Less

Submitted 23 September, 2021; originally announced September 2021.

Comments: Seventh international workshop on computational linguistics of Uralic languages (IWCLUL)

arXiv:2109.08702 [pdf, other]

When a Computer Cracks a Joke: Automated Generation of Humorous Headlines

Authors: Khalid Alnajjar, Mika Hämäläinen

Abstract: Automated news generation has become a major interest for new agencies in the past. Oftentimes headlines for such automatically generated news articles are unimaginative as they have been generated with ready-made templates. We present a computationally creative approach for headline generation that can generate humorous versions of existing headlines. We evaluate our system with human judges and… ▽ More Automated news generation has become a major interest for new agencies in the past. Oftentimes headlines for such automatically generated news articles are unimaginative as they have been generated with ready-made templates. We present a computationally creative approach for headline generation that can generate humorous versions of existing headlines. We evaluate our system with human judges and compare the results to human authored humorous titles. The headlines produced by the system are considered funny 36\% of the time by human evaluators. △ Less

Submitted 17 September, 2021; originally announced September 2021.

Comments: Proceedings of the 12th International Conference on Computational Creativity (ICCC 2021)

arXiv:2108.09546 [pdf, ps, other]

How Cute is Pikachu? Gathering and Ranking Pokémon Properties from Data with Pokémon Word Embeddings

Authors: Mika Hämäläinen, Khalid Alnajjar, Niko Partanen

Abstract: We present different methods for obtaining descriptive properties automatically for the 151 original Pokémon. We train several different word embeddings models on a crawled Pokémon corpus, and use them to rank automatically English adjectives based on how characteristic they are to a given Pokémon. Based on our experiments, it is better to train a model with domain specific data than to use a pret… ▽ More We present different methods for obtaining descriptive properties automatically for the 151 original Pokémon. We train several different word embeddings models on a crawled Pokémon corpus, and use them to rank automatically English adjectives based on how characteristic they are to a given Pokémon. Based on our experiments, it is better to train a model with domain specific data than to use a pretrained model. Word2Vec produces less noise in the results than fastText model. Furthermore, we expand the list of properties for each Pokémon automatically. However, none of the methods is spot on and there is a considerable amount of noise in the different semantic models. Our models have been released on Zenodo. △ Less

Submitted 21 August, 2021; originally announced August 2021.

Comments: English translation of Hämäläinen, M., Alnajjar, K. \& Partanen, N. (2021). Nettikorpuksen avulla tuotettuja sanavektorimalleja Pokémonien ominaisuuksien kuvaamiseksi. In Saarikivi, T. \& Saarikivi, J. (eds.) \textit{Turhan tiedon kirja -- Tutkimuksista pois jätettyjä sivuja}

arXiv:2108.00308 [pdf, ps, other]

Human Evaluation of Creative NLG Systems: An Interdisciplinary Survey on Recent Papers

Authors: Mika Hämäläinen, Khalid Alnajjar

Abstract: We survey human evaluation in papers presenting work on creative natural language generation that have been published in INLG 2020 and ICCC 2020. The most typical human evaluation method is a scaled survey, typically on a 5 point scale, while many other less common methods exist. The most commonly evaluated parameters are meaning, syntactic correctness, novelty, relevance and emotional value, amon… ▽ More We survey human evaluation in papers presenting work on creative natural language generation that have been published in INLG 2020 and ICCC 2020. The most typical human evaluation method is a scaled survey, typically on a 5 point scale, while many other less common methods exist. The most commonly evaluated parameters are meaning, syntactic correctness, novelty, relevance and emotional value, among many others. Our guidelines for future evaluation include clearly defining the goal of the generative system, asking questions as concrete as possible, testing the evaluation setup, using multiple different evaluation setups, reporting the entire evaluation process and potential biases clearly, and finally analyzing the evaluation results in a more profound way than merely reporting the most typical statistics. △ Less

Submitted 31 July, 2021; originally announced August 2021.

Comments: Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021)

arXiv:2107.03266 [pdf, ps, other]

Lemmatization of Historical Old Literary Finnish Texts in Modern Orthography

Authors: Mika Hämäläinen, Niko Partanen, Khalid Alnajjar

Abstract: Texts written in Old Literary Finnish represent the first literary work ever written in Finnish starting from the 16th century. There have been several projects in Finland that have digitized old publications and made them available for research use. However, using modern NLP methods in such data poses great challenges. In this paper we propose an approach for simultaneously normalizing and lemmat… ▽ More Texts written in Old Literary Finnish represent the first literary work ever written in Finnish starting from the 16th century. There have been several projects in Finland that have digitized old publications and made them available for research use. However, using modern NLP methods in such data poses great challenges. In this paper we propose an approach for simultaneously normalizing and lemmatizing Old Literary Finnish into modern spelling. Our best model reaches to 96.3\% accuracy in texts written by Agricola and 87.7\% accuracy in other contemporary out-of-domain text. Our method has been made freely available on Zenodo and Github. △ Less

Submitted 7 July, 2021; originally announced July 2021.

Comments: la 28e Conférence sur le Traitement Automatique des Langues Naturelles (TALN)

arXiv:2106.03389 [pdf, other]

Never guess what I heard... Rumor Detection in Finnish News: a Dataset and a Baseline

Authors: Mika Hämäläinen, Khalid Alnajjar, Niko Partanen, Jack Rueter

Abstract: This study presents a new dataset on rumor detection in Finnish language news headlines. We have evaluated two different LSTM based models and two different BERT models, and have found very significant differences in the results. A fine-tuned FinBERT reaches the best overall accuracy of 94.3% and rumor label accuracy of 96.0% of the time. However, a model fine-tuned on Multilingual BERT reaches th… ▽ More This study presents a new dataset on rumor detection in Finnish language news headlines. We have evaluated two different LSTM based models and two different BERT models, and have found very significant differences in the results. A fine-tuned FinBERT reaches the best overall accuracy of 94.3% and rumor label accuracy of 96.0% of the time. However, a model fine-tuned on Multilingual BERT reaches the best factual label accuracy of 97.2%. Our results suggest that the performance difference is due to a difference in the original training data. Furthermore, we find that a regular LSTM model works better than one trained with a pretrained word2vec model. These findings suggest that more work needs to be done for pretrained models in Finnish language as they have been trained on small and biased corpora. △ Less

Submitted 7 June, 2021; originally announced June 2021.

Comments: 2021 Workshop on NLP4IF: Censorship, Disinformation, and Propaganda

arXiv:2105.12428 [pdf, other]

Neural Morphology Dataset and Models for Multiple Languages, from the Large to the Endangered

Authors: Mika Hämäläinen, Niko Partanen, Jack Rueter, Khalid Alnajjar

Abstract: We train neural models for morphological analysis, generation and lemmatization for morphologically rich languages. We present a method for automatically extracting substantially large amount of training data from FSTs for 22 languages, out of which 17 are endangered. The neural models follow the same tagset as the FSTs in order to make it possible to use them as fallback systems together with the… ▽ More We train neural models for morphological analysis, generation and lemmatization for morphologically rich languages. We present a method for automatically extracting substantially large amount of training data from FSTs for 22 languages, out of which 17 are endangered. The neural models follow the same tagset as the FSTs in order to make it possible to use them as fallback systems together with the FSTs. The source code, models and datasets have been released on Zenodo. △ Less

Submitted 26 May, 2021; originally announced May 2021.

Comments: The 23rd Nordic Conference on Computational Linguistics (NoDaLiDa 2021)

arXiv:2105.05542 [pdf, other]

!Qué maravilla! Multimodal Sarcasm Detection in Spanish: a Dataset and a Baseline

Authors: Khalid Alnajjar, Mika Hämäläinen

Abstract: We construct the first ever multimodal sarcasm dataset for Spanish. The audiovisual dataset consists of sarcasm annotated text that is aligned with video and audio. The dataset represents two varieties of Spanish, a Latin American variety and a Peninsular Spanish variety, which ensures a wider dialectal coverage for this global language. We present several models for sarcasm detection that will se… ▽ More We construct the first ever multimodal sarcasm dataset for Spanish. The audiovisual dataset consists of sarcasm annotated text that is aligned with video and audio. The dataset represents two varieties of Spanish, a Latin American variety and a Peninsular Spanish variety, which ensures a wider dialectal coverage for this global language. We present several models for sarcasm detection that will serve as baselines in the future research. Our results show that results with text only (89%) are worse than when combining text with audio (91.9%). Finally, the best results are obtained when combining all the modalities: text, audio and video (93.1%). △ Less

Submitted 12 May, 2021; originally announced May 2021.

Comments: Accepted to The Third Workshop on Multimodal Artificial Intelligence (MAI-Workshop)

arXiv:2104.05361 [pdf, ps, other]

The Great Misalignment Problem in Human Evaluation of NLP Methods

Authors: Mika Hämäläinen, Khalid Alnajjar

Abstract: We outline the Great Misalignment Problem in natural language processing research, this means simply that the problem definition is not in line with the method proposed and the human evaluation is not in line with the definition nor the method. We study this misalignment problem by surveying 10 randomly sampled papers published in ACL 2020 that report results with human evaluation. Our results sho… ▽ More We outline the Great Misalignment Problem in natural language processing research, this means simply that the problem definition is not in line with the method proposed and the human evaluation is not in line with the definition nor the method. We study this misalignment problem by surveying 10 randomly sampled papers published in ACL 2020 that report results with human evaluation. Our results show that only one paper was fully in line in terms of problem definition, method and evaluation. Only two papers presented a human evaluation that was in line with what was modeled in the method. These results highlight that the Great Misalignment Problem is a major one and it affects the validity and reproducibility of results obtained by a human evaluation. △ Less

Submitted 12 April, 2021; originally announced April 2021.

Comments: Workshop on Human Evaluation of NLP Systems at EACL 2021

arXiv:2103.13275 [pdf, other]

doi 10.31885/9789515150257.24

When Word Embeddings Become Endangered

Authors: Khalid Alnajjar

Abstract: Big languages such as English and Finnish have many natural language processing (NLP) resources and models, but this is not the case for low-resourced and endangered languages as such resources are so scarce despite the great advantages they would provide for the language communities. The most common types of resources available for low-resourced and endangered languages are translation dictionari… ▽ More Big languages such as English and Finnish have many natural language processing (NLP) resources and models, but this is not the case for low-resourced and endangered languages as such resources are so scarce despite the great advantages they would provide for the language communities. The most common types of resources available for low-resourced and endangered languages are translation dictionaries and universal dependencies. In this paper, we present a method for constructing word embeddings for endangered languages using existing word embeddings of different resource-rich languages and the translation dictionaries of resource-poor languages. Thereafter, the embeddings are fine-tuned using the sentences in the universal dependencies and aligned to match the semantic spaces of the big languages; resulting in cross-lingual embeddings. The endangered languages we work with here are Erzya, Moksha, Komi-Zyrian and Skolt Sami. Furthermore, we build a universal sentiment analysis model for all the languages that are part of this study, whether endangered or not, by utilizing cross-lingual word embeddings. The evaluation conducted shows that our word embeddings for endangered languages are well-aligned with the resource-rich languages, and they are suitable for training task-specific models as demonstrated by our sentiment analysis model which achieved a high accuracy. All our cross-lingual word embeddings and the sentiment analysis model have been released openly via an easy-to-use Python library. △ Less

Submitted 24 March, 2021; originally announced March 2021.

Journal ref: In M. Hämäläinen, N. Partanen, & K. Alnajjar (Eds.), Multilingual Facilitation (pp. 275-288). University of Helsinki (2021)

arXiv:2012.05318 [pdf, other]

doi 10.1145/3423337.3429435

Normalization of Different Swedish Dialects Spoken in Finland

Authors: Mika Hämäläinen, Niko Partanen, Khalid Alnajjar

Abstract: Our study presents a dialect normalization method for different Finland Swedish dialects covering six regions. We tested 5 different models, and the best model improved the word error rate from 76.45 to 28.58. Contrary to results reported in earlier research on Finnish dialects, we found that training the model with one word at a time gave best results. We believe this is due to the size of the tr… ▽ More Our study presents a dialect normalization method for different Finland Swedish dialects covering six regions. We tested 5 different models, and the best model improved the word error rate from 76.45 to 28.58. Contrary to results reported in earlier research on Finnish dialects, we found that training the model with one word at a time gave best results. We believe this is due to the size of the training data available for the model. Our models are accessible as a Python package. The study provides important information about the adaptability of these methods in different contexts, and gives important baselines for further study. △ Less

Submitted 9 December, 2020; originally announced December 2020.

Comments: In Proceedings of the 4th ACM SIGSPATIAL Workshop on Geospatial Humanities (GeoHumanities'20)

arXiv:2012.02578 [pdf, other]

Ve'rdd. Narrowing the Gap between Paper Dictionaries, Low-Resource NLP and Community Involvement

Authors: Khalid Alnajjar, Mika Hämäläinen, Jack Rueter, Niko Partanen

Abstract: We present an open-source online dictionary editing system, Ve'rdd, that offers a chance to re-evaluate and edit grassroots dictionaries that have been exposed to multiple amateur editors. The idea is to incorporate community activities into a state-of-the-art finite-state language description of a seriously endangered minority language, Skolt Sami. Problems involve getting the community to take p… ▽ More We present an open-source online dictionary editing system, Ve'rdd, that offers a chance to re-evaluate and edit grassroots dictionaries that have been exposed to multiple amateur editors. The idea is to incorporate community activities into a state-of-the-art finite-state language description of a seriously endangered minority language, Skolt Sami. Problems involve getting the community to take part in things above the pencil-and-paper level. At times, it seems that the native speakers and the dictionary oriented are lacking technical understanding to utilize the infrastructures which might make their work more meaningful in the future, i.e. multiple reuse of all of their input. Therefore, our system integrates with the existing tools and infrastructures for Uralic language masking the technical complexities behind a user-friendly UI. △ Less

Submitted 4 December, 2020; originally announced December 2020.

Comments: Proceedings of the 28th International Conference on Computational Linguistics: System Demonstrations

arXiv:2010.05269 [pdf, other]

Automated Prediction of Medieval Arabic Diacritics

Authors: Khalid Alnajjar, Mika Hämäläinen, Niko Partanen, Jack Rueter

Abstract: This study uses a character level neural machine translation approach trained on a long short-term memory-based bi-directional recurrent neural network architecture for diacritization of Medieval Arabic. The results improve from the online tool used as a baseline. A diacritization model have been published openly through an easy to use Python package available on PyPi and Zenodo. We have found tha… ▽ More This study uses a character level neural machine translation approach trained on a long short-term memory-based bi-directional recurrent neural network architecture for diacritization of Medieval Arabic. The results improve from the online tool used as a baseline. A diacritization model have been published openly through an easy to use Python package available on PyPi and Zenodo. We have found that context size should be considered when optimizing a feasible prediction model. △ Less

Submitted 11 October, 2020; originally announced October 2020.

arXiv:2009.02685 [pdf, ps, other]

Automatic Dialect Adaptation in Finnish and its Effect on Perceived Creativity

Authors: Mika Hämäläinen, Niko Partanen, Khalid Alnajjar, Jack Rueter, Thierry Poibeau

Abstract: We present a novel approach for adapting text written in standard Finnish to different dialects. We experiment with character level NMT models both by using a multi-dialectal and transfer learning approaches. The models are tested with over 20 different dialects. The results seem to favor transfer learning, although not strongly over the multi-dialectal approach. We study the influence dialectal a… ▽ More We present a novel approach for adapting text written in standard Finnish to different dialects. We experiment with character level NMT models both by using a multi-dialectal and transfer learning approaches. The models are tested with over 20 different dialects. The results seem to favor transfer learning, although not strongly over the multi-dialectal approach. We study the influence dialectal adaptation has on perceived creativity of computer generated poetry. Our results suggest that the more the dialect deviates from the standard Finnish, the lower scores people tend to give on an existing evaluation metric. However, on a word association test, people associate creativity and originality more with dialect and fluency more with standard Finnish. △ Less

Submitted 6 September, 2020; originally announced September 2020.

Comments: In proceedings of the Eleventh International Conference on Computational Creativity

arXiv:1910.13946 [pdf, other]

Let's FACE it. Finnish Poetry Generation with Aesthetics and Framing

Authors: Mika Hämäläinen, Khalid Alnajjar

Abstract: We present a creative poem generator for the morphologically rich Finnish language. Our method falls into the master-apprentice paradigm, where a computationally creative genetic algorithm teaches a BRNN model to generate poetry. We model several parts of poetic aesthetics in the fitness function of the genetic algorithm, such as sonic features, semantic coherence, imagery and metaphor. Furthermor… ▽ More We present a creative poem generator for the morphologically rich Finnish language. Our method falls into the master-apprentice paradigm, where a computationally creative genetic algorithm teaches a BRNN model to generate poetry. We model several parts of poetic aesthetics in the fitness function of the genetic algorithm, such as sonic features, semantic coherence, imagery and metaphor. Furthermore, we justify the creativity of our method based on the FACE theory on computational creativity and take additional care in evaluating our system by automatic metrics for concepts together with human evaluation for aesthetics, framing and expressions. △ Less

Submitted 30 October, 2019; originally announced October 2019.

Journal ref: Proceedings of the 12th International Conference on Natural Language Generation (INLG 2019)

arXiv:1907.04954 [pdf, other]

Modelling the Socialization of Creative Agents in a Master-Apprentice Setting: The Case of Movie Title Puns

Authors: Mika Hämäläinen, Khalid Alnajjar

Abstract: This paper presents work on modelling the social psychological aspect of socialization in the case of a computationally creative master-apprentice system. In each master-apprentice pair, the master, a genetic algorithm, is seen as a parent for its apprentice, which is an NMT based sequence-to-sequence model. The effect of different parenting styles on the creative output of each pair is in the foc… ▽ More This paper presents work on modelling the social psychological aspect of socialization in the case of a computationally creative master-apprentice system. In each master-apprentice pair, the master, a genetic algorithm, is seen as a parent for its apprentice, which is an NMT based sequence-to-sequence model. The effect of different parenting styles on the creative output of each pair is in the focus of this study. This approach brings a novel view point to computational social creativity, which has mainly focused in the past on computationally creative agents being on a socially equal level, whereas our approach studies the phenomenon in the context of a social hierarchy. △ Less

Submitted 10 July, 2019; originally announced July 2019.

Journal ref: Proceedings of the 10th International Conference on Computational Creativity. Grace, K., Cook, M., Ventura, D. & Maher, M. L. (eds.). Association for Computational Creativity, p. 266-273 (2019)

Showing 1–29 of 29 results for author: Alnajjar, K