Showing 1–2 of 2 results for author: Hazim, R
-
The SAMER Arabic Text Simplification Corpus
Authors:
Bashar Alhafni,
Reem Hazim,
Juan PiƱeros Liberato,
Muhamed Al Khalil,
Nizar Habash
Abstract:
We present the SAMER Corpus, the first manually annotated Arabic parallel corpus for text simplification targeting school-aged learners. Our corpus comprises texts of 159K words selected from 15 publicly available Arabic fiction novels most of which were published between 1865 and 1955. Our corpus includes readability level annotations at both the document and word levels, as well as two simplifie…
▽ More
We present the SAMER Corpus, the first manually annotated Arabic parallel corpus for text simplification targeting school-aged learners. Our corpus comprises texts of 159K words selected from 15 publicly available Arabic fiction novels most of which were published between 1865 and 1955. Our corpus includes readability level annotations at both the document and word levels, as well as two simplified parallel versions for each text targeting learners at two different readability levels. We describe the corpus selection process, and outline the guidelines we followed to create the annotations and ensure their quality. Our corpus is publicly available to support and encourage research on Arabic text simplification, Arabic automatic readability assessment, and the development of Arabic pedagogical language technologies.
△ Less
Submitted 29 April, 2024;
originally announced April 2024.
-
Arabic Word-level Readability Visualization for Assisted Text Simplification
Authors:
Reem Hazim,
Hind Saddiki,
Bashar Alhafni,
Muhamed Al Khalil,
Nizar Habash
Abstract:
This demo paper presents a Google Docs add-on for automatic Arabic word-level readability visualization. The add-on includes a lemmatization component that is connected to a five-level readability lexicon and Arabic WordNet-based substitution suggestions. The add-on can be used for assessing the reading difficulty of a text and identifying difficult words as part of the task of manual text simplif…
▽ More
This demo paper presents a Google Docs add-on for automatic Arabic word-level readability visualization. The add-on includes a lemmatization component that is connected to a five-level readability lexicon and Arabic WordNet-based substitution suggestions. The add-on can be used for assessing the reading difficulty of a text and identifying difficult words as part of the task of manual text simplification. We make our add-on and its code publicly available.
△ Less
Submitted 19 October, 2022;
originally announced October 2022.