-
Grammar Assistance Using Syntactic Structures (GAUSS)
Authors:
Olga Zamaraeva,
Lorena S. Allegue,
Carlos Gómez-Rodríguez,
Anastasiia Ogneva,
Margarita Alonso-Ramos
Abstract:
Automatic grammar coaching serves an important purpose of advising on standard grammar varieties while not imposing social pressures or reinforcing established social roles. Such systems already exist but most of them are for English and few of them offer meaningful feedback. Furthermore, they typically rely completely on neural methods and require huge computational resources which most of the wo…
▽ More
Automatic grammar coaching serves an important purpose of advising on standard grammar varieties while not imposing social pressures or reinforcing established social roles. Such systems already exist but most of them are for English and few of them offer meaningful feedback. Furthermore, they typically rely completely on neural methods and require huge computational resources which most of the world cannot afford. We propose a grammar coaching system for Spanish that relies on (i) a rich linguistic formalism capable of giving informative feedback; and (ii) a faster parsing algorithm which makes using this formalism practical in a real-world application. The approach is feasible for any language for which there is a computerized grammar and is less reliant on expensive and environmentally costly neural methods. We seek to contribute to Greener AI and to address global education challenges by raising the standards of inclusivity and engagement in grammar coaching.
△ Less
Submitted 26 June, 2024;
originally announced June 2024.
-
Spanish Resource Grammar version 2023
Authors:
Olga Zamaraeva,
Lorena S. Allegue,
Carlos Gómez-Rodríguez
Abstract:
We present the latest version of the Spanish Resource Grammar (SRG), a grammar of Spanish implemented in the HPSG formalism. Such grammars encode a complex set of hypotheses about syntax making them a resource for empirical testing of linguistic theory. They also encode a strict notion of grammaticality which makes them a resource for natural language processing applications in computer-assisted l…
▽ More
We present the latest version of the Spanish Resource Grammar (SRG), a grammar of Spanish implemented in the HPSG formalism. Such grammars encode a complex set of hypotheses about syntax making them a resource for empirical testing of linguistic theory. They also encode a strict notion of grammaticality which makes them a resource for natural language processing applications in computer-assisted language learning. This version of the SRG uses the recent version of the Freeling morphological analyzer and is released along with an automatically created, manually verified treebank of 2,291 sentences. We explain the treebanking process, emphasizing how it is different from treebanking with manual annotation and how it contributes to empirically-driven development of syntactic theory. The treebanks' high level of consistency and detail makes them a resource for training high-quality semantic parsers and generally systems that benefit from precise and detailed semantics. Finally, we present the grammar's coverage and overgeneration on 100 sentences from a learner corpus, a new research line related to develo** methodologies for robust empirical evaluation of hypotheses in second language acquisition.
△ Less
Submitted 26 March, 2024; v1 submitted 23 September, 2023;
originally announced September 2023.
-
Revisiting Supertagging for HPSG
Authors:
Olga Zamaraeva,
Carlos Gómez-Rodríguez
Abstract:
We present new supertaggers trained on HPSG-based treebanks. These treebanks feature high-quality annotation based on a well-developed linguistic theory and include diverse and challenging test datasets, beyond the usual WSJ section 23 and Wikipedia data. HPSG supertagging has previously relied on MaxEnt-based models. We use SVM and neural CRF- and BERT-based methods and show that both SVM and neu…
▽ More
We present new supertaggers trained on HPSG-based treebanks. These treebanks feature high-quality annotation based on a well-developed linguistic theory and include diverse and challenging test datasets, beyond the usual WSJ section 23 and Wikipedia data. HPSG supertagging has previously relied on MaxEnt-based models. We use SVM and neural CRF- and BERT-based methods and show that both SVM and neural supertaggers achieve considerably higher accuracy compared to the baseline. Our fine-tuned BERT-based tagger achieves 97.26% accuracy on 1000 sentences from WSJ23 and 93.88% on the completely out-of-domain The Cathedral and the Bazaar (cb)). We conclude that it therefore makes sense to integrate these new supertaggers into modern HPSG parsers, and we also hope that the diverse and difficult datasets we used here will gain more popularity in the field. We contribute the complete dataset reformatted for token classification.
△ Less
Submitted 14 September, 2023;
originally announced September 2023.
-
A Summary of the First Workshop on Language Technology for Language Documentation and Revitalization
Authors:
Graham Neubig,
Shruti Rijhwani,
Alexis Palmer,
Jordan MacKenzie,
Hilaria Cruz,
Xinjian Li,
Matthew Lee,
Aditi Chaudhary,
Luke Gessler,
Steven Abney,
Shirley Anugrah Hayati,
Antonios Anastasopoulos,
Olga Zamaraeva,
Emily Prud'hommeaux,
Jennette Child,
Sara Child,
Rebecca Knowles,
Sarah Moeller,
Jeffrey Micher,
Yiyuan Li,
Sydney Zink,
Mengzhou Xia,
Roshan S Sharma,
Patrick Littell
Abstract:
Despite recent advances in natural language processing and other language technology, the application of such technology to language documentation and conservation has been limited. In August 2019, a workshop was held at Carnegie Mellon University in Pittsburgh to attempt to bring together language community members, documentary linguists, and technologists to discuss how to bridge this gap and cr…
▽ More
Despite recent advances in natural language processing and other language technology, the application of such technology to language documentation and conservation has been limited. In August 2019, a workshop was held at Carnegie Mellon University in Pittsburgh to attempt to bring together language community members, documentary linguists, and technologists to discuss how to bridge this gap and create prototypes of novel and practical language revitalization technologies. This paper reports the results of this workshop, including issues discussed, and various conceived and implemented technologies for nine languages: Arapaho, Cayuga, Inuktitut, Irish Gaelic, Kidaw'ida, Kwak'wala, Ojibwe, San Juan Quiahije Chatino, and Seneca.
△ Less
Submitted 27 April, 2020;
originally announced April 2020.