Skip to main content

Showing 1–21 of 21 results for author: Langlais, P

.
  1. arXiv:2405.15110  [pdf, other

    cs.CL

    CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems

    Authors: Abbas Ghaddar, David Alfonso-Hermelo, Philippe Langlais, Mehdi Rezagholizadeh, Boxing Chen, Prasanna Parthasarathi

    Abstract: In this work, we dive deep into one of the popular knowledge-grounded dialogue benchmarks that focus on faithfulness, FaithDial. We show that a significant portion of the FaithDial data contains annotation artifacts, which may bias models towards completely ignoring the conversation history. We therefore introduce CHARP, a diagnostic test set, designed for an improved evaluation of hallucinations… ▽ More

    Submitted 23 May, 2024; originally announced May 2024.

    Comments: To appear in Findings ACL 2024

  2. arXiv:2403.00252  [pdf, other

    cs.CL cs.AI cs.LG

    EUROPA: A Legal Multilingual Keyphrase Generation Dataset

    Authors: Olivier Salaün, Frédéric Piedboeuf, Guillaume Le Berre, David Alfonso Hermelo, Philippe Langlais

    Abstract: Keyphrase generation has primarily been explored within the context of academic research articles, with a particular focus on scientific domains and the English language. In this work, we present EUROPA, a dataset for multilingual keyphrase generation in the legal domain. It is derived from legal judgments from the Court of Justice of the European Union (EU), and contains instances in all 24 EU of… ▽ More

    Submitted 14 June, 2024; v1 submitted 29 February, 2024; originally announced March 2024.

    Comments: 19 pages, 2 figures, accepted at ACL 2024

  3. arXiv:2402.14895  [pdf, other

    cs.CL cs.AI cs.LG

    Data Augmentation is Dead, Long Live Data Augmentation

    Authors: Frédéric Piedboeuf, Philippe Langlais

    Abstract: Textual data augmentation (DA) is a prolific field of study where novel techniques to create artificial data are regularly proposed, and that has demonstrated great efficiency on small data settings, at least for text classification tasks. In this paper, we challenge those results, showing that classical data augmentation is simply a way of performing better fine-tuning, and that spending more tim… ▽ More

    Submitted 22 February, 2024; originally announced February 2024.

    Comments: 8 pages

  4. arXiv:2401.07760  [pdf, other

    cs.CL

    On the importance of Data Scale in Pretraining Arabic Language Models

    Authors: Abbas Ghaddar, Philippe Langlais, Mehdi Rezagholizadeh, Boxing Chen

    Abstract: Pretraining monolingual language models have been proven to be vital for performance in Arabic Natural Language Processing (NLP) tasks. In this paper, we conduct a comprehensive study on the role of data in Arabic Pretrained Language Models (PLMs). More precisely, we reassess the performance of a suite of state-of-the-art Arabic PLMs by retraining them on massive-scale, high-quality Arabic corpora… ▽ More

    Submitted 15 January, 2024; originally announced January 2024.

  5. arXiv:2311.11140  [pdf, other

    cs.DL

    The state of OAI-PMH repositories in Canadian Universities

    Authors: Frédéric Piedboeuf, Guillaume Le Berre, David Alfonso-Hermelo, Olivier Charbonneau, Philippe Langlais

    Abstract: This article presents a study of the current state of Universities Institutional Repositories (UIRs) in Canada. UIRs are vital to sharing information and documents, mainly Electronic Thesis and Dissertation (ETDs), and theoretically allow anyone, anywhere, to access the documents contained within the repository. Despite calls for consistent and shareable metadata in these repositories, our literat… ▽ More

    Submitted 18 November, 2023; originally announced November 2023.

    Comments: Published at DCMI -- International conference on dublin core and metadata applications, 2023

  6. arXiv:2307.09706  [pdf, other

    cs.CL cs.AI cs.LG

    RaTE: a Reproducible automatic Taxonomy Evaluation by Filling the Gap

    Authors: Tianjian Gao, Phillipe Langlais

    Abstract: Taxonomies are an essential knowledge representation, yet most studies on automatic taxonomy construction (ATC) resort to manual evaluation to score proposed algorithms. We argue that automatic taxonomy evaluation (ATE) is just as important as taxonomy construction. We propose RaTE, an automatic label-free taxonomy scoring procedure, which relies on a large pre-trained language model. We apply our… ▽ More

    Submitted 18 July, 2023; originally announced July 2023.

    Comments: 15th International Conference on Computational Semantics (IWCS), Association for Computational Linguistics (ACL)

  7. arXiv:2305.04971  [pdf, other

    cs.LG cs.CL

    LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization

    Authors: Peng Lu, Ahmad Rashid, Ivan Kobyzev, Mehdi Rezagholizadeh, Philippe Langlais

    Abstract: Regularization techniques are crucial to improving the generalization performance and training efficiency of deep neural networks. Many deep learning algorithms rely on weight decay, dropout, batch/layer normalization to converge faster and generalize. Label Smoothing (LS) is another simple, versatile and efficient regularization which can be applied to various supervised classification tasks. Con… ▽ More

    Submitted 8 May, 2023; originally announced May 2023.

    Comments: Accepted at ACL2023 (Findings)

  8. arXiv:2212.05956  [pdf, other

    cs.CL cs.LG

    Improving Generalization of Pre-trained Language Models via Stochastic Weight Averaging

    Authors: Peng Lu, Ivan Kobyzev, Mehdi Rezagholizadeh, Ahmad Rashid, Ali Ghodsi, Philippe Langlais

    Abstract: Knowledge Distillation (KD) is a commonly used technique for improving the generalization of compact Pre-trained Language Models (PLMs) on downstream tasks. However, such methods impose the additional burden of training a separate teacher model for every new dataset. Alternatively, one may directly work on the improvement of the optimization procedure of the compact model toward better generalizat… ▽ More

    Submitted 16 December, 2022; v1 submitted 12 December, 2022; originally announced December 2022.

    Comments: Published at EMNLP 2022 (Findings)

  9. arXiv:2205.10687  [pdf, other

    cs.CL

    Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Understanding

    Authors: Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Duan Xinyu, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Philippe Langlais

    Abstract: There is a growing body of work in recent years to develop pre-trained language models (PLMs) for the Arabic language. This work concerns addressing two major problems in existing Arabic PLMs which constraint progress of the Arabic NLU and NLG fields.First, existing Arabic PLMs are not well-explored and their pre-trainig can be improved significantly using a more methodical approach. Second, there… ▽ More

    Submitted 21 May, 2022; originally announced May 2022.

  10. arXiv:2204.07674  [pdf, other

    cs.CL

    CILDA: Contrastive Data Augmentation using Intermediate Layer Knowledge Distillation

    Authors: Md Akmal Haidar, Mehdi Rezagholizadeh, Abbas Ghaddar, Khalil Bibi, Philippe Langlais, Pascal Poupart

    Abstract: Knowledge distillation (KD) is an efficient framework for compressing large-scale pre-trained language models. Recent years have seen a surge of research aiming to improve KD by leveraging Contrastive Learning, Intermediate Layer Distillation, Data Augmentation, and Adversarial Training. In this work, we propose a learning based data augmentation technique tailored for knowledge distillation, call… ▽ More

    Submitted 15 April, 2022; originally announced April 2022.

  11. arXiv:2112.04329  [pdf, other

    cs.CL

    JABER and SABER: Junior and Senior Arabic BERt

    Authors: Abbas Ghaddar, Yimeng Wu, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Duan Xinyu, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Philippe Langlais

    Abstract: Language-specific pre-trained models have proven to be more accurate than multilingual ones in a monolingual evaluation setting, Arabic is no exception. However, we found that previously released Arabic BERT models were significantly under-trained. In this technical report, we present JABER and SABER, Junior and Senior Arabic BERt respectively, our pre-trained language model prototypes dedicated f… ▽ More

    Submitted 9 January, 2022; v1 submitted 8 December, 2021; originally announced December 2021.

    Comments: Technical Report; v2: add SABER and CAMeLBERT evaluation; v3: fix minor typos and grammatical errors

  12. arXiv:2111.05196  [pdf, other

    cs.CL

    NATURE: Natural Auxiliary Text Utterances for Realistic Spoken Language Evaluation

    Authors: David Alfonso-Hermelo, Ahmad Rashid, Abbas Ghaddar, Philippe Langlais, Mehdi Rezagholizadeh

    Abstract: Slot-filling and intent detection are the backbone of conversational agents such as voice assistants, and are active areas of research. Even though state-of-the-art techniques on publicly available benchmarks show impressive performance, their ability to generalize to realistic scenarios is yet to be demonstrated. In this work, we present NATURE, a set of simple spoken-language oriented transforma… ▽ More

    Submitted 28 January, 2022; v1 submitted 9 November, 2021; originally announced November 2021.

    Comments: 20 pages, 4 figures, accepted to NeurIPS 2021 Track Datasets and Benchmarks

  13. arXiv:2109.10164  [pdf, other

    cs.CL

    RAIL-KD: RAndom Intermediate Layer Map** for Knowledge Distillation

    Authors: Md Akmal Haidar, Nithin Anchuri, Mehdi Rezagholizadeh, Abbas Ghaddar, Philippe Langlais, Pascal Poupart

    Abstract: Intermediate layer knowledge distillation (KD) can improve the standard KD technique (which only targets the output of teacher and student models) especially over large pre-trained language models. However, intermediate layer distillation suffers from excessive computational burdens and engineering efforts required for setting up a proper layer map**. To address these problems, we propose a RAnd… ▽ More

    Submitted 1 October, 2021; v1 submitted 21 September, 2021; originally announced September 2021.

  14. arXiv:2109.10147  [pdf, other

    cs.CL

    Knowledge Distillation with Noisy Labels for Natural Language Understanding

    Authors: Shivendra Bhardwaj, Abbas Ghaddar, Ahmad Rashid, Khalil Bibi, Chengyang Li, Ali Ghodsi, Philippe Langlais, Mehdi Rezagholizadeh

    Abstract: Knowledge Distillation (KD) is extensively used to compress and deploy large pre-trained language models on edge devices for real-world applications. However, one neglected area of research is the impact of noisy (corrupted) labels on KD. We present, to the best of our knowledge, the first study on KD with noisy labels in Natural Language Understanding (NLU). We document the scope of the problem a… ▽ More

    Submitted 21 September, 2021; originally announced September 2021.

  15. End-to-End Self-Debiasing Framework for Robust NLU Training

    Authors: Abbas Ghaddar, Philippe Langlais, Mehdi Rezagholizadeh, Ahmad Rashid

    Abstract: Existing Natural Language Understanding (NLU) models have been shown to incorporate dataset biases leading to strong performance on in-distribution (ID) test sets but poor performance on out-of-distribution (OOD) ones. We introduce a simple yet effective debiasing framework whereby the shallow representations of the main model are used to derive a bias model and both models are trained simultaneou… ▽ More

    Submitted 5 September, 2021; originally announced September 2021.

    Comments: Findings ACL 2021

    Journal ref: Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021; August; 2021; pages 1923--1929

  16. Context-aware Adversarial Training for Name Regularity Bias in Named Entity Recognition

    Authors: Abbas Ghaddar, Philippe Langlais, Ahmad Rashid, Mehdi Rezagholizadeh

    Abstract: In this work, we examine the ability of NER models to use contextual information when predicting the type of an ambiguous entity. We introduce NRB, a new testbed carefully designed to diagnose Name Regularity Bias of NER models. Our results indicate that all state-of-the-art models we tested show such a bias; BERT fine-tuned models significantly outperforming feature-based (LSTM-CRF) ones on NRB,… ▽ More

    Submitted 24 July, 2021; originally announced July 2021.

    Comments: MIT Press\TACL 2021\Presented at ACL 2021 This is the exact same content of the TACL version, except the figures and tables are better aligned

    Journal ref: journal={Transactions of the Association for Computational Linguistics}, volume={9}, pages={586--604}, year={2021},

  17. arXiv:2006.02679  [pdf

    cs.DL

    Digital interfaces of historical newspapers: opportunities, restrictions and recommendations

    Authors: Eva Pfanzelter, Sarah Oberbichler, Jani Marjanen, Pierre-Carl Langlais, Stefan Hechl

    Abstract: Many libraries offer free access to digitised historical newspapers via user interfaces. After an initial period of search and filter options as the only features, the availability of more advanced tools and the desire for more options among users has ushered in a period of interface development. However, this raises a number of open questions and challenges. For example, how can we provide interf… ▽ More

    Submitted 4 June, 2020; originally announced June 2020.

  18. arXiv:1809.08962  [pdf, other

    cs.CL cs.AI

    WiRe57 : A Fine-Grained Benchmark for Open Information Extraction

    Authors: William Léchelle, Fabrizio Gotti, Philippe Langlais

    Abstract: We build a reference for the task of Open Information Extraction, on five documents. We tentatively resolve a number of issues that arise, including inference and granularity. We seek to better pinpoint the requirements for the task. We produce our annotation guidelines specifying what is correct to extract and what is not. In turn, we use this reference to score existing Open IE systems. We addre… ▽ More

    Submitted 1 August, 2019; v1 submitted 24 September, 2018; originally announced September 2018.

  19. arXiv:1806.05559  [pdf, other

    cs.CL cs.LG stat.ML

    Extracting Parallel Sentences with Bidirectional Recurrent Neural Networks to Improve Machine Translation

    Authors: Francis Grégoire, Philippe Langlais

    Abstract: Parallel sentence extraction is a task addressing the data sparsity problem found in multilingual natural language processing applications. We propose a bidirectional recurrent neural network based approach to extract parallel sentences from collections of multilingual texts. Our experiments with noisy parallel corpora show that we can achieve promising results against a competitive baseline by re… ▽ More

    Submitted 24 August, 2018; v1 submitted 13 June, 2018; originally announced June 2018.

    Comments: 12 pages, 7 figures, COLING 2018. arXiv admin note: text overlap with arXiv:1709.09783

  20. arXiv:1806.03489  [pdf, other

    cs.CL

    Robust Lexical Features for Improved Neural Network Named-Entity Recognition

    Authors: Abbas Ghaddar, Philippe Langlais

    Abstract: Neural network approaches to Named-Entity Recognition reduce the need for carefully hand-crafted features. While some features do remain in state-of-the-art systems, lexical features have been mostly discarded, with the exception of gazetteers. In this work, we show that this is unfair: lexical features are actually quite useful. We propose to embed words and entity types into a low-dimensional ve… ▽ More

    Submitted 9 June, 2018; originally announced June 2018.

    Comments: 12 pages, to appear in COLING 2018

  21. arXiv:1709.09783  [pdf, other

    cs.CL

    A Deep Neural Network Approach To Parallel Sentence Extraction

    Authors: Francis Grégoire, Philippe Langlais

    Abstract: Parallel sentence extraction is a task addressing the data sparsity problem found in multilingual natural language processing applications. We propose an end-to-end deep neural network approach to detect translational equivalence between sentences in two different languages. In contrast to previous approaches, which typically rely on multiples models and various word alignment features, by leverag… ▽ More

    Submitted 27 September, 2017; originally announced September 2017.

    Comments: 9 pages, 5 figures