Falcon 7b for Software Mention Detection in Scholarly Documents

Khan, AmeerAli; Ramadan, Qusai; Yang, Cong; Boukhers, Zeyd

Computer Science > Machine Learning

arXiv:2405.08514 (cs)

[Submitted on 14 May 2024]

Title:Falcon 7b for Software Mention Detection in Scholarly Documents

Authors:AmeerAli Khan, Qusai Ramadan, Cong Yang, Zeyd Boukhers

View PDF HTML (experimental)

Abstract:This paper aims to tackle the challenge posed by the increasing integration of software tools in research across various disciplines by investigating the application of Falcon-7b for the detection and classification of software mentions within scholarly texts. Specifically, the study focuses on solving Subtask I of the Software Mention Detection in Scholarly Publications (SOMD), which entails identifying and categorizing software mentions from academic literature. Through comprehensive experimentation, the paper explores different training strategies, including a dual-classifier approach, adaptive sampling, and weighted loss scaling, to enhance detection accuracy while overcoming the complexities of class imbalance and the nuanced syntax of scholarly writing. The findings highlight the benefits of selective labelling and adaptive sampling in improving the model's performance. However, they also indicate that integrating multiple strategies does not necessarily result in cumulative improvements. This research offers insights into the effective application of large language models for specific tasks such as SOMD, underlining the importance of tailored approaches to address the unique challenges presented by academic text analysis.

Comments:	Accepted for publication by the first Workshop on Natural Scientific Language Processing and Research Knowledge Graphs - NSLP (@ ESCAI)
Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL); Digital Libraries (cs.DL)
Cite as:	arXiv:2405.08514 [cs.LG]
	(or arXiv:2405.08514v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2405.08514

Submission history

From: Zeyd Boukhers [view email]
[v1] Tue, 14 May 2024 11:37:26 UTC (251 KB)

Computer Science > Machine Learning

Title:Falcon 7b for Software Mention Detection in Scholarly Documents

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Falcon 7b for Software Mention Detection in Scholarly Documents

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators