Unsupervised Domain-agnostic Fake News Detection using Multi-modal Weak Signals

Silva, Amila; Luo, Ling; Karunasekera, Shanika; Leckie, Christopher

Computer Science > Machine Learning

arXiv:2305.11349 (cs)

COVID-19 e-print

Important: e-prints posted on arXiv are not peer-reviewed by arXiv; they should not be relied upon without context to guide clinical practice or health-related behavior and should not be reported in news media as established information without consulting multiple experts in the field.

[Submitted on 18 May 2023]

Title:Unsupervised Domain-agnostic Fake News Detection using Multi-modal Weak Signals

Authors:Amila Silva, Ling Luo, Shanika Karunasekera, Christopher Leckie

View PDF

Abstract:The emergence of social media as one of the main platforms for people to access news has enabled the wide dissemination of fake news. This has motivated numerous studies on automating fake news detection. Although there have been limited attempts at unsupervised fake news detection, their performance suffers due to not exploiting the knowledge from various modalities related to news records and due to the presence of various latent biases in the existing news datasets. To address these limitations, this work proposes an effective framework for unsupervised fake news detection, which first embeds the knowledge available in four modalities in news records and then proposes a novel noise-robust self-supervised learning technique to identify the veracity of news records from the multi-modal embeddings. Also, we propose a novel technique to construct news datasets minimizing the latent biases in existing news datasets. Following the proposed approach for dataset construction, we produce a Large-scale Unlabelled News Dataset consisting 419,351 news articles related to COVID-19, acronymed as LUND-COVID. We trained the proposed unsupervised framework using LUND-COVID to exploit the potential of large datasets, and evaluate it using a set of existing labelled datasets. Our results show that the proposed unsupervised framework largely outperforms existing unsupervised baselines for different tasks such as multi-modal fake news detection, fake news early detection and few-shot fake news detection, while yielding notable improvements for unseen domains during training.

Comments:	15 pages
Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as:	arXiv:2305.11349 [cs.LG]
	(or arXiv:2305.11349v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2305.11349

Submission history

From: Amila Silva [view email]
[v1] Thu, 18 May 2023 23:49:31 UTC (9,120 KB)

Computer Science > Machine Learning

Title:Unsupervised Domain-agnostic Fake News Detection using Multi-modal Weak Signals

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Unsupervised Domain-agnostic Fake News Detection using Multi-modal Weak Signals

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators