Improving word mover's distance by leveraging self-attention matrix

Yamagiwa, Hiroaki; Yokoi, Sho; Shimodaira, Hidetoshi

Computer Science > Computation and Language

arXiv:2211.06229 (cs)

[Submitted on 11 Nov 2022 (v1), last revised 2 Nov 2023 (this version, v2)]

Title:Improving word mover's distance by leveraging self-attention matrix

Authors:Hiroaki Yamagiwa, Sho Yokoi, Hidetoshi Shimodaira

View PDF

Abstract:Measuring the semantic similarity between two sentences is still an important task. The word mover's distance (WMD) computes the similarity via the optimal alignment between the sets of word embeddings. However, WMD does not utilize word order, making it challenging to distinguish sentences with significant overlaps of similar words, even if they are semantically very different. Here, we attempt to improve WMD by incorporating the sentence structure represented by BERT's self-attention matrix (SAM). The proposed method is based on the Fused Gromov-Wasserstein distance, which simultaneously considers the similarity of the word embedding and the SAM for calculating the optimal transport between two sentences. Experiments demonstrate the proposed method enhances WMD and its variants in paraphrase identification with near-equivalent performance in semantic textual similarity. Our code is available at \url{this https URL}.

Comments:	24 pages, accepted to EMNLP 2023 Findings
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2211.06229 [cs.CL]
	(or arXiv:2211.06229v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2211.06229

Submission history

From: Hiroaki Yamagiwa [view email]
[v1] Fri, 11 Nov 2022 14:25:08 UTC (920 KB)
[v2] Thu, 2 Nov 2023 15:58:47 UTC (1,595 KB)

Computer Science > Computation and Language

Title:Improving word mover's distance by leveraging self-attention matrix

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Improving word mover's distance by leveraging self-attention matrix

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators