SiMa: Effective and Efficient Matching Across Data Silos Using Graph Neural Networks

Koutras, Christos; Hai, Rihan; Psarakis, Kyriakos; Fragkoulis, Marios; Katsifodimos, Asterios

Computer Science > Databases

arXiv:2206.12733 (cs)

[Submitted on 25 Jun 2022 (v1), last revised 3 Mar 2024 (this version, v2)]

Title:SiMa: Effective and Efficient Matching Across Data Silos Using Graph Neural Networks

Authors:Christos Koutras, Rihan Hai, Kyriakos Psarakis, Marios Fragkoulis, Asterios Katsifodimos

View PDF HTML (experimental)

Abstract:How can we leverage existing column relationships within silos, to predict similar ones across silos? Can we do this efficiently and effectively? Existing matching approaches do not exploit prior knowledge, relying on prohibitively expensive similarity computations. In this paper we present the first technique for matching columns across data silos, called SiMa, which leverages Graph Neural Networks (GNNs) to learn from existing column relationships within data silos, and dataset-specific profiles. The main novelty of SiMa is its ability to be trained incrementally on column relationships within each silo individually, without requiring the consolidation of all datasets in a single place. Our experiments show that SiMa is more effective than the - otherwise inapplicable to the setting of silos - state-of-the-art matching methods, while requiring orders of magnitude less computational resources. Moreover, we demonstrate that SiMa considerably outperforms other state-of-the-art column representation learning methods.

Subjects:	Databases (cs.DB)
Cite as:	arXiv:2206.12733 [cs.DB]
	(or arXiv:2206.12733v2 [cs.DB] for this version)
	https://doi.org/10.48550/arXiv.2206.12733

Submission history

From: Christos Koutras [view email]
[v1] Sat, 25 Jun 2022 21:18:08 UTC (4,423 KB)
[v2] Sun, 3 Mar 2024 07:27:07 UTC (4,823 KB)

Computer Science > Databases

Title:SiMa: Effective and Efficient Matching Across Data Silos Using Graph Neural Networks

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Databases

Title:SiMa: Effective and Efficient Matching Across Data Silos Using Graph Neural Networks

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators