Search | arXiv e-print repository

doi 10.1109/TBDATA.2024.3407573

CompanyKG: A Large-Scale Heterogeneous Graph for Company Similarity Quantification

Authors: Lele Cao, Vilhelm von Ehrenheim, Mark Granroth-Wilding, Richard Anselmo Stahl, Andrew McCornack, Armin Catovic, Dhiana Deva Cavacanti Rocha

Abstract: In the investment industry, it is often essential to carry out fine-grained company similarity quantification for a range of purposes, including market map**, competitor analysis, and mergers and acquisitions. We propose and publish a knowledge graph, named CompanyKG, to represent and learn diverse company features and relations. Specifically, 1.17 million companies are represented as nodes enri… ▽ More In the investment industry, it is often essential to carry out fine-grained company similarity quantification for a range of purposes, including market map**, competitor analysis, and mergers and acquisitions. We propose and publish a knowledge graph, named CompanyKG, to represent and learn diverse company features and relations. Specifically, 1.17 million companies are represented as nodes enriched with company description embeddings; and 15 different inter-company relations result in 51.06 million weighted edges. To enable a comprehensive assessment of methods for company similarity quantification, we have devised and compiled three evaluation tasks with annotated test sets: similarity prediction, competitor retrieval and similarity ranking. We present extensive benchmarking results for 11 reproducible predictive methods categorized into three groups: node-only, edge-only, and node+edge. To the best of our knowledge, CompanyKG is the first large-scale heterogeneous graph dataset originating from a real-world investment platform, tailored for quantifying inter-company similarity. △ Less

Submitted 9 June, 2024; v1 submitted 18 June, 2023; originally announced June 2023.

Comments: CompanyKG (version 1.x). Published by IEEE Transactions on Big Data (12 pages, 10 figures and 2 tables) + Appendix (9 pages, 1 figures and 4 tables). Code: https://github.com/EQTPartners/CompanyKG ; Data: https://zenodo.org/record/8010239

Report number: CompanyKG-V01 MSC Class: 05C85; 05C12; 68T07; 68T50; 05C90 ACM Class: E.0; I.2.1; I.2.6; H.4.0; J.0; I.2.8; I.2.7

arXiv:2208.01296 [pdf, other]

Silo NLP's Participation at WAT2022

Authors: Shantipriya Parida, Subhadarshi Panda, Stig-Arne Grönroos, Mark Granroth-Wilding, Mika Koistinen

Abstract: This paper provides the system description of "Silo NLP's" submission to the Workshop on Asian Translation (WAT2022). We have participated in the Indic Multimodal tasks (English->Hindi, English->Malayalam, and English->Bengali Multimodal Translation). For text-only translation, we trained Transformers from scratch and fine-tuned mBART-50 models. For multimodal translation, we used the same mBART a… ▽ More This paper provides the system description of "Silo NLP's" submission to the Workshop on Asian Translation (WAT2022). We have participated in the Indic Multimodal tasks (English->Hindi, English->Malayalam, and English->Bengali Multimodal Translation). For text-only translation, we trained Transformers from scratch and fine-tuned mBART-50 models. For multimodal translation, we used the same mBART architecture and extracted object tags from the images to use as visual features concatenated with the text sequence. Our submission tops many tasks including English->Hindi multimodal translation (evaluation test), English->Malayalam text-only and multimodal translation (evaluation test), English->Bengali multimodal translation (challenge test), and English->Bengali text-only translation (evaluation test). △ Less

Submitted 2 August, 2022; originally announced August 2022.

Comments: Submitted to Workshop on Asian Translation (WAT2022)

arXiv:1912.05320 [pdf, other]

CoSimLex: A Resource for Evaluating Graded Word Similarity in Context

Authors: Carlos Santos Armendariz, Matthew Purver, Matej Ulčar, Senja Pollak, Nikola Ljubešić, Marko Robnik-Šikonja, Mark Granroth-Wilding, Kristiina Vaik

Abstract: State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Standard tasks and datasets for intrinsic evaluation of embeddings are based on judgements of similarity, but ignore context; standard tasks for word sense disambiguation take account of context but do not provide continuous… ▽ More State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Standard tasks and datasets for intrinsic evaluation of embeddings are based on judgements of similarity, but ignore context; standard tasks for word sense disambiguation take account of context but do not provide continuous measures of meaning similarity. This paper describes an effort to build a new dataset, CoSimLex, intended to fill this gap. Building on the standard pairwise similarity task of SimLex-999, it provides context-dependent similarity measures; covers not only discrete differences in word sense but more subtle, graded changes in meaning; and covers not only a well-resourced language (English) but a number of less-resourced languages. We define the task and evaluation metrics, outline the dataset collection methodology, and describe the status of the dataset so far. △ Less

Submitted 29 October, 2020; v1 submitted 11 December, 2019; originally announced December 2019.

ACM Class: I.2.7

Journal ref: Proceedings of the 12th Language Resources and Evaluation Conference (2020) 5878-5886

Showing 1–3 of 3 results for author: Granroth-Wilding, M