LongEval-Retrieval: French-English Dynamic Test Collection for Continuous Web Search Evaluation

Deveaud, Petra Galuščáková Romain; Gonzalez-Saez, Gabriela; Mulhem, Philippe; Goeuriot, Lorraine; Piroi, Florina; Popel, Martin

Computer Science > Information Retrieval

arXiv:2303.03229 (cs)

[Submitted on 6 Mar 2023 (v1), last revised 27 Apr 2023 (this version, v2)]

Title:LongEval-Retrieval: French-English Dynamic Test Collection for Continuous Web Search Evaluation

Authors:Petra Galuščáková Romain Deveaud, Gabriela Gonzalez-Saez, Philippe Mulhem, Lorraine Goeuriot, Florina Piroi, Martin Popel

View PDF

Abstract:LongEval-Retrieval is a Web document retrieval benchmark that focuses on continuous retrieval evaluation. This test collection is intended to be used to study the temporal persistence of Information Retrieval systems and will be used as the test collection in the Longitudinal Evaluation of Model Performance Track (LongEval) at CLEF 2023. This benchmark simulates an evolving information system environment - such as the one a Web search engine operates in - where the document collection, the query distribution, and relevance all move continuously, while following the Cranfield paradigm for offline evaluation. To do that, we introduce the concept of a dynamic test collection that is composed of successive sub-collections each representing the state of an information system at a given time step. In LongEval-Retrieval, each sub-collection contains a set of queries, documents, and soft relevance assessments built from click models. The data comes from Qwant, a privacy-preserving Web search engine that primarily focuses on the French market. LongEval-Retrieval also provides a 'mirror' collection: it is initially constructed in the French language to benefit from the majority of Qwant's traffic, before being translated to English. This paper presents the creation process of LongEval-Retrieval and provides baseline runs and analysis.

Subjects:	Information Retrieval (cs.IR)
Cite as:	arXiv:2303.03229 [cs.IR]
	(or arXiv:2303.03229v2 [cs.IR] for this version)
	https://doi.org/10.48550/arXiv.2303.03229

Submission history

From: Petra Galuščáková [view email]
[v1] Mon, 6 Mar 2023 15:47:45 UTC (219 KB)
[v2] Thu, 27 Apr 2023 09:51:56 UTC (219 KB)

Computer Science > Information Retrieval

Title:LongEval-Retrieval: French-English Dynamic Test Collection for Continuous Web Search Evaluation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Information Retrieval

Title:LongEval-Retrieval: French-English Dynamic Test Collection for Continuous Web Search Evaluation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators