Search | arXiv e-print repository

doi 10.1145/3626772.3657911

RLStop: A Reinforcement Learning Stop** Method for TAR

Abstract: We present RLStop, a novel Technology Assisted Review (TAR) stop** rule based on reinforcement learning that helps minimise the number of documents that need to be manually reviewed within TAR applications. RLStop is trained on example rankings using a reward function to identify the optimal point to stop examining documents. Experiments at a range of target recall levels on multiple benchmark d… ▽ More We present RLStop, a novel Technology Assisted Review (TAR) stop** rule based on reinforcement learning that helps minimise the number of documents that need to be manually reviewed within TAR applications. RLStop is trained on example rankings using a reward function to identify the optimal point to stop examining documents. Experiments at a range of target recall levels on multiple benchmark datasets (CLEF e-Health, TREC Total Recall, and Reuters RCV1) demonstrated that RLStop substantially reduces the workload required to screen a document collection for relevance. RLStop outperforms a wide range of alternative approaches, achieving performance close to the maximum possible for the task under some circumstances. △ Less

Submitted 7 June, 2024; v1 submitted 3 May, 2024; originally announced May 2024.

Comments: Accepted at SIGIR 2024

arXiv:2312.03171 [pdf, other]

Combining Counting Processes and Classification Improves a Stop** Rule for Technology Assisted Review

Authors: Reem Bin-Hezam, Mark Stevenson

Abstract: Technology Assisted Review (TAR) stop** rules aim to reduce the cost of manually assessing documents for relevance by minimising the number of documents that need to be examined to ensure a desired level of recall. This paper extends an effective stop** rule using information derived from a text classifier that can be trained without the need for any additional annotation. Experiments on multi… ▽ More Technology Assisted Review (TAR) stop** rules aim to reduce the cost of manually assessing documents for relevance by minimising the number of documents that need to be examined to ensure a desired level of recall. This paper extends an effective stop** rule using information derived from a text classifier that can be trained without the need for any additional annotation. Experiments on multiple data sets (CLEF e-Health, TREC Total Recall, TREC Legal and RCV1) showed that the proposed approach consistently improves performance and outperforms several alternative methods. △ Less

Submitted 5 December, 2023; originally announced December 2023.

Comments: Accepted at EMNLP 2023 Findings

arXiv:2311.08597 [pdf, other]

doi 10.1145/3631990

Stop** Methods for Technology Assisted Reviews based on Point Processes

Authors: Mark Stevenson, Reem Bin-Hezam

Abstract: Technology Assisted Review (TAR), which aims to reduce the effort required to screen collections of documents for relevance, is used to develop systematic reviews of medical evidence and identify documents that must be disclosed in response to legal proceedings. Stop** methods are algorithms which determine when to stop screening documents during the TAR process, hel** to ensure that workload… ▽ More Technology Assisted Review (TAR), which aims to reduce the effort required to screen collections of documents for relevance, is used to develop systematic reviews of medical evidence and identify documents that must be disclosed in response to legal proceedings. Stop** methods are algorithms which determine when to stop screening documents during the TAR process, hel** to ensure that workload is minimised while still achieving a high level of recall. This paper proposes a novel stop** method based on point processes, which are statistical models that can be used to represent the occurrence of random events. The approach uses rate functions to model the occurrence of relevant documents in the ranking and compares four candidates, including one that has not previously been used for this purpose (hyperbolic). Evaluation is carried out using standard datasets (CLEF e-Health, TREC Total Recall, TREC Legal), and this work is the first to explore stop** method robustness by reporting performance on a range of rankings of varying effectiveness. Results show that the proposed method achieves the desired level of recall without requiring an excessive number of documents to be examined in the majority of cases and also compares well against multiple alternative approaches. △ Less

Submitted 14 November, 2023; originally announced November 2023.

Comments: Accepted by ACM Transactions on Information Systems (TOIS)

Showing 1–3 of 3 results for author: Bin-Hezam, R