-
Finding a Second Wind: Speeding Up Graph Traversal Queries in RDBMSs Using Column-Oriented Processing
Authors:
Mikhail Firsov,
Michael Polyntsov,
Kirill Smirnov,
George Chernishev
Abstract:
Recursive queries and recursive derived tables constitute an important part of the SQL standard. Their efficient processing is important for many real-life applications that rely on graph or hierarchy traversal. Position-enabled column-stores offer a novel opportunity to improve run times for this type of queries. Such systems allow the engine to explicitly use data positions (row ids) inside its…
▽ More
Recursive queries and recursive derived tables constitute an important part of the SQL standard. Their efficient processing is important for many real-life applications that rely on graph or hierarchy traversal. Position-enabled column-stores offer a novel opportunity to improve run times for this type of queries. Such systems allow the engine to explicitly use data positions (row ids) inside its core and thus, enable novel efficient implementations of query plan operators.
In this paper, we present an approach that significantly speeds up recursive query processing inside RDBMSes. Its core idea is to employ a particular aspect of column-store technology (late materialization) which enables the query engine to manipulate data positions during query execution. Based on it, we propose two sets of Volcano-style operators intended to process different query cases.
In order validate our ideas, we have implemented the proposed approach in PosDB, an RDBMS column-store with SQL support. We experimentally demonstrate the viability of our approach by providing a comparison with PostgreSQL. Experiments show that for breadth-first search: 1) our position-based approach yields up to 6x better results than PostgreSQL, 2) our tuple-based one results in only 3x improvement when using a special rewriting technique, but it can work in a larger number of cases, and 3) both approaches can't be emulated in row-stores efficiently.
△ Less
Submitted 16 August, 2023;
originally announced August 2023.
-
Solving Data Quality Problems with Desbordante: a Demo
Authors:
George Chernishev,
Michael Polyntsov,
Anton Chizhov,
Kirill Stupakov,
Ilya Shchuckin,
Alexander Smirnov,
Maxim Strutovsky,
Alexey Shlyonskikh,
Mikhail Firsov,
Stepan Manannikov,
Nikita Bobrov,
Daniil Goncharov,
Ilia Barutkin,
Vladislav Shalnev,
Kirill Muraviev,
Anna Rakhmukova,
Dmitriy Shcheka,
Anton Chernikov,
Mikhail Vyrodov,
Yaroslav Kurbatov,
Maxim Fofanov,
Sergei Belokonnyi,
Pavel Anosov,
Arthur Saliou,
Eduard Gaisin
, et al. (1 additional authors not shown)
Abstract:
Data profiling is an essential process in modern data-driven industries. One of its critical components is the discovery and validation of complex statistics, including functional dependencies, data constraints, association rules, and others.
However, most existing data profiling systems that focus on complex statistics do not provide proper integration with the tools used by contemporary data s…
▽ More
Data profiling is an essential process in modern data-driven industries. One of its critical components is the discovery and validation of complex statistics, including functional dependencies, data constraints, association rules, and others.
However, most existing data profiling systems that focus on complex statistics do not provide proper integration with the tools used by contemporary data scientists. This creates a significant barrier to the adoption of these tools in the industry. Moreover, existing systems were not created with industrial-grade workloads in mind. Finally, they do not aim to provide descriptive explanations, i.e. why a given pattern is not found. It is a significant issue as it is essential to understand the underlying reasons for a specific pattern's absence to make informed decisions based on the data.
Because of that, these patterns are effectively rest in thin air: their application scope is rather limited, they are rarely used by the broader public. At the same time, as we are going to demonstrate in this presentation, complex statistics can be efficiently used to solve many classic data quality problems.
Desbordante is an open-source data profiler that aims to close this gap. It is built with emphasis on industrial application: it is efficient, scalable, resilient to crashes, and provides explanations. Furthermore, it provides seamless Python integration by offloading various costly operations to the C++ core, not only mining.
In this demonstration, we show several scenarios that allow end users to solve different data quality problems. Namely, we showcase typo detection, data deduplication, and data anomaly detection scenarios.
△ Less
Submitted 28 July, 2023; v1 submitted 27 July, 2023;
originally announced July 2023.
-
Desbordante: from benchmarking suite to high-performance science-intensive data profiler (preprint)
Authors:
George Chernishev,
Michael Polyntsov,
Anton Chizhov,
Kirill Stupakov,
Ilya Shchuckin,
Alexander Smirnov,
Maxim Strutovsky,
Alexey Shlyonskikh,
Mikhail Firsov,
Stepan Manannikov,
Nikita Bobrov,
Daniil Goncharov,
Ilia Barutkin,
Vladislav Shalnev,
Kirill Muraviev,
Anna Rakhmukova,
Dmitriy Shcheka,
Anton Chernikov,
Dmitrii Mandelshtam,
Mikhail Vyrodov,
Arthur Saliou,
Eduard Gaisin,
Kirill Smirnov
Abstract:
Pioneering data profiling systems such as Metanome and OpenClean brought public attention to science-intensive data profiling. This type of profiling aims to extract complex patterns (primitives) such as functional dependencies, data constraints, association rules, and others. However, these tools are research prototypes rather than production-ready systems.
The following work presents Desbordan…
▽ More
Pioneering data profiling systems such as Metanome and OpenClean brought public attention to science-intensive data profiling. This type of profiling aims to extract complex patterns (primitives) such as functional dependencies, data constraints, association rules, and others. However, these tools are research prototypes rather than production-ready systems.
The following work presents Desbordante - a high-performance science-intensive data profiler with open source code. Unlike similar systems, it is built with emphasis on industrial application in a multi-user environment. It is efficient, resilient to crashes, and scalable. Its efficiency is ensured by implementing discovery algorithms in C++, resilience is achieved by extensive use of containerization, and scalability is based on replication of containers.
Desbordante aims to open industrial-grade primitive discovery to a broader public, focusing on domain experts who are not IT professionals. Aside from the discovery of various primitives, Desbordante offers primitive validation, which not only reports whether a given instance of primitive holds or not, but also points out what prevents it from holding via the use of special screens. Next, Desbordante supports pipelines - ready-to-use functionality implemented using the discovered primitives, for example, typo detection. We provide built-in pipelines, and the users can construct their own via provided Python bindings. Unlike other profilers, Desbordante works not only with tabular data, but with graph and transactional data as well.
In this paper, we present Desbordante, the vision behind it and its use-cases. To provide a more in-depth perspective, we discuss its current state, architecture, and design decisions it is built on. Additionally, we outline our future plans.
△ Less
Submitted 14 January, 2023;
originally announced January 2023.
-
Implementing the Comparison-Based External Sort
Authors:
Michael Polyntsov,
Valentin Grigorev,
Kirill Smirnov,
George Chernishev
Abstract:
In the age of big data, sorting is an indispensable operation for DBMSes and similar systems. Having data sorted can help produce query plans with significantly lower run times. It also can provide other benefits like having non-blocking operators which will produce data steadily (without bursts), or operators with reduced memory footprint.
Sorting may be required on any step of query processing…
▽ More
In the age of big data, sorting is an indispensable operation for DBMSes and similar systems. Having data sorted can help produce query plans with significantly lower run times. It also can provide other benefits like having non-blocking operators which will produce data steadily (without bursts), or operators with reduced memory footprint.
Sorting may be required on any step of query processing, i.e., be it source data or intermediate results. At the same time, the data to be sorted may not fit into main memory. In this case, an external sort operator, which writes intermediate results to disk, should be used.
In this paper we consider an external sort operator of the comparison-based sort type. We discuss its implementation and describe related design decisions. Our aim is to study the impact on performance of a data structure used on the merge step. For this, we have experimentally evaluated three data structures implemented inside a DBMS.
Results have shown that it is worthwhile to make an effort to implement an efficient data structure for run merging, even on modern commodity computers which are usually disk-bound. Moreover, we demonstrated that using a loser tree is a more efficient approach than both the naive approach and the heap-based one.
△ Less
Submitted 26 July, 2022;
originally announced July 2022.