The Influence of Dimensions on the Complexity of Computing Decision Trees

Kobourov, Stephen G.; Löffler, Maarten; Montecchiani, Fabrizio; Pilipczuk, Marcin; Rutter, Ignaz; Seidel, Raimund; Sorge, Manuel; Wulms, Jules

Computer Science > Computational Complexity

arXiv:2205.07756 (cs)

[Submitted on 16 May 2022 (v1), last revised 2 Jun 2022 (this version, v2)]

Title:The Influence of Dimensions on the Complexity of Computing Decision Trees

Authors:Stephen G. Kobourov, Maarten Löffler, Fabrizio Montecchiani, Marcin Pilipczuk, Ignaz Rutter, Raimund Seidel, Manuel Sorge, Jules Wulms

View PDF

Abstract:A decision tree recursively splits a feature space $\mathbb{R}^{d}$ and then assigns class labels based on the resulting partition. Decision trees have been part of the basic machine-learning toolkit for decades. A large body of work treats heuristic algorithms to compute a decision tree from training data, usually aiming to minimize in particular the size of the resulting tree. In contrast, little is known about the complexity of the underlying computational problem of computing a minimum-size tree for the given training data. We study this problem with respect to the number $d$ of dimensions of the feature space. We show that it can be solved in $O(n^{2d + 1}d)$ time, but under reasonable complexity-theoretic assumptions it is not possible to achieve $f(d) \cdot n^{o(d / \log d)}$ running time, where $n$ is the number of training examples. The problem is solvable in $(dR)^{O(dR)} \cdot n^{1+o(1)}$ time, if there are exactly two classes and $R$ is an upper bound on the number of tree leaves labeled with the first~class.

Comments:	13 pages, 8 figures
Subjects:	Computational Complexity (cs.CC)
Cite as:	arXiv:2205.07756 [cs.CC]
	(or arXiv:2205.07756v2 [cs.CC] for this version)
	https://doi.org/10.48550/arXiv.2205.07756

Submission history

From: Manuel Sorge [view email]
[v1] Mon, 16 May 2022 15:30:56 UTC (179 KB)
[v2] Thu, 2 Jun 2022 14:24:33 UTC (179 KB)

Computer Science > Computational Complexity

Title:The Influence of Dimensions on the Complexity of Computing Decision Trees

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computational Complexity

Title:The Influence of Dimensions on the Complexity of Computing Decision Trees

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators