Sample-Optimal Low-Rank Approximation of Distance Matrices

Indyk, Piotr; Vakilian, Ali; Wagner, Tal; Woodruff, David

Computer Science > Data Structures and Algorithms

arXiv:1906.00339 (cs)

[Submitted on 2 Jun 2019]

Title:Sample-Optimal Low-Rank Approximation of Distance Matrices

Authors:Piotr Indyk, Ali Vakilian, Tal Wagner, David Woodruff

View PDF

Abstract:A distance matrix $A \in \mathbb R^{n \times m}$ represents all pairwise distances, $A_{ij}=\mathrm{d}(x_i,y_j)$, between two point sets $x_1,...,x_n$ and $y_1,...,y_m$ in an arbitrary metric space $(\mathcal Z, \mathrm{d})$. Such matrices arise in various computational contexts such as learning image manifolds, handwriting recognition, and multi-dimensional unfolding.
In this work we study algorithms for low-rank approximation of distance matrices. Recent work by Bakshi and Woodruff (NeurIPS 2018) showed it is possible to compute a rank-$k$ approximation of a distance matrix in time $O((n+m)^{1+\gamma}) \cdot \mathrm{poly}(k,1/\epsilon)$, where $\epsilon>0$ is an error parameter and $\gamma>0$ is an arbitrarily small constant. Notably, their bound is sublinear in the matrix size, which is unachievable for general matrices.
We present an algorithm that is both simpler and more efficient. It reads only $O((n+m) k/\epsilon)$ entries of the input matrix, and has a running time of $O(n+m) \cdot \mathrm{poly}(k,1/\epsilon)$. We complement the sample complexity of our algorithm with a matching lower bound on the number of entries that must be read by any algorithm. We provide experimental results to validate the approximation quality and running time of our algorithm.

Comments:	COLT 2019
Subjects:	Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG)
Cite as:	arXiv:1906.00339 [cs.DS]
	(or arXiv:1906.00339v1 [cs.DS] for this version)
	https://doi.org/10.48550/arXiv.1906.00339

Submission history

From: Tal Wagner [view email]
[v1] Sun, 2 Jun 2019 04:15:17 UTC (530 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.DS

< prev | next >

new | recent | 2019-06

Change to browse by:

cs
cs.LG

References & Citations

DBLP - CS Bibliography

listing | bibtex

Piotr Indyk
Ali Vakilian
Tal Wagner
David P. Woodruff

export BibTeX citation

Computer Science > Data Structures and Algorithms

Title:Sample-Optimal Low-Rank Approximation of Distance Matrices

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Data Structures and Algorithms

Title:Sample-Optimal Low-Rank Approximation of Distance Matrices

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators