Skip to main content

Showing 1–47 of 47 results for author: Duchi, J C

.
  1. arXiv:2403.16336  [pdf, other

    stat.ML cs.LG math.ST stat.ME

    Predictive Inference in Multi-environment Scenarios

    Authors: John C. Duchi, Suyash Gupta, Kuanhao Jiang, Pragya Sur

    Abstract: We address the challenge of constructing valid confidence intervals and sets in problems of prediction across multiple environments. We investigate two types of coverage suitable for these problems, extending the jackknife and split-conformal methods to show how to obtain distribution-free coverage in such non-traditional, hierarchical data-generating scenarios. Our contributions also include exte… ▽ More

    Submitted 24 March, 2024; originally announced March 2024.

  2. arXiv:2402.08794  [pdf, ps, other

    cs.IT math.ST

    An information-theoretic lower bound in time-uniform estimation

    Authors: John C. Duchi, Saminul Haque

    Abstract: We present an information-theoretic lower bound for the problem of parameter estimation with time-uniform coverage guarantees. Via a new a reduction to sequential testing, we obtain stronger lower bounds that capture the hardness of the time-uniform setting. In the case of location model estimation, logistic regression, and exponential family models, our $Ω(\sqrt{n^{-1}\log \log n})$ lower bound i… ▽ More

    Submitted 11 June, 2024; v1 submitted 13 February, 2024; originally announced February 2024.

    Comments: 16 pages

  3. arXiv:2311.01453  [pdf, other

    stat.ML cs.LG stat.ME

    PPI++: Efficient Prediction-Powered Inference

    Authors: Anastasios N. Angelopoulos, John C. Duchi, Tijana Zrnic

    Abstract: We present PPI++: a computationally lightweight methodology for estimation and inference based on a small labeled dataset and a typically much larger dataset of machine-learning predictions. The methods automatically adapt to the quality of available predictions, yielding easy-to-compute confidence sets -- for parameters of any dimensionality -- that always improve on classical intervals using onl… ▽ More

    Submitted 25 March, 2024; v1 submitted 2 November, 2023; originally announced November 2023.

    Comments: Code available at https://github.com/aangelopoulos/ppi_py

  4. arXiv:2202.04166  [pdf, other

    stat.ME stat.ML

    The Lifecycle of a Statistical Model: Model Failure Detection, Identification, and Refitting

    Authors: Alnur Ali, Maxime Cauchois, John C. Duchi

    Abstract: The statistical machine learning community has demonstrated considerable resourcefulness over the years in develo** highly expressive tools for estimation, prediction, and inference. The bedrock assumptions underlying these developments are that the data comes from a fixed population and displays little heterogeneity. But reality is significantly more complex: statistical models now routinely fa… ▽ More

    Submitted 8 February, 2022; originally announced February 2022.

  5. arXiv:2101.02696  [pdf, other

    math.OC cs.LG stat.ML

    Accelerated, Optimal, and Parallel: Some Results on Model-Based Stochastic Optimization

    Authors: Karan Chadha, Gary Cheng, John C. Duchi

    Abstract: We extend the Approximate-Proximal Point (aProx) family of model-based methods for solving stochastic convex optimization problems, including stochastic subgradient, proximal point, and bundle methods, to the minibatch and accelerated setting. To do so, we propose specific model-based algorithms and an acceleration scheme for which we provide non-asymptotic convergence guarantees, which are order-… ▽ More

    Submitted 7 January, 2021; originally announced January 2021.

    Comments: 24 pages, 17 figures

  6. arXiv:2010.05893  [pdf, other

    math.OC cs.LG stat.ML

    Large-Scale Methods for Distributionally Robust Optimization

    Authors: Daniel Levy, Yair Carmon, John C. Duchi, Aaron Sidford

    Abstract: We propose and analyze algorithms for distributionally robust optimization of convex losses with conditional value at risk (CVaR) and $χ^2$ divergence uncertainty sets. We prove that our algorithms require a number of gradient evaluations independent of training set size and number of parameters, making them suitable for large-scale applications. For $χ^2$ uncertainty sets these are the first such… ▽ More

    Submitted 10 December, 2020; v1 submitted 12 October, 2020; originally announced October 2020.

    Comments: 63 pages, NeurIPS 2020

  7. arXiv:2008.04267  [pdf, other

    stat.ML cs.LG stat.ME

    Robust Validation: Confident Predictions Even When Distributions Shift

    Authors: Maxime Cauchois, Suyash Gupta, Alnur Ali, John C. Duchi

    Abstract: While the traditional viewpoint in machine learning and statistics assumes training and testing samples come from the same population, practice belies this fiction. One strategy -- coming from robust statistics and optimization -- is thus to build a model robust to distributional perturbations. In this paper, we take a different approach to describe procedures for robust predictive inference, wher… ▽ More

    Submitted 4 July, 2024; v1 submitted 10 August, 2020; originally announced August 2020.

    Comments: Published in the Journal of the American Statistical Association (JASA 2024)

  8. arXiv:2006.13476  [pdf, other

    cs.LG math.OC stat.ML

    Second-Order Information in Non-Convex Stochastic Optimization: Power and Limitations

    Authors: Yossi Arjevani, Yair Carmon, John C. Duchi, Dylan J. Foster, Ayush Sekhari, Karthik Sridharan

    Abstract: We design an algorithm which finds an $ε$-approximate stationary point (with $\|\nabla F(x)\|\le ε$) using $O(ε^{-3})$ stochastic gradient and Hessian-vector products, matching guarantees that were previously available only under a stronger assumption of access to multiple queries with the same random seed. We prove a lower bound which establishes that this rate is optimal and---surprisingly---tha… ▽ More

    Submitted 24 June, 2020; originally announced June 2020.

    Comments: Accepted to CONFERENCE ON LEARNING THEORY (COLT) 2020

  9. arXiv:2005.10630  [pdf, other

    cs.CR cs.LG stat.ML

    Near Instance-Optimality in Differential Privacy

    Authors: Hilal Asi, John C. Duchi

    Abstract: We develop two notions of instance optimality in differential privacy, inspired by classical statistical theory: one by defining a local minimax risk and the other by considering unbiased mechanisms and analogizing the Cramer-Rao bound, and we show that the local modulus of continuity of the estimand of interest completely determines these quantities. We also develop a complementary collection mec… ▽ More

    Submitted 16 May, 2020; originally announced May 2020.

  10. First-Order Methods for Nonconvex Quadratic Minimization

    Authors: Yair Carmon, John C. Duchi

    Abstract: We consider minimization of indefinite quadratics with either trust-region (norm) constraints or cubic regularization. Despite the nonconvexity of these problems we prove that, under mild assumptions, gradient descent converges to their global solutions, and give a non-asymptotic rate of convergence for the cubic variant. We also consider Krylov subspace solutions and establish sharp convergence g… ▽ More

    Submitted 10 March, 2020; originally announced March 2020.

    Comments: This is a SIAM Review preprint covering our papers "Gradient Descent Finds the Cubic-Regularized Non-Convex Newton Step" (SIOPT, 2019) and "Analysis of Krylov Subspace Solutions of Regularized Nonconvex Quadratic Problems" (NeurIPS, 2018); some materials in Section 6 are new

    Journal ref: SIAM Review, 62(2), 2020, pp. 395--436

  11. arXiv:1912.02365  [pdf, other

    math.OC cs.IT cs.LG stat.ML

    Lower Bounds for Non-Convex Stochastic Optimization

    Authors: Yossi Arjevani, Yair Carmon, John C. Duchi, Dylan J. Foster, Nathan Srebro, Blake Woodworth

    Abstract: We lower bound the complexity of finding $ε$-stationary points (with gradient norm at most $ε$) using stochastic first-order methods. In a well-studied model where algorithms access smooth, potentially non-convex functions through queries to an unbiased stochastic gradient oracle with bounded variance, we prove that (in the worst case) any algorithm requires at least $ε^{-4}$ queries to find an… ▽ More

    Submitted 27 February, 2022; v1 submitted 4 December, 2019; originally announced December 2019.

    Comments: Correction to hard instance dimensions in Theorem 3

  12. arXiv:1909.10455  [pdf, other

    math.OC cs.IT cs.LG stat.ML

    Necessary and Sufficient Geometries for Gradient Methods

    Authors: Daniel Levy, John C. Duchi

    Abstract: We study the impact of the constraint set and gradient geometry on the convergence of online and stochastic methods for convex optimization, providing a characterization of the geometries for which stochastic gradient and adaptive gradient methods are (minimax) optimal. In particular, we show that when the constraint set is quadratically convex, diagonally pre-conditioned stochastic gradient metho… ▽ More

    Submitted 28 October, 2019; v1 submitted 23 September, 2019; originally announced September 2019.

    Comments: 23 pages. To appear at NeurIPS 2019

  13. arXiv:1906.06032  [pdf, other

    cs.LG stat.ML

    Adversarial Training Can Hurt Generalization

    Authors: Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John C. Duchi, Percy Liang

    Abstract: While adversarial training can improve robust accuracy (against an adversary), it sometimes hurts standard accuracy (when there is no adversary). Previous work has studied this tradeoff between standard and robust accuracy, but only in the setting where no predictor performs well on both objectives in the infinite data limit. In this paper, we show that even when the optimal predictor with infinit… ▽ More

    Submitted 26 August, 2019; v1 submitted 14 June, 2019; originally announced June 2019.

  14. arXiv:1905.13736  [pdf, other

    stat.ML cs.CV cs.LG

    Unlabeled Data Improves Adversarial Robustness

    Authors: Yair Carmon, Aditi Raghunathan, Ludwig Schmidt, Percy Liang, John C. Duchi

    Abstract: We demonstrate, theoretically and empirically, that adversarial robustness can significantly benefit from semisupervised learning. Theoretically, we revisit the simple Gaussian model of Schmidt et al. that shows a sample complexity gap between standard and robust classification. We prove that unlabeled data bridges this gap: a simple semisupervised learning procedure (self-training) achieves high… ▽ More

    Submitted 13 January, 2022; v1 submitted 31 May, 2019; originally announced May 2019.

    Comments: Corrected some math typos in the proof of Lemma 1

  15. The importance of better models in stochastic optimization

    Authors: Hilal Asi, John C. Duchi

    Abstract: Standard stochastic optimization methods are brittle, sensitive to stepsize choices and other algorithmic parameters, and they exhibit instability outside of well-behaved families of objectives. To address these challenges, we investigate models for stochastic minimization and learning problems that exhibit better robustness to problem families and algorithmic parameters. With appropriately accura… ▽ More

    Submitted 20 March, 2019; originally announced March 2019.

  16. arXiv:1903.02675  [pdf, other

    cs.LG cs.DS math.OC stat.ML

    A Rank-1 Sketch for Matrix Multiplicative Weights

    Authors: Yair Carmon, John C. Duchi, Aaron Sidford, Kevin Tian

    Abstract: We show that a simple randomized sketch of the matrix multiplicative weight (MMW) update enjoys (in expectation) the same regret bounds as MMW, up to a small constant factor. Unlike MMW, where every step requires full matrix exponentiation, our steps require only a single product of the form $e^A b$, which the Lanczos method approximates efficiently. Our key technique is to view the sketch as a… ▽ More

    Submitted 12 August, 2019; v1 submitted 6 March, 2019; originally announced March 2019.

  17. Mean Estimation from One-Bit Measurements

    Authors: Alon Kipnis, John C. Duchi

    Abstract: We consider the problem of estimating the mean of a symmetric log-concave distribution under the constraint that only a single bit per sample from this distribution is available to the estimator. We study the mean squared error as a function of the sample size (and hence the number of bits). We consider three settings: first, a centralized setting, where an encoder may release $n$ bits given a sam… ▽ More

    Submitted 9 May, 2022; v1 submitted 10 January, 2019; originally announced January 2019.

    Comments: Accepted for publication in the IEEE Transactions on Information Theory

    Journal ref: IEEE Transactions on Information Theory ( Volume: 68, Issue: 9, September 2022)

  18. arXiv:1810.05633  [pdf, other

    math.OC stat.ML

    Stochastic (Approximate) Proximal Point Methods: Convergence, Optimality, and Adaptivity

    Authors: Hilal Asi, John C. Duchi

    Abstract: We develop model-based methods for solving stochastic convex optimization problems, introducing the approximate-proximal point, or aProx, family, which includes stochastic subgradient, proximal point, and bundle methods. When the modeling approaches we propose are appropriately accurate, the methods enjoy stronger convergence and robustness guarantees than classical approaches, even though the mod… ▽ More

    Submitted 3 July, 2019; v1 submitted 12 October, 2018; originally announced October 2018.

    Comments: To appear in SIAM Journal on Optimization

    Journal ref: SIAM Journal on Optimization 29(3), pp. 2257--2290, 2019

  19. arXiv:1806.09222  [pdf, other

    math.OC

    Analysis of Krylov Subspace Solutions of Regularized Nonconvex Quadratic Problems

    Authors: Yair Carmon, John C. Duchi

    Abstract: We provide convergence rates for Krylov subspace solutions to the trust-region and cubic-regularized (nonconvex) quadratic problems. Such solutions may be efficiently computed by the Lanczos method and have long been used in practice. We prove error bounds of the form $1/t^2$ and $e^{-4t/\sqrtκ}$, where $κ$ is a condition number for the problem, and $t$ is the Krylov subspace order (number of Lanc… ▽ More

    Submitted 2 January, 2019; v1 submitted 24 June, 2018; originally announced June 2018.

  20. arXiv:1806.05756  [pdf, other

    math.ST cs.IT

    The Right Complexity Measure in Locally Private Estimation: It is not the Fisher Information

    Authors: John C. Duchi, Feng Ruan

    Abstract: We identify fundamental tradeoffs between statistical utility and privacy under local models of privacy in which data is kept private even from the statistician, providing instance-specific bounds for private estimation and learning problems by develo** the \emph{local minimax risk}. In contrast to approaches based on worst-case (minimax) error, which are conservative, this allows us to evaluate… ▽ More

    Submitted 29 September, 2020; v1 submitted 14 June, 2018; originally announced June 2018.

  21. arXiv:1804.08116  [pdf, ps, other

    math.ST

    A constrained risk inequality for general losses

    Authors: John C. Duchi, Feng Ruan

    Abstract: We provide a general constrained risk inequality that applies to arbitrary non-decreasing losses, extending a result of Brown and Low [Ann. Stat. 1996]. Given two distributions $P_0$ and $P_1$, we find a lower bound for the risk of estimating a parameter $θ(P_1)$ under $P_1$ given an upper bound on the risk of estimating the parameter $θ(P_0)$ under $P_0$. The inequality is a useful pedagogical to… ▽ More

    Submitted 16 April, 2020; v1 submitted 22 April, 2018; originally announced April 2018.

    Comments: 10 pages. This version (v2) adds applications to efficient nonparametric estimation and some pedagogical comments

  22. arXiv:1804.03761  [pdf, other

    stat.ML cs.LG

    Derivative free optimization via repeated classification

    Authors: Tatsunori B. Hashimoto, Steve Yadlowsky, John C. Duchi

    Abstract: We develop an algorithm for minimizing a function using $n$ batched function value measurements at each of $T$ rounds by using classifiers to identify a function's sublevel set. We show that sufficiently accurate classifiers can achieve linear convergence rates, and show that the convergence rate is tied to the difficulty of active learning sublevel sets. Further, we show that the bootstrap is a c… ▽ More

    Submitted 10 April, 2018; originally announced April 2018.

    Comments: At AISTATS2018

  23. arXiv:1711.02226  [pdf, other

    stat.ML

    Unsupervised Transformation Learning via Convex Relaxations

    Authors: Tatsunori B. Hashimoto, John C. Duchi, Percy Liang

    Abstract: Our goal is to extract meaningful transformations from raw images, such as varying the thickness of lines in handwriting or the lighting in a portrait. We propose an unsupervised approach to learn such transformations by attempting to reconstruct an image from a linear combination of transformations of its nearest neighbors. On handwritten digits and celebrity portraits, we show that even with lin… ▽ More

    Submitted 6 November, 2017; originally announced November 2017.

    Comments: To appear at NIPS 2017

  24. arXiv:1711.00841  [pdf, other

    math.OC

    Lower Bounds for Finding Stationary Points II: First-Order Methods

    Authors: Yair Carmon, John C. Duchi, Oliver Hinder, Aaron Sidford

    Abstract: We establish lower bounds on the complexity of finding $ε$-stationary points of smooth, non-convex high-dimensional functions using first-order methods. We prove that deterministic first-order methods, even applied to arbitrarily smooth functions, cannot achieve convergence rates in $ε$ better than $ε^{-8/5}$, which is within $ε^{-1/15}\log\frac{1}ε$ of the best known rate for such methods. Moreov… ▽ More

    Submitted 2 November, 2017; originally announced November 2017.

  25. arXiv:1710.11606  [pdf, other

    math.OC

    Lower Bounds for Finding Stationary Points I

    Authors: Yair Carmon, John C. Duchi, Oliver Hinder, Aaron Sidford

    Abstract: We prove lower bounds on the complexity of finding $ε$-stationary points (points $x$ such that $\|\nabla f(x)\| \le ε$) of smooth, high-dimensional, and potentially non-convex functions $f$. We consider oracle-based complexity measures, where an algorithm is given access to the value and all derivatives of $f$ at a query point $x$. We show that for any (potentially randomized) algorithm… ▽ More

    Submitted 15 August, 2019; v1 submitted 31 October, 2017; originally announced October 2017.

  26. arXiv:1708.00952  [pdf, other

    math.ST

    Mean Estimation from Adaptive One-bit Measurements

    Authors: Alon Kipnis, John C. Duchi

    Abstract: We consider the problem of estimating the mean of a normal distribution under the following constraint: the estimator can access only a single bit from each sample from this distribution. We study the squared error risk in this estimation as a function of the number of samples and one-bit measurements $n$. We consider an adaptive estimation setting where the single-bit sent at step $n$ is a functi… ▽ More

    Submitted 10 October, 2017; v1 submitted 2 August, 2017; originally announced August 2017.

  27. arXiv:1705.02766  [pdf, other

    math.OC

    "Convex Until Proven Guilty": Dimension-Free Acceleration of Gradient Descent on Non-Convex Functions

    Authors: Yair Carmon, Oliver Hinder, John C. Duchi, Aaron Sidford

    Abstract: We develop and analyze a variant of Nesterov's accelerated gradient descent (AGD) for minimization of smooth non-convex functions. We prove that one of two cases occurs: either our AGD variant converges quickly, as if the function was convex, or we produce a certificate that the function is "guilty" of being non-convex. This non-convexity certificate allows us to exploit negative curvature and obt… ▽ More

    Submitted 8 May, 2017; originally announced May 2017.

  28. arXiv:1705.02356  [pdf, other

    math.ST cs.IT math.OC

    Solving (most) of a set of quadratic equalities: Composite optimization for robust phase retrieval

    Authors: John C. Duchi, Feng Ruan

    Abstract: We develop procedures, based on minimization of the composition $f(x) = h(c(x))$ of a convex function $h$ and smooth function $c$, for solving random collections of quadratic equalities, applying our methodology to phase retrieval problems. We show that the prox-linear algorithm we develop can solve phase retrieval problems---even with adversarially faulty measurements---with high probability as s… ▽ More

    Submitted 22 April, 2018; v1 submitted 5 May, 2017; originally announced May 2017.

    Comments: 55 pages, 9 figures

  29. arXiv:1612.00547  [pdf, other

    math.OC cs.DS

    Gradient Descent Finds the Cubic-Regularized Non-Convex Newton Step

    Authors: Yair Carmon, John C. Duchi

    Abstract: We consider the minimization of non-convex quadratic forms regularized by a cubic term, which exhibit multiple saddle points and poor local minima. Nonetheless, we prove that, under mild assumptions, gradient descent approximates the $\textit{global minimum}$ to within $\varepsilon$ accuracy in $O(\varepsilon^{-1}\log(1/\varepsilon))$ steps for large $\varepsilon$ and $O(\log(1/\varepsilon))$ step… ▽ More

    Submitted 30 August, 2022; v1 submitted 1 December, 2016; originally announced December 2016.

    Comments: Corrected Lemma 4.6(iii) and changed the title and some notation to match the journal version of the paper

  30. arXiv:1611.00756  [pdf, other

    math.OC cs.DS

    Accelerated Methods for Non-Convex Optimization

    Authors: Yair Carmon, John C. Duchi, Oliver Hinder, Aaron Sidford

    Abstract: We present an accelerated gradient method for non-convex optimization problems with Lipschitz continuous first and second derivatives. The method requires time $O(ε^{-7/4} \log(1/ ε) )$ to find an $ε$-stationary point, meaning a point $x$ such that $\|\nabla f(x)\| \le ε$. The method improves upon the $O(ε^{-2} )$ complexity of gradient descent and provides the additional second-order guarantee th… ▽ More

    Submitted 2 February, 2017; v1 submitted 2 November, 2016; originally announced November 2016.

  31. arXiv:1603.00126  [pdf, ps, other

    math.ST cs.IT

    Multiclass Classification, Information, Divergence, and Surrogate Risk

    Authors: John C. Duchi, Khashayar Khosravi, Feng Ruan

    Abstract: We provide a unifying view of statistical information measures, multi-way Bayesian hypothesis testing, loss functions for multi-class classification problems, and multi-distribution $f$-divergences, elaborating equivalence results between all of these objects, and extending existing results for binary outcome spaces to more general ones. We consider a generalization of $f$-divergences to multiple… ▽ More

    Submitted 10 September, 2017; v1 submitted 29 February, 2016; originally announced March 2016.

  32. arXiv:1508.00882  [pdf, ps, other

    math.OC stat.ML

    Asynchronous stochastic convex optimization

    Authors: John C. Duchi, Sorathan Chaturapruek, Christopher Ré

    Abstract: We show that asymptotically, completely asynchronous stochastic gradient procedures achieve optimal (even to constant factors) convergence rates for the solution of convex optimization problems under nearly the same conditions required for asymptotic optimality of standard stochastic gradient procedures. Roughly, the noise inherent to the stochastic approximation scheme dominates any noise from as… ▽ More

    Submitted 4 August, 2015; originally announced August 2015.

    Comments: 38 pages, 8 figures

  33. arXiv:1412.4451  [pdf, ps, other

    math.ST cs.IT

    Privacy and Statistical Risk: Formalisms and Minimax Bounds

    Authors: Rina Foygel Barber, John C. Duchi

    Abstract: We explore and compare a variety of definitions for privacy and disclosure limitation in statistical estimation and data analysis, including (approximate) differential privacy, testing-based definitions of privacy, and posterior guarantees on disclosure risk. We give equivalence results between the definitions, shedding light on the relationships between different formalisms for privacy. We also t… ▽ More

    Submitted 14 December, 2014; originally announced December 2014.

    Comments: 29 pages

  34. arXiv:1405.0782  [pdf, ps, other

    cs.IT cs.LG math.ST

    Optimality guarantees for distributed statistical estimation

    Authors: John C. Duchi, Michael I. Jordan, Martin J. Wainwright, Yuchen Zhang

    Abstract: Large data sets often require performing distributed statistical estimation, with a full data set split across multiple machines and limited communication between machines. To study such scenarios, we define and study some refinements of the classical minimax risk that apply to distributed settings, comparing to the performance of estimators with access to the entire data. Lower bounds on these qu… ▽ More

    Submitted 20 June, 2014; v1 submitted 5 May, 2014; originally announced May 2014.

    Comments: 34 pages, 1 figure. Preliminary version appearing in Neural Information Processing Systems 2013 (http://papers.nips.cc/paper/4902-information-theoretic-lower-bounds-for-distributed-statistical-estimation-with-communication-constraints)

  35. arXiv:1312.2139  [pdf, ps, other

    math.OC cs.IT stat.ML

    Optimal rates for zero-order convex optimization: the power of two function evaluations

    Authors: John C. Duchi, Michael I. Jordan, Martin J. Wainwright, Andre Wibisono

    Abstract: We consider derivative-free algorithms for stochastic and non-stochastic convex optimization problems that use only function values rather than gradients. Focusing on non-asymptotic bounds on convergence rates, we show that if pairs of function values are available, algorithms for $d$-dimensional optimization that use gradient estimates based on random perturbations suffer a factor of at most… ▽ More

    Submitted 20 August, 2014; v1 submitted 7 December, 2013; originally announced December 2013.

    Comments: 34 pages

  36. arXiv:1311.2669  [pdf, ps, other

    cs.IT math.ST

    Distance-based and continuum Fano inequalities with applications to statistical estimation

    Authors: John C. Duchi, Martin J. Wainwright

    Abstract: In this technical note, we give two extensions of the classical Fano inequality in information theory. The first extends Fano's inequality to the setting of estimation, providing lower bounds on the probability that an estimator of a discrete quantity is within some distance $t$ of the quantity. The second inequality extends our bound to a continuum setting and provides a volume-based bound. We il… ▽ More

    Submitted 30 December, 2013; v1 submitted 11 November, 2013; originally announced November 2013.

    Comments: 16 pages, 1 figure

  37. arXiv:1305.6000  [pdf, ps, other

    math.ST cs.CR cs.IT

    Local Privacy and Minimax Bounds: Sharp Rates for Probability Estimation

    Authors: John C. Duchi, Michael I. Jordan, Martin J. Wainwright

    Abstract: We provide a detailed study of the estimation of probability distributions---discrete and continuous---in a stringent setting in which data is kept private even from the statistician. We give sharp minimax rates of convergence for estimation in these locally private settings, exhibiting fundamental tradeoffs between privacy and convergence rate, as well as providing tools to allow movement along t… ▽ More

    Submitted 26 May, 2013; originally announced May 2013.

    Comments: 27 pages, 1 figure

  38. arXiv:1305.5029  [pdf, ps, other

    math.ST cs.LG stat.ML

    Divide and Conquer Kernel Ridge Regression: A Distributed Algorithm with Minimax Optimal Rates

    Authors: Yuchen Zhang, John C. Duchi, Martin J. Wainwright

    Abstract: We establish optimal convergence rates for a decomposition-based scalable approach to kernel ridge regression. The method is simple to describe: it randomly partitions a dataset of size N into m subsets of equal size, computes an independent kernel ridge regression estimator for each subset, then averages the local solutions into a global predictor. This partitioning leads to a substantial reducti… ▽ More

    Submitted 29 April, 2014; v1 submitted 22 May, 2013; originally announced May 2013.

  39. arXiv:1302.3203  [pdf, ps, other

    math.ST cs.CR cs.IT

    Local Privacy, Data Processing Inequalities, and Statistical Minimax Rates

    Authors: John C. Duchi, Michael I. Jordan, Martin J. Wainwright

    Abstract: Working under a model of privacy in which data remains private even from the statistician, we study the tradeoff between privacy guarantees and the utility of the resulting statistical estimators. We prove bounds on information-theoretic quantities, including mutual information and Kullback-Leibler divergence, that depend on the privacy guarantees. When combined with standard minimax techniques, i… ▽ More

    Submitted 27 August, 2014; v1 submitted 13 February, 2013; originally announced February 2013.

    Comments: 59 pages, 3 figures

  40. arXiv:1210.2085  [pdf, ps, other

    stat.ML cs.IT cs.LG

    Privacy Aware Learning

    Authors: John C. Duchi, Michael I. Jordan, Martin J. Wainwright

    Abstract: We study statistical risk minimization problems under a privacy model in which the data is kept confidential even from the learner. In this local privacy framework, we establish sharp upper and lower bounds on the convergence rates of statistical estimation procedures. As a consequence, we exhibit a precise tradeoff between the amount of privacy the data preserves and the utility, as measured by c… ▽ More

    Submitted 10 October, 2013; v1 submitted 7 October, 2012; originally announced October 2012.

    Comments: 60 pages

  41. arXiv:1209.4129  [pdf, ps, other

    stat.ML cs.LG stat.CO

    Comunication-Efficient Algorithms for Statistical Optimization

    Authors: Yuchen Zhang, John C. Duchi, Martin Wainwright

    Abstract: We analyze two communication-efficient algorithms for distributed statistical optimization on large-scale data sets. The first algorithm is a standard averaging method that distributes the $N$ data samples evenly to $\nummac$ machines, performs separate minimization on each subset, and then averages the estimates. We provide a sharp analysis of this average mixture algorithm, showing that under a… ▽ More

    Submitted 11 October, 2013; v1 submitted 18 September, 2012; originally announced September 2012.

    Comments: 44 pages, to appear in Journal of Machine Learning Research (JMLR)

  42. arXiv:1208.0129  [pdf, ps, other

    stat.ML cs.LG

    Oracle inequalities for computationally adaptive model selection

    Authors: Alekh Agarwal, Peter L. Bartlett, John C. Duchi

    Abstract: We analyze general model selection procedures using penalized empirical loss minimization under computational constraints. While classical model selection approaches do not consider computational aspects of performing model selection, we argue that any practical model selection procedure must not only trade off estimation and approximation error, but also the computational effort required to compu… ▽ More

    Submitted 1 August, 2012; originally announced August 2012.

  43. arXiv:1204.1688  [pdf, ps, other

    math.ST cs.LG stat.ML

    The asymptotics of ranking algorithms

    Authors: John C. Duchi, Lester Mackey, Michael I. Jordan

    Abstract: We consider the predictive problem of supervised ranking, where the task is to rank sets of candidate items returned in response to queries. Although there exist statistical procedures that come with guarantees of consistency in this setting, these procedures require that individuals provide a complete ranking of all items, which is rarely feasible in practice. Instead, individuals routinely provi… ▽ More

    Submitted 26 November, 2013; v1 submitted 7 April, 2012; originally announced April 2012.

    Comments: Published in at http://dx.doi.org/10.1214/13-AOS1142 the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)

    Report number: IMS-AOS-AOS1142

    Journal ref: Annals of Statistics 2013, Vol. 41, No. 5, 2292-2323

  44. arXiv:1110.2529  [pdf, ps, other

    stat.ML cs.LG math.OC

    The Generalization Ability of Online Algorithms for Dependent Data

    Authors: Alekh Agarwal, John C. Duchi

    Abstract: We study the generalization performance of online learning algorithms trained on samples coming from a dependent source of data. We show that the generalization error of any stable online algorithm concentrates around its regret--an easily computable statistic of the online performance of the algorithm--when the underlying ergodic process is $β$- or $φ$-mixing. We show high probability error bound… ▽ More

    Submitted 6 June, 2012; v1 submitted 11 October, 2011; originally announced October 2011.

    Comments: 26 pages, 1 figure

  45. arXiv:1105.4681  [pdf, ps, other

    math.OC stat.ML

    Ergodic Mirror Descent

    Authors: John C. Duchi, Alekh Agarwal, Mikael Johansson, Michael I. Jordan

    Abstract: We generalize stochastic subgradient descent methods to situations in which we do not receive independent samples from the distribution over which we optimize, but instead receive samples that are coupled over time. We show that as long as the source of randomness is suitably ergodic---it converges quickly enough to a stationary distribution---the method enjoys strong convergence guarantees, both… ▽ More

    Submitted 1 August, 2012; v1 submitted 24 May, 2011; originally announced May 2011.

    Comments: 35 pages, 2 figures

  46. arXiv:1104.5525  [pdf, ps, other

    math.OC stat.ML

    Distributed Delayed Stochastic Optimization

    Authors: Alekh Agarwal, John C. Duchi

    Abstract: We analyze the convergence of gradient-based optimization algorithms that base their updates on delayed stochastic gradient information. The main application of our results is to the development of gradient-based distributed optimization algorithms where a master node performs parameter updates while worker nodes compute stochastic gradients based on local information in parallel, which may give r… ▽ More

    Submitted 28 April, 2011; originally announced April 2011.

    Comments: 27 pages, 4 figures

  47. arXiv:1103.4296  [pdf, ps, other

    math.OC stat.ML

    Randomized Smoothing for Stochastic Optimization

    Authors: John C. Duchi, Peter L. Bartlett, Martin J. Wainwright

    Abstract: We analyze convergence rates of stochastic optimization procedures for non-smooth convex optimization problems. By combining randomized smoothing techniques with accelerated gradient methods, we obtain convergence rates of stochastic optimization procedures, both in expectation and with high probability, that have optimal dependence on the variance of the gradient estimates. To the best of our kno… ▽ More

    Submitted 7 April, 2012; v1 submitted 22 March, 2011; originally announced March 2011.

    Comments: 39 pages, 3 figures