-
Nearest Neighbor Sampling for Covariate Shift Adaptation
Authors:
François Portier,
Lionel Truquet,
Ikko Yamane
Abstract:
Many existing covariate shift adaptation methods estimate sample weights given to loss values to mitigate the gap between the source and the target distribution. However, estimating the optimal weights typically involves computationally expensive matrix inversion and hyper-parameter tuning. In this paper, we propose a new covariate shift adaptation method which avoids estimating the weights. The b…
▽ More
Many existing covariate shift adaptation methods estimate sample weights given to loss values to mitigate the gap between the source and the target distribution. However, estimating the optimal weights typically involves computationally expensive matrix inversion and hyper-parameter tuning. In this paper, we propose a new covariate shift adaptation method which avoids estimating the weights. The basic idea is to directly work on unlabeled target data, labeled according to the $k$-nearest neighbors in the source dataset. Our analysis reveals that setting $k = 1$ is an optimal choice. This property removes the necessity of tuning the only hyper-parameter $k$ and leads to a running time quasi-linear in the sample size. Our results include sharp rates of convergence for our estimator, with a tight control of the mean square error and explicit constants. In particular, the variance of our estimators has the same rate of convergence as for standard parametric estimation despite their non-parametric nature. The proposed estimator shares similarities with some matching-based treatment effect estimators used, e.g., in biostatistics, econometrics, and epidemiology. Our experiments show that it achieves drastic reduction in the running time with remarkable accuracy.
△ Less
Submitted 28 June, 2024; v1 submitted 15 December, 2023;
originally announced December 2023.
-
Time series models on the simplex, with an application to dynamic modeling of relative abundance data in Ecology
Authors:
Guillaume Franchi,
Lionel Truquet
Abstract:
Motivated by the dynamic modeling of relative abundance data in ecology, we introduce a general approach to model time series on the simplex. Our approach is based on a general construction of infinite memory models, called chains with complete connections. Simple conditions ensuring the existence of stationary paths are given for the transition kernel that defines the dynamic. We then study in de…
▽ More
Motivated by the dynamic modeling of relative abundance data in ecology, we introduce a general approach to model time series on the simplex. Our approach is based on a general construction of infinite memory models, called chains with complete connections. Simple conditions ensuring the existence of stationary paths are given for the transition kernel that defines the dynamic. We then study in details two specific examples with a Dirichlet and a multivariate logistic-normal conditional distribution. Inference methods can be based on either likelihood maximization or on some convex criteria that can be used to initialize likelihood optimization. We also give an interpretation of our models in term of additive perturbations on the simplex and relative risk ratios which are useful to analyze abundance data in ecosystems. An illustration concerning the evolution of the distribution of three species of Scandinavian birds is provided.
△ Less
Submitted 1 February, 2023;
originally announced February 2023.
-
Multivariate time series models for mixed data
Authors:
Zinsou Max Debaly,
Lionel Truquet
Abstract:
We introduce a general approach for modeling the dynamic of multivariate time series when the data are of mixed type (binary/count/continuous). Our method is quite flexible and conditionally on past values, each coordinate at time $t$ can have a distribution compatible with a standard univariate time series model such as GARCH, ARMA, INGARCH or logistic models whereas past values of the other coor…
▽ More
We introduce a general approach for modeling the dynamic of multivariate time series when the data are of mixed type (binary/count/continuous). Our method is quite flexible and conditionally on past values, each coordinate at time $t$ can have a distribution compatible with a standard univariate time series model such as GARCH, ARMA, INGARCH or logistic models whereas past values of the other coordinates play the role of exogenous covariates in the dynamic. The simultaneous dependence in the multivariate time series can be modeled with a copula. Additional exogenous covariates are also allowed in the dynamic. We first study usual stability properties of these models and then show that autoregressive parameters can be consistently estimated equation-by-equation using a pseudo-maximum likelihood method, leading to a fast implementation even when the number of time series is large. Moreover, we prove consistency results when a parametric copula model is fitted to the time series and in the case of Gaussian copulas, we show that the likelihood estimator of the correlation matrix is strongly consistent. We carefully check all our assumptions for two prototypical examples: a GARCH/INGARCH model and logistic/log-linear INGARCH model. Our results are illustrated with numerical experiments as well as two real data sets.
△ Less
Submitted 2 April, 2021;
originally announced April 2021.