Search | arXiv e-print repository

FAM: Relative Flatness Aware Minimization

Authors: Linara Adilova, Amr Abourayya, Jianning Li, Amin Dada, Henning Petzka, Jan Egger, Jens Kleesiek, Michael Kamp

Abstract: Flatness of the loss curve around a model at hand has been shown to empirically correlate with its generalization ability. Optimizing for flatness has been proposed as early as 1994 by Hochreiter and Schmidthuber, and was followed by more recent successful sharpness-aware optimization techniques. Their widespread adoption in practice, though, is dubious because of the lack of theoretically grounde… ▽ More Flatness of the loss curve around a model at hand has been shown to empirically correlate with its generalization ability. Optimizing for flatness has been proposed as early as 1994 by Hochreiter and Schmidthuber, and was followed by more recent successful sharpness-aware optimization techniques. Their widespread adoption in practice, though, is dubious because of the lack of theoretically grounded connection between flatness and generalization, in particular in light of the reparameterization curse - certain reparameterizations of a neural network change most flatness measures but do not change generalization. Recent theoretical work suggests that a particular relative flatness measure can be connected to generalization and solves the reparameterization curse. In this paper, we derive a regularizer based on this relative flatness that is easy to compute, fast, efficient, and works with arbitrary loss functions. It requires computing the Hessian only of a single layer of the network, which makes it applicable to large neural networks, and with it avoids an expensive map** of the loss surface in the vicinity of the model. In an extensive empirical evaluation we show that this relative flatness aware minimization (FAM) improves generalization in a multitude of applications and models, both in finetuning and standard training. We make the code available at github. △ Less

Submitted 5 July, 2023; originally announced July 2023.

Comments: Proceedings of the 2nd Annual Workshop on Topology, Algebra, and Geometry in Machine Learning (TAG-ML) at the 40 th International Conference on Machine Learning, Honolulu, Hawaii, USA. 2023

arXiv:2203.01035 [pdf, other]

Discriminating Against Unrealistic Interpolations in Generative Adversarial Networks

Authors: Henning Petzka, Ted Kronvall, Cristian Sminchisescu

Abstract: Interpolations in the latent space of deep generative models is one of the standard tools to synthesize semantically meaningful mixtures of generated samples. As the generator function is non-linear, commonly used linear interpolations in the latent space do not yield the shortest paths in the sample space, resulting in non-smooth interpolations. Recent work has therefore equipped the latent space… ▽ More Interpolations in the latent space of deep generative models is one of the standard tools to synthesize semantically meaningful mixtures of generated samples. As the generator function is non-linear, commonly used linear interpolations in the latent space do not yield the shortest paths in the sample space, resulting in non-smooth interpolations. Recent work has therefore equipped the latent space with a suitable metric to enforce shortest paths on the manifold of generated samples. These are often, however, susceptible of veering away from the manifold of real samples, resulting in smooth but unrealistic generation that requires an additional method to assess the sample quality along paths. Generative Adversarial Networks (GANs), by construction, measure the sample quality using its discriminator network. In this paper, we establish that the discriminator can be used effectively to avoid regions of low sample quality along shortest paths. By reusing the discriminator network to modify the metric on the latent space, we propose a lightweight solution for improved interpolations in pre-trained GANs. △ Less

Submitted 2 March, 2022; originally announced March 2022.

Comments: The first two authors made equal contribution

arXiv:2001.00939 [pdf, other]

Relative Flatness and Generalization

Authors: Henning Petzka, Michael Kamp, Linara Adilova, Cristian Sminchisescu, Mario Boley

Abstract: Flatness of the loss curve is conjectured to be connected to the generalization ability of machine learning models, in particular neural networks. While it has been empirically observed that flatness measures consistently correlate strongly with generalization, it is still an open theoretical problem why and under which circumstances flatness is connected to generalization, in particular in light… ▽ More Flatness of the loss curve is conjectured to be connected to the generalization ability of machine learning models, in particular neural networks. While it has been empirically observed that flatness measures consistently correlate strongly with generalization, it is still an open theoretical problem why and under which circumstances flatness is connected to generalization, in particular in light of reparameterizations that change certain flatness measures but leave generalization unchanged. We investigate the connection between flatness and generalization by relating it to the interpolation from representative data, deriving notions of representativeness, and feature robustness. The notions allow us to rigorously connect flatness and generalization and to identify conditions under which the connection holds. Moreover, they give rise to a novel, but natural relative flatness measure that correlates strongly with generalization, simplifies to ridge regression for ordinary least squares, and solves the reparameterization issue. △ Less

Submitted 4 November, 2021; v1 submitted 3 January, 2020; originally announced January 2020.

Comments: The first two authors made equal contribution; Accepted for publication at NeurIPS 2021; arXiv admin note: substantial text overlap with arXiv:1912.00058

arXiv:1912.00058 [pdf, other]

A Reparameterization-Invariant Flatness Measure for Deep Neural Networks

Authors: Henning Petzka, Linara Adilova, Michael Kamp, Cristian Sminchisescu

Abstract: The performance of deep neural networks is often attributed to their automated, task-related feature construction. It remains an open question, though, why this leads to solutions with good generalization, even in cases where the number of parameters is larger than the number of samples. Back in the 90s, Hochreiter and Schmidhuber observed that flatness of the loss surface around a local minimum c… ▽ More The performance of deep neural networks is often attributed to their automated, task-related feature construction. It remains an open question, though, why this leads to solutions with good generalization, even in cases where the number of parameters is larger than the number of samples. Back in the 90s, Hochreiter and Schmidhuber observed that flatness of the loss surface around a local minimum correlates with low generalization error. For several flatness measures, this correlation has been empirically validated. However, it has recently been shown that existing measures of flatness cannot theoretically be related to generalization due to a lack of invariance with respect to reparameterizations. We propose a natural modification of existing flatness measures that results in invariance to reparameterization. △ Less

Submitted 29 November, 2019; originally announced December 2019.

Comments: 14 pages; accepted at Workshop "Science meets Engineering of Deep Learning", 33rd Conference on Neural Information Processing Systems (NeurIPS 2019)

arXiv:1812.06486 [pdf, other]

Non-attracting Regions of Local Minima in Deep and Wide Neural Networks

Authors: Henning Petzka, Cristian Sminchisescu

Abstract: Understanding the loss surface of neural networks is essential for the design of models with predictable performance and their success in applications. Experimental results suggest that sufficiently deep and wide neural networks are not negatively impacted by suboptimal local minima. Despite recent progress, the reason for this outcome is not fully understood. Could deep networks have very few, if… ▽ More Understanding the loss surface of neural networks is essential for the design of models with predictable performance and their success in applications. Experimental results suggest that sufficiently deep and wide neural networks are not negatively impacted by suboptimal local minima. Despite recent progress, the reason for this outcome is not fully understood. Could deep networks have very few, if at all, suboptimal local optima? or could all of them be equally good? We provide a construction to show that suboptimal local minima (i.e., non-global ones), even though degenerate, exist for fully connected neural networks with sigmoid activation functions. The local minima obtained by our construction belong to a connected set of local solutions that can be escaped from via a non-increasing path on the loss curve. For extremely wide neural networks of decreasing width after the wide layer, we prove that every suboptimal local minimum belongs to such a connected set. This provides a partial explanation for the successful application of deep neural networks. In addition, we also characterize under what conditions the same construction leads to saddle points instead of local minima for deep neural networks. △ Less

Submitted 31 August, 2020; v1 submitted 16 December, 2018; originally announced December 2018.

arXiv:1709.08894 [pdf, other]

On the regularization of Wasserstein GANs

Authors: Henning Petzka, Asja Fischer, Denis Lukovnicov

Abstract: Since their invention, generative adversarial networks (GANs) have become a popular approach for learning to model a distribution of real (unlabeled) data. Convergence problems during training are overcome by Wasserstein GANs which minimize the distance between the model and the empirical distribution in terms of a different metric, but thereby introduce a Lipschitz constraint into the optimizatio… ▽ More Since their invention, generative adversarial networks (GANs) have become a popular approach for learning to model a distribution of real (unlabeled) data. Convergence problems during training are overcome by Wasserstein GANs which minimize the distance between the model and the empirical distribution in terms of a different metric, but thereby introduce a Lipschitz constraint into the optimization problem. A simple way to enforce the Lipschitz constraint on the class of functions, which can be modeled by the neural network, is weight clip**. It was proposed that training can be improved by instead augmenting the loss by a regularization term that penalizes the deviation of the gradient of the critic (as a function of the network's input) from one. We present theoretical arguments why using a weaker regularization term enforcing the Lipschitz constraint is preferable. These arguments are supported by experimental results on toy data sets. △ Less

Submitted 5 March, 2018; v1 submitted 26 September, 2017; originally announced September 2017.

Comments: Published as a conference paper at ICLR 2018. * Henning Petzka and Asja Fischer contributed equally to this work (11 pages +13 pages appendix)

arXiv:1604.06718 [pdf, ps, other]

Perforation conditions and almost algebraic order in Cuntz semigroups

Authors: Ramon Antoine, Francesc Perera, Henning Petzka

Abstract: For a C$^*$-algebra $A$, it is an important problem to determine the Cuntz semigroup $\mathrm{Cu}(A\otimes\mathcal{Z})$ in terms of $\mathrm{Cu}(A)$. We approach this problem from the point of view of semigroup tensor products in the category of abstract Cuntz semigroups, by analysing the passage of significant properties from $\mathrm{Cu}(A)$ to… ▽ More For a C$^*$-algebra $A$, it is an important problem to determine the Cuntz semigroup $\mathrm{Cu}(A\otimes\mathcal{Z})$ in terms of $\mathrm{Cu}(A)$. We approach this problem from the point of view of semigroup tensor products in the category of abstract Cuntz semigroups, by analysing the passage of significant properties from $\mathrm{Cu}(A)$ to $\mathrm{Cu}(A)\otimes_\mathrm{Cu}\mathrm{Cu}(\mathcal{Z})$. We describe the effect of the natural map $\mathrm{Cu}(A)\to\mathrm{Cu}(A)\otimes_\mathrm{Cu}\mathrm{Cu}(\mathcal{Z})$ in the order of $\mathrm{Cu}(A)$, and show that, if $A$ has real rank zero and no elementary subquotients, $\mathrm{Cu}(A)\otimes_\mathrm{Cu}\mathrm{Cu}(\mathcal{Z})$ enjoys the corresponding property of having a dense set of (equivalence classes of) projections. In the simple, nonelementary, real rank zero and stable rank one situation, our investigations lead us to identify almost unperforation for projections with the fact that tensoring with $\mathcal{Z}$ is inert at the level of the Cuntz semigroup. △ Less

Submitted 19 October, 2016; v1 submitted 22 April, 2016; originally announced April 2016.

Comments: 30 pages; minor bugs fixed. This paper has been accepted for publication in Proceedings of the Royal Society of Edinburgh Section A Mathematics and will appear in a revised form subsequent to editorial input by the ICMS/Royal Soc. of Edinburgh. Material on these pages is copyright Cambridge University Press. http://journals.cambridge.org/action/displayJournal?jid=PRM

MSC Class: 06B35; 15A69; 46L05; 46L35; 46L80

arXiv:1603.07199 [pdf, ps, other]

Comparison Properties of the Cuntz semigroup and applications to C*-algebras

Authors: Joan Bosa, Henning Petzka

Abstract: We study comparison properties in the category Cu aiming to lift results to the C*-algebraic setting. We introduce a new comparison property and relate it to both the CFP and $ω$-comparison. We show differences of all properties by providing examples, which suggest that the corona factorization property for C*-algebras might allow for both finite and infinite projections. In addition, we show that… ▽ More We study comparison properties in the category Cu aiming to lift results to the C*-algebraic setting. We introduce a new comparison property and relate it to both the CFP and $ω$-comparison. We show differences of all properties by providing examples, which suggest that the corona factorization property for C*-algebras might allow for both finite and infinite projections. In addition, we show that Rørdam's simple, nuclear C*-algebra with a finite and an infinite projection does not have the CFP. △ Less

Submitted 29 March, 2016; v1 submitted 23 March, 2016; originally announced March 2016.

Comments: 25 pages

MSC Class: 46L05; 46L35

arXiv:1508.02211 [pdf, ps, other]

doi 10.1016/j.jfa.2016.01.013

Corrigendum to "Regularity for stably projectionless, simple C*-algebras"

Authors: Henning Petzka, Aaron Tikuisis

Abstract: An error is identified and corrected in the construction of a non-Z-stable, stably projectionless, simple, nuclear C*-algebra carried out in a paper by the second author. An error is identified and corrected in the construction of a non-Z-stable, stably projectionless, simple, nuclear C*-algebra carried out in a paper by the second author. △ Less

Submitted 10 August, 2015; originally announced August 2015.

Comments: 4 pages

Report number: SOAR-GMJT-01

Journal ref: Journal of Functional Analysis 270 (2016), 2376-2380

arXiv:1305.7495 [pdf, ps, other]

Geometric Structure of Dimension Functions of Certain Continuous Fields

Authors: Ramon Antoine, Joan Bosa, Francesc Perera, Henning Petzka

Abstract: In this paper we study structural properties of the Cuntz semigroup and its functionals for continuous fields of C*-algebras over finite dimensional spaces. In a variety of cases, this leads to an answer to a conjecture posed by Blackadar and Handelman. Enroute to our results, we determine when the stable rank of continuous fields of C*-algebras over one dimensional spaces is one. In this paper we study structural properties of the Cuntz semigroup and its functionals for continuous fields of C*-algebras over finite dimensional spaces. In a variety of cases, this leads to an answer to a conjecture posed by Blackadar and Handelman. Enroute to our results, we determine when the stable rank of continuous fields of C*-algebras over one dimensional spaces is one. △ Less

Submitted 6 June, 2013; v1 submitted 31 May, 2013; originally announced May 2013.

Comments: Minor corrections and additions; Improvement of proposition 1.8

MSC Class: 46L35; 46L80; 06F05

arXiv:1209.4547 [pdf, ps, other]

The Blackadar-Handelman theorem for non-unital C*-algebras

Authors: Henning Petzka

Abstract: A well-known theorem of Blackadar and Handelman states that every unital stably finite C*-algebra has a bounded quasitrace. Rather strong generalizations of stable finiteness to the non-unital case can be obtained by either requiring the multiplier algebra to be stably finite, or alternatively requiring it to be at least stably not properly infinite. This paper deals with the question whether the… ▽ More A well-known theorem of Blackadar and Handelman states that every unital stably finite C*-algebra has a bounded quasitrace. Rather strong generalizations of stable finiteness to the non-unital case can be obtained by either requiring the multiplier algebra to be stably finite, or alternatively requiring it to be at least stably not properly infinite. This paper deals with the question whether the Blackadar-Handelman result can be extended to the non-unital case with respect to these generalizations of stable finiteness. For suitably well-behaved C*-algebras there is a positive result, but none of the non-unital versions holds in full generality. Two examples of C*-algebras are constructed. The first one is a non-unital, stably commutative C*-algebra $A$ that contradicts the weakest possible generalization of the Blackadar-Handelman theorem: The multiplier algebras of all matrix algebras over $A$ are finite, while $A$ has no bounded quasitrace. The second example is a non-unital, simple C*-algebra $B$ that is stably non-stable, i.e. no matrix algebra over $B$ is a stable C*-algebra. In fact, the multiplier algebras over all matrix algebras of this C*-algebra are not properly infinite. Moreover, the C*-algebra $B$ has no bounded quasitrace and therefore gives a simple counterexample to a possible generalization of the Blackadar-Handelman theorem. △ Less

Submitted 21 September, 2012; v1 submitted 20 September, 2012; originally announced September 2012.

Comments: 14 pages

arXiv:1209.4545 [pdf, ps, other]

On certain multiplier projections

Authors: Henning Petzka

Abstract: Let $\MCZK$, denote the multiplier algebra over $\CZK$, the algebra of continuous functions into the compact operators with spectrum the infinite product of two-spheres. We consider multiplier projections in $\MCZK$ of a certain diagonal form. We show that, while for each multiplier projection $Q$ of the special form, we have that $Q(x)\in\BH\setminus \KK$ for all $x\in \prod_{j=1}^\infty S^2$, th… ▽ More Let $\MCZK$, denote the multiplier algebra over $\CZK$, the algebra of continuous functions into the compact operators with spectrum the infinite product of two-spheres. We consider multiplier projections in $\MCZK$ of a certain diagonal form. We show that, while for each multiplier projection $Q$ of the special form, we have that $Q(x)\in\BH\setminus \KK$ for all $x\in \prod_{j=1}^\infty S^2$, the ideal generated by $Q$ in $\MCZK$ might be proper. We further show that the ideal generated by a multiplier projection of the special form is proper if and only if the projection is stably finite. △ Less

Submitted 21 September, 2012; v1 submitted 20 September, 2012; originally announced September 2012.

Comments: 15 pages

Showing 1–12 of 12 results for author: Petzka, H