Loss Symmetry and Noise Equilibrium of Stochastic Gradient Descent

Ziyin, Liu; Wang, Mingze; Li, Hongchao; Wu, Lei

Computer Science > Machine Learning

arXiv:2402.07193 (cs)

[Submitted on 11 Feb 2024 (v1), last revised 3 Jun 2024 (this version, v2)]

Title:Loss Symmetry and Noise Equilibrium of Stochastic Gradient Descent

Authors:Liu Ziyin, Mingze Wang, Hongchao Li, Lei Wu

View PDF HTML (experimental)

Abstract:Symmetries exist abundantly in the loss function of neural networks. We characterize the learning dynamics of stochastic gradient descent (SGD) when exponential symmetries, a broad subclass of continuous symmetries, exist in the loss function. We establish that when gradient noises do not balance, SGD has the tendency to move the model parameters toward a point where noises from different directions are balanced. Here, a special type of fixed point in the constant directions of the loss function emerges as a candidate for solutions for SGD. As the main theoretical result, we prove that every parameter $\theta$ connects without loss function barrier to a unique noise-balanced fixed point $\theta^*$. The theory implies that the balancing of gradient noise can serve as a novel alternative mechanism for relevant phenomena such as progressive sharpening and flattening and can be applied to understand common practical problems such as representation normalization, matrix factorization, warmup, and formation of latent representations.

Comments:	preprint
Subjects:	Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
Cite as:	arXiv:2402.07193 [cs.LG]
	(or arXiv:2402.07193v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2402.07193

Submission history

From: Liu Ziyin [view email]
[v1] Sun, 11 Feb 2024 13:00:04 UTC (1,106 KB)
[v2] Mon, 3 Jun 2024 17:49:41 UTC (1,404 KB)

Computer Science > Machine Learning

Title:Loss Symmetry and Noise Equilibrium of Stochastic Gradient Descent

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Loss Symmetry and Noise Equilibrium of Stochastic Gradient Descent

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators