DoGE: Domain Reweighting with Generalization Estimation

Fan, Simin; Pagliardini, Matteo; Jaggi, Martin

Computer Science > Machine Learning

arXiv:2310.15393 (cs)

[Submitted on 23 Oct 2023 (v1), last revised 5 Feb 2024 (this version, v2)]

Title:DoGE: Domain Reweighting with Generalization Estimation

Authors:Simin Fan, Matteo Pagliardini, Martin Jaggi

View PDF

Abstract:The coverage and composition of the pretraining data significantly impacts the generalization ability of Large Language Models (LLMs). Despite its importance, recent LLMs still rely on heuristics and trial and error to increase or reduce the influence of data-domains. We propose DOmain reweighting with Generalization Estimation (DoGE), which optimizes the probability of sampling from each domain (domain weights) in a principled way. Our approach is a two-stage process consisting of (i) training a proxy model to obtain domain weights using a bi-level optimization algorithm; (ii) training a larger base model by sampling training domains according to the learned domain weights. In our experiments, we extensively show how DoGE improves the generalization of the base model to any target data mixture. On the SlimPajama dataset, our base model gets better perplexity and few-shot reasoning accuracies across $6$ tasks compared to baseline methods. Moreover, aiming to generalize to out-of-domain target tasks, which is unseen in the pretraining corpus (OOD domain), DoGE can effectively identify inter-domain dependencies, and consistently achieves better test perplexity on the target domain.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2310.15393 [cs.LG]
	(or arXiv:2310.15393v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2310.15393

Submission history

From: Simin Fan [view email]
[v1] Mon, 23 Oct 2023 22:51:58 UTC (2,595 KB)
[v2] Mon, 5 Feb 2024 16:33:05 UTC (7,719 KB)

Computer Science > Machine Learning

Title:DoGE: Domain Reweighting with Generalization Estimation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:DoGE: Domain Reweighting with Generalization Estimation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators