Prompt Generation Networks for Input-based Adaptation of Frozen Vision Transformers

Loedeman, Jochem; Stol, Maarten C.; Han, Tengda; Asano, Yuki M.

Computer Science > Computer Vision and Pattern Recognition

arXiv:2210.06466 (cs)

[Submitted on 12 Oct 2022 (v1), last revised 19 Apr 2023 (this version, v2)]

Title:Prompt Generation Networks for Input-based Adaptation of Frozen Vision Transformers

Authors:Jochem Loedeman, Maarten C. Stol, Tengda Han, Yuki M. Asano

View PDF

Abstract:With the introduction of the transformer architecture in computer vision, increasing model scale has been demonstrated as a clear path to achieving performance and robustness gains. However, with model parameter counts reaching the billions, classical finetuning approaches are becoming increasingly limiting and even unfeasible when models become hosted as inference APIs, as in NLP. To this end, visual prompt learning, whereby a model is adapted by learning additional inputs, has emerged as a potential solution for adapting frozen and cloud-hosted models: During inference, this neither requires access to the internals of models' forward pass function, nor requires any post-processing. In this work, we propose the Prompt Generation Network (PGN) that generates high performing, input-dependent prompts by sampling from an end-to-end learned library of tokens. We further introduce the "prompt inversion" trick, with which PGNs can be efficiently trained in a latent space but deployed as strictly input-only prompts for inference. We show the PGN is effective in adapting pre-trained models to various new datasets: It surpasses previous methods by a large margin on 12/12 datasets and even outperforms full-finetuning on 5/12, while requiring 100x less parameters.

Comments:	Tech report, 12 pages. Code: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2210.06466 [cs.CV]
	(or arXiv:2210.06466v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2210.06466

Submission history

From: Tengda Han [view email]
[v1] Wed, 12 Oct 2022 17:59:58 UTC (8,433 KB)
[v2] Wed, 19 Apr 2023 15:48:45 UTC (4,279 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Prompt Generation Networks for Input-based Adaptation of Frozen Vision Transformers

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Prompt Generation Networks for Input-based Adaptation of Frozen Vision Transformers

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators