DearKD: Data-Efficient Early Knowledge Distillation for Vision Transformers

Chen, Xianing; Cao, Qiong; Zhong, Yujie; Zhang, **g; Gao, Shenghua; Tao, Dacheng

Computer Science > Computer Vision and Pattern Recognition

arXiv:2204.12997 (cs)

[Submitted on 27 Apr 2022 (v1), last revised 28 Apr 2022 (this version, v2)]

Title:DearKD: Data-Efficient Early Knowledge Distillation for Vision Transformers

Authors:Xianing Chen, Qiong Cao, Yujie Zhong, **g Zhang, Shenghua Gao, Dacheng Tao

View PDF

Abstract:Transformers are successfully applied to computer vision due to their powerful modeling capacity with self-attention. However, the excellent performance of transformers heavily depends on enormous training images. Thus, a data-efficient transformer solution is urgently needed. In this work, we propose an early knowledge distillation framework, which is termed as DearKD, to improve the data efficiency required by transformers. Our DearKD is a two-stage framework that first distills the inductive biases from the early intermediate layers of a CNN and then gives the transformer full play by training without distillation. Further, our DearKD can be readily applied to the extreme data-free case where no real images are available. In this case, we propose a boundary-preserving intra-divergence loss based on DeepInversion to further close the performance gap against the full-data counterpart. Extensive experiments on ImageNet, partial ImageNet, data-free setting and other downstream tasks prove the superiority of DearKD over its baselines and state-of-the-art methods.

Comments:	CVPR 2022
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2204.12997 [cs.CV]
	(or arXiv:2204.12997v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2204.12997

Submission history

From: Qiong Cao [view email]
[v1] Wed, 27 Apr 2022 15:11:04 UTC (2,155 KB)
[v2] Thu, 28 Apr 2022 14:36:21 UTC (2,155 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:DearKD: Data-Efficient Early Knowledge Distillation for Vision Transformers

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:DearKD: Data-Efficient Early Knowledge Distillation for Vision Transformers

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators