Learning Better Features for Face Detection with Feature Fusion and Segmentation Supervision

Tian, Wanxin; Wang, Zixuan; Shen, Haifeng; Deng, Weihong; Meng, Yi**; Chen, Binghui; Zhang, Xiubao; Zhao, Yuan; Huang, Xiehe

Computer Science > Computer Vision and Pattern Recognition

arXiv:1811.08557 (cs)

[Submitted on 20 Nov 2018 (v1), last revised 25 Apr 2019 (this version, v3)]

Title:Learning Better Features for Face Detection with Feature Fusion and Segmentation Supervision

Authors:Wanxin Tian, Zixuan Wang, Haifeng Shen, Weihong Deng, Yi** Meng, Binghui Chen, Xiubao Zhang, Yuan Zhao, Xiehe Huang

View PDF

Abstract:The performance of face detectors has been largely improved with the development of convolutional neural network. However, it remains challenging for face detectors to detect tiny, occluded or blurry faces. Besides, most face detectors can't locate face's position precisely and can't achieve high Intersection-over-Union (IoU) scores. We assume that problems inside are inadequate use of supervision information and imbalance between semantics and details at all level feature maps in CNN even with Feature Pyramid Networks (FPN). In this paper, we present a novel single-shot face detection network, named DF$^2$S$^2$ (Detection with Feature Fusion and Segmentation Supervision), which introduces a more effective feature fusion pyramid and a more efficient segmentation branch on ResNet-50 to handle mentioned problems. Specifically, inspired by FPN and SENet, we apply semantic information from higher-level feature maps as contextual cues to augment low-level feature maps via a spatial and channel-wise attention style, preventing details from being covered by too much semantics and making semantics and details complement each other. We further propose a semantic segmentation branch to best utilize detection supervision information meanwhile applying attention mechanism in a self-supervised manner. The segmentation branch is supervised by weak segmentation ground-truth (no extra annotation is required) in a hierarchical manner, deprecated in the inference time so it wouldn't compromise the inference speed. We evaluate our model on WIDER FACE dataset and achieved state-of-art results.

Comments:	10 pages, 4 figures. arXiv admin note: text overlap with arXiv:1711.07246, arXiv:1712.00721 by other authors
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1811.08557 [cs.CV]
	(or arXiv:1811.08557v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1811.08557

Submission history

From: Wanxin Tian [view email]
[v1] Tue, 20 Nov 2018 07:31:03 UTC (4,882 KB)
[v2] Wed, 24 Apr 2019 03:27:57 UTC (4,938 KB)
[v3] Thu, 25 Apr 2019 10:10:12 UTC (5,088 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Learning Better Features for Face Detection with Feature Fusion and Segmentation Supervision

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Learning Better Features for Face Detection with Feature Fusion and Segmentation Supervision

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators