Protect, Show, Attend and Tell: Empowering Image Captioning Models with Ownership Protection

Lim, Jian Han; Chan, Chee Seng; Ng, Kam Woh; Fan, Lixin; Yang, Qiang

Computer Science > Computer Vision and Pattern Recognition

arXiv:2008.11009 (cs)

[Submitted on 25 Aug 2020 (v1), last revised 31 Aug 2021 (this version, v2)]

Title:Protect, Show, Attend and Tell: Empowering Image Captioning Models with Ownership Protection

Authors:Jian Han Lim, Chee Seng Chan, Kam Woh Ng, Lixin Fan, Qiang Yang

View PDF

Abstract:By and large, existing Intellectual Property (IP) protection on deep neural networks typically i) focus on image classification task only, and ii) follow a standard digital watermarking framework that was conventionally used to protect the ownership of multimedia and video content. This paper demonstrates that the current digital watermarking framework is insufficient to protect image captioning tasks that are often regarded as one of the frontiers AI problems. As a remedy, this paper studies and proposes two different embedding schemes in the hidden memory state of a recurrent neural network to protect the image captioning model. From empirical points, we prove that a forged key will yield an unusable image captioning model, defeating the purpose of infringement. To the best of our knowledge, this work is the first to propose ownership protection on image captioning task. Also, extensive experiments show that the proposed method does not compromise the original image captioning performance on all common captioning metrics on Flickr30k and MS-COCO datasets, and at the same time it is able to withstand both removal and ambiguity attacks. Code is available at this https URL

Comments:	Accepted at Pattern Recognition, 17 pages
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)
Cite as:	arXiv:2008.11009 [cs.CV]
	(or arXiv:2008.11009v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2008.11009

Submission history

From: Chee Seng Chan [view email]
[v1] Tue, 25 Aug 2020 13:48:35 UTC (1,780 KB)
[v2] Tue, 31 Aug 2021 09:36:59 UTC (11,330 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Protect, Show, Attend and Tell: Empowering Image Captioning Models with Ownership Protection

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Protect, Show, Attend and Tell: Empowering Image Captioning Models with Ownership Protection

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators