MTLE: A Multitask Learning Encoder of Visual Feature Representations for Video and Movie Description

Nina, Oliver; Garcia, Washington; Clouse, Scott; Yilmaz, Alper

Computer Science > Machine Learning

arXiv:1809.07257 (cs)

[Submitted on 19 Sep 2018]

Title:MTLE: A Multitask Learning Encoder of Visual Feature Representations for Video and Movie Description

Authors:Oliver Nina, Washington Garcia, Scott Clouse, Alper Yilmaz

View PDF

Abstract:Learning visual feature representations for video analysis is a daunting task that requires a large amount of training samples and a proper generalization framework. Many of the current state of the art methods for video captioning and movie description rely on simple encoding mechanisms through recurrent neural networks to encode temporal visual information extracted from video data. In this paper, we introduce a novel multitask encoder-decoder framework for automatic semantic description and captioning of video sequences. In contrast to current approaches, our method relies on distinct decoders that train a visual encoder in a multitask fashion. Our system does not depend solely on multiple labels and allows for a lack of training data working even with datasets where only one single annotation is viable per video. Our method shows improved performance over current state of the art methods in several metrics on multi-caption and single-caption datasets. To the best of our knowledge, our method is the first method to use a multitask approach for encoding video features. Our method demonstrates its robustness on the Large Scale Movie Description Challenge (LSMDC) 2017 where our method won the movie description task and its results were ranked among other competitors as the most helpful for the visually impaired.

Comments:	This is a pre-print version of our soon to be released paper
Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
Cite as:	arXiv:1809.07257 [cs.LG]
	(or arXiv:1809.07257v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1809.07257

Submission history

From: Oliver Nina [view email]
[v1] Wed, 19 Sep 2018 15:50:18 UTC (1,875 KB)

Computer Science > Machine Learning

Title:MTLE: A Multitask Learning Encoder of Visual Feature Representations for Video and Movie Description

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:MTLE: A Multitask Learning Encoder of Visual Feature Representations for Video and Movie Description

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators