$C^3$: Compositional Counterfactual Contrastive Learning for Video-grounded Dialogues

Le, Hung; Chen, Nancy F.; Hoi, Steven C. H.

Computer Science > Machine Learning

arXiv:2106.08914 (cs)

[Submitted on 16 Jun 2021 (v1), last revised 5 Aug 2023 (this version, v2)]

Title:$C^3$: Compositional Counterfactual Contrastive Learning for Video-grounded Dialogues

Authors:Hung Le, Nancy F. Chen, Steven C.H. Hoi

View PDF

Abstract:Video-grounded dialogue systems aim to integrate video understanding and dialogue understanding to generate responses that are relevant to both the dialogue and video context. Most existing approaches employ deep learning models and have achieved remarkable performance, given the relatively small datasets available. However, the results are partly accomplished by exploiting biases in the datasets rather than develo** multimodal reasoning, resulting in limited generalization. In this paper, we propose a novel approach of Compositional Counterfactual Contrastive Learning ($C^3$) to develop contrastive training between factual and counterfactual samples in video-grounded dialogues. Specifically, we design factual/counterfactual sampling based on the temporal steps in videos and tokens in dialogues and propose contrastive loss functions that exploit object-level or action-level variance. Different from prior approaches, we focus on contrastive hidden state representations among compositional output tokens to optimize the representation space in a generation setting. We achieved promising performance gains on the Audio-Visual Scene-Aware Dialogues (AVSD) benchmark and showed the benefits of our approach in grounding video and dialogue context.

Comments:	24th Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL)
Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2106.08914 [cs.LG]
	(or arXiv:2106.08914v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2106.08914

Submission history

From: Hung Le [view email]
[v1] Wed, 16 Jun 2021 16:05:27 UTC (8,330 KB)
[v2] Sat, 5 Aug 2023 08:04:15 UTC (8,905 KB)

Computer Science > Machine Learning

Title:$C^3$: Compositional Counterfactual Contrastive Learning for Video-grounded Dialogues

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:$C^3$: Compositional Counterfactual Contrastive Learning for Video-grounded Dialogues

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators