Representing visual classification as a linear combination of words

Agarwal, Shobhit; Semenov, Yevgeniy R.; Lotter, William

Computer Science > Artificial Intelligence

arXiv:2311.10933v1 (cs)

[Submitted on 18 Nov 2023]

Title:Representing visual classification as a linear combination of words

Authors:Shobhit Agarwal, Yevgeniy R. Semenov, William Lotter

View PDF

Abstract:Explainability is a longstanding challenge in deep learning, especially in high-stakes domains like healthcare. Common explainability methods highlight image regions that drive an AI model's decision. Humans, however, heavily rely on language to convey explanations of not only "where" but "what". Additionally, most explainability approaches focus on explaining individual AI predictions, rather than describing the features used by an AI model in general. The latter would be especially useful for model and dataset auditing, and potentially even knowledge generation as AI is increasingly being used in novel tasks. Here, we present an explainability strategy that uses a vision-language model to identify language-based descriptors of a visual classification task. By leveraging a pre-trained joint embedding space between images and text, our approach estimates a new classification task as a linear combination of words, resulting in a weight for each word that indicates its alignment with the vision-based classifier. We assess our approach using two medical imaging classification tasks, where we find that the resulting descriptors largely align with clinical knowledge despite a lack of domain-specific language training. However, our approach also identifies the potential for 'shortcut connections' in the public datasets used. Towards a functional measure of explainability, we perform a pilot reader study where we find that the AI-identified words can enable non-expert humans to perform a specialized medical task at a non-trivial level. Altogether, our results emphasize the potential of using multimodal foundational models to deliver intuitive, language-based explanations of visual tasks.

Comments:	To be published in the Proceedings of the 3rd Machine Learning for Health symposium, Proceedings of Machine Learning Research (PMLR)
Subjects:	Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
Cite as:	arXiv:2311.10933 [cs.AI]
	(or arXiv:2311.10933v1 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2311.10933

Submission history

From: William Lotter [view email]
[v1] Sat, 18 Nov 2023 02:00:20 UTC (12,235 KB)

Computer Science > Artificial Intelligence

Title:Representing visual classification as a linear combination of words

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:Representing visual classification as a linear combination of words

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators