Skip to main content

Showing 1–1 of 1 results for author: Burda-Lassen, O

.
  1. arXiv:2405.17475  [pdf, other

    cs.CV cs.AI cs.CL cs.LG

    How Culturally Aware are Vision-Language Models?

    Authors: Olena Burda-Lassen, Aman Chadha, Shashank Goswami, Vinija Jain

    Abstract: An image is often said to be worth a thousand words, and certain images can tell rich and insightful stories. Can these stories be told via image captioning? Images from folklore genres, such as mythology, folk dance, cultural signs, and symbols, are vital to every culture. Our research compares the performance of four popular vision-language models (GPT-4V, Gemini Pro Vision, LLaVA, and OpenFlami… ▽ More

    Submitted 24 May, 2024; originally announced May 2024.