Skip to content

Opening book details…

About this document

Quantifying Image Caption Concreteness For Multimodal Dataset Curation by Anugrah Aidin Yotolembah is a document available to read on EtoBox.

The document introduces a new metric called Image Caption Concreteness (ICC) for evaluating the visual concreteness of image captions without needing an image reference. This metric aims to improve the quality of multimodal datasets by filtering out abstract or subjective captions that can hinder the training of vision-language models. The authors demonstrate that using ICC for dataset curation enhances performance in image captioning and representation learning, especially in resource-constrained settings.

Author
Anugrah Aidin Yotolembah
Language
EN