Opening book details…
Can I read BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs on EtoBox?
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs by Yang, Zhantao; Feng, Ruili; Yan, Keyu; Wang, Huangji; Wang, Zhicai; Zhu, Shangwen; Zhang, Han; Xiao, Jie; Wu, Pingyu; Zhu, Kai; Chen, Jixuan; Xie, Chen-Wei; Yang, Yue; Zhang, Hongyang; Liu, Yu; Cheng, Fan is a scholarly article available to read on EtoBox.
What is BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs about?
Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions often carry lengthy, intertwined contexts that are difficult to parse and frequently overlook essential cues, posing a great barrier for models like GroundingDINO and SDXL, which lack the strong text encoding and syntax analysis needed to fully leverage dense captions. To address this, we propose BACON, a prompting method that breaks down VLM-generated captions into disentangled, structured elements such as objects, relationships, styles, and themes. This approach not only minimizes confusion from handling complex contexts but also allows for efficient transfer into a JSON dictionary, enabling models without linguistic processing capabilities to easily access key information. We annotated 100,000 image-caption pairs using BACON with GPT-4V and trained an LLaVA captioner on this dataset, enabling it to produce BACON-style captions without relying on costly GPT-4V. Evaluations of overall quality, precision, and recall-as well as user studies-demonstrate that the resulting caption model consistently outperf
- Author
- Yang, Zhantao; Feng, Ruili; Yan, Keyu; Wang, Huangji; Wang, Zhicai; Zhu, Shangwen; Zhang, Han; Xiao, Jie; Wu, Pingyu; Zhu, Kai; Chen, Jixuan; Xie, Chen-Wei; Yang, Yue; Zhang, Hongyang; Liu, Yu; Cheng, Fan
- Published
- 2024
- Language
- EN