Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Text-only Synthesis for Image Captioning

May 28, 2024

Qing Zhou, Junlin Huang, Qiang Li, Junyu Gao, Qi Wang

Figure 1 for Text-only Synthesis for Image Captioning

Figure 2 for Text-only Synthesis for Image Captioning

Figure 3 for Text-only Synthesis for Image Captioning

Figure 4 for Text-only Synthesis for Image Captioning

Share this with someone who'll enjoy it:

Abstract:From paired image-text training to text-only training for image captioning, the pursuit of relaxing the requirements for high-cost and large-scale annotation of good quality data remains consistent. In this paper, we propose Text-only Synthesis for Image Captioning (ToCa), which further advances this relaxation with fewer human labor and less computing time. Specifically, we deconstruct caption text into structures and lexical words, which serve as the fundamental components of the caption. By combining different structures and lexical words as inputs to the large language model, massive captions that contain various patterns of lexical words are generated. This method not only approaches the target domain but also surpasses it by generating new captions, thereby enhancing the zero-shot generalization ability of the model. Considering the different levels of data access in the real world, we define three synthesis scenarios: cross-domain synthesis, in-domain synthesis, and data-efficient synthesis. Experiments in these scenarios demonstrate the generalizability, transferability and practicability of ToCa with a nearly 5 CIDEr improvement for zero-shot cross-domain captioning and a maximum increase of over 20 CIDEr for data-efficient captioning.

View paper on

Share this with someone who'll enjoy it:

Title:Text-only Synthesis for Image Captioning

Paper and Code