Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification

Dec 26, 2023

Qinying Liu, Kecheng Zheng, Wei Wu, Zhan Tong, Yu Liu, Wei Chen, Zilei Wang, Yujun Shen

Figure 1 for TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification

Figure 2 for TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification

Figure 3 for TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification

Figure 4 for TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification

Share this with someone who'll enjoy it:

Abstract:The crux of learning vision-language models is to extract semantically aligned information from visual and linguistic data. Existing attempts usually face the problem of coarse alignment, e.g., the vision encoder struggles in localizing an attribute-specified object. In this work, we propose an embarrassingly simple approach to better align image and text features with no need of additional data formats other than image-text pairs. Concretely, given an image and its paired text, we manage to parse objects (e.g., cat) and attributes (e.g., black) from the description, which are highly likely to exist in the image. It is noteworthy that the parsing pipeline is fully automatic and thus enjoys good scalability. With these parsed semantics as supervision signals, we can complement the commonly used image-text contrastive loss with the multi-tag classification loss. Extensive experimental results on a broad suite of semantic segmentation datasets substantiate the average 3.65\% improvement of our framework over existing alternatives. Furthermore, the visualization results indicate that attribute supervision makes vision-language models accurately localize attribute-specified objects. Project page and code can be found at https://qinying-liu.github.io/Tag-Align.

View paper on

Share this with someone who'll enjoy it:

Title:TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification

Paper and Code