Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers

Jan 31, 2021

Lisa Anne Hendricks, John Mellor, Rosalia Schneider, Jean-Baptiste Alayrac, Aida Nematzadeh

Figure 1 for Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers

Figure 2 for Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers

Figure 3 for Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers

Figure 4 for Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers

Share this with someone who'll enjoy it:

Abstract:Recently multimodal transformer models have gained popularity because their performance on language and vision tasks suggest they learn rich visual-linguistic representations. Focusing on zero-shot image retrieval tasks, we study three important factors which can impact the quality of learned representations: pretraining data, the attention mechanism, and loss functions. By pretraining models on six datasets, we observe that dataset noise and language similarity to our downstream task are important indicators of model performance. Through architectural analysis, we learn that models with a multimodal attention mechanism can outperform deeper models with modality specific attention mechanisms. Finally, we show that successful contrastive losses used in the self-supervised learning literature do not yield similar performance gains when used in multimodal transformers

* pre-print of MIT Press Publication version

View paper on

Share this with someone who'll enjoy it:

Title:Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers

Paper and Code