Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Jax Law

Universal Sentence Representation Learning with Conditional Masked Language Model

Dec 29, 2020

Ziyi Yang, Yinfei Yang, Daniel Cer, Jax Law, Eric Darve

Figure 1 for Universal Sentence Representation Learning with Conditional Masked Language Model

Figure 2 for Universal Sentence Representation Learning with Conditional Masked Language Model

Figure 3 for Universal Sentence Representation Learning with Conditional Masked Language Model

Figure 4 for Universal Sentence Representation Learning with Conditional Masked Language Model

Abstract:This paper presents a novel training method, Conditional Masked Language Modeling (CMLM), to effectively learn sentence representations on large scale unlabeled corpora. CMLM integrates sentence representation learning into MLM training by conditioning on the encoded vectors of adjacent sentences. Our English CMLM model achieves state-of-the-art performance on SentEval, even outperforming models learned using (semi-)supervised signals. As a fully unsupervised learning method, CMLM can be conveniently extended to a broad range of languages and domains. We find that a multilingual CMLM model co-trained with bitext retrieval~(BR) and natural language inference~(NLI) tasks outperforms the previous state-of-the-art multilingual models by a large margin. We explore the same language bias of the learned representations, and propose a principle component based approach to remove the language identifying information from the representation while still retaining sentence semantics.

* preprint, updated license

Via

Access Paper or Ask Questions

Multilingual Universal Sentence Encoder for Semantic Retrieval

Jul 09, 2019

Yinfei Yang, Daniel Cer, Amin Ahmad, Mandy Guo, Jax Law, Noah Constant, Gustavo Hernandez Abrego, Steve Yuan, Chris Tar, Yun-Hsuan Sung(+2 more)

Abstract:We introduce two pre-trained retrieval focused multilingual sentence encoding models, respectively based on the Transformer and CNN model architectures. The models embed text from 16 languages into a single semantic space using a multi-task trained dual-encoder that learns tied representations using translation based bridge tasks (Chidambaram al., 2018). The models provide performance that is competitive with the state-of-the-art on: semantic retrieval (SR), translation pair bitext retrieval (BR) and retrieval question answering (ReQA). On English transfer learning tasks, our sentence-level embeddings approach, and in some cases exceed, the performance of monolingual, English only, sentence embedding models. Our models are made available for download on TensorFlow Hub.

* 6 pages, 6 tables, 2 listings, and 1 figure

Via

Access Paper or Ask Questions