Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Learning Cross-lingual Embeddings from Twitter via Distant Supervision

May 17, 2019

Jose Camacho-Collados, Yerai Doval, Eugenio Martínez-Cámara, Luis Espinosa-Anke, Francesco Barbieri, Steven Schockaert

Figure 1 for Learning Cross-lingual Embeddings from Twitter via Distant Supervision

Figure 2 for Learning Cross-lingual Embeddings from Twitter via Distant Supervision

Figure 3 for Learning Cross-lingual Embeddings from Twitter via Distant Supervision

Figure 4 for Learning Cross-lingual Embeddings from Twitter via Distant Supervision

Share this with someone who'll enjoy it:

Abstract:Cross-lingual embeddings represent the meaning of words from different languages in the same vector space. Recent work has shown that it is possible to construct such representations by aligning independently learned monolingual embedding spaces, and that accurate alignments can be obtained even without external bilingual data. In this paper we explore a research direction which has been surprisingly neglected in the literature: leveraging noisy user-generated text to learn cross-lingual embeddings particularly tailored towards social media applications. While the noisiness and informal nature of the social media genre poses additional challenges to cross-lingual embedding methods, we find that it also provides key opportunities due to the abundance of code-switching and the existence of a shared vocabulary of emoji and named entities. Our contribution consists in a very simple post-processing step that exploits these phenomena to significantly improve the performance of state-of-the-art alignment methods.

* 11 pages, 5 tables, 1 figure, 1 appendix

View paper on

Share this with someone who'll enjoy it:

Title:Learning Cross-lingual Embeddings from Twitter via Distant Supervision

Paper and Code