Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Iulia Turc

Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Oct 07, 2022

Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu, Julian Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, Kristina Toutanova

Figure 1 for Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Figure 2 for Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Figure 3 for Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Figure 4 for Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Abstract:Visually-situated language is ubiquitous -- sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domain-specific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, image captioning. In addition to the novel pretraining strategy, we introduce a variable-resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state-of-the-art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images.

Via

Access Paper or Ask Questions

Measuring Attribution in Natural Language Generation Models

Dec 23, 2021

Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, David Reitter

Figure 1 for Measuring Attribution in Natural Language Generation Models

Figure 2 for Measuring Attribution in Natural Language Generation Models

Figure 3 for Measuring Attribution in Natural Language Generation Models

Figure 4 for Measuring Attribution in Natural Language Generation Models

Abstract:With recent improvements in natural language generation (NLG) models for various applications, it has become imperative to have the means to identify and evaluate whether NLG output is only sharing verifiable information about the external world. In this work, we present a new evaluation framework entitled Attributable to Identified Sources (AIS) for assessing the output of natural language generation models, when such output pertains to the external world. We first define AIS and introduce a two-stage annotation pipeline for allowing annotators to appropriately evaluate model output according to AIS guidelines. We empirically validate this approach on three generation datasets (two in the conversational QA domain and one in summarization) via human evaluation studies that suggest that AIS could serve as a common framework for measuring whether model-generated statements are supported by underlying sources. We release guidelines for the human evaluation studies.

Via

Access Paper or Ask Questions

Revisiting the Primacy of English in Zero-shot Cross-lingual Transfer

Jun 30, 2021

Iulia Turc, Kenton Lee, Jacob Eisenstein, Ming-Wei Chang, Kristina Toutanova

Figure 1 for Revisiting the Primacy of English in Zero-shot Cross-lingual Transfer

Figure 2 for Revisiting the Primacy of English in Zero-shot Cross-lingual Transfer

Figure 3 for Revisiting the Primacy of English in Zero-shot Cross-lingual Transfer

Figure 4 for Revisiting the Primacy of English in Zero-shot Cross-lingual Transfer

Abstract:Despite their success, large pre-trained multilingual models have not completely alleviated the need for labeled data, which is cumbersome to collect for all target languages. Zero-shot cross-lingual transfer is emerging as a practical solution: pre-trained models later fine-tuned on one transfer language exhibit surprising performance when tested on many target languages. English is the dominant source language for transfer, as reinforced by popular zero-shot benchmarks. However, this default choice has not been systematically vetted. In our study, we compare English against other transfer languages for fine-tuning, on two pre-trained multilingual models (mBERT and mT5) and multiple classification and question answering tasks. We find that other high-resource languages such as German and Russian often transfer more effectively, especially when the set of target languages is diverse or unknown a priori. Unexpectedly, this can be true even when the training sets were automatically translated from English. This finding can have immediate impact on multilingual zero-shot systems, and should inform future benchmark designs.

Via

Access Paper or Ask Questions

The MultiBERTs: BERT Reproductions for Robustness Analysis

Jun 30, 2021

Thibault Sellam, Steve Yadlowsky, Jason Wei, Naomi Saphra, Alexander D'Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das(+2 more)

Figure 1 for The MultiBERTs: BERT Reproductions for Robustness Analysis

Figure 2 for The MultiBERTs: BERT Reproductions for Robustness Analysis

Figure 3 for The MultiBERTs: BERT Reproductions for Robustness Analysis

Figure 4 for The MultiBERTs: BERT Reproductions for Robustness Analysis

Abstract:Experiments with pretrained models such as BERT are often based on a single checkpoint. While the conclusions drawn apply to the artifact (i.e., the particular instance of the model), it is not always clear whether they hold for the more general procedure (which includes the model architecture, training data, initialization scheme, and loss function). Recent work has shown that re-running pretraining can lead to substantially different conclusions about performance, suggesting that alternative evaluations are needed to make principled statements about procedures. To address this question, we introduce MultiBERTs: a set of 25 BERT-base checkpoints, trained with similar hyper-parameters as the original BERT model but differing in random initialization and data shuffling. The aim is to enable researchers to draw robust and statistically justified conclusions about pretraining procedures. The full release includes 25 fully trained checkpoints, as well as statistical guidelines and a code library implementing our recommended hypothesis testing methods. Finally, for five of these models we release a set of 28 intermediate checkpoints in order to support research on learning dynamics.

* Checkpoints and example analyses: http://goo.gle/multiberts

Via

Access Paper or Ask Questions

CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation

Mar 31, 2021

Jonathan H. Clark, Dan Garrette, Iulia Turc, John Wieting

Figure 1 for CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation

Figure 2 for CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation

Figure 3 for CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation

Figure 4 for CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation

Abstract:Pipelined NLP systems have largely been superseded by end-to-end neural modeling, yet nearly all commonly-used models still require an explicit tokenization step. While recent tokenization approaches based on data-derived subword lexicons are less brittle than manually engineered tokenizers, these techniques are not equally suited to all languages, and the use of any fixed vocabulary may limit a model's ability to adapt. In this paper, we present CANINE, a neural encoder that operates directly on character sequences, without explicit tokenization or vocabulary, and a pre-training strategy that operates either directly on characters or optionally uses subwords as a soft inductive bias. To use its finer-grained input effectively and efficiently, CANINE combines downsampling, which reduces the input sequence length, with a deep transformer stack, which encodes context. CANINE outperforms a comparable mBERT model by 2.8 F1 on TyDi QA, a challenging multilingual benchmark, despite having 28% fewer model parameters.

Via

Access Paper or Ask Questions

Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Sep 25, 2019

Iulia Turc, Ming-Wei Chang, Kenton Lee, Kristina Toutanova

Figure 1 for Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Figure 2 for Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Figure 3 for Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Figure 4 for Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Abstract:Recent developments in natural language representations have been accompanied by large and expensive models that leverage vast amounts of general-domain text through self-supervised pre-training. Due to the cost of applying such models to down-stream tasks, several model compression techniques on pre-trained language representations have been proposed (Sun et al., 2019; Sanh, 2019). However, surprisingly, the simple baseline of just pre-training and fine-tuning compact models has been overlooked. In this paper, we first show that pre-training remains important in the context of smaller architectures, and fine-tuning pre-trained compact models can be competitive to more elaborate methods proposed in concurrent work. Starting with pre-trained compact models, we then explore transferring task knowledge from large fine-tuned models through standard knowledge distillation. The resulting simple, yet effective and general algorithm, Pre-trained Distillation, brings further improvements. Through extensive experiments, we more generally explore the interaction between pre-training and distillation under two variables that have been under-studied: model size and properties of unlabeled task data. One surprising observation is that they have a compound effect even when sequentially applied on the same data. To accelerate future research, we will make our 24 pre-trained miniature BERT models publicly available.

* Added comparison to concurrent work

Via

Access Paper or Ask Questions