Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Shengqi Zhu

Imagen 3

Aug 13, 2024

Imagen-Team-Google, :, Jason Baldridge, Jakob Bauer, Mukul Bhutani, Nicole Brichtova, Andrew Bunner, Kelvin Chan, Yichang Chen, Sander Dieleman(+240 more)

Abstract:We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred over other state-of-the-art (SOTA) models at the time of evaluation. In addition, we discuss issues around safety and representation, as well as methods we used to minimize the potential harm of our models.

Via

Access Paper or Ask Questions

What We Talk About When We Talk About LMs: Implicit Paradigm Shifts and the Ship of Language Models

Jul 02, 2024

Shengqi Zhu, Jeffrey M. Rzeszotarski

Abstract:The term Language Models (LMs), as a time-specific collection of models of interest, is constantly reinvented, with its referents updated much like the $\textit{Ship of Theseus}$ replaces its parts but remains the same ship in essence. In this paper, we investigate this $\textit{Ship of Language Models}$ problem, wherein scientific evolution takes the form of continuous, implicit retrofits of key existing terms. We seek to initiate a novel perspective of scientific progress, in addition to the more well-studied emergence of new terms. To this end, we construct the data infrastructure based on recent NLP publications. Then, we perform a series of text-based analyses toward a detailed, quantitative understanding of the use of Language Models as a term of art. Our work highlights how systems and theories influence each other in scientific discourse, and we call for attention to the transformation of this Ship that we all are contributing to.

Via

Access Paper or Ask Questions

More than Classification: A Unified Framework for Event Temporal Relation Extraction

May 28, 2023

Quzhe Huang, Yutong Hu, Shengqi Zhu, Yansong Feng, Chang Liu, Dongyan Zhao

Figure 1 for More than Classification: A Unified Framework for Event Temporal Relation Extraction

Figure 2 for More than Classification: A Unified Framework for Event Temporal Relation Extraction

Figure 3 for More than Classification: A Unified Framework for Event Temporal Relation Extraction

Figure 4 for More than Classification: A Unified Framework for Event Temporal Relation Extraction

Abstract:Event temporal relation extraction~(ETRE) is usually formulated as a multi-label classification task, where each type of relation is simply treated as a one-hot label. This formulation ignores the meaning of relations and wipes out their intrinsic dependency. After examining the relation definitions in various ETRE tasks, we observe that all relations can be interpreted using the start and end time points of events. For example, relation \textit{Includes} could be interpreted as event 1 starting no later than event 2 and ending no earlier than event 2. In this paper, we propose a unified event temporal relation extraction framework, which transforms temporal relations into logical expressions of time points and completes the ETRE by predicting the relations between certain time point pairs. Experiments on TB-Dense and MATRES show significant improvements over a strong baseline and outperform the state-of-the-art model by 0.3\% on both datasets. By representing all relations in a unified framework, we can leverage the relations with sufficient data to assist the learning of other relations, thus achieving stable improvement in low-data scenarios. When the relation definitions are changed, our method can quickly adapt to the new ones by simply modifying the logic expressions that map time points to new event relations. The code is released at \url{https://github.com/AndrewZhe/A-Unified-Framework-for-ETRE}.

* ACL 2023 Main Conference

Via

Access Paper or Ask Questions

Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocRED

Apr 17, 2022

Quzhe Huang, Shibo Hao, Yuan Ye, Shengqi Zhu, Yansong Feng, Dongyan Zhao

Figure 1 for Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocRED

Figure 2 for Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocRED

Figure 3 for Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocRED

Figure 4 for Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocRED

Abstract:DocRED is a widely used dataset for document-level relation extraction. In the large-scale annotation, a \textit{recommend-revise} scheme is adopted to reduce the workload. Within this scheme, annotators are provided with candidate relation instances from distant supervision, and they then manually supplement and remove relational facts based on the recommendations. However, when comparing DocRED with a subset relabeled from scratch, we find that this scheme results in a considerable amount of false negative samples and an obvious bias towards popular entities and relations. Furthermore, we observe that the models trained on DocRED have low recall on our relabeled dataset and inherit the same bias in the training data. Through the analysis of annotators' behaviors, we figure out the underlying reason for the problems above: the scheme actually discourages annotators from supplementing adequate instances in the revision phase. We appeal to future research to take into consideration the issues with the recommend-revise scheme when designing new models and annotation schemes. The relabeled dataset is released at \url{https://github.com/AndrewZhe/Revisit-DocRED}, to serve as a more reliable test set of document RE models.

* ACL 2022 Main Conference

Via

Access Paper or Ask Questions

Exploring Distantly-Labeled Rationales in Neural Network Models

Jun 03, 2021

Quzhe Huang, Shengqi Zhu, Yansong Feng, Dongyan Zhao

Figure 1 for Exploring Distantly-Labeled Rationales in Neural Network Models

Figure 2 for Exploring Distantly-Labeled Rationales in Neural Network Models

Figure 3 for Exploring Distantly-Labeled Rationales in Neural Network Models

Figure 4 for Exploring Distantly-Labeled Rationales in Neural Network Models

Abstract:Recent studies strive to incorporate various human rationales into neural networks to improve model performance, but few pay attention to the quality of the rationales. Most existing methods distribute their models' focus to distantly-labeled rationale words entirely and equally, while ignoring the potential important non-rationale words and not distinguishing the importance of different rationale words. In this paper, we propose two novel auxiliary loss functions to make better use of distantly-labeled rationales, which encourage models to maintain their focus on important words beyond labeled rationales (PINs) and alleviate redundant training on non-helpful rationales (NoIRs). Experiments on two representative classification tasks show that our proposed methods can push a classification model to effectively learn crucial clues from non-perfect rationales while maintaining the ability to spread its focus to other unlabeled important words, thus significantly outperform existing methods.

* ACL 2021

Via

Access Paper or Ask Questions

Three Sentences Are All You Need: Local Path Enhanced Document Relation Extraction

Jun 03, 2021

Quzhe Huang, Shengqi Zhu, Yansong Feng, Yuan Ye, Yuxuan Lai, Dongyan Zhao

Figure 1 for Three Sentences Are All You Need: Local Path Enhanced Document Relation Extraction

Figure 2 for Three Sentences Are All You Need: Local Path Enhanced Document Relation Extraction

Figure 3 for Three Sentences Are All You Need: Local Path Enhanced Document Relation Extraction

Figure 4 for Three Sentences Are All You Need: Local Path Enhanced Document Relation Extraction

Abstract:Document-level Relation Extraction (RE) is a more challenging task than sentence RE as it often requires reasoning over multiple sentences. Yet, human annotators usually use a small number of sentences to identify the relationship between a given entity pair. In this paper, we present an embarrassingly simple but effective method to heuristically select evidence sentences for document-level RE, which can be easily combined with BiLSTM to achieve good performance on benchmark datasets, even better than fancy graph neural network based methods. We have released our code at https://github.com/AndrewZhe/Three-Sentences-Are-All-You-Need.

* ACL 2021

Via

Access Paper or Ask Questions