Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Sagnik Ray Choudhury

FutureGen: LLM-RAG Approach to Generate the Future Work of Scientific Article

Mar 20, 2025

Ibrahim Al Azher, Miftahul Jannat Mokarrama, Zhishuai Guo, Sagnik Ray Choudhury, Hamed Alhoori

Abstract:The future work section of a scientific article outlines potential research directions by identifying gaps and limitations of a current study. This section serves as a valuable resource for early-career researchers seeking unexplored areas and experienced researchers looking for new projects or collaborations. In this study, we generate future work suggestions from key sections of a scientific article alongside related papers and analyze how the trends have evolved. We experimented with various Large Language Models (LLMs) and integrated Retrieval-Augmented Generation (RAG) to enhance the generation process. We incorporate a LLM feedback mechanism to improve the quality of the generated content and propose an LLM-as-a-judge approach for evaluation. Our results demonstrated that the RAG-based approach with LLM feedback outperforms other methods evaluated through qualitative and quantitative metrics. Moreover, we conduct a human evaluation to assess the LLM as an extractor and judge. The code and dataset for this project are here, code: HuggingFace

* 19 pages, 5 figures

Via

Access Paper or Ask Questions

Implications of Annotation Artifacts in Edge Probing Test Datasets

Oct 20, 2023

Sagnik Ray Choudhury, Jushaan Kalra

Figure 1 for Implications of Annotation Artifacts in Edge Probing Test Datasets

Figure 2 for Implications of Annotation Artifacts in Edge Probing Test Datasets

Figure 3 for Implications of Annotation Artifacts in Edge Probing Test Datasets

Figure 4 for Implications of Annotation Artifacts in Edge Probing Test Datasets

Abstract:Edge probing tests are classification tasks that test for grammatical knowledge encoded in token representations coming from contextual encoders such as large language models (LLMs). Many LLM encoders have shown high performance in EP tests, leading to conjectures about their ability to encode linguistic knowledge. However, a large body of research claims that the tests necessarily do not measure the LLM's capacity to encode knowledge, but rather reflect the classifiers' ability to learn the problem. Much of this criticism stems from the fact that often the classifiers have very similar accuracy when an LLM vs a random encoder is used. Consequently, several modifications to the tests have been suggested, including information theoretic probes. We show that commonly used edge probing test datasets have various biases including memorization. When these biases are removed, the LLM encoders do show a significant difference from the random ones, even with the simple non-information theoretic probes.

* Accepted CoNLL 2023, code: https://github.com/Josh1108/EPtest.git

Via

Access Paper or Ask Questions

Explaining Interactions Between Text Spans

Oct 20, 2023

Sagnik Ray Choudhury, Pepa Atanasova, Isabelle Augenstein

Abstract:Reasoning over spans of tokens from different parts of the input is essential for natural language understanding (NLU) tasks such as fact-checking (FC), machine reading comprehension (MRC) or natural language inference (NLI). However, existing highlight-based explanations primarily focus on identifying individual important tokens or interactions only between adjacent tokens or tuples of tokens. Most notably, there is a lack of annotations capturing the human decision-making process w.r.t. the necessary interactions for informed decision-making in such tasks. To bridge this gap, we introduce SpanEx, a multi-annotator dataset of human span interaction explanations for two NLU tasks: NLI and FC. We then investigate the decision-making processes of multiple fine-tuned large language models in terms of the employed connections between spans in separate parts of the input and compare them to the human reasoning processes. Finally, we present a novel community detection based unsupervised method to extract such interaction explanations from a model's inner workings.

* code: https://github.com/copenlu/spanex , dataset: https://huggingface.co/datasets/copenlu/spanex. Accepted EMNLP 2023

Via

Access Paper or Ask Questions

Machine Reading, Fast and Slow: When Do Models "Understand" Language?

Sep 15, 2022

Sagnik Ray Choudhury, Anna Rogers, Isabelle Augenstein

Figure 1 for Machine Reading, Fast and Slow: When Do Models "Understand" Language?

Figure 2 for Machine Reading, Fast and Slow: When Do Models "Understand" Language?

Figure 3 for Machine Reading, Fast and Slow: When Do Models "Understand" Language?

Figure 4 for Machine Reading, Fast and Slow: When Do Models "Understand" Language?

Abstract:Two of the most fundamental challenges in Natural Language Understanding (NLU) at present are: (a) how to establish whether deep learning-based models score highly on NLU benchmarks for the 'right' reasons; and (b) to understand what those reasons would even be. We investigate the behavior of reading comprehension models with respect to two linguistic 'skills': coreference resolution and comparison. We propose a definition for the reasoning steps expected from a system that would be 'reading slowly', and compare that with the behavior of five models of the BERT family of various sizes, observed through saliency scores and counterfactual explanations. We find that for comparison (but not coreference) the systems based on larger encoders are more likely to rely on the 'right' information, but even they struggle with generalization, suggesting that they still learn specific lexical patterns rather than the general principles of comparison.

* Accepted COLING 2022

Via

Access Paper or Ask Questions

Can Edge Probing Tasks Reveal Linguistic Knowledge in QA Models?

Sep 18, 2021

Sagnik Ray Choudhury, Nikita Bhutani, Isabelle Augenstein

Figure 1 for Can Edge Probing Tasks Reveal Linguistic Knowledge in QA Models?

Figure 2 for Can Edge Probing Tasks Reveal Linguistic Knowledge in QA Models?

Figure 3 for Can Edge Probing Tasks Reveal Linguistic Knowledge in QA Models?

Figure 4 for Can Edge Probing Tasks Reveal Linguistic Knowledge in QA Models?

Abstract:There have been many efforts to try to understand what gram-matical knowledge (e.g., ability to understand the part of speech of a token) is encoded in large pre-trained language models (LM). This is done through 'Edge Probing' (EP) tests: simple ML models that predict the grammatical properties ofa span (whether it has a particular part of speech) using only the LM's token representations. However, most NLP applications use fine-tuned LMs. Here, we ask: if a LM is fine-tuned, does the encoding of linguistic information in it change, as measured by EP tests? Conducting experiments on multiple question-answering (QA) datasets, we answer that question negatively: the EP test results do not change significantly when the fine-tuned QA model performs well or in adversarial situations where the model is forced to learn wrong correlations. However, a critical analysis of the EP task datasets reveals that EP models may rely on spurious correlations to make predictions. This indicates even if fine-tuning changes the encoding of such knowledge, the EP tests might fail to measure it.

Via

Access Paper or Ask Questions

Intent Features for Rich Natural Language Understanding

Apr 21, 2021

Brian Lester, Sagnik Ray Choudhury, Rashmi Prasad, Srinivas Bangalore

Figure 1 for Intent Features for Rich Natural Language Understanding

Figure 2 for Intent Features for Rich Natural Language Understanding

Figure 3 for Intent Features for Rich Natural Language Understanding

Figure 4 for Intent Features for Rich Natural Language Understanding

Abstract:Complex natural language understanding modules in dialog systems have a richer understanding of user utterances, and thus are critical in providing a better user experience. However, these models are often created from scratch, for specific clients and use cases, and require the annotation of large datasets. This encourages the sharing of annotated data across multiple clients. To facilitate this we introduce the idea of intent features: domain and topic agnostic properties of intents that can be learned from the syntactic cues only, and hence can be shared. We introduce a new neural network architecture, the Global-Local model, that shows significant improvement over strong baselines for identifying these features in a deployed, multi-intent natural language understanding module, and, more generally, in a classification setting where a part of an utterance has to be classified utilizing the whole context.

* Camera-ready for NAACL 2021 Industry Track

Via

Access Paper or Ask Questions

Quantifying Gender Bias Towards Politicians in Cross-Lingual Language Models

Apr 15, 2021

Karolina Stańczak, Sagnik Ray Choudhury, Tiago Pimentel, Ryan Cotterell, Isabelle Augenstein

Figure 1 for Quantifying Gender Bias Towards Politicians in Cross-Lingual Language Models

Figure 2 for Quantifying Gender Bias Towards Politicians in Cross-Lingual Language Models

Figure 3 for Quantifying Gender Bias Towards Politicians in Cross-Lingual Language Models

Figure 4 for Quantifying Gender Bias Towards Politicians in Cross-Lingual Language Models

Abstract:While the prevalence of large pre-trained language models has led to significant improvements in the performance of NLP systems, recent research has demonstrated that these models inherit societal biases extant in natural language. In this paper, we explore a simple method to probe pre-trained language models for gender bias, which we use to effect a multi-lingual study of gender bias towards politicians. We construct a dataset of 250k politicians from most countries in the world and quantify adjective and verb usage around those politicians' names as a function of their gender. We conduct our study in 7 languages across 6 different language modeling architectures. Our results demonstrate that stance towards politicians in pre-trained language models is highly dependent on the language used. Finally, contrary to previous findings, our study suggests that larger language models do not tend to be significantly more gender-biased than smaller ones.

Via

Access Paper or Ask Questions

Multiple Word Embeddings for Increased Diversity of Representation

Oct 09, 2020

Brian Lester, Daniel Pressel, Amy Hemmeter, Sagnik Ray Choudhury, Srinivas Bangalore

Figure 1 for Multiple Word Embeddings for Increased Diversity of Representation

Figure 2 for Multiple Word Embeddings for Increased Diversity of Representation

Figure 3 for Multiple Word Embeddings for Increased Diversity of Representation

Figure 4 for Multiple Word Embeddings for Increased Diversity of Representation

Abstract:Most state-of-the-art models in natural language processing (NLP) are neural models built on top of large, pre-trained, contextual language models that generate representations of words in context and are fine-tuned for the task at hand. The improvements afforded by these "contextual embeddings" come with a high computational cost. In this work, we explore a simple technique that substantially and consistently improves performance over a strong baseline with negligible increase in run time. We concatenate multiple pre-trained embeddings to strengthen our representation of words. We show that this concatenation technique works across many tasks, datasets, and model types. We analyze aspects of pre-trained embedding similarity and vocabulary coverage and find that the representational diversity between different pre-trained embeddings is the driving force of why this technique works. We provide open source implementations of our models in both TensorFlow and PyTorch.

* arXiv admin note: text overlap with arXiv:2001.01167

Via

Access Paper or Ask Questions

Constrained Decoding for Computationally Efficient Named Entity Recognition Taggers

Oct 09, 2020

Brian Lester, Daniel Pressel, Amy Hemmeter, Sagnik Ray Choudhury, Srinivas Bangalore

Figure 1 for Constrained Decoding for Computationally Efficient Named Entity Recognition Taggers

Figure 2 for Constrained Decoding for Computationally Efficient Named Entity Recognition Taggers

Figure 3 for Constrained Decoding for Computationally Efficient Named Entity Recognition Taggers

Figure 4 for Constrained Decoding for Computationally Efficient Named Entity Recognition Taggers

Abstract:Current state-of-the-art models for named entity recognition (NER) are neural models with a conditional random field (CRF) as the final layer. Entities are represented as per-token labels with a special structure in order to decode them into spans. Current work eschews prior knowledge of how the span encoding scheme works and relies on the CRF learning which transitions are illegal and which are not to facilitate global coherence. We find that by constraining the output to suppress illegal transitions we can train a tagger with a cross-entropy loss twice as fast as a CRF with differences in F1 that are statistically insignificant, effectively eliminating the need for a CRF. We analyze the dynamics of tag co-occurrence to explain when these constraints are most effective and provide open source implementations of our tagger in both PyTorch and TensorFlow.

* Findings of EMNLP 2020

Via

Access Paper or Ask Questions

Computationally Efficient NER Taggers with Combined Embeddings and Constrained Decoding

Jan 05, 2020

Brian Lester, Daniel Pressel, Amy Hemmeter, Sagnik Ray Choudhury

Figure 1 for Computationally Efficient NER Taggers with Combined Embeddings and Constrained Decoding

Figure 2 for Computationally Efficient NER Taggers with Combined Embeddings and Constrained Decoding

Figure 3 for Computationally Efficient NER Taggers with Combined Embeddings and Constrained Decoding

Abstract:Current State-of-the-Art models in Named Entity Recognition (NER) are neural models with a Conditional Random Field (CRF) as the final network layer, and pre-trained "contextual embeddings". The CRF layer is used to facilitate global coherence between labels, and the contextual embeddings provide a better representation of words in context. However, both of these improvements come at a high computational cost. In this work, we explore two simple techniques that substantially improve NER performance over a strong baseline with negligible cost. First, we use multiple pre-trained embeddings as word representations via concatenation. Second, we constrain the tagger, trained using a cross-entropy loss, during decoding to eliminate illegal transitions. While training a tagger on CoNLL 2003 we find a $786$\% speed-up over a contextual embeddings-based tagger without sacrificing strong performance. We also show that the concatenation technique works across multiple tasks and datasets. We analyze aspects of similarity and coverage between pre-trained embeddings and the dynamics of tag co-occurrence to explain why these techniques work. We provide an open source implementation of our tagger using these techniques in three popular deep learning frameworks --- TensorFlow, Pytorch, and DyNet.

Via

Access Paper or Ask Questions