Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Mohammad Reza Modarres

NormXLogit: The Head-on-Top Never Lies

Nov 25, 2024

Sina Abbasi, Mohammad Reza Modarres, Mohammad Taher Pilehvar

Figure 1 for NormXLogit: The Head-on-Top Never Lies

Figure 2 for NormXLogit: The Head-on-Top Never Lies

Figure 3 for NormXLogit: The Head-on-Top Never Lies

Figure 4 for NormXLogit: The Head-on-Top Never Lies

Abstract:The Transformer architecture has emerged as the dominant choice for building large language models (LLMs). However, with new LLMs emerging on a frequent basis, it is important to consider the potential value of architecture-agnostic approaches that can provide interpretability across a variety of architectures. Despite recent successes in the interpretability of LLMs, many existing approaches rely on complex methods that are often tied to a specific model design and come with a significant computational cost. To address these limitations, we propose a novel technique, called NormXLogit, for assessing the significance of individual input tokens. This method operates based on the input and output representations associated with each token. First, we demonstrate that during the pre-training of LLMs, the norms of word embeddings capture the importance of input tokens. Second, we reveal a significant relationship between a token's importance and the extent to which its representation can resemble the model's final prediction. Through extensive analysis, we show that our approach consistently outperforms existing gradient-based methods in terms of faithfulness. Additionally, our method achieves better performance in layer-wise explanations compared to the most prominent architecture-specific methods.

Via

Access Paper or Ask Questions

RepMatch: Quantifying Cross-Instance Similarities in Representation Space

Oct 12, 2024

Mohammad Reza Modarres, Sina Abbasi, Mohammad Taher Pilehvar

Figure 1 for RepMatch: Quantifying Cross-Instance Similarities in Representation Space

Figure 2 for RepMatch: Quantifying Cross-Instance Similarities in Representation Space

Figure 3 for RepMatch: Quantifying Cross-Instance Similarities in Representation Space

Figure 4 for RepMatch: Quantifying Cross-Instance Similarities in Representation Space

Abstract:Advances in dataset analysis techniques have enabled more sophisticated approaches to analyzing and characterizing training data instances, often categorizing data based on attributes such as ``difficulty''. In this work, we introduce RepMatch, a novel method that characterizes data through the lens of similarity. RepMatch quantifies the similarity between subsets of training instances by comparing the knowledge encoded in models trained on them, overcoming the limitations of existing analysis methods that focus solely on individual instances and are restricted to within-dataset analysis. Our framework allows for a broader evaluation, enabling similarity comparisons across arbitrary subsets of instances, supporting both dataset-to-dataset and instance-to-dataset analyses. We validate the effectiveness of RepMatch across multiple NLP tasks, datasets, and models. Through extensive experimentation, we demonstrate that RepMatch can effectively compare datasets, identify more representative subsets of a dataset (that lead to better performance than randomly selected subsets of equivalent size), and uncover heuristics underlying the construction of some challenge datasets.

Via

Access Paper or Ask Questions