Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Srinath Srinivasan

Extracting Usable Predictions from Quantized Networks through Uncertainty Quantification for OOD Detection

Mar 02, 2024

Rishi Singhal, Srinath Srinivasan

Abstract:OOD detection has become more pertinent with advances in network design and increased task complexity. Identifying which parts of the data a given network is misclassifying has become as valuable as the network's overall performance. We can compress the model with quantization, but it suffers minor performance loss. The loss of performance further necessitates the need to derive the confidence estimate of the network's predictions. In line with this thinking, we introduce an Uncertainty Quantification(UQ) technique to quantify the uncertainty in the predictions from a pre-trained vision model. We subsequently leverage this information to extract valuable predictions while ignoring the non-confident predictions. We observe that our technique saves up to 80% of ignored samples from being misclassified. The code for the same is available here.

Via

Access Paper or Ask Questions

The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models

Dec 01, 2023

Satya Sai Srinath Namburi, Makesh Sreedhar, Srinath Srinivasan, Frederic Sala

Figure 1 for The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models

Figure 2 for The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models

Figure 3 for The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models

Figure 4 for The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models

Abstract:Compressing large language models (LLMs), often consisting of billions of parameters, provides faster inference, smaller memory footprints, and enables local deployment. Two standard compression techniques are pruning and quantization, with the former eliminating redundant connections in model layers and the latter representing model parameters with fewer bits. The key tradeoff is between the degree of compression and the impact on the quality of the compressed model. Existing research on LLM compression primarily focuses on performance in terms of general metrics like perplexity or downstream task accuracy. More fine-grained metrics, such as those measuring parametric knowledge, remain significantly underexplored. To help bridge this gap, we present a comprehensive analysis across multiple model families (ENCODER, ENCODER-DECODER, and DECODER) using the LAMA and LM-HARNESS benchmarks in order to systematically quantify the effect of commonly employed compression techniques on model performance. A particular focus is on tradeoffs involving parametric knowledge, with the goal of providing practitioners with practical insights to help make informed decisions on compression. We release our codebase1 to enable further research.

* Accepted to EMNLP 2023 Findings

Via

Access Paper or Ask Questions