Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Feb 16, 2024

Dayou Du, Yijia Zhang, Shijie Cao, Jiaqi Guo, Ting Cao, Xiaowen Chu, Ningyi Xu

Figure 1 for BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Figure 2 for BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Figure 3 for BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Figure 4 for BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Share this with someone who'll enjoy it:

Abstract:The upscaling of Large Language Models (LLMs) has yielded impressive advances in natural language processing, yet it also poses significant deployment challenges. Weight quantization has emerged as a widely embraced solution to reduce memory and computational demands. This paper introduces BitDistiller, a framework that synergizes Quantization-Aware Training (QAT) with Knowledge Distillation (KD) to boost the performance of LLMs at ultra-low precisions (sub-4-bit). Specifically, BitDistiller first incorporates a tailored asymmetric quantization and clipping technique to maximally preserve the fidelity of quantized weights, and then proposes a novel Confidence-Aware Kullback-Leibler Divergence (CAKLD) objective, which is employed in a self-distillation manner to enable faster convergence and superior model performance. Empirical evaluations demonstrate that BitDistiller significantly surpasses existing methods in both 3-bit and 2-bit configurations on general language understanding and complex reasoning benchmarks. Notably, BitDistiller is shown to be more cost-effective, demanding fewer data and training resources. The code is available at https://github.com/DD-DuDa/BitDistiller.

View paper on

Share this with someone who'll enjoy it:

Title:BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Paper and Code