Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:S-STE: Continuous Pruning Function for Efficient 2:4 Sparse Pre-training

Sep 13, 2024

Yuezhou Hu, Jun Zhu, Jianfei Chen

Figure 1 for S-STE: Continuous Pruning Function for Efficient 2:4 Sparse Pre-training

Figure 2 for S-STE: Continuous Pruning Function for Efficient 2:4 Sparse Pre-training

Figure 3 for S-STE: Continuous Pruning Function for Efficient 2:4 Sparse Pre-training

Figure 4 for S-STE: Continuous Pruning Function for Efficient 2:4 Sparse Pre-training

Share this with someone who'll enjoy it:

Abstract:Training deep neural networks (DNNs) is costly. Fortunately, Nvidia Ampere and Hopper GPUs can accelerate matrix multiplications twice as fast as a dense equivalent by implementing 2:4 sparsity. However, previous STE-based 2:4 pre-training methods (e.g. STE with hard-thresholding, SR-STE) suffer from optimization difficulties because of discontinuous pruning function. In this study, we comprehensively analyse the bottleneck of traditional N:M sparse training and recognize three drawbacks with discontinuity: incorrect descending direction, inability to predict the amount of descent and sparse mask oscillation. In the light of this statement, we propose S-STE, a simple yet powerful 2:4 training method that contains two parts: to continuously project weights to be 2:4 sparse, and to rescale sparse weights with a per-tensor fixed scaling factor. Besides, we adopt minimum-variance unbiased estimation for activation gradient and FP8 quantization for whole process. Results show that our method surpass previous 2:4 pre-training recipes and is comparable even with full parameter models.

View paper on

Share this with someone who'll enjoy it:

Title:S-STE: Continuous Pruning Function for Efficient 2:4 Sparse Pre-training

Paper and Code