Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Approximate attention with MLP: a pruning strategy for attention-based model in multivariate time series forecasting

Oct 31, 2024

Suhan Guo, Jiahong Deng, Yi Wei, Hui Dou, Furao Shen, Jian Zhao

Figure 1 for Approximate attention with MLP: a pruning strategy for attention-based model in multivariate time series forecasting

Figure 2 for Approximate attention with MLP: a pruning strategy for attention-based model in multivariate time series forecasting

Figure 3 for Approximate attention with MLP: a pruning strategy for attention-based model in multivariate time series forecasting

Figure 4 for Approximate attention with MLP: a pruning strategy for attention-based model in multivariate time series forecasting

Share this with someone who'll enjoy it:

Abstract:Attention-based architectures have become ubiquitous in time series forecasting tasks, including spatio-temporal (STF) and long-term time series forecasting (LTSF). Yet, our understanding of the reasons for their effectiveness remains limited. This work proposes a new way to understand self-attention networks: we have shown empirically that the entire attention mechanism in the encoder can be reduced to an MLP formed by feedforward, skip-connection, and layer normalization operations for temporal and/or spatial modeling in multivariate time series forecasting. Specifically, the Q, K, and V projection, the attention score calculation, the dot-product between the attention score and the V, and the final projection can be removed from the attention-based networks without significantly degrading the performance that the given network remains the top-tier compared to other SOTA methods. For spatio-temporal networks, the MLP-replace-attention network achieves a reduction in FLOPS of $62.579\%$ with a loss in performance less than $2.5\%$; for LTSF, a reduction in FLOPs of $42.233\%$ with a loss in performance less than $2\%$.

View paper on

Share this with someone who'll enjoy it:

Title:Approximate attention with MLP: a pruning strategy for attention-based model in multivariate time series forecasting

Paper and Code