Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:RepViT: Revisiting Mobile CNN From ViT Perspective

Jul 27, 2023

Ao Wang, Hui Chen, Zijia Lin, Hengjun Pu, Guiguang Ding

Figure 1 for RepViT: Revisiting Mobile CNN From ViT Perspective

Figure 2 for RepViT: Revisiting Mobile CNN From ViT Perspective

Figure 3 for RepViT: Revisiting Mobile CNN From ViT Perspective

Figure 4 for RepViT: Revisiting Mobile CNN From ViT Perspective

Share this with someone who'll enjoy it:

Abstract:Recently, lightweight Vision Transformers (ViTs) demonstrate superior performance and lower latency compared with lightweight Convolutional Neural Networks (CNNs) on resource-constrained mobile devices. This improvement is usually attributed to the multi-head self-attention module, which enables the model to learn global representations. However, the architectural disparities between lightweight ViTs and lightweight CNNs have not been adequately examined. In this study, we revisit the efficient design of lightweight CNNs and emphasize their potential for mobile devices. We incrementally enhance the mobile-friendliness of a standard lightweight CNN, specifically MobileNetV3, by integrating the efficient architectural choices of lightweight ViTs. This ends up with a new family of pure lightweight CNNs, namely RepViT. Extensive experiments show that RepViT outperforms existing state-of-the-art lightweight ViTs and exhibits favorable latency in various vision tasks. On ImageNet, RepViT achieves over 80\% top-1 accuracy with nearly 1ms latency on an iPhone 12, which is the first time for a lightweight model, to the best of our knowledge. Our largest model, RepViT-M3, obtains 81.4\% accuracy with only 1.3ms latency. The code and trained models are available at \url{https://github.com/jameslahm/RepViT}.

* 9 pages, 7 figures

View paper on

Share this with someone who'll enjoy it:

Title:RepViT: Revisiting Mobile CNN From ViT Perspective

Paper and Code