Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Exploring Plain Vision Transformer Backbones for Object Detection

Mar 30, 2022

Yanghao Li, Hanzi Mao, Ross Girshick, Kaiming He

Figure 1 for Exploring Plain Vision Transformer Backbones for Object Detection

Figure 2 for Exploring Plain Vision Transformer Backbones for Object Detection

Figure 3 for Exploring Plain Vision Transformer Backbones for Object Detection

Figure 4 for Exploring Plain Vision Transformer Backbones for Object Detection

Share this with someone who'll enjoy it:

Abstract:We explore the plain, non-hierarchical Vision Transformer (ViT) as a backbone network for object detection. This design enables the original ViT architecture to be fine-tuned for object detection without needing to redesign a hierarchical backbone for pre-training. With minimal adaptations for fine-tuning, our plain-backbone detector can achieve competitive results. Surprisingly, we observe: (i) it is sufficient to build a simple feature pyramid from a single-scale feature map (without the common FPN design) and (ii) it is sufficient to use window attention (without shifting) aided with very few cross-window propagation blocks. With plain ViT backbones pre-trained as Masked Autoencoders (MAE), our detector, named ViTDet, can compete with the previous leading methods that were all based on hierarchical backbones, reaching up to 61.3 box AP on the COCO dataset using only ImageNet-1K pre-training. We hope our study will draw attention to research on plain-backbone detectors. Code will be made available.

* Tech report

View paper on

Share this with someone who'll enjoy it:

Title:Exploring Plain Vision Transformer Backbones for Object Detection

Paper and Code