Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Yunpeng Shen

LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection

Jun 05, 2024

Qiang Chen, Xiangbo Su, Xinyu Zhang, Jian Wang, Jiahui Chen, Yunpeng Shen, Chuchu Han, Ziliang Chen, Weixiang Xu, Fanrong Li(+5 more)

Figure 1 for LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection

Figure 2 for LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection

Figure 3 for LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection

Figure 4 for LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection

Abstract:In this paper, we present a light-weight detection transformer, LW-DETR, which outperforms YOLOs for real-time object detection. The architecture is a simple stack of a ViT encoder, a projector, and a shallow DETR decoder. Our approach leverages recent advanced techniques, such as training-effective techniques, e.g., improved loss and pretraining, and interleaved window and global attentions for reducing the ViT encoder complexity. We improve the ViT encoder by aggregating multi-level feature maps, and the intermediate and final feature maps in the ViT encoder, forming richer feature maps, and introduce window-major feature map organization for improving the efficiency of interleaved attention computation. Experimental results demonstrate that the proposed approach is superior over existing real-time detectors, e.g., YOLO and its variants, on COCO and other benchmark datasets. Code and models are available at (https://github.com/Atten4Vis/LW-DETR).

Via

Access Paper or Ask Questions