Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:GSTO: Gated Scale-Transfer Operation for Multi-Scale Feature Learning in Pixel Labeling

Jun 28, 2020

Zhuoying Wang, Yongtao Wang, Zhi Tang, Yangyan Li, Ying Chen, Haibin Ling, Weisi Lin

Figure 1 for GSTO: Gated Scale-Transfer Operation for Multi-Scale Feature Learning in Pixel Labeling

Figure 2 for GSTO: Gated Scale-Transfer Operation for Multi-Scale Feature Learning in Pixel Labeling

Figure 3 for GSTO: Gated Scale-Transfer Operation for Multi-Scale Feature Learning in Pixel Labeling

Figure 4 for GSTO: Gated Scale-Transfer Operation for Multi-Scale Feature Learning in Pixel Labeling

Share this with someone who'll enjoy it:

Abstract:Existing CNN-based methods for pixel labeling heavily depend on multi-scale features to meet the requirements of both semantic comprehension and detail preservation. State-of-the-art pixel labeling neural networks widely exploit conventional scale-transfer operations, i.e., up-sampling and down-sampling to learn multi-scale features. In this work, we find that these operations lead to scale-confused features and suboptimal performance because they are spatial-invariant and directly transit all feature information cross scales without spatial selection. To address this issue, we propose the Gated Scale-Transfer Operation (GSTO) to properly transit spatial-filtered features to another scale. Specifically, GSTO can work either with or without extra supervision. Unsupervised GSTO is learned from the feature itself while the supervised one is guided by the supervised probability matrix. Both forms of GSTO are lightweight and plug-and-play, which can be flexibly integrated into networks or modules for learning better multi-scale features. In particular, by plugging GSTO into HRNet, we get a more powerful backbone (namely GSTO-HRNet) for pixel labeling, and it achieves new state-of-the-art results on the COCO benchmark for human pose estimation and other benchmarks for semantic segmentation including Cityscapes, LIP and Pascal Context, with negligible extra computational cost. Moreover, experiment results demonstrate that GSTO can also significantly boost the performance of multi-scale feature aggregation modules like PPM and ASPP. Code will be made available at https://github.com/VDIGPKU/GSTO.

View paper on

Share this with someone who'll enjoy it:

Title:GSTO: Gated Scale-Transfer Operation for Multi-Scale Feature Learning in Pixel Labeling

Paper and Code