Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:UNesT: Local Spatial Representation Learning with Hierarchical Transformer for Efficient Medical Segmentation

Sep 28, 2022

Xin Yu, Qi Yang, Yinchi Zhou, Leon Y. Cai, Riqiang Gao, Ho Hin Lee, Thomas Li, Shunxing Bao, Zhoubing Xu, Thomas A. Lasko(+5 more)

Figure 1 for UNesT: Local Spatial Representation Learning with Hierarchical Transformer for Efficient Medical Segmentation

Figure 2 for UNesT: Local Spatial Representation Learning with Hierarchical Transformer for Efficient Medical Segmentation

Figure 3 for UNesT: Local Spatial Representation Learning with Hierarchical Transformer for Efficient Medical Segmentation

Figure 4 for UNesT: Local Spatial Representation Learning with Hierarchical Transformer for Efficient Medical Segmentation

Share this with someone who'll enjoy it:

Abstract:Transformer-based models, capable of learning better global dependencies, have recently demonstrated exceptional representation learning capabilities in computer vision and medical image analysis. Transformer reformats the image into separate patches and realize global communication via the self-attention mechanism. However, positional information between patches is hard to preserve in such 1D sequences, and loss of it can lead to sub-optimal performance when dealing with large amounts of heterogeneous tissues of various sizes in 3D medical image segmentation. Additionally, current methods are not robust and efficient for heavy-duty medical segmentation tasks such as predicting a large number of tissue classes or modeling globally inter-connected tissues structures. Inspired by the nested hierarchical structures in vision transformer, we proposed a novel 3D medical image segmentation method (UNesT), employing a simplified and faster-converging transformer encoder design that achieves local communication among spatially adjacent patch sequences by aggregating them hierarchically. We extensively validate our method on multiple challenging datasets, consisting anatomies of 133 structures in brain, 14 organs in abdomen, 4 hierarchical components in kidney, and inter-connected kidney tumors). We show that UNesT consistently achieves state-of-the-art performance and evaluate its generalizability and data efficiency. Particularly, the model achieves whole brain segmentation task complete ROI with 133 tissue classes in single network, outperforms prior state-of-the-art method SLANT27 ensembled with 27 network tiles, our model performance increases the mean DSC score of the publicly available Colin and CANDI dataset from 0.7264 to 0.7444 and from 0.6968 to 0.7025, respectively.

* 19 pages, 17 figures. arXiv admin note: text overlap with arXiv:2203.02430

View paper on

Share this with someone who'll enjoy it:

Title:UNesT: Local Spatial Representation Learning with Hierarchical Transformer for Efficient Medical Segmentation

Paper and Code