Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Disentangled Pre-training for Human-Object Interaction Detection

Apr 02, 2024

Zhuolong Li, Xingao Li, Changxing Ding, Xiangmin Xu

Figure 1 for Disentangled Pre-training for Human-Object Interaction Detection

Figure 2 for Disentangled Pre-training for Human-Object Interaction Detection

Figure 3 for Disentangled Pre-training for Human-Object Interaction Detection

Figure 4 for Disentangled Pre-training for Human-Object Interaction Detection

Share this with someone who'll enjoy it:

Abstract:Detecting human-object interaction (HOI) has long been limited by the amount of supervised data available. Recent approaches address this issue by pre-training according to pseudo-labels, which align object regions with HOI triplets parsed from image captions. However, pseudo-labeling is tricky and noisy, making HOI pre-training a complex process. Therefore, we propose an efficient disentangled pre-training method for HOI detection (DP-HOI) to address this problem. First, DP-HOI utilizes object detection and action recognition datasets to pre-train the detection and interaction decoder layers, respectively. Then, we arrange these decoder layers so that the pre-training architecture is consistent with the downstream HOI detection task. This facilitates efficient knowledge transfer. Specifically, the detection decoder identifies reliable human instances in each action recognition dataset image, generates one corresponding query, and feeds it into the interaction decoder for verb classification. Next, we combine the human instance verb predictions in the same image and impose image-level supervision. The DP-HOI structure can be easily adapted to the HOI detection task, enabling effective model parameter initialization. Therefore, it significantly enhances the performance of existing HOI detection models on a broad range of rare categories. The code and pre-trained weight are available at https://github.com/xingaoli/DP-HOI.

* Accepted by CVPR2024

View paper on

Share this with someone who'll enjoy it:

Title:Disentangled Pre-training for Human-Object Interaction Detection

Paper and Code