Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Diana Borza

An Attention-Based Deep Learning Model for Multiple Pedestrian Attributes Recognition

Apr 02, 2020

Ehsan Yaghoubi, Diana Borza, João Neves, Aruna Kumar, Hugo Proença

Figure 1 for An Attention-Based Deep Learning Model for Multiple Pedestrian Attributes Recognition

Figure 2 for An Attention-Based Deep Learning Model for Multiple Pedestrian Attributes Recognition

Figure 3 for An Attention-Based Deep Learning Model for Multiple Pedestrian Attributes Recognition

Figure 4 for An Attention-Based Deep Learning Model for Multiple Pedestrian Attributes Recognition

Abstract:The automatic characterization of pedestrians in surveillance footage is a tough challenge, particularly when the data is extremely diverse with cluttered backgrounds, and subjects are captured from varying distances, under multiple poses, with partial occlusion. Having observed that the state-of-the-art performance is still unsatisfactory, this paper provides a novel solution to the problem, with two-fold contributions: 1) considering the strong semantic correlation between the different full-body attributes, we propose a multi-task deep model that uses an element-wise multiplication layer to extract more comprehensive feature representations. In practice, this layer serves as a filter to remove irrelevant background features, and is particularly important to handle complex, cluttered data; and 2) we introduce a weighted-sum term to the loss function that not only relativizes the contribution of each task (kind of attributed) but also is crucial for performance improvement in multiple-attribute inference settings. Our experiments were performed on two well-known datasets (RAP and PETA) and point for the superiority of the proposed method with respect to the state-of-the-art. The code is available at https://github.com/Ehsan-Yaghoubi/MAN-PAR-.

* Submitted to Image and Vision Computing journal

Via

Access Paper or Ask Questions

An Implicit Attention Mechanism for Deep Learning Pedestrian Re-identification Frameworks

Jan 31, 2020

Ehsan Yaghoubi, Diana Borza, Aruna Kumar, Hugo Proença

Figure 1 for An Implicit Attention Mechanism for Deep Learning Pedestrian Re-identification Frameworks

Figure 2 for An Implicit Attention Mechanism for Deep Learning Pedestrian Re-identification Frameworks

Figure 3 for An Implicit Attention Mechanism for Deep Learning Pedestrian Re-identification Frameworks

Figure 4 for An Implicit Attention Mechanism for Deep Learning Pedestrian Re-identification Frameworks

Abstract:Attention is defined as the preparedness for the mental selection of certain aspects in a physical environment. In the computer vision domain, this mechanism is of most interest, as it helps to define the segments of an image/video that are critical for obtaining a specific decision. This paper introduces one 'implicit' attentional mechanism for deep learning frameworks, that provides simultaneously: 1) masks-free; and 2) foreground-focused samples for the inference phase. The main idea is to generate synthetic data composed of interleaved segments from the original learning set, while using class information only from specific segments. During the learning phase, the newly generated samples feed the network, keeping their label exclusively consistent with the identity from where the region-of-interest was cropped. Hence, as the model receives images of each identity with inconsistent unwanted areas, it naturally pays the most attention to the label consistent consistent regions, which we observed to be equivalent to learn an effective receptive field. During the test phase, samples are provided without any mask, and the network naturally disregards the detrimental information, which is the insight for the observed improvements in performance. As a proof-of-concept, we consider the challenging problem of pedestrian re-identification and compare the effectiveness of our solution to the state-of-the-art techniques in the well known Richly Annotated Pedestrian (RAP) dataset. The code is available at https://github.com/Ehsan-Yaghoubi/reid-strong-baseline.

* To be published in IEEE international conference on image processing 2020 (ICIP 2020). Codes are available at https://github.com/Ehsan-Yaghoubi/reid-strong-baseline

Via

Access Paper or Ask Questions