Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Reducing Inference Latency with Concurrent Architectures for Image Recognition

Nov 13, 2020

Ramyad Hadidi, Jiashen Cao, Michael S. Ryoo, Hyesoon Kim

Figure 1 for Reducing Inference Latency with Concurrent Architectures for Image Recognition

Figure 2 for Reducing Inference Latency with Concurrent Architectures for Image Recognition

Figure 3 for Reducing Inference Latency with Concurrent Architectures for Image Recognition

Figure 4 for Reducing Inference Latency with Concurrent Architectures for Image Recognition

Share this with someone who'll enjoy it:

Abstract:Satisfying the high computation demand of modern deep learning architectures is challenging for achieving low inference latency. The current approaches in decreasing latency only increase parallelism within a layer. This is because architectures typically capture a single-chain dependency pattern that prevents efficient distribution with a higher concurrency (i.e., simultaneous execution of one inference among devices). Such single-chain dependencies are so widespread that even implicitly biases recent neural architecture search (NAS) studies. In this visionary paper, we draw attention to an entirely new space of NAS that relaxes the single-chain dependency to provide higher concurrency and distribution opportunities. To quantitatively compare these architectures, we propose a score that encapsulates crucial metrics such as communication, concurrency, and load balancing. Additionally, we propose a new generator and transformation block that consistently deliver superior architectures compared to current state-of-the-art methods. Finally, our preliminary results show that these new architectures reduce the inference latency and deserve more attention.

View paper on

Share this with someone who'll enjoy it:

Title:Reducing Inference Latency with Concurrent Architectures for Image Recognition

Paper and Code