Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Ternary MobileNets via Per-Layer Hybrid Filter Banks

Nov 04, 2019

Dibakar Gope, Jesse Beu, Urmish Thakker, Matthew Mattina

Figure 1 for Ternary MobileNets via Per-Layer Hybrid Filter Banks

Figure 2 for Ternary MobileNets via Per-Layer Hybrid Filter Banks

Figure 3 for Ternary MobileNets via Per-Layer Hybrid Filter Banks

Figure 4 for Ternary MobileNets via Per-Layer Hybrid Filter Banks

Share this with someone who'll enjoy it:

Abstract:MobileNets family of computer vision neural networks have fueled tremendous progress in the design and organization of resource-efficient architectures in recent years. New applications with stringent real-time requirements on highly constrained devices require further compression of MobileNets-like already compute-efficient networks. Model quantization is a widely used technique to compress and accelerate neural network inference and prior works have quantized MobileNets to 4-6 bits albeit with a modest to significant drop in accuracy. While quantization to sub-byte values (i.e. precision less than or equal to 8 bits) has been valuable, even further quantization of MobileNets to binary or ternary values is necessary to realize significant energy savings and possibly runtime speedups on specialized hardware, such as ASICs and FPGAs. Under the key observation that convolutional filters at each layer of a deep neural network may respond differently to ternary quantization, we propose a novel quantization method that generates per-layer hybrid filter banks consisting of full-precision and ternary weight filters for MobileNets. The layer-wise hybrid filter banks essentially combine the strengths of full-precision and ternary weight filters to derive a compact, energy-efficient architecture for MobileNets. Using this proposed quantization method, we quantized a substantial portion of weight filters of MobileNets to ternary values resulting in 27.98% savings in energy, and a 51.07% reduction in the model size, while achieving comparable accuracy and no degradation in throughput on specialized hardware in comparison to the baseline full-precision MobileNets.

View paper on

OpenReview

Share this with someone who'll enjoy it:

Title:Ternary MobileNets via Per-Layer Hybrid Filter Banks

Paper and Code