Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Georgios Mentzos

Accelerating TinyML Inference on Microcontrollers through Approximate Kernels

Sep 25, 2024

Giorgos Armeniakos, Georgios Mentzos, Dimitrios Soudris

Figure 1 for Accelerating TinyML Inference on Microcontrollers through Approximate Kernels

Figure 2 for Accelerating TinyML Inference on Microcontrollers through Approximate Kernels

Figure 3 for Accelerating TinyML Inference on Microcontrollers through Approximate Kernels

Figure 4 for Accelerating TinyML Inference on Microcontrollers through Approximate Kernels

Abstract:The rapid growth of microcontroller-based IoT devices has opened up numerous applications, from smart manufacturing to personalized healthcare. Despite the widespread adoption of energy-efficient microcontroller units (MCUs) in the Tiny Machine Learning (TinyML) domain, they still face significant limitations in terms of performance and memory (RAM, Flash). In this work, we combine approximate computing and software kernel design to accelerate the inference of approximate CNN models on MCUs. Our kernel-based approximation framework firstly unpacks the operands of each convolution layer and then conducts an offline calculation to determine the significance of each operand. Subsequently, through a design space exploration, it employs a computation skipping approximation strategy based on the calculated significance. Our evaluation on an STM32-Nucleo board and 2 popular CNNs trained on the CIFAR-10 dataset shows that, compared to state-of-the-art exact inference, our Pareto optimal solutions can feature on average 21% latency reduction with no degradation in Top-1 classification accuracy, while for lower accuracy requirements, the corresponding reduction becomes even more pronounced.

Via

Access Paper or Ask Questions