Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Matthew J. Marinella

Analog fast Fourier transforms for scalable and efficient signal processing

Sep 27, 2024

T. Patrick Xiao, Ben Feinberg, David K. Richardson, Matthew Cannon, Harsha Medu, Vineet Agrawal, Matthew J. Marinella, Sapan Agarwal, Christopher H. Bennett

Figure 1 for Analog fast Fourier transforms for scalable and efficient signal processing

Figure 2 for Analog fast Fourier transforms for scalable and efficient signal processing

Figure 3 for Analog fast Fourier transforms for scalable and efficient signal processing

Figure 4 for Analog fast Fourier transforms for scalable and efficient signal processing

Abstract:Edge devices are being deployed at increasing volumes to sense and act on information from the physical world. The discrete Fourier transform (DFT) is often necessary to make this sensed data suitable for further processing $\unicode{x2013}$ such as by artificial intelligence (AI) algorithms $\unicode{x2013}$ and for transmission over communication networks. Analog in-memory computing has been shown to be a fast and energy-efficient solution for processing edge AI workloads, but not for Fourier transforms. This is because of the existence of the fast Fourier transform (FFT) algorithm, which enormously reduces the complexity of the DFT but has so far belonged only to digital processors. Here, we show that the FFT can be mapped to analog in-memory computing systems, enabling them to efficiently scale to arbitrarily large Fourier transforms without requiring large sizes or large numbers of non-volatile memory arrays. We experimentally demonstrate analog FFTs on 1D audio and 2D image signals, using a large-scale charge-trapping memory array with precisely tunable, low-conductance analog states. The scalability of both the new analog FFT approach and the charge-trapping memory device is leveraged to compute a 65,536-point analog DFT, a scale that is otherwise inaccessible by analog systems and which is $>$1000$\times$ larger than any previous analog DFT demonstration. The analog FFT also provides more numerically precise DFTs with greater tolerance to device and circuit non-idealities than a direct matrix-vector multiplication approach. We show that the extension of the FFT algorithm to analog in-memory processors leads to design considerations that differ markedly from digital implementations, and that analog Fourier transforms have a substantial power efficiency advantage at all size scales over FFTs implemented on state-of-the-art digital hardware.

Via

Access Paper or Ask Questions

Shape-Dependent Multi-Weight Magnetic Artificial Synapses for Neuromorphic Computing

Nov 22, 2021

Thomas Leonard, Samuel Liu, Mahshid Alamdar, Can Cui, Otitoaleke G. Akinola, Lin Xue, T. Patrick Xiao, Joseph S. Friedman, Matthew J. Marinella, Christopher H. Bennett(+1 more)

Figure 1 for Shape-Dependent Multi-Weight Magnetic Artificial Synapses for Neuromorphic Computing

Figure 2 for Shape-Dependent Multi-Weight Magnetic Artificial Synapses for Neuromorphic Computing

Figure 3 for Shape-Dependent Multi-Weight Magnetic Artificial Synapses for Neuromorphic Computing

Figure 4 for Shape-Dependent Multi-Weight Magnetic Artificial Synapses for Neuromorphic Computing

Abstract:In neuromorphic computing, artificial synapses provide a multi-weight conductance state that is set based on inputs from neurons, analogous to the brain. Additional properties of the synapse beyond multiple weights can be needed, and can depend on the application, requiring the need for generating different synapse behaviors from the same materials. Here, we measure artificial synapses based on magnetic materials that use a magnetic tunnel junction and a magnetic domain wall. By fabricating lithographic notches in a domain wall track underneath a single magnetic tunnel junction, we achieve 4-5 stable resistance states that can be repeatably controlled electrically using spin orbit torque. We analyze the effect of geometry on the synapse behavior, showing that a trapezoidal device has asymmetric weight updates with high controllability, while a straight device has higher stochasticity, but with stable resistance levels. The device data is input into neuromorphic computing simulators to show the usefulness of application-specific synaptic functions. Implementing an artificial neural network applied on streamed Fashion-MNIST data, we show that the trapezoidal magnetic synapse can be used as a metaplastic function for efficient online learning. Implementing a convolutional neural network for CIFAR-100 image recognition, we show that the straight magnetic synapse achieves near-ideal inference accuracy, due to the stability of its resistance levels. This work shows multi-weight magnetic synapses are a feasible technology for neuromorphic computing and provides design guidelines for emerging artificial synapse technologies.

* 27 pages 6 figures 1 table

Via

Access Paper or Ask Questions

On the Accuracy of Analog Neural Network Inference Accelerators

Sep 12, 2021

T. Patrick Xiao, Ben Feinberg, Christopher H. Bennett, Venkatraman Prabhakar, Prashant Saxena, Vineet Agrawal, Sapan Agarwal, Matthew J. Marinella

Figure 1 for On the Accuracy of Analog Neural Network Inference Accelerators

Figure 2 for On the Accuracy of Analog Neural Network Inference Accelerators

Figure 3 for On the Accuracy of Analog Neural Network Inference Accelerators

Figure 4 for On the Accuracy of Analog Neural Network Inference Accelerators

Abstract:Specialized accelerators have recently garnered attention as a method to reduce the power consumption of neural network inference. A promising category of accelerators utilizes nonvolatile memory arrays to both store weights and perform $\textit{in situ}$ analog computation inside the array. While prior work has explored the design space of analog accelerators to optimize performance and energy efficiency, there is seldom a rigorous evaluation of the accuracy of these accelerators. This work shows how architectural design decisions, particularly in mapping neural network parameters to analog memory cells, influence inference accuracy. When evaluated using ResNet50 on ImageNet, the resilience of the system to analog non-idealities - cell programming errors, analog-to-digital converter resolution, and array parasitic resistances - all improve when analog quantities in the hardware are made proportional to the weights in the network. Moreover, contrary to the assumptions of prior work, nearly equivalent resilience to cell imprecision can be achieved by fully storing weights as analog quantities, rather than spreading weight bits across multiple devices, often referred to as bit slicing. By exploiting proportionality, analog system designers have the freedom to match the precision of the hardware to the needs of the algorithm, rather than attempting to guarantee the same level of precision in the intermediate results as an equivalent digital accelerator. This ultimately results in an analog accelerator that is more accurate, more robust to analog errors, and more energy-efficient.

* Changes in v2: corrected typos that caused mislabeled results in Fig. 15, corrected a label in Fig. 6, minor text edits for clarity. No content changes

Via

Access Paper or Ask Questions

High-Speed CMOS-Free Purely Spintronic Asynchronous Recurrent Neural Network

Jul 05, 2021

Pranav O. Mathews, Christian B. Duffee, Abel Thayil, Ty E. Stovall, Christopher H. Bennett, Felipe Garcia-Sanchez, Matthew J. Marinella, Jean Anne C. Incorvia, Naimul Hassan, Xuan Hu(+1 more)

Figure 1 for High-Speed CMOS-Free Purely Spintronic Asynchronous Recurrent Neural Network

Figure 2 for High-Speed CMOS-Free Purely Spintronic Asynchronous Recurrent Neural Network

Figure 3 for High-Speed CMOS-Free Purely Spintronic Asynchronous Recurrent Neural Network

Figure 4 for High-Speed CMOS-Free Purely Spintronic Asynchronous Recurrent Neural Network

Abstract:Neuromorphic computing systems overcome the limitations of traditional von Neumann computing architectures. These computing systems can be further improved upon by using emerging technologies that are more efficient than CMOS for neural computation. Recent research has demonstrated memristors and spintronic devices in various neural network designs boost efficiency and speed. This paper presents a biologically inspired fully spintronic neuron used in a fully spintronic Hopfield RNN. The network is used to solve tasks, and the results are compared against those of current Hopfield neuromorphic architectures which use emerging technologies.

Via

Access Paper or Ask Questions

Controllable reset behavior in domain wall-magnetic tunnel junction artificial neurons for task-adaptable computation

Jan 08, 2021

Samuel Liu, Christopher H. Bennett, Joseph S. Friedman, Matthew J. Marinella, David Paydarfar, Jean Anne C. Incorvia

Figure 1 for Controllable reset behavior in domain wall-magnetic tunnel junction artificial neurons for task-adaptable computation

Figure 2 for Controllable reset behavior in domain wall-magnetic tunnel junction artificial neurons for task-adaptable computation

Figure 3 for Controllable reset behavior in domain wall-magnetic tunnel junction artificial neurons for task-adaptable computation

Figure 4 for Controllable reset behavior in domain wall-magnetic tunnel junction artificial neurons for task-adaptable computation

Abstract:Neuromorphic computing with spintronic devices has been of interest due to the limitations of CMOS-driven von Neumann computing. Domain wall-magnetic tunnel junction (DW-MTJ) devices have been shown to be able to intrinsically capture biological neuron behavior. Edgy-relaxed behavior, where a frequently firing neuron experiences a lower action potential threshold, may provide additional artificial neuronal functionality when executing repeated tasks. In this study, we demonstrate that this behavior can be implemented in DW-MTJ artificial neurons via three alternative mechanisms: shape anisotropy, magnetic field, and current-driven soft reset. Using micromagnetics and analytical device modeling to classify the Optdigits handwritten digit dataset, we show that edgy-relaxed behavior improves both classification accuracy and classification rate for ordered datasets while sacrificing little to no accuracy for a randomized dataset. This work establishes methods by which artificial spintronic neurons can be flexibly adapted to datasets.

* 5 pages, 5 figures

Via

Access Paper or Ask Questions

Domain Wall Leaky Integrate-and-Fire Neurons with Shape-Based Configurable Activation Functions

Nov 11, 2020

Wesley H. Brigner, Naimul Hassan, Xuan Hu, Christopher H. Bennett, Felipe Garcia-Sanchez, Can Cui, Alvaro Velasquez, Matthew J. Marinella, Jean Anne C. Incorvia, Joseph S. Friedman

Figure 1 for Domain Wall Leaky Integrate-and-Fire Neurons with Shape-Based Configurable Activation Functions

Figure 2 for Domain Wall Leaky Integrate-and-Fire Neurons with Shape-Based Configurable Activation Functions

Figure 3 for Domain Wall Leaky Integrate-and-Fire Neurons with Shape-Based Configurable Activation Functions

Figure 4 for Domain Wall Leaky Integrate-and-Fire Neurons with Shape-Based Configurable Activation Functions

Abstract:Complementary metal oxide semiconductor (CMOS) devices display volatile characteristics, and are not well suited for analog applications such as neuromorphic computing. Spintronic devices, on the other hand, exhibit both non-volatile and analog features, which are well-suited to neuromorphic computing. Consequently, these novel devices are at the forefront of beyond-CMOS artificial intelligence applications. However, a large quantity of these artificial neuromorphic devices still require the use of CMOS, which decreases the efficiency of the system. To resolve this, we have previously proposed a number of artificial neurons and synapses that do not require CMOS for operation. Although these devices are a significant improvement over previous renditions, their ability to enable neural network learning and recognition is limited by their intrinsic activation functions. This work proposes modifications to these spintronic neurons that enable configuration of the activation functions through control of the shape of a magnetic domain wall track. Linear and sigmoidal activation functions are demonstrated in this work, which can be extended through a similar approach to enable a wide variety of activation functions.

Via

Access Paper or Ask Questions

Device-aware inference operations in SONOS nonvolatile memory arrays

Apr 02, 2020

Christopher H. Bennett, T. Patrick Xiao, Ryan Dellana, Vineet Agrawal, Ben Feinberg, Venkatraman Prabhakar, Krishnaswamy Ramkumar, Long Hinh, Swatilekha Saha, Vijay Raghavan(+3 more)

Figure 1 for Device-aware inference operations in SONOS nonvolatile memory arrays

Figure 2 for Device-aware inference operations in SONOS nonvolatile memory arrays

Figure 3 for Device-aware inference operations in SONOS nonvolatile memory arrays

Figure 4 for Device-aware inference operations in SONOS nonvolatile memory arrays

Abstract:Non-volatile memory arrays can deploy pre-trained neural network models for edge inference. However, these systems are affected by device-level noise and retention issues. Here, we examine damage caused by these effects, introduce a mitigation strategy, and demonstrate its use in fabricated array of SONOS (Silicon-Oxide-Nitride-Oxide-Silicon) devices. On MNIST, fashion-MNIST, and CIFAR-10 tasks, our approach increases resilience to synaptic noise and drift. We also show strong performance can be realized with ADCs of 5-8 bits precision.

* To be presented at IEEE International Physics Reliability Symposium (IRPS) 2020

Via

Access Paper or Ask Questions

Unsupervised Competitive Hardware Learning Rule for Spintronic Clustering Architecture

Mar 24, 2020

Alvaro Velasquez, Christopher H. Bennett, Naimul Hassan, Wesley H. Brigner, Otitoaleke G. Akinola, Jean Anne C. Incorvia, Matthew J. Marinella, Joseph S. Friedman

Figure 1 for Unsupervised Competitive Hardware Learning Rule for Spintronic Clustering Architecture

Figure 2 for Unsupervised Competitive Hardware Learning Rule for Spintronic Clustering Architecture

Figure 3 for Unsupervised Competitive Hardware Learning Rule for Spintronic Clustering Architecture

Figure 4 for Unsupervised Competitive Hardware Learning Rule for Spintronic Clustering Architecture

Abstract:We propose a hardware learning rule for unsupervised clustering within a novel spintronic computing architecture. The proposed approach leverages the three-terminal structure of domain-wall magnetic tunnel junction devices to establish a feedback loop that serves to train such devices when they are used as synapses in a neuromorphic computing architecture.

Via

Access Paper or Ask Questions

Plasticity-Enhanced Domain-Wall MTJ Neural Networks for Energy-Efficient Online Learning

Mar 04, 2020

Christopher H. Bennett, T. Patrick Xiao, Can Cui, Naimul Hassan, Otitoaleke G. Akinola, Jean Anne C. Incorvia, Alvaro Velasquez, Joseph S. Friedman, Matthew J. Marinella

Figure 1 for Plasticity-Enhanced Domain-Wall MTJ Neural Networks for Energy-Efficient Online Learning

Figure 2 for Plasticity-Enhanced Domain-Wall MTJ Neural Networks for Energy-Efficient Online Learning

Figure 3 for Plasticity-Enhanced Domain-Wall MTJ Neural Networks for Energy-Efficient Online Learning

Figure 4 for Plasticity-Enhanced Domain-Wall MTJ Neural Networks for Energy-Efficient Online Learning

Abstract:Machine learning implements backpropagation via abundant training samples. We demonstrate a multi-stage learning system realized by a promising non-volatile memory device, the domain-wall magnetic tunnel junction (DW-MTJ). The system consists of unsupervised (clustering) as well as supervised sub-systems, and generalizes quickly (with few samples). We demonstrate interactions between physical properties of this device and optimal implementation of neuroscience-inspired plasticity learning rules, and highlight performance on a suite of tasks. Our energy analysis confirms the value of the approach, as the learning budget stays below 20 $\mu J$ even for large tasks used typically in machine learning.

Via

Access Paper or Ask Questions

Evaluating complexity and resilience trade-offs in emerging memory inference machines

Feb 25, 2020

Christopher H. Bennett, Ryan Dellana, T. Patrick Xiao, Ben Feinberg, Sapan Agarwal, Suma Cardwell, Matthew J. Marinella, William Severa, Brad Aimone

Figure 1 for Evaluating complexity and resilience trade-offs in emerging memory inference machines

Figure 2 for Evaluating complexity and resilience trade-offs in emerging memory inference machines

Figure 3 for Evaluating complexity and resilience trade-offs in emerging memory inference machines

Figure 4 for Evaluating complexity and resilience trade-offs in emerging memory inference machines

Abstract:Neuromorphic-style inference only works well if limited hardware resources are maximized properly, e.g. accuracy continues to scale with parameters and complexity in the face of potential disturbance. In this work, we use realistic crossbar simulations to highlight that compact implementations of deep neural networks are unexpectedly susceptible to collapse from multiple system disturbances. Our work proposes a middle path towards high performance and strong resilience utilizing the Mosaics framework, and specifically by re-using synaptic connections in a recurrent neural network implementation that possesses a natural form of noise-immunity.

Via

Access Paper or Ask Questions