Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Andre Viebke

Performance Modelling of Deep Learning on Intel Many Integrated Core Architectures

Jun 04, 2019

Andre Viebke, Sabri Pllana, Suejb Memeti, Joanna Kolodziej

Figure 1 for Performance Modelling of Deep Learning on Intel Many Integrated Core Architectures

Figure 2 for Performance Modelling of Deep Learning on Intel Many Integrated Core Architectures

Figure 3 for Performance Modelling of Deep Learning on Intel Many Integrated Core Architectures

Figure 4 for Performance Modelling of Deep Learning on Intel Many Integrated Core Architectures

Abstract:Many complex problems, such as natural language processing or visual object detection, are solved using deep learning. However, efficient training of complex deep convolutional neural networks for large data sets is computationally demanding and requires parallel computing resources. In this paper, we present two parameterized performance models for estimation of execution time of training convolutional neural networks on the Intel many integrated core architecture. While for the first performance model we minimally use measurement techniques for parameter value estimation, in the second model we estimate more parameters based on measurements. We evaluate the prediction accuracy of performance models in the context of training three different convolutional neural network architectures on the Intel Xeon Phi. The achieved average performance prediction accuracy is about 15% for the first model and 11% for second model.

* Preprint, HPCS. arXiv admin note: substantial text overlap with arXiv:1702.07908

Via

Access Paper or Ask Questions

CHAOS: A Parallelization Scheme for Training Convolutional Neural Networks on Intel Xeon Phi

Feb 25, 2017

Andre Viebke, Suejb Memeti, Sabri Pllana, Ajith Abraham

Figure 1 for CHAOS: A Parallelization Scheme for Training Convolutional Neural Networks on Intel Xeon Phi

Figure 2 for CHAOS: A Parallelization Scheme for Training Convolutional Neural Networks on Intel Xeon Phi

Figure 3 for CHAOS: A Parallelization Scheme for Training Convolutional Neural Networks on Intel Xeon Phi

Figure 4 for CHAOS: A Parallelization Scheme for Training Convolutional Neural Networks on Intel Xeon Phi

Abstract:Deep learning is an important component of big-data analytic tools and intelligent applications, such as, self-driving cars, computer vision, speech recognition, or precision medicine. However, the training process is computationally intensive, and often requires a large amount of time if performed sequentially. Modern parallel computing systems provide the capability to reduce the required training time of deep neural networks. In this paper, we present our parallelization scheme for training convolutional neural networks (CNN) named Controlled Hogwild with Arbitrary Order of Synchronization (CHAOS). Major features of CHAOS include the support for thread and vector parallelism, non-instant updates of weight parameters during back-propagation without a significant delay, and implicit synchronization in arbitrary order. CHAOS is tailored for parallel computing systems that are accelerated with the Intel Xeon Phi. We evaluate our parallelization approach empirically using measurement techniques and performance modeling for various numbers of threads and CNN architectures. Experimental results for the MNIST dataset of handwritten digits using the total number of threads on the Xeon Phi show speedups of up to 103x compared to the execution on one thread of the Xeon Phi, 14x compared to the sequential execution on Intel Xeon E5, and 58x compared to the sequential execution on Intel Core i5.

* The Journal of Supercomputing, 2017

Via

Access Paper or Ask Questions