Publication

Distributed Learning of CNNs on Heterogeneous CPU/GPU Architectures

José Pedro Rebelo Ferreira Marques, Gabriel Falcao Paiva Fernandes
2018
Article

Résumé

The convolutional neural networks (CNNs) have proven to be powerful classification tools in tasks that range from check reading to medical diagnosis, reaching close to human perception, and in some cases surpassing it. However, the problems to solve are becoming larger and more complex, which translates to larger CNNs, leading to longer training times that not even the adoption of Graphics Processing Units (GPUs) could keep up to. This problem is partially solved by using more processing units and distributed training methods that are offered by several frameworks dedicated to neural network training, such as Caffe, Torch, or TensorFlow. However, these techniques do not take full advantage of the possible parallelization offered by CNNs and the cooperative use of heterogeneous devices with different processing capabilities, clock speeds, memory size, among others. This paper presents a new method for the parallel training of CNNs where only the convolutional layer is distributed. The paper analyzes the influence of network size, bandwidth, batch size, number of devices, including their processing capabilities, and other parameters. Results show that this technique is capable of diminishing the training time without affecting the classification performance for both CPUs and GPUs. For the CIFAR-10 dataset, using a CNN with two convolutional layers, and 500 and 1500 kernels, respectively, best speedups achieve 3.28 x using four CPUs and 2.45 x with three GPUs. Larger datasets will certainly require more than 60-90% of processing time calculating convolutions, and speedups will tend to increase accordingly.

Source officielle

https://infoscience.epfl.ch/record/262517?ln=fr

À propos de ce résultat

Cette page est générée automatiquement et peut contenir des informations qui ne sont pas correctes, complètes, à jour ou pertinentes par rapport à votre recherche. Il en va de même pour toutes les autres pages de ce site. Veillez à vérifier les informations auprès des sources officielles de l'EPFL.

Graph Chatbot

Chattez avec Graph Search

Posez n’importe quelle question sur les cours, conférences, exercices, recherches, actualités, etc. de l’EPFL ou essayez les exemples de questions ci-dessous.

AVERTISSEMENT : Le chatbot Graph n'est pas programmé pour fournir des réponses explicites ou catégoriques à vos questions. Il transforme plutôt vos questions en demandes API qui sont distribuées aux différents services informatiques officiellement administrés par l'EPFL. Son but est uniquement de collecter et de recommander des références pertinentes à des contenus que vous pouvez explorer pour vous aider à répondre à vos questions.

Connectez-vous pour utiliser Chat avec Graph Search

José Pedro Rebelo Ferreira Marques, Gabriel Falcao Paiva Fernandes
2018
Article

Résumé

Source officielle

https://infoscience.epfl.ch/record/262517?ln=fr

À propos de ce résultat

Proximité ontologique

Génie informatique

Calcul intensif: Parallélisme (informatique)

Information engineering

Apprentissage automatique: Réseau de neurones artificiels

Concepts associés (35)

Publications associées (48)

MOOCs associés (8)

Distributed Learning of CNNs on Heterogeneous CPU/GPU Architectures

Graph Chatbot

Chattez avec Graph Search

Exploring High-Performance and Energy-Efficient Architectures for Edge AI-Enabled Applications

DBFS: Dynamic Bitwidth-Frequency Scaling for Efficient Software-defined SIMD

Compilation and Design Space Exploration of Dataflow Programs for Heterogeneous CPU-GPU Platforms

Compilation and Design Space Exploration of Dataflow Programs for Heterogeneous CPU-GPU Platforms

DBFS: Dynamic Bitwidth-Frequency Scaling for Efficient Software-defined SIMD

Exploring High-Performance and Energy-Efficient Architectures for Edge AI-Enabled Applications