Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates

Chattez avec Graph Search

Posez n’importe quelle question sur les cours, conférences, exercices, recherches, actualités, etc. de l’EPFL ou essayez les exemples de questions ci-dessous.

AVERTISSEMENT : Le chatbot Graph n'est pas programmé pour fournir des réponses explicites ou catégoriques à vos questions. Il transforme plutôt vos questions en demandes API qui sont distribuées aux différents services informatiques officiellement administrés par l'EPFL. Son but est uniquement de collecter et de recommander des références pertinentes à des contenus que vous pouvez explorer pour vous aider à répondre à vos questions.

It has been experimentally observed that the efficiency of distributed training with stochastic gradient (SGD) depends decisively on the batch size and—in asynchronous implementations—on the gradient staleness. Especially, it has been observed that the speedup saturates beyond a certain batch size and/or when the delays grow too large. We identify a data-dependent parameter that explains the speedup saturation in both these settings. Our comprehensive theoretical analysis, for strongly convex, convex and non-convex settings, unifies and generalized prior work directions that often focused on only one of these two aspects. In particular, our approach allows us to derive improved speedup results under frequently considered sparsity assumptions. Our insights give rise to theoretically based guidelines on how the learning rates can be adjusted in practice. We show that our results are tight and illustrate key findings in numerical experiments.

Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates

Graph Chatbot

Chattez avec Graph Search

Understanding generalization and robustness in modern deep learning

Optimization Algorithms for Decentralized, Distributed and Collaborative Machine Learning

On the Generalization of Stochastic Gradient Descent with Momentum

Optimization Algorithms for Decentralized, Distributed and Collaborative Machine Learning

Understanding generalization and robustness in modern deep learning

On the Generalization of Stochastic Gradient Descent with Momentum