Publication

Don't cry over spilled records: Memory elasticity of data-parallel applications and its application to cluster scheduling

Willy Zwaenepoel, Calin Iorgulescu, Syed Mohammad Aunn Raza, Florin Dinu, Wajih Ul Hassan
2017
Article de conférence

Résumé

Understanding the performance of data-parallel workloads when resource-constrained has significant practical importance but unfortunately has received only limited attention. This paper identifies, quantifies and demonstrates memory elasticity, an intrinsic property of data-parallel tasks. Memory elasticity allows tasks to run with significantly less memory that they would ideally want while only paying a moderate performance penalty. For example, we find that given as little as 10% of ideal memory, PageRank and NutchIndexing Hadoop reducers become only 1.2x/1.75x and 1.08x slower. We show that memory elasticity is prevalent in the Hadoop, Spark, Tez and Flink frameworks. We also show that memory elasticity is predictable in nature by building simple models for Hadoop and extending them to Tez and Spark. To demonstrate the potential benefits of leveraging memory elasticity, this paper further explores its application to cluster scheduling. In this setting, we observe that the resource vs. time trade-off enabled by memory elasticity becomes a task queuing time vs task runtime trade-off. Tasks may complete faster when scheduled with less memory because their waiting time is reduced. We show that a scheduler can turn this task-level trade-off into improved job completion time and cluster-wide memory utilization. We have integrated memory elasticity into Apache YARN. We show gains of up to 60% in average job completion time on a 50-node Hadoop cluster. Extensive simulations show similar improvements over a large number of scenarios.

Source officielle

https://infoscience.epfl.ch/record/255993?ln=fr

À propos de ce résultat

Cette page est générée automatiquement et peut contenir des informations qui ne sont pas correctes, complètes, à jour ou pertinentes par rapport à votre recherche. Il en va de même pour toutes les autres pages de ce site. Veillez à vérifier les informations auprès des sources officielles de l'EPFL.

Graph Chatbot

Chattez avec Graph Search

Posez n’importe quelle question sur les cours, conférences, exercices, recherches, actualités, etc. de l’EPFL ou essayez les exemples de questions ci-dessous.

AVERTISSEMENT : Le chatbot Graph n'est pas programmé pour fournir des réponses explicites ou catégoriques à vos questions. Il transforme plutôt vos questions en demandes API qui sont distribuées aux différents services informatiques officiellement administrés par l'EPFL. Son but est uniquement de collecter et de recommander des références pertinentes à des contenus que vous pouvez explorer pour vous aider à répondre à vos questions.

Connectez-vous pour utiliser Chat avec Graph Search

Willy Zwaenepoel, Calin Iorgulescu, Syed Mohammad Aunn Raza, Florin Dinu, Wajih Ul Hassan
2017
Article de conférence

Résumé

Source officielle

https://infoscience.epfl.ch/record/255993?ln=fr

À propos de ce résultat

Proximité ontologique

Génie informatique

Calcul intensif: Parallélisme (informatique)

Concepts associés (34)

Publications associées (32)

MOOCs associés (4)

Don't cry over spilled records: Memory elasticity of data-parallel applications and its application to cluster scheduling

Graph Chatbot

Chattez avec Graph Search

Manycore clique enumeration with fast set intersections

Improving Main-memory Database System Performance through Cooperative Multitasking

An Associativity-Agnostic in-Cache Computing Architecture Optimized for Multiplication

Manycore clique enumeration with fast set intersections

Improving Main-memory Database System Performance through Cooperative Multitasking

An Associativity-Agnostic in-Cache Computing Architecture Optimized for Multiplication