Publication

CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning

Anastasia Ailamaki, Benjamin Cyrille Damien Gaidioz, Manolis Karpathiotakis, Styliani Asimina Giannakopoulou
2017
Article de conférence

Résumé

Data cleaning has become an indispensable part of data analysis due to the increasing amount of dirty data. Data scientists spend most of their time preparing dirty data before it can be used for data analysis. At the same time, the existing tools that attempt to automate the data cleaning procedure typically focus on a specific use case and operation. Still, even such specialized tools exhibit long running times or fail to process large datasets. Therefore, from a user’s perspective, one is forced to use a different, potentially inefficient tool for each category of errors. This paper addresses the coverage and efficiency problems of data cleaning. It introduces CleanM ( pronounced clean’em), a language which can express multiple types of cleaning operations. CleanM goes through a three-level translation process for optimiza- tion purposes; a different family of optimizations is applied in each abstraction level. Thus, CleanM can express complex data cleaning tasks, optimize them in a unified way, and deploy them in a scaleout fashion. We validate the applicability of CleanM by using it on top of CleanDB, a newly designed and implemented framework which can query heterogeneous data. When compared to existing data cleaning solutions, CleanDB a) covers more data corruption cases, b) scales better, and can handle cases for which its competitors are unable to terminate, and c) uses a single interface for querying and for data cleaning

Source officielle

https://infoscience.epfl.ch/record/230431?ln=fr

À propos de ce résultat

Cette page est générée automatiquement et peut contenir des informations qui ne sont pas correctes, complètes, à jour ou pertinentes par rapport à votre recherche. Il en va de même pour toutes les autres pages de ce site. Veillez à vérifier les informations auprès des sources officielles de l'EPFL.

Graph Chatbot

Chattez avec Graph Search

Posez n’importe quelle question sur les cours, conférences, exercices, recherches, actualités, etc. de l’EPFL ou essayez les exemples de questions ci-dessous.

AVERTISSEMENT : Le chatbot Graph n'est pas programmé pour fournir des réponses explicites ou catégoriques à vos questions. Il transforme plutôt vos questions en demandes API qui sont distribuées aux différents services informatiques officiellement administrés par l'EPFL. Son but est uniquement de collecter et de recommander des références pertinentes à des contenus que vous pouvez explorer pour vous aider à répondre à vos questions.

Connectez-vous pour utiliser Chat avec Graph Search

Anastasia Ailamaki, Benjamin Cyrille Damien Gaidioz, Manolis Karpathiotakis, Styliani Asimina Giannakopoulou
2017
Article de conférence

Résumé

Source officielle

https://infoscience.epfl.ch/record/230431?ln=fr

À propos de ce résultat

Proximité ontologique

Business

Administration des affaires: Gestion des données

Concepts associés (31)

Publications associées (41)

MOOCs associés (7)

CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning

Graph Chatbot

Chattez avec Graph Search

Dataset for "Elastocapillary menisci mediate interaction of neighboring structures at the surface of a compliant solid"

Continuous corrected particle number concentration data in 10 sec resolution measured in the Swiss aerosol container using a whole air inlet during MOSAiC 2019/2020

Continuous corrected particle number concentration data in 10 sec resolution, measured in the Swiss aerosol container during MOSAiC 2019/2020

Dataset for "Elastocapillary menisci mediate interaction of neighboring structures at the surface of a compliant solid"

Continuous corrected particle number concentration data in 10 sec resolution measured in the Swiss aerosol container using a whole air inlet during MOSAiC 2019/2020

Continuous corrected particle number concentration data in 10 sec resolution, measured in the Swiss aerosol container during MOSAiC 2019/2020