Publication

Machine learning for cross-gazetteer matching of natural features

Elise Acheson
2019
Journal paper

Abstract

Defining and identifying duplicate records in a dataset is a challenging task which grows more complex when the modeled entities themselves are hard to delineate. In the geospatial domain, it may not be clear where a mountain, stream, or valley ends and begins, a problem carried over when such entities are catalogued in gazetteers. In this paper, we take two gazetteers, GeoNames and SwissNames3D, and perform matching - identifying records in each that are about the same entity - across a sample of natural feature records. We first perform rule-based matching, establishing competitive results, then apply machine learning using Random Forests, a method well-suited to the matching task. We report on the performance of a wider array of matching features than has been previously studied, including domain-specific ones such as feature type, land cover class, and elevation. Our results show an increase in performance using machine learning over rules, with a notable performance gain from considering feature types, but negligible gains from other specialized matching features. We argue that future work in this area should strive to be more reproducible and report results on a realistic testing pipeline including candidate selection, feature extraction, and classification.

Official source

https://infoscience.epfl.ch/record/267566?ln=en

About this result

This page is automatically generated and may contain information that is not correct, complete, up-to-date, or relevant to your search query. The same applies to every other page on this website. Please make sure to verify the information with EPFL's official sources.

Machine learning for cross-gazetteer matching of natural features

Graph Chatbot

Chat with Graph Search

Data-driven approaches for non-invasive cuffless blood pressure monitoring

The Virtue of Complexity in Return Prediction

Measurement in the Age of Information

The Virtue of Complexity in Return Prediction

Measurement in the Age of Information

Data-driven approaches for non-invasive cuffless blood pressure monitoring