A pragmatic guide to geoparsing evaluation
Publication Date
2019-09-19Journal Title
Language Resources and Evaluation
ISSN
1574-020X
Publisher
Springer Netherlands
Volume
54
Issue
3
Pages
683-712
Language
en
Type
Article
This Version
VoR
Metadata
Show full item recordCitation
Gritta, M., Pilehvar, M. T., & Collier, N. (2019). A pragmatic guide to geoparsing evaluation. Language Resources and Evaluation, 54 (3), 683-712. https://doi.org/10.1007/s10579-019-09475-3
Abstract
Abstract: Empirical methods in geoparsing have thus far lacked a standard evaluation framework describing the task, metrics and data used to compare state-of-the-art systems. Evaluation is further made inconsistent, even unrepresentative of real world usage by the lack of distinction between the different types of toponyms, which necessitates new guidelines, a consolidation of metrics and a detailed toponym taxonomy with implications for Named Entity Recognition (NER) and beyond. To address these deficiencies, our manuscript introduces a new framework in three parts. (Part 1) Task Definition: clarified via corpus linguistic analysis proposing a fine-grained Pragmatic Taxonomy of Toponyms. (Part 2) Metrics: discussed and reviewed for a rigorous evaluation including recommendations for NER/Geoparsing practitioners. (Part 3) Evaluation data: shared via a new dataset called GeoWebNews to provide test/train examples and enable immediate use of our contributions. In addition to fine-grained Geotagging and Toponym Resolution (Geocoding), this dataset is also suitable for prototyping and evaluating machine learning NLP models.
Keywords
Original Paper, Geoparsing, Toponym resolution, Geotagging, Geocoding, Named Entity Recognition, Machine learning, Evaluation framework, Geonames, Toponyms, Natural language understanding, Pragmatics
Sponsorship
Natural Environment Research Council (NE/M009009/1)
Medical Research Council (MR/M025160/1)
Engineering and Physical Sciences Research Council (EP/M005089/1)
Identifiers
s10579-019-09475-3, 9475
External DOI: https://doi.org/10.1007/s10579-019-09475-3
This record's URL: https://www.repository.cam.ac.uk/handle/1810/308852
Rights
Licence:
https://creativecommons.org/licenses/by/4.0/