Instance Segmentation, Body Part Parsing, and Pose Estimation of Human Figures in Pictorial Maps

In recent years, convolutional neural networks (CNNs) have been applied successfully to recognise persons, their body parts and pose keypoints in photos and videos. The transfer of these techniques to artificially created images is rather unexplored, though challenging since these images are drawn in different styles, body proportions, and levels of abstraction. In this work, we study these problems on the basis of pictorial maps where we identify included human figures with two consecutive CNNs: We first segment individual figures with Mask R-CNN, and then parse their body parts and estimate their poses simultaneously with four different UNet++ versions. We train the CNNs with a mixture of real persons and synthetic figures and compare the results with manually annotated test datasets consisting of pictorial figures. By varying the training datasets and the CNN configurations, we were able to improve the original Mask R-CNN model and we achieved moderately satisfying results with the UNet++ versions. The extracted figures may be used for animation and storytelling and may be relevant for the analysis of historic and contemporary maps.

Keywords

4013 Geomatic Engineering, 40 Engineering, Generic health relevance

Journal Title

International Journal of Cartography

Journal ISSN

2372-9333
2372-9341

Volume Title

8

Publisher

Taylor & Francis

Publisher DOI

https://doi.org/10.1080/23729333.2021.1949087

Rights and licensing

Except where otherwised noted, this item's license is described as Attribution 4.0 International

Sponsorship

MRC (MR/T043229/1)
MRC (MR/Y033884/1)

Collections

University of Cambridge Research Outputs (Articles and Conferences)