Repository logo
 

Machine learning classification of entrepreneurs in British historical census data

Accepted version
Peer-reviewed

Type

Article

Change log

Authors

Montebruno, P 
Bennett, RJ 
Smith, H 
Lieshout, CV 

Abstract

This paper presents a binary classification of entrepreneurs in British historical data based on the recent availability of big data from the I-CeM dataset. The main task of the paper is to attribute an employment status to individuals that did not fully report entrepreneur status in earlier censuses (1851-1881). The paper assesses the accuracy of different classifiers and machine learning algorithms, including Deep Learning, for this classification problem. We first adopt a ground-truth dataset from the later censuses to train the computer with a Logistic Regression (which is standard in the literature for this kind of binary classification) to recognize entrepreneurs distinct from non-entrepreneurs (i.e. workers). Our initial accuracy for this base-line method is 0.74. We compare the Logistic Regression with ten optimized machine learning algorithms: Nearest Neighbors, Linear and Radial Support Vector Machine, Gaussian Process, Decision Tree, Random Forest, Neural Network, AdaBoost, Naive Bayes, and Quadratic Discriminant Analysis. The best results are boosting and ensemble methods. AdaBoost achieves an accuracy of 0.95. Deep-Learning, as a standalone category of algorithms, further improves accuracy to 0.96 without using the rich text-data that characterizes the OccString feature, a string of up to 500 characters with the full occupational statement of each individual collected in the earlier censuses. Finally, and now using this OccString feature, we implement both shallow (bag-of-words algorithm) learning and Deep Learning (Recurrent Neural Network with a Long Short-Term Memory layer) algorithms. These methods all achieve accuracies above 0.99 with Deep Learning Recurrent Neural Network as the best model with an accuracy of 0.9978. The results show that standard algorithms for classification can be outperformed by machine learning algorithms. This confirms the value of extending the techniques traditionally used in the literature for this type of classification problem.

Description

Keywords

Machine learning, Deep learning, Logistic regression, Classification, Big data, Census

Journal Title

Information Processing and Management

Conference Name

Journal ISSN

0306-4573
1873-5371

Volume Title

57

Publisher

Elsevier BV
Sponsorship
Leverhulme Trust (EM-2012-008/7)
Economic and Social Research Council (ES/M010953/1)
Isaac Newton Trust (17.07(d))
Isaac Newton Trust (18.40(g))
ESRC Leverhulme Trust Isaac Newton Trust