Sandwich boosting for accurate estimation in partially linear models for grouped data

Young, Elliot H; Shah, Rajen D

doi:https://doi.org/10.17863/CAM.107026

Sandwich boosting for accurate estimation in partially linear models for grouped data

Accepted version

Peer-reviewed

Repository URI

https://www.repository.cam.ac.uk/handle/1810/365972

Repository DOI

https://doi.org/10.17863/CAM.107026

Files

Primary Accepted version (900.27 KB)

Type

Article

Authors

Young, Elliot H

Shah, Rajen D

https://orcid.org/0000-0001-9073-3782

Abstract

jats:titleAbstract</jats:title> jats:pWe study partially linear models in settings where observations are arranged in independent groups but may exhibit within-group dependence. Existing approaches estimate linear model parameters through weighted least squares, with optimal weights (given by the inverse covariance of the response, conditional on the covariates) typically estimated by maximizing a (restricted) likelihood from random effects modelling or by using generalized estimating equations. We introduce a new ‘sandwich loss’ whose population minimizer coincides with the weights of these approaches when the parametric forms for the conditional covariance are well-specified, but can yield arbitrarily large improvements in linear parameter estimation accuracy when they are not. Under relatively mild conditions, our estimated coefficients are asymptotically Gaussian and enjoy minimal variance among estimators with weights restricted to a given class of functions, when user-chosen regression methods are used to estimate nuisance functions. We further expand the class of functional forms for the weights that may be fitted beyond parametric models by leveraging the flexibility of modern machine learning methods within a new gradient boosting scheme for minimizing the sandwich loss. We demonstrate the effectiveness of both the sandwich loss and what we call ‘sandwich boosting’ in a variety of settings with simulated and real-world data.</jats:p>

Keywords

49 Mathematical Sciences, 4905 Statistics

Journal Title

Journal of the Royal Statistical Society Series B: Statistical Methodology

Journal ISSN

1369-7412
1467-9868

Publisher

Oxford University Press (OUP)

Publisher DOI

https://doi.org/10.1093/jrsssb/qkae032

Rights

Attribution 4.0 International

Sponsorship

Engineering and Physical Sciences Research Council (EP/N031938/1)

Collections

University of Cambridge Research Outputs (Articles and Conferences)

Sandwich boosting for accurate estimation in partially linear models for grouped data

Accepted version

Peer-reviewed

Repository URI

Repository DOI

Files

Type

Change log

Authors

Abstract

Description

Keywords

Journal Title

Conference Name

Journal ISSN

Volume Title

Publisher

Publisher DOI

Rights

Sponsorship

Collections