Advancing nanobody discovery and engineering with sequence-based machine learning
Repository URI
Repository DOI
Change log
Authors
Abstract
Nanobodies, also known as single-domain antibodies, are small antibody fragments naturally produced in camelids. Notable for their high expressibility, solubility, and stability, they exhibit binding affinities comparable to those of conventional antibodies. Since the approval of the first nanobody-based drug in 2019, they have gained significant momentum as therapeutics. Yet the development of these biologics remains a challenge. While in vitro directed evolution technologies are relatively fast and cheap to deploy, and in silico tools rapidly advancing, the gold standard for generating therapeutic antibodies with favourable properties in vivo remains discovery from animal immunisation or patients. This raises the question of whether a computational design strategy will ever meet the challenge of generating nanobodies with immune-system-like properties, including long half-life, high stability, and low toxicity. Despite recent advances, computational approaches are still hindered by limited data and a lack of predictive models tailored to nanobodies. This thesis seeks to bridge these gaps by generating new nanobody datasets and developing nanobody-specific computational frameworks. In particular, I introduce three novel sequence-based computational strategies to optimise the developability and functionality of engineered nanobodies. First, I present AbNatiV, a deep-learning model for assessing the nativeness of antibodies. It enables the generation of antibodies and nanobodies undistinguishable from immune-system derived ones. With AbNatiV, I develop an experimentally validated automated humanisation pipeline for nanobodies that optimises humanness while preserving nanobody nativeness. Second, I introduce AbNatiV2, an updated version of AbNatiV that integrates new nanobody immune repertoires with more efficient training objectives and enhanced transformer features. In conjunction, I develop p-AbNatiV2, a paired VQ-VAE that leverages the unpaired AbNatiV2 models to jointly evaluate the humanness of paired antibody sequences and predict their pairing compatibility. Third, I present NanoMelt, a nanobody thermostability predictor trained on a dataset of 640 apparent melting temperatures, which contains 129 new measurements. NanoMelt serves as a case study in learning protein biophysical traits from limited data. I demonstrate that NanoMelt effectively guides the selection of highly stable nanobody candidates. Supported by experimental applications, both within this work and through external collaborations, this thesis contributes to the foundation of the in silico nanobody discovery platform of tomorrow.
