Repository logo
 

Interpretable Toxicity Predictions and Exploring Chemical Space: Developing and Testing Cheminformatics and Explainable AI Tools


Loading...
Thumbnail Image

Type

Change log

Abstract

This thesis studies chemical space exploration and explainability for toxicity prediction. The four chapters, show how molecules can be predicted, tested and explained in ways that are both chemically meaningful and computationally robust. Chapter 1 offers an introduction to the cheminformatics and its applications.

Chapter 2 focuses on the application of machine learning and cheminformatics tools to environmental toxicity prediction, in collaboration with researchers at DTU. This work includes contributions to two manuscripts: one addressing ML-based data gap filling in ecotoxicity, and the second developing and explaining uncertainty-aware models for predicting non-cancer and developmental/reproductive human toxicity using both Bayesian and frequentist methods. I developed chemical space mapping in the context of molecular structure concepts and explored the use of descriptor attribution for extracting relevant substructures.

In Chapter 3, the use of molecular counterfactuals to recover meaningful structural alerts is investigated. Using skin sensitisation and mutagenicity datasets, it is demonstrated how counterfactuals can serve as a bridge between traditional toxicological knowledge and modern XAI techniques.

Finally, Chapter 4 uses GNNExplainer, a graph-agnostic explainability method applied to a range of toxicity classification models to evaluate how graph-based explanations align with toxicologically meaningful features across well-studied human endpoints, highlighting the importance and limitations of model interpretability in cheminformatics.

Together, these studies advance the interface between prediction and explanation, particularly in the context of toxicity, where transparency and trust are critical. The thesis is structured to reflect a progression from predictive modelling and uncertainty (Chapter 1), to two complementary explainability strategies (Chapters 2 and 3). Together these contribute to the broader goal of developing interpretable, reliable, and chemically-grounded machine learning tools for molecular science.

Description

Date

2025-10-30

Advisors

Goodman, Jonathan

Qualification

Doctor of Philosophy (PhD)

Awarding Institution

University of Cambridge

Rights and licensing

Except where otherwised noted, this item's license is described as All rights reserved