Interpretable Toxicity Predictions and Exploring Chemical Space: Developing and Testing Cheminformatics and Explainable AI Tools
Repository URI
Repository DOI
Change log
Authors
Abstract
This thesis studies chemical space exploration and explainability for toxicity prediction. The four chapters, show how molecules can be predicted, tested and explained in ways that are both chemically meaningful and computationally robust. Chapter 1 offers an introduction to the cheminformatics and its applications.
Chapter 2 focuses on the application of machine learning and cheminformatics tools to environmental toxicity prediction, in collaboration with researchers at DTU. This work includes contributions to two manuscripts: one addressing ML-based data gap filling in ecotoxicity, and the second developing and explaining uncertainty-aware models for predicting non-cancer and developmental/reproductive human toxicity using both Bayesian and frequentist methods. I developed chemical space mapping in the context of molecular structure concepts and explored the use of descriptor attribution for extracting relevant substructures.
In Chapter 3, the use of molecular counterfactuals to recover meaningful structural alerts is investigated. Using skin sensitisation and mutagenicity datasets, it is demonstrated how counterfactuals can serve as a bridge between traditional toxicological knowledge and modern XAI techniques.
Finally, Chapter 4 uses GNNExplainer, a graph-agnostic explainability method applied to a range of toxicity classification models to evaluate how graph-based explanations align with toxicologically meaningful features across well-studied human endpoints, highlighting the importance and limitations of model interpretability in cheminformatics.
Together, these studies advance the interface between prediction and explanation, particularly in the context of toxicity, where transparency and trust are critical. The thesis is structured to reflect a progression from predictive modelling and uncertainty (Chapter 1), to two complementary explainability strategies (Chapters 2 and 3). Together these contribute to the broader goal of developing interpretable, reliable, and chemically-grounded machine learning tools for molecular science.
