Repository logo
 

Selective contrastive explanations using knowledge graphs for commonsense question answering


Loading...
Thumbnail Image

Type

Change log

Abstract

Trust in an AI system requires not only knowing that the system chose the right answer, but also knowing that it arrived at this decision for the right reasons. In this thesis, I will investigate how to generate and evaluate explanations for the decisions taken by commonsense reasoning systems for question answering. For this analysis to be valid, the model explanation must be a true representation of the facts used by the model, that is, it must be faithful. Psychological research shows that humans prefer explanations that are minimal and focus on a specific aspect of the question, but these properties have so far been under-appreciated in NLP work. I elicit such explanations from humans, using school-level science questions as my domain, and compare them to models' explanations.

I use models that produce explanations by selecting facts from a knowledge graph given as input. My first contribution is an evaluation method that quantifies the faithfulness of these explanations. My experiments reveal that the explanations given by two pre-existing such models are unfaithful to a high degree. I also show that faithfulness can be increased by modifying the architecture. My second contribution is a set of methods for selecting the input knowledge graph from a larger one. I propose a weighting scheme that scores facts according to whether they appear in comparable explanations, and find that the resulting graphs achieve up to a 39% higher accuracy over the standard method. My third contribution is a knowledge graph schema that is designed to reduce the complexity, redundancy, and ambiguity of edges and nodes within it. I populate the schema by translating facts from the WorldTree dataset, resulting in explanations for 1778 questions. The overall graph has 14,000 edges. Compared to a popular commonsense knowledge graph, it contains 2.7 times fewer redundant concepts and 1.7 times fewer complex concepts. My final contribution is a human annotation procedure for building minimal and focussed explanations. The results show that annotators achieve agreement of κ=0.34 on which facts are necessary in the explanations. I then compare faithful explanations from two models with this human gold standard, and find a normalised graph edit distance of 0.81 on average. My results indicate that models use facts dissimilar to those used by humans to answer commonsense questions. In sum, this thesis performs and advocates systematic comparisons between humans and model behaviour based on the central ideas of explanation clarity and faithfulness. Explanations are useful insofar as they clearly express ideas, and if faithfulness is not considered, we risk creating misleading explanations.

Description

Date

2025-08-18

Advisors

Teufel, Simone

Qualification

Doctor of Philosophy (PhD)

Awarding Institution

University of Cambridge

Rights and licensing

Except where otherwised noted, this item's license is described as All rights reserved
Sponsorship
Homerton College Charter Postgraduate Award