Repository logo
 

Adapting Audio Foundation Models for Heart Sound Analysis

Accepted version
Peer-reviewed

Change log

Abstract

Foundation models - large pretrained neural networks - have shown potential for heart sound classification tasks. However, a key question is still how to adapt a general audio foundation model best to these tasks. This work systematically studies three domain adaptation techniques, freezing the foundation model and training a linear layer on top (linear probing, LP), fine-tuning (FT), and continued pretraining (CP), on two audio foundation models using four public heart sound databases. Our findings demonstrate that LP alone is insufficient for heart sound analysis tasks. While FT improves performance over LP, it yields models that generalise poorly to unseen datasets. To overcome this limitation, we introduce CP as a novel method for heart sounds. We find that further pretraining a model on all datasets together produces a heart sound-specific yet task-agnostic foundation model, which boosts LP and FT performance by up to 13%. Furthermore, two CP variants are studied, and we find that using the downstream dataset only for CP improves the learned representations and boosts LP and FT performance the most. These findings underscore that choosing the correct adaptation strategy is critical for heart sound analysis tasks.

Description

Keywords

Journal Title

Conference Name

Computing in Cardiology (CinC 2025)

Journal ISSN

Volume Title

Publisher

Publisher DOI

Publisher URL

Rights and licensing

Except where otherwised noted, this item's license is described as All Rights Reserved
Sponsorship
European Commission Horizon 2020 (H2020) ERC (833296)
ERC Project 833296