Data-Driven Derivation of an "Informer Compound Set" for Improved Selection of Active Compounds in High-Throughput Screening.

Despite the usefulness of high-throughput screening (HTS) in drug discovery, for some systems, low assay throughput or high screening cost can prohibit the screening of large numbers of compounds. In such cases, iterative cycles of screening involving active learning (AL) are employed, creating the need for smaller "informer sets" that can be routinely screened to build predictive models for selecting compounds from the screening collection for follow-up screens. Here, we present a data-driven derivation of an informer compound set with improved predictivity of active compounds in HTS, and we validate its benefit over randomly selected training sets on 46 PubChem assays comprising at least 300,000 compounds and covering a wide range of assay biology. The informer compound set showed improvement in BEDROC($\alpha$ = 100), PRAUC, and ROCAUC values averaged over all assays of 0.024, 0.014, and 0.016, respectively, compared to randomly selected training sets, all with paired $t$-test p-values <10$^{-15}$. A per-assay assessment showed that the BEDROC($\alpha$ = 100), which is of particular relevance for early retrieval of actives, improved for 38 out of 46 assays, increasing the success rate of smaller follow-up screens. Overall, we showed that an informer set derived from historical HTS activity data can be employed for routine small-scale exploratory screening in an assay-agnostic fashion. This approach led to a consistent improvement in hit rates in follow-up screens without compromising scaffold retrieval. The informer set is adjustable in size depending on the number of compounds one intends to screen, as performance gains are realized for sets with more than 3,000 compounds, and this set is therefore applicable to a variety of situations. Finally, our results indicate that random sampling may not adequately cover descriptor space, drawing attention to the importance of the composition of the training set for predicting actives.

Keywords

Drug Evaluation, Preclinical, High-Throughput Screening Assays, Informatics, Machine Learning

Journal Title

Journal of Chemical Information and Modeling

Journal ISSN

1549-9596
1549-960X

Volume Title

56

Publisher

American Chemical Society

Publisher DOI

https://doi.org/10.1021/acs.jcim.6b00244

Rights and licensing

Except where otherwised noted, this item's license is described as http://www.rioxx.net/licenses/all-rights-reserved

Sponsorship

The Netherlands Organisation for Scientific Research (Grant ID: NWO-017.009-065), Novartis Institutes for BioMedical Research, Prins Bernhard Cultuurfonds, European Research Commission

Collections

Scholarly Works - Chemistry
Symplectic mapped items for data match