Investigating the reliability and interpretability of machine learning frameworks for chemical retrosynthesis

Friedrich Hastedt; Rowan M. Bailey; Klaus Hellgardt; Sophia N. Yaliraki; Ehecatl Antonio del Rio Chanona; Dongda Zhang

doi:10.1039/D4DD00007B

Investigating the reliability and interpretability of machine learning frameworks for chemical retrosynthesis†

Friedrich Hastedt,

*^a Rowan M. Bailey,

^b Klaus Hellgardt,

^a Sophia N. Yaliraki,

^b Ehecatl Antonio del Rio Chanona

*^a and Dongda Zhang

*^c

Author affiliations

* Corresponding authors

^a Department of Chemical Engineering, Imperial College London, London, UK
E-mail: friedrich.hastedt18@ic.ac.uk, k.hellgardt@imperial.ac.uk, a.del-rio-chanona@imperial.ac.uk

^b Department of Chemistry, Imperial College London, London, UK
E-mail: r.bailey22@imperial.ac.uk, s.yaliraki@imperial.ac.uk

^c Department of Chemical Engineering, University of Manchester, Manchester, UK
E-mail: dongda.zhang@manchester.ac.uk

Abstract

Machine learning models for chemical retrosynthesis have attracted substantial interest in recent years. Unaddressed challenges, particularly the absence of robust evaluation metrics for performance comparison, and the lack of black-box interpretability, obscure model limitations and impede progress in the field. We present an automated benchmarking pipeline designed for effective model performance comparisons. With an emphasis on user-friendly design, we aim to streamline accessibility and facilitate utilisation within the research community. Additionally, we suggest and perform a new interpretability study to uncover the degree of chemical understanding acquired by retrosynthesis models. Our results reveal that frameworks based on chemical reaction rules yield the most diverse, chemically valid, and feasible reactions, whereas purely data-driven frameworks suffer from unfeasible and invalid predictions. The interpretability study emphasises that incorporating reaction rules not only enhances model performance but also improves interpretability. For simple molecules, we show that Graph Neural Networks identify relevant functional groups in the product molecule, offering model interpretability. Sequence-to-sequence Transformers are not found to provide such an explanation. As the molecule and reaction mechanism grow more complex, both data-driven models propose unfeasible disconnections without offering a chemical rationale. We stress the importance of incorporating chemically meaningful descriptors within deep-learning models. Our study provides valuable guidance for the future development of retrosynthesis frameworks.

This article is part of the themed collection: Celebrating the 200th Anniversary of the University of Manchester

Supplementary files

Transparent peer review

To support increased transparency, we offer authors the option to publish the peer review history alongside their article.

View this article’s peer review history

Article information

DOI: https://doi.org/10.1039/D4DD00007B
Article type: Paper
Submitted: 13 jan 2024
Accepted: 23 mai 2024
First published: 23 mai 2024
This article is Open Access

Download Citation

Digital Discovery, 2024,3, 1194-1212

Permissions

Request permissions

Investigating the reliability and interpretability of machine learning frameworks for chemical retrosynthesis

F. Hastedt, R. M. Bailey, K. Hellgardt, S. N. Yaliraki, E. A. del Rio Chanona and D. Zhang, Digital Discovery, 2024, 3, 1194 DOI: 10.1039/D4DD00007B

This article is licensed under a Creative Commons Attribution-NonCommercial 3.0 Unported Licence. You can use material from this article in other publications, without requesting further permission from the RSC, provided that the correct acknowledgement is given and it is not used for commercial purposes.

To request permission to reproduce material from this article in a commercial publication, please go to the Copyright Clearance Center request page.

If you are an author contributing to an RSC publication, you do not need to request permission provided correct acknowledgement is given.

If you are the author of this article, you do not need to request permission to reproduce figures and diagrams provided correct acknowledgement is given. If you want to reproduce the whole article in a third-party commercial publication (excluding your thesis/dissertation for which permission is not required) please go to the Copyright Clearance Center request page.

Digital Discovery

Investigating the reliability and interpretability of machine learning frameworks for chemical retrosynthesis†

Abstract

Supplementary files

Transparent peer review

Article information

Download Citation

Permissions

Investigating the reliability and interpretability of machine learning frameworks for chemical retrosynthesis

Social activity

Search articles by author

Spotlight

Advertisements