Featurization strategies for polymer sequence or composition design by machine learning

Roshan A. Patel; Carlos H. Borca; Michael A. Webb

doi:10.1039/D1ME00160D

Featurization strategies for polymer sequence or composition design by machine learning†

Roshan A. Patel,^a Carlos H. Borca^a and Michael A. Webb

*^a

Author affiliations

* Corresponding authors

^a Department of Chemical and Biological Engineering, Princeton University, Princeton, NJ, USA
E-mail: mawebb@princeton.edu

Abstract

The emergence of data-intensive scientific discovery and machine learning has dramatically changed the way in which scientists and engineers approach materials design. Nevertheless, for designing macromolecules or polymers, one limitation is the lack of appropriate methods or standards for converting systems into chemically informed, machine-readable representations. This featurization process is critical to building predictive models that can guide polymer discovery. Although standard molecular featurization techniques have been deployed on homopolymers, such approaches capture neither the multiscale nature nor topological complexity of copolymers, and they have limited application to systems that cannot be characterized by a single repeat unit. Herein, we present, evaluate, and analyze a series of featurization strategies suitable for copolymer systems. These strategies are systematically examined in diverse prediction tasks sourced from four distinct datasets that enable understanding of how featurization can impact copolymer property prediction. Based on this comparative analysis, we suggest directly encoding polymer size in polymer representations when possible, adopting topological descriptors or convolutional neural networks when the precise polymer sequence is known, and using chemically informed unit representations when developing extrapolative models. These results provide guidance and future directions regarding polymer featurization for copolymer design by machine learning.

This article is part of the themed collections: Machine Learning and Artificial Intelligence: A cross-journal collection and Emerging Investigator Series

Article information

https://doi.org/10.1039/D1ME00160D

Article type

Paper

Submitted

03 Nov 2021

Accepted

10 Mar 2022

First published

21 Mar 2022

Download Citation

Mol. Syst. Des. Eng., 2022,7, 661-676

Permissions

Request permissions

Featurization strategies for polymer sequence or composition design by machine learning

R. A. Patel, C. H. Borca and M. A. Webb, Mol. Syst. Des. Eng., 2022, 7, 661 DOI: 10.1039/D1ME00160D

To request permission to reproduce material from this article, please go to the Copyright Clearance Center request page.

If you are an author contributing to an RSC publication, you do not need to request permission provided correct acknowledgement is given.

If you are the author of this article, you do not need to request permission to reproduce figures and diagrams provided correct acknowledgement is given. If you want to reproduce the whole article in a third-party publication (excluding your thesis/dissertation for which permission is not required) please go to the Copyright Clearance Center request page.

Molecular Systems Design & Engineering

Featurization strategies for polymer sequence or composition design by machine learning†

Abstract

Supplementary files

Article information

Download Citation

Author version available

Permissions

Featurization strategies for polymer sequence or composition design by machine learning

Social activity

Search articles by author

Spotlight

Advertisements