Developing large language models for quantum chemistry simulation input generation

Pieter Floris Jacobs; Robert Pollice

doi:10.1039/D4DD00366G

Developing large language models for quantum chemistry simulation input generation†

Pieter Floris Jacobs

^a and Robert Pollice

*^a

Author affiliations

* Corresponding authors

^a Stratingh Institute for Chemistry, Nijenborgh 3, 9747 AG Groningen, The Netherlands
E-mail: r.pollice@rug.nl

Abstract

Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise. While large language models (LLMs) have shown impressive capabilities in synthesizing code from natural language prompts, they often struggle with DSLs, likely due to their limited exposure during training. In this work, we investigate the potential of foundational LLMs for generating input files for the quantum chemistry package ORCA by establishing a general framework that can be adapted to other DSLs. To improve upon Image ID:d4dd00366g-u1.gif as our base model, we explore the impact of prompt engineering, retrieval-augmented generation, and finetuning via synthetically generated datasets. We find that finetuning, even with synthetic datasets as small as 500 samples, significantly improves performance. Additionally, we observe that finetuning shows synergism with advanced prompt engineering such as chain-of-thought prompting. Consequently, our best finetuned models outperform the formally much more powerful Image ID:d4dd00366g-u2.gif model. In turn, finetuning GPT-4o with the same small synthetic dataset leads to a further substantial performance improvement, suggesting our approach to be more general rather than limited to LLMs with poor base proficiency. All tools and datasets are made openly available for future research. We believe that this research lays the groundwork for a wider adoption of LLMs for DSLs in chemistry and beyond.

Supplementary files

Transparent peer review

To support increased transparency, we offer authors the option to publish the peer review history alongside their article.

View this article’s peer review history

Article information

DOI: https://doi.org/10.1039/D4DD00366G
Article type: Paper
Submitted: 12 Nov 2024
Accepted: 04 Feb 2025
First published: 05 Feb 2025
This article is Open Access

Download Citation

Digital Discovery, 2025,4, 762-775

Permissions

Request permissions

Developing large language models for quantum chemistry simulation input generation

P. F. Jacobs and R. Pollice, Digital Discovery, 2025, 4, 762 DOI: 10.1039/D4DD00366G

This article is licensed under a Creative Commons Attribution 3.0 Unported Licence. You can use material from this article in other publications without requesting further permissions from the RSC, provided that the correct acknowledgement is given.

Digital Discovery

Developing large language models for quantum chemistry simulation input generation†

Abstract

Supplementary files

Transparent peer review

Article information

Download Citation

Permissions

Developing large language models for quantum chemistry simulation input generation

Social activity

Search articles by author

Spotlight

Advertisements