Machine Learning-Based Prediction of Phase Equilibria and Interfacial Properties of Multicomponent CO2 Mixtures with Impurities

Application to CO2 Transportation

Journal Article (2026)
Author(s)

Darshan Raju (TU Delft - Mechanical Engineering)

Roar Skartlien (Institute for Energy Technology)

Mahinder Ramdin (TU Delft - Mechanical Engineering)

Thijs J.H. Vlugt (TU Delft - Mechanical Engineering)

Research Group
Engineering Thermodynamics
DOI related publication
https://doi.org/10.1021/acs.iecr.6c03388 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Research Group
Engineering Thermodynamics
Journal title
Industrial and Engineering Chemistry Research
Issue number
35
Volume number
65
Pages (from-to)
18863-18890
Page Views
23
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Impurities in captured CO2 broaden the two-phase envelope, decrease the interfacial tension (IFT), and affect safe dense-phase transport. These properties are difficult to measure and expensive to compute using molecular simulations. We present the first systematic machine learning (ML) surrogate framework for phase equilibria and interfacial properties of impure CO2, trained on Perturbed-Chain Polar Statistical Associating Fluid Theory (PCP-SAFT) equation of state (EoS) + cDFT data. Within this framework, the hyperparameter-optimized TabPFN most accurately predicts Pbubble, Pdew, IFT, and interfacial thickness. For IFT, we combined the Winterfeld–Scriven–Davis correlation with an ML-based residual correction. The hybrid model with stochastic variational Gaussian process regression reduces the root mean squared error (RMSE) of IFT to 0.02 mN m–1. Symbolic regression additionally provides an interpretable expression for this correction. For sampling CO2-rich compositions, active learning is the most data-efficient strategy. Trained surrogates predict each state point in ca. 1–5 ms, compared to ca. 0.18–1.8 s for cDFT, enabling rapid property estimation for computational fluid dynamics (CFD), inverse design, and the planning of simulations and experiments.