We present ProcGNN, a physics-informed graph neural network for stream-level chemical process surrogate modeling. Implemented directly in PyTorch, ProcGNN uses shared unit and stream representations to predict temperature, pressure, mass flow rate, and component mole fractions across different process flowsheets.
The overall architecture of ProcGNN.
Rigorous process simulation is widely used for process design and optimization, but repeated simulations can be computationally expensive. Conventional surrogate models are usually developed for a fixed process structure and a predefined set of input and output variables. Reusing them becomes difficult when unit operations, material-stream connections, or prediction targets change. ProcGNN addresses this problem by representing unit operations as nodes and material streams as directed edges. A shared prediction function is applied to each stream, allowing the prediction set to follow the streams represented in the supplied flowsheet.
ProcGNN combines three main components. First, it encodes unit-operation features, stream features, and feed conditions using common definitions across flowsheets. Feed-ratio features represent the air-to-methane and water-to-methane relationships, while a relation gate incorporates heat-transfer interactions between related hot- and cold-side streams. After four bidirectional GNN layers, each stream prediction combines its source and destination unit representations, stream features, a Set2Set graph representation, and directly encoded feed conditions. This gives each prediction access to both information from connected units and information from the entire flowsheet.
Second, flow-aware message passing treats the two propagation directions separately. Forward propagation aggregates source-unit information at each destination through incoming streams; backward propagation aggregates destination-unit information at each source through outgoing streams. In both directions, attention is normalized over the outgoing streams of the same source unit, allowing the model to compare branches that lead to different destinations. Differential encoding provides each update with a learned feature of the difference between the aggregated message and the unit’s previous representation. Residual fusion also retains the previous and initial unit representations as messages propagate through the graph.
Third, physics-informed training penalizes conservation violations using predicted inlet and outlet stream properties. Mass balances apply to valid units with represented inlet and outlet streams, component balances apply to non-reactive units, and atom balances apply to reactive units. These unit-level rules can be reused across flowsheets wherever the corresponding units and streams are represented. The conservation penalties are combined with supervised stream-property losses after an initial supervised training period. Separate prediction heads estimate temperature and pressure, composition, and mass flow rate; softmax ensures that the predicted component mole fractions are nonnegative and sum to one.
Experiments cover ten steam methane reforming (SMR) flowsheets under single-process prediction, multi-process prediction, zero-shot transfer, and limited-data fine-tuning. Compared with the strongest competing model in each setting, ProcGNN reduces average sMAPE from 0.1140 to 0.0681 in single-process prediction and from 0.6773 to 0.2196 in multi-process prediction, corresponding to relative error reductions of 40.2% and 67.6%, respectively. In zero-shot transfer to an unseen flowsheet, it achieves the lowest sMAPE for eight of the ten stream properties. Ablations further show that the architectural components and the three conservation terms contribute to prediction accuracy and conservation satisfaction.
A short description of ProcGNN:
- Stream-level prediction across flowsheets: Unit operations and material streams are represented as nodes and directed edges. Shared prediction heads estimate temperature, pressure, mass flow rate, and the mole fractions of H₂O, H₂, CH₄, CO₂, CO, O₂, and N₂ for the streams represented in each flowsheet. Set2Set readout and direct feed encoding supplement the endpoint unit and stream features.
- Flow-aware bidirectional message passing: Forward and backward propagation aggregate information over incoming and outgoing neighborhoods, respectively. Attention is normalized across each source unit’s outgoing streams to compare branches. Differential encoding and residual fusion combine neighborhood information with the unit’s own representation.
- Unit-level conservation regularization: Predicted inlet and outlet streams are used to compute mass, component, and atom conservation penalties at applicable units. The same balance rules are reused across flowsheets to encourage physical consistency during supervised training.
- Evaluation on unseen flowsheets: A source model is trained on nine SMR flowsheets and evaluated on the remaining unseen flowsheet before and after fine-tuning. These experiments test whether the shared model transfers to a process structure absent from training.
ProcGNN is available at:
Cite “ProcGNN” as:
@misc{procgnn,
title = {A Physics-Informed Stream-Level Graph Neural Network for Multi-Flowsheet Chemical Process Surrogate Modeling},
author = {Jun Hee Cho and Jun Ho Eom and Huu-Tuong Ho and Jun Young Kim and O-Joun Lee},
year = {},
eprint = {},
archivePrefix = {},
primaryClass = {},
doi = {},
url = {},
}