TITLE:
Finite-Volume-Aligned Physics Regularisation for Reservoir Pressure Surrogates under Distribution Shift: Robustness, Residual Monitoring, and Resolution-Transfer Limits
AUTHORS:
Franck-Hilaire Essiagne, Kouassi Louis Kra, Moussa Camara
KEYWORDS:
Physics-Informed Machine Learning, Surrogate Modelling, Reservoir Simulation, Darcy Flow, Finite Volume, Distribution Shift, Residual Monitoring, Selective Simulation, Resolution Transfer
JOURNAL NAME:
Engineering,
Vol.18 No.10,
October
8,
2026
ABSTRACT: Fast reservoir surrogates are attractive for uncertainty quantification, optimisation, and repeated forecasting, yet data-only networks can lose accuracy and local conservation when geology or well controls leave the training distribution. This study isolates the marginal value of a finite-volume-aligned physics loss and then examines whether its residual can support trustworthy deployment. A 32 × 32 benchmark was generated for steady single-phase Darcy flow over heterogeneous permeability fields and variable injector-producer configurations. Identical 8089-parameter dilated convolutional networks were trained from 60 labelled simulations using pressure error alone (Data-60) or the same objective plus the normalized residual of the finite-volume stencil that generated the labels (PI-60). For the three-seed ensemble on fixed tests, PI-60 reduced mean relative pressure error by 5.3% on Test-ID and 10.4% on Fixed Compound-OOD (Test-OOD), with a 14.1% reduction on the independently regenerated Independent Combined-Shift stress test: P90 and P95 Fixed Compound-OOD errors decreased by 13.0% and 14.7%. The raw residual separated the designed ID and OOD populations extremely well (AUC = 0.997) but ranked case-wise OOD pressure error weakly (Spearman ρ = 0.148; worst-20%-error AUROC = 0.648). Stability-aware weighting materially improved this diagnostic role: a five-step Jacobi-preconditioned conjugate-gradient inverse-action score increased ρ to 0.382 and AUROC to 0.790, while the exact inverse-weighted energy score provided an analysis ceiling of ρ = 0.466 and AUROC = 0.838. In selective-simulation analysis, the preconditioned score reduced risk-coverage AURC by 4.6% relative to the raw residual; combining it with seed disagreement did not help. Continuous severity sweeps further showed that PI-60 reduced integrated pressure error by 5.7% for shortening correlation length and 6.9% for increasing well rate, while reducing linear degradation slopes by 48.3% and 27.0%. Finally, a zero-shot 16 × 16/32 × 32/64 × 64 test exposed an important limit: the finite-volume reference showed decreasing adjacent-grid discrepancy and an empirical order near 2.38, whereas both learned surrogates became less coherent under refinement. PI-60 reduced the 32-to-64 discrepancy by 4.5% relative to Data-60 but did not restore convergence. These results distinguish fixed-grid finite-volume alignment from genuine discretization consistency and support residual monitoring only when stability, calibration, and resolution limits are made explicit. Batched CPU inference remained approximately 51 times faster than the reference finite-volume solve.