Repository navigation
Reject truncated restart files and incompatible boundary decompositions - #1948
sbryngelson wants to merge 6 commits into
Conversation
A job killed while writing restart_data/lustre_<step>.dat leaves a short file. MPI reads past its end return short without an error, so the next restart silently read the missing tail as garbage (Inf/NaN pressure). Check the file size against the bytes about to be read and abort with a clear message, for both the shared file and file_per_process.
There was a problem hiding this comment.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Copilot review overview
4 open findings
n_MOKis computed fromm_glb_readinstead ofn_glb_read. This can make the computed… · New The expression2*nb*nnodeis evaluated in default integer arithmetic before being added to… · NewierrfromMPI_FILE_GET_SIZEis not checked. If this MPI call fails,file_bytesmay be… · New The phrase'holds <file_bytes> of <expected_bytes> expected bytes'is a bit grammatically… · New
What changed in this PR
Adds a restart-file size sanity check to prevent silent short MPI reads from truncated restart files, aborting with a clear error message instead of proceeding with garbage data.
Changes:
- Compute the number of variables per cell to be read (including optional QBMM data) as an MPI-offset-sized integer.
- Check restart file size via
MPI_FILE_GET_SIZEin bothfile_per_processand shared-file branches before reading. - Introduce
s_check_restart_file_sizehelper to centralize truncation detection and abort messaging.
| File | Description |
|---|---|
| src/simulation/m_start_up.fpp | Adds pre-read restart file size validation and a helper routine to abort on truncated restart files. |
🧠 Review effort: Lite
Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #1948 +/- ##
==========================================
- Coverage 62.64% 61.79% -0.85%
==========================================
Files 86 86
Lines 22425 22783 +358
Branches 3325 3355 +30
==========================================
+ Hits 14048 14079 +31
- Misses 6119 6214 +95
- Partials 2258 2490 +232 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Lines of Code
|
4612c60 to
d0c59cd
Compare
pre_process writes restart_data/boundary_conditions/bc_<rank>.dat per rank, and simulation reads them by its own rank index. Run on a different rank count, each rank silently reads another rank's boundary slab, and a Dirichlet (-17) inflow then blows up within a few steps. pre_process now records the decomposition in decomposition.dat. Simulation aborts with a clear message when it differs; post_process, which only uses the data for boundary ghost cells, warns. Files from before this change are not checked. The per-rank MPI_File_open calls are now checked as well. (cherry picked from commit 8e17a3c)
…not aborting (cherry picked from commit 3ed989a)



Summary
Truncated restart files and per-rank boundary files written for another MPI decomposition could be read silently as invalid simulation inputs. This PR combines file-size validation with boundary-decomposition validation and checked boundary-file opens. Post-processing retains the documented warning behavior for decomposition mismatches, and legacy boundary files without decomposition metadata remain compatible.
Consolidation
This existing PR retains its original commits and discussion. Changes from #1947 are added as separate commits with original authors/messages and
cherry-pick -xprovenance. The complete source PR descriptions, verification records, and limitations are reproduced below. Their original verification claims are historical records, not fresh runs on this combined head.Verification of the combined branch
git diff --checkpasses excluding generated golden metadata; its original formatting is retained../mfc.sh precheckpasses all seven gates: formatting, spelling, toolchain lint/tests, source lint, documentation references, parameter documentation, and example case validation.This consolidation was performed with OpenAI Codex.
Contribution Policy
We do not accept pull requests generated primarily by AI without genuine understanding or real-world usage context.
All contributions are expected to demonstrate:
If these expectations are not met, we would prefer to implement the changes ourselves rather than spend time reviewing low-effort submissions.
Acknowledgement
PR template credit: junegunn
Original PR documentation
#1948: Restart: abort on a truncated restart file instead of reading garbage
Source: #1948
Original head:
d0c59cd90ed0fd26ae4230ff5aa3b86f272897a3Complete original PR description
Description
A restart file truncated by a job killed while writing it (
restart_data/lustre_<step>.dat) is read without any check, and the run either starts from garbage (in production: Inf/NaN pressure) or dies without saying why.Root cause.
s_read_parallel_data_files(src/simulation/m_start_up.fpp) opens the restart file and readssys_sizevariables (plus the qbmmpb/mvwhen present) at offsets computed from the global grid. It never compares the file size to what it reads. An MPI read past end-of-file returns short without raising an error, so the missing tail comes back as whatever was in the buffer.Fix. The new helper
s_check_restart_file_sizecomparesMPI_FILE_GET_SIZEagainst the bytes about to be read, and aborts with a message naming the file and both sizes. It is called in both the shared-file and thefile_per_processbranches. It only checks for a file that is too short, so files with extra trailing data (e.g. Lagrangianbeta) still read. Intact files are unaffected.Verification
Production (OLCF Frontier). A job killed during a save left a short
lustre_<step>.dat. The next restart read it as Inf pressure. The workaround was to pick the latest restart whose size equals nx·ny·nz·sys_size·8 bytes.This branch. I ran this on a Frontier CPU compute node (GNU 12.3 + Cray MPICH, Release) with a 2D 50x40 case on 2 ranks. I ran to step 20 with saves every 10 steps, truncated
lustre_10.datfrom 80000 to 40000 bytes, and restarted from step 10.Restart file ./restart_data/lustre_10.dat holds 40000 of 80000 expected bytes. It is truncated, e.g. by a job killed while writing it; restart from an earlier step.masterRestarting from the intact file on this branch runs normally.
simulationcompiles with CCE 19 (CPU) and GNU 12.3. Precheck passes, apart from twotest_thermochemcases that fail on the Frontier login node because they compile with the system/usr/bin/gfortran(addressed by #1943). Existing goldens are unaffected, because they read intact files. No regression test is added: the harness cannot express an expected abort.Contribution Policy
We do not accept pull requests generated primarily by AI without genuine understanding or real-world usage context.
All contributions are expected to demonstrate:
If these expectations are not met, we would prefer to implement the changes ourselves rather than spend time reviewing low-effort submissions.
Acknowledgement
The problem was hit in production runs on Frontier; the fix was exercised on the small case above.
PR template credit: junegunn
#1947: BC I/O: refuse per-rank boundary files written for another decomposition
Source: #1947
Original head:
3ed989ac96da521d9d210c1c3de084034b5d0d10Complete original PR description
Description
Restarting (or just running
simulation) on a different rank count frompre_processsilently corrupts Dirichlet (-17) inflow and boundary-patch data.Root cause. When
bc_iois on (any-17boundary ornum_bc_patches > 0),pre_processwrites the boundary types and buffers per rank, torestart_data/boundary_conditions/bc_<rank>.dat(s_write_parallel_boundary_condition_files).simulationandpost_processopenbc_<proc_rank>.datby their own rank index and read it unchecked (s_read_parallel_boundary_condition_files). The restart file itself is in a global layout and reads on any rank count, so the only thing tying a run to the pre_process decomposition is this directory, and nothing checks it. On another rank count each rank gets another rank's boundary slab, with the wrong position and possibly the wrong size. The per-rankMPI_File_openwas not checked either, so a missing file went unnoticed too.Fix (minimal).
pre_processrank 0 also writesrestart_data/boundary_conditions/decomposition.dat, containingnum_procs num_procs_x num_procs_y num_procs_z.simulationaborts on a mismatch, with a message that names both decompositions.post_processonly uses the data for boundary ghost cells, so it warns instead of aborting (new optionalstrictargument). Workflows that post-process on a different rank count keep working, but are told that boundary ghost values are wrong.decomposition.datand are read as before, unchecked.MPI_File_opencalls now go throughs_check_mpi_file_open.Not done here: the full fix. The full fix would write and read these arrays in a global parallel-I/O layout, like the restart file, so that any rank count works. I did not do it in this PR because it is not contained:
-buff_size:m+buff_sizehalos, which overlap between ranks and include corner cells, andbuff_sizecan differ between targets;Verification
Production (OLCF Frontier, CCE 19, OpenACC).
pre_processran on 192 ranks and the restart ran on 64 ranks at step 15714. The inflow cellj = 0reached ICFL 1.015 at step 15723. In a second run (360 ranks, then 240), p = -4e47 appeared at the inflow face at step 2. Runs that kept the rank count were unaffected.This branch. I ran these on a Frontier CPU compute node (GNU 12.3 + Cray MPICH, Release). The case is a 2D uniform channel, 50x40, with
bc_x%beg = -17.decomposition.dat=2 2 1 1./restart_data/boundary_conditions was written for 2 ranks (2x1x1) but this run uses 1 ranks (1x1x1). These per-rank boundary files only work on the decomposition that wrote them: run on that rank count, or rerun pre_process on this one.mastermastermaster(nodecomposition.dat)Builds:
pre_process,simulationandpost_processcompile with CCE 19 on CPU, andpre_processandsimulationwith GNU 12.3. The post_process warning path was compiled but not run. Precheck passes, apart from twotest_thermochemcases that fail on the Frontier login node because they compile with the system/usr/bin/gfortran(addressed by #1943).No regression test is added. The test harness compares goldens and has no way to expect an abort, and a test that restarts on a different rank count would need harness support.
Contribution Policy
We do not accept pull requests generated primarily by AI without genuine understanding or real-world usage context.
All contributions are expected to demonstrate:
If these expectations are not met, we would prefer to implement the changes ourselves rather than spend time reviewing low-effort submissions.
Acknowledgement
The bug was found in production runs on Frontier; the fix was exercised on the small cases above.
PR template credit: junegunn