Skip to content

Fixed per-thread array bounds on all GPU compilers, with a generated state-vector layout - #1941

Open
sbryngelson wants to merge 6 commits into
MFlowCode:masterfrom
sbryngelson:fixed-bounds
Open

sbryngelson wants to merge 6 commits into
MFlowCode:masterfrom
sbryngelson:fixed-bounds

Conversation

@sbryngelson

Copy link
Copy Markdown
Member

Summary

Per-thread arrays in GPU kernels now get compile-time extents in every GPU simulation build, from one place, and case optimization supplies exact extents including sys_size.

Runtime-sized private arrays (dimension(num_fluids), dimension(sys_size)) cannot live in registers and spill to scratch. The USING_AMD guards fixed this for amdflang only; profiling on AFAR 24.3 (rocprofv3) showed WENO 11x and HLLC 4.9x slower without them. #1938 turned the same fixed bounds on for every compiler, and CI showed CCE 12-36% and NVHPC OpenMP 9-62% faster (AMD unchanged), but GNU CPU up to 39% slower. So this PR applies them to GPU builds only.

Changes (four commits)

  1. BOUND() replaces the USING_AMD guards. The 83 copy-pasted #:if [not MFC_CASE_OPTIMIZATION and] USING_AMD blocks become single declarations such as dimension(${BOUND('num_fluids')}$). BOUND returns the fixed maximum from one table (num_fluids 3, nb 3, sys_size 70 or 10 + species, WENO and QBMM extents, and so on) or the runtime extent. A converter checked every existing literal against that table. MFC_FIXED_BOUNDS (CMake, default ON) is decided per target: only the offloaded simulation uses fixed bounds, so pre_process/post_process callers always match the routines they call. s_check_amd becomes s_check_fixed_bounds, and the sys_size limit is checked right after the layout is computed (the old input-time check ran before sys_size existed).
  2. The state-vector layout is generated from one table. toolchain/mfc/params/eqn_layout.py lists each model's fields with their size and condition as small Python expressions. The params codegen turns them into generated_eqn_idx.fpp, replacing about 140 hand-written lines of s_initialize_eqn_idx. Checked against master's routine over all 7,962,624 combinations of model, the 13 layout flags, 1D/multi-D, and the counts: every eqn_idx field and sys_size matches.
  3. Case optimization bakes in exact extents. It evaluates the same table for the case and writes CASE_OPT_SIZES (sys_size and the hypoelastic stress count) into case.fpp, so BOUND gives exact compile-time extents under case optimization on every compiler. The sys_size variable stays a runtime variable, so nothing that declares or updates it on the device changes. Simulation checks at startup that the runtime layout matches the baked-in values. Since CASE_OPT_SIZES is in case.fpp, it is part of the build fingerprint.
  4. Fixed bounds are on for every GPU compiler (NVHPC and CCE, OpenACC and OpenMP); CPU builds keep runtime extents. Docs describe BOUND().

Behavior change

GPU simulation builds without --case-optimization now accept at most 3 fluids, nb <= 3, and sys_size <= 70 (or 10 + species) on every GPU compiler, as amdflang already did; larger cases rebuild with --case-optimization, and the error message says so. CPU builds and pre_process/post_process are unaffected.

Testing

  • Generated code: for amdflang simulation, and for gfortran simulation/pre_process/post_process, the preprocessed Fortran matches master apart from the checker rename and messages.
  • HPCFund MI210, AFAR 24.3, --gpu mp, release: full test suite, 718/718 passed.
  • gfortran CPU: full test suite, 718/718 passed.
  • Case-optimized 5eq_rk3_weno3_hllc on GPU: runs with CASE_OPT_SIZES = {'sys_size': 8} baked in (3D, two fluids), HLLC and CBC arrays sized exactly 8, startup check passes; grind 1.93 against 2.50 without case optimization (--mem 2).
  • AMD benchmark against master, same node (--mem 2): geometric-mean ratio 1.000, every case within ±2%. AMD already had fixed bounds, so no change is expected; the gains are on NVHPC and CCE.
  • NVHPC and CCE: the CI Test and Bench jobs on this PR.

This PR was prepared with Claude Code (AI-assisted).

Acknowledgement

  • I confirm this PR meets the above expectations and reflects my own understanding and real-world context.

…itch

Per-thread arrays in GPU kernels get compile-time extents because
runtime-sized private arrays spill to scratch (4-5x on amdflang). The 83
copy-pasted `#:if [not MFC_CASE_OPTIMIZATION and] USING_AMD` blocks become
single declarations using ${BOUND('name')}$, which returns the fixed maximum
from one table (num_fluids 3, nb 3, sys_size 70 or 10 + species, ...) or the
runtime extent. A converter checked every literal against that table.

MFC_FIXED_BOUNDS (CMake, default ON) is decided per target: only the
offloaded simulation uses fixed bounds, so pre/post_process callers always
match the routines they call. This commit keeps it amdflang-only, and the
generated code for amdflang simulation and for every non-fixed build matches
master. s_check_amd becomes s_check_fixed_bounds and absorbs the CBC limits.
…table

toolchain/mfc/params/eqn_layout.py lists each model's fields in order with
their size and condition as small Python expressions. The params codegen
translates them into generated_eqn_idx.fpp, which s_initialize_eqn_idx now
includes in place of 140 hand-written lines (the hypoelastic shear tables
stay hand-written). evaluate_layout() computes the same layout for a case,
which case optimization will use for a compile-time sys_size.

Checked against master's routine over all 7,962,624 combinations of model,
the 13 layout flags, 1D/multi-D, and the fluid/velocity/dimension/bubble/
species counts: every eqn_idx field and sys_size match, and evaluate_layout
agrees on sys_size for a 300,000-case sample.
Case optimization now evaluates the layout table for the case and writes
CASE_OPT_SIZES (sys_size and the hypoelastic stress count) into case.fpp.
BOUND() returns those exact values, or the case-optimized parameters, under
case optimization on any compiler; before, sys_size arrays fell back to the
runtime value or, on amdflang, to the fixed maximum of 70.

The sys_size variable itself stays a runtime variable filled by the
generated layout, so nothing that declares or updates it on the device
changes. Simulation checks at startup that the runtime layout matches the
baked-in values. CASE_OPT_SIZES is part of case.fpp, so it is already in the
build fingerprint: a case with a different layout gets its own build.
MFC_FIXED_BOUNDS now applies to NVHPC and CCE (OpenACC and OpenMP), not
just amdflang. CPU builds keep runtime extents: with GNU the fixed bounds
cost up to 39% on the 5-equation cases. In MFlowCode#1938's CI benchmarks the fixed
bounds sped up CCE by 12-36% and NVHPC OpenMP by 9-62%, with AMD unchanged.

GPU simulation builds without case optimization now take at most 3 fluids,
nb <= 3, and sys_size <= 70 (or 10 + species); larger cases rebuild with
--case-optimization, and s_check_fixed_bounds says so. Docs describe BOUND()
and why runtime-sized per-thread arrays are slow on GPUs.
Copilot AI balanced review requested due to automatic review settings October 3, 2026 06:59

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Fixed-size GPU dummy arguments conflict with shorter host-side simulation arrays in common one- and two-fluid paths.

Review effort: Balanced
Findings: 2 High severity · 2 Low severity

Open (4)
What changed in this PR

Centralizes GPU per-thread array bounds and generates equation-state layouts for consistent runtime and case-optimized sizing.

Changes:

  • Introduces BOUND() with GPU maxima and case-specific extents.
  • Generates eqn_idx and sys_size from a shared Python layout.
  • Updates validation, tests, CMake code generation, and documentation.
File Description
toolchain/​mfc/​params/​generators/​fortran_gen.py Generates equation-index includes.
toolchain/​mfc/​params/​eqn_layout.py Defines and evaluates state-vector layouts.
toolchain/​mfc/​params_tests/​test_fortran_gen.py Updates generated-file tests.
toolchain/​mfc/​params_tests/​test_eqn_layout.py Tests layout evaluation and translation.
toolchain/​mfc/​lint_source.py Renames the fixed-bound checker exemption.
toolchain/​mfc/​case.py Computes case-optimized extents.
src/​simulation/​m_weno.fpp Applies fixed WENO scratch bounds.
src/​simulation/​m_viscous.fpp Applies fluid and dimension bounds.
src/​simulation/​m_time_steppers.fpp Bounds timestep scratch arrays.
src/​simulation/​m_surface_tension.fpp Bounds capillary tensors.
src/​simulation/​m_start_up.fpp Validates baked and maximum sizes.
src/​simulation/​m_sim_helpers.fpp Bounds helper arrays.
src/​simulation/​m_riemann_state.fpp Bounds Riemann-state scratch tensors.
src/​simulation/​m_riemann_solver_lf.fpp Bounds LF solver arrays.
src/​simulation/​m_riemann_solver_hypo_hlld.fpp Bounds hypoelastic HLLD arrays.
src/​simulation/​m_riemann_solver_hlld.fpp Bounds HLLD fluid arrays.
src/​simulation/​m_riemann_solver_hllc.fpp Bounds HLLC state arrays.
src/​simulation/​m_riemann_solver_hll.fpp Bounds HLL fluid arrays.
src/​simulation/​m_qbmm.fpp Bounds QBMM scratch arrays.
src/​simulation/​m_pressure_relaxation.fpp Bounds relaxation arrays.
src/​simulation/​m_ibm.fpp Bounds immersed-boundary scratch arrays.
src/​simulation/​m_data_output.fpp Bounds runtime diagnostic arrays.
src/​simulation/​m_compute_cbc.fpp Bounds CBC helper arguments.
src/​simulation/​m_cbc.fpp Bounds CBC working arrays.
src/​simulation/​m_bubbles_EL.fpp Bounds Lagrangian-bubble arrays.
src/​simulation/​m_bubbles_EE.fpp Bounds Eulerian-bubble arrays.
src/​simulation/​m_acoustic_src.fpp Bounds acoustic-source fluid arrays.
src/​common/​m_variables_conversion.fpp Bounds conversion-kernel arrays.
src/​common/​m_phase_change.fpp Bounds and initializes phase arrays.
src/​common/​m_global_parameters_common.fpp Consumes generated equation layouts.
src/​common/​m_eos.fpp Bounds EOS helper arguments.
src/​common/​m_checker_common.fpp Enforces fixed-bound limits.
src/​common/​include/​shared_parallel_macros.fpp Defines centralized bound selection.
docs/​documentation/​gpuParallelization.md Documents BOUND() usage.
docs/​documentation/​contributing.md Updates generated-file documentation.
CMakeLists.txt Adds the fixed-bounds option.
cmake/​ParamsCodegen.cmake Registers generated layout files.
cmake/​Fypp.cmake Enables fixed bounds per target.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/common/m_eos.fpp
Comment on lines +645 to +648
real(wp), intent(in) :: pres, rho, gamma, pi_inf
real(wp), dimension(${BOUND('num_fluids')}$), intent(in) :: adv
real(wp), intent(out) :: c
real(wp), dimension(${BOUND('num_fluids')}$), intent(in), optional :: alpha_rho

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 41a73ae. The probe-output locals (s_write_probe_files) and s_report_icfl_violation's locals are now ${BOUND('num_fluids')}$. The ICFL path was worse than nonconforming: it reaches the conversion kernel via s_compute_cell_state, whose whole-array alpha_K normalization wrote past a one- or two-fluid host array under mpp_lim (see the other thread).

Comment on lines +190 to +191
real(wp), dimension(${BOUND('num_fluids')}$), intent(inout) :: alpha_rho_K, alpha_K
real(wp), optional, dimension(${BOUND('num_fluids')}$), intent(in) :: G

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 41a73ae. The wrapper's alpha_K/alpha_rho_K locals and both wrappers' optional G now use ${BOUND('num_fluids')}$; every caller that passes G passes fluid_pp(:)%G (sized num_fluids_max), so that stays conforming. The kernel also now normalizes only alpha_K(1:num_fluids): the whole-array form wrote past a short host actual with mpp_lim, and divided the uninitialized padding even with a conforming one.

Comment thread docs/documentation/contributing.md Outdated
registers a single ninja-tracked `add_custom_command` (DEPENDS all `params/*.py`) that
invokes `cmake_gen.py` and writes the 18 generated includes under
invokes `cmake_gen.py` and writes the 21 generated includes under
`build/include/<target>/`. There is no configure-time generation: all 18 files are build

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 41a73ae: 21 throughout (matches the list in cmake/ParamsCodegen.cmake).

Comment thread src/simulation/m_start_up.fpp Outdated
@:PROHIBIT(hypoelasticity .and. eqn_idx%stress%end - eqn_idx%stress%beg + 1 /= ${CASE_OPT_SIZES['n_stress']}$, &
& "stress count differs from the case-optimized build; rebuild it")
#:elif MFC_FIXED_BOUNDS
@:PROHIBIT(sys_size > ${SYS_SIZE_MAX}$, "sys_size <= ${SYS_SIZE_MAX}$ in GPU builds")

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 41a73ae: the message now ends with "; rebuild with --case-optimization", like the num_fluids/nb checks.

@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.00000% with 7 lines in your changes missing coverage. Please review.
✅ Project coverage is 62.77%. Comparing base (f5ef7e7) to head (41a73ae).
⚠️ Report is 5 commits behind head on master.

Files with missing lines Patch % Lines
src/common/m_checker_common.fpp 0.00% 4 Missing ⚠️
src/common/m_global_parameters_common.fpp 87.50% 0 Missing and 2 partials ⚠️
src/simulation/m_data_output.fpp 83.33% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##           master    #1941      +/-   ##
==========================================
- Coverage   62.80%   62.77%   -0.04%     
==========================================
  Files          86       86              
  Lines       22387    22295      -92     
  Branches     3306     3291      -15     
==========================================
- Hits        14061    13995      -66     
+ Misses       6071     6060      -11     
+ Partials     2255     2240      -15     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

# Conflicts:
#	docs/documentation/gpuParallelization.md
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

Lines of Code

File Lines Diff
src/simulation/m_compute_cbc.fpp 172 -102
src/common/m_global_parameters_common.fpp 159 -91
src/simulation/m_riemann_state.fpp 1152 -33
src/simulation/m_qbmm.fpp 850 -29
src/common/m_eos.fpp 490 -28
src/common/m_variables_conversion.fpp 936 -24
src/simulation/m_riemann_solver_hllc.fpp 1285 -23
src/simulation/m_sim_helpers.fpp 238 -17
src/simulation/m_ibm.fpp 1423 -16
src/common/include/shared_parallel_macros.fpp 176 +15
src/simulation/m_cbc.fpp 1101 -13
src/simulation/m_riemann_solver_hypo_hlld.fpp 773 -10
src/simulation/m_viscous.fpp 1012 -9
src/simulation/m_weno.fpp 1312 -9
src/simulation/m_pressure_relaxation.fpp 208 -8
src/simulation/m_surface_tension.fpp 255 -8
src/simulation/m_riemann_solver_lf.fpp 502 -7
src/simulation/m_riemann_solver_hll.fpp 601 -6
src/simulation/m_bubbles_EE.fpp 299 -5
src/simulation/m_data_output.fpp 1537 -5
src/simulation/m_time_steppers.fpp 876 -5
src/common/m_phase_change.fpp 281 -4
src/simulation/m_acoustic_src.fpp 512 -4
src/simulation/m_bubbles_EL.fpp 1635 -4
src/simulation/m_riemann_solver_hlld.fpp 188 -4
src/simulation/m_start_up.fpp 1241 -3
Directory Lines Diff
common 10284 -132
simulation 27656 -320
total 46468 -452

In GPU simulation builds BOUND('num_fluids') is 3, but three host paths still
passed num_fluids-sized arrays to routines with BOUND dummies:
s_convert_species_to_mixture_variables (its alpha_K/alpha_rho_K locals and the
forwarded G), s_report_icfl_violation, and s_write_probe_files. With one or two
fluids and mpp_lim, the kernel's whole-array normalization of alpha_K wrote past
the end of the host local. Declare those with BOUND (every G caller passes
fluid_pp(:)%G, sized num_fluids_max), and normalize only alpha_K(1:num_fluids).

Also: fix the generated-file count in contributing.md (21, not 18), and tell
users to rebuild with --case-optimization when sys_size exceeds the GPU maximum.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants