Skip to content

Retire the WENO7 amdflang workaround; record which ones AFAR 24.3 still needs - #1939

Merged
sbryngelson merged 1 commit into
MFlowCode:masterfrom
sbryngelson:afar24-retire-workarounds
Oct 5, 2026
Merged

sbryngelson merged 1 commit into
MFlowCode:masterfrom
sbryngelson:afar24-retire-workarounds

Conversation

@sbryngelson

Copy link
Copy Markdown
Member

Summary

After the move to AFAR 24.3.0 (#1920), this checks which amdflang workarounds are still needed, removes the one that isn't, and records the results.

Workaround AFAR 24.3 Action
WENO7 coefficients written element-wise (#1660) the false-nuw fix, llvm/llvm-project#198014, is in the drop's source; the WENO7 tests pass at -O3 removed: back to reversed-stride sections
-attributor-max-pi-accesses=16384 (#1759) the cap is unchanged in 24.3 (cl::init(512)), and the GPU image has 486 kernels against a threshold of about 512 (ROCm/llvm-project#4070, fix proposed in #4094) kept; comment notes 24.3
HLLC split GPU_PARALLEL_LOOP call site merging it faults every HLLC case on 24.3 (28/28, memory access fault in the merged kernel) kept; comment notes 24.3
USING_AMD fixed-bound guards on 24.3, turning them off gives correct results but is 4-5x slower: private arrays sized by runtime globals spill to scratch (WENO 11x, HLLC 4.9x; rocprofv3) kept; docs corrected to say why
No GPU kernels inside BLOCK a minimal reproducer runs correctly on 24.3; not verified inside MFC rule kept; docs note it
Re_size host copies (#1588) the amdflang bug is fixed (ROCm/llvm-project#2890), but the scalar copies also satisfy the Cray OpenACC array-element trap (#1815) that source lint enforces kept, unchanged

Testing

On HPCFund MI210 (gfx90a), AFAR 24.3.0, amdflang --gpu mp release build (-O3), no case optimization:

  • 161 tests: all viscous (63), hypoelastic (62), and WENO7 (12) tests, 30 IBM tests, and samples of HLL, HLLC, and LF. 161/161 passed.
  • ./mfc.sh bench against master on the same node: geometric-mean speed ratio 1.006, every case within ±2%.

These runs used a build that also read Re_size directly. That change was dropped afterwards for the Cray reason above, so the Riemann solvers are back to master's code and this PR's only code change is the WENO7 revert. The WENO7 revert affects every compiler, so CI covers it on each one.

Results are unchanged: every test matches its golden file.

This PR was prepared with Claude Code (AI-assisted).

Acknowledgement

  • I confirm this PR meets the above expectations and reflects my own understanding and real-world context.

…ll needs

- Restore the reversed-stride sections in the WENO7 coefficients
  (MFlowCode#1660). The false-nuw fix (llvm/llvm-project#198014) is in AFAR 24.3.
- Note that the attributor cap flag (MFlowCode#1759) and HLLC's split call site
  are still needed on 24.3 (cap unchanged, ROCm/llvm-project#4070;
  merging the call sites faults every HLLC case).
- Correct the docs: on 24.3 the USING_AMD guards are needed for
  performance (4-5x), not correctness; note the BLOCK reproducer.
Copilot AI balanced review requested due to automatic review settings October 3, 2026 02:55
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Lines of Code

File Lines Diff
src/simulation/m_weno.fpp 1321 -19
Directory Lines Diff
simulation 27976 -19
total 46920 -19

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The restored expressions are equivalent to the element-wise assignments, and the documentation accurately reflects the retained workarounds.

Review effort: Balanced
Findings: None

What changed in this PR

Retires the obsolete WENO7 amdflang workaround and documents the remaining AFAR 24.3 constraints.

Changes:

  • Restores concise reversed-stride WENO7 coefficient assignments.
  • Records retained HLLC and linker workarounds.
  • Updates AMD performance and block guidance.
File Description
src/​simulation/​m_weno.fpp Restores reversed-stride WENO7 sections.
src/​simulation/​m_riemann_solver_hllc.fpp Notes the HLLC workaround remains necessary.
docs/​documentation/​gpuParallelization.md Documents current AFAR 24.3 findings.
cmake/​MFCTargets.cmake Records why the Attributor override remains.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@codecov

codecov Bot commented Oct 3, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 62.77%. Comparing base (c7feed2) to head (cd2550c).

Additional details and impacted files
@@            Coverage Diff             @@
##           master    #1939      +/-   ##
==========================================
- Coverage   62.80%   62.77%   -0.04%     
==========================================
  Files          86       86              
  Lines       22385    22366      -19     
  Branches     3304     3304              
==========================================
- Hits        14060    14041      -19     
  Misses       6073     6073              
  Partials     2252     2252              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@sbryngelson
sbryngelson merged commit bee66c2 into MFlowCode:master Oct 5, 2026
148 of 150 checks passed
@sbryngelson
sbryngelson deleted the afar24-retire-workarounds branch October 5, 2026 00:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants