Add -fopenmp-assume-no-thread-state to the amdflang offload flags (~1.2x on AFAR 24.3) - #1940
Merged
sbryngelson merged 2 commits intoOct 3, 2026
Merged
Conversation
No MFC kernel changes OpenMP ICVs on the device, so the per-kernel thread-state bookkeeping can go. On AFAR 24.3 (MI210) every benchmark case gets 1.17-1.35x faster (geomean 1.21x); 207 tests pass. This is the only part of -fopenmp-target-fast (dropped in MFlowCode#1450) not already set.
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The isolated compiler-flag change matches MFC’s device runtime usage and has targeted correctness and performance validation.
Review effort: Balanced
Findings: None
What changed in this PR
Adds a verified LLVMFlang OpenMP-offload optimization that removes unnecessary device thread-state bookkeeping.
Changes:
- Adds
-fopenmp-assume-no-thread-statefor LLVMFlang GPU builds. - Documents the assumption and measured performance benefit.
| File | Description |
|---|---|
cmake/MFCTargets.cmake |
Adds the optimization flag to LLVMFlang offload compilation. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #1940 +/- ##
=======================================
Coverage 62.80% 62.80%
=======================================
Files 86 86
Lines 22385 22385
Branches 3304 3304
=======================================
Hits 14060 14060
Misses 6073 6073
Partials 2252 2252 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
-fopenmp-assume-no-thread-stateto the amdflang (LLVMFlang) OpenMP-offload compile flags. This is the one part of-fopenmp-target-fastthat MFC did not already pass: on AFAR 24.3,-fopenmp-target-fastexpands to-O3 -fopenmp-assume-no-nested-parallelism -fopenmp-assume-no-thread-state, and the first two are already in the flags. #1450 dropped-fopenmp-target-fastbecause it was in the reproducer for #1449.The flag promises that no kernel changes OpenMP ICVs (the runtime's internal settings, such as thread counts) inside a target region, which lets the compiler drop the per-kernel thread-state bookkeeping. MFC's only runtime-setting call is
omp_set_default_device, made on the host before any kernel launches. The GPU macros emit nonum_threads,omp_set_*, or nested parallelism.Results
HPCFund MI210 (gfx90a), AFAR 24.3.0, release build (
-O3), no case optimization,./mfc.sh bench --mem 2, same node as the master baseline. Grind time (lower is better):Geometric mean: 1.21x. Same-node run-to-run noise in these measurements is about 1-2%. The Frontier (AMD) Bench job on this PR gives an independent MI250X number.
Testing
207 tests on the same build: all viscous, hypoelastic, WENO7, IBM (where #1449's corruption appeared), and chemistry tests, plus samples of HLL, HLLC, and LF. 207/207 passed.
Only amdflang GPU builds are affected; other compilers do not see this flag.
This PR was prepared with Claude Code (AI-assisted).
Acknowledgement