Skip to content

Add -fopenmp-assume-no-thread-state to the amdflang offload flags (~1.2x on AFAR 24.3) - #1940

Merged
sbryngelson merged 2 commits into
MFlowCode:masterfrom
sbryngelson:amdflang-no-thread-state
Oct 3, 2026
Merged

sbryngelson merged 2 commits into
MFlowCode:masterfrom
sbryngelson:amdflang-no-thread-state

Conversation

@sbryngelson

Copy link
Copy Markdown
Member

Summary

Adds -fopenmp-assume-no-thread-state to the amdflang (LLVMFlang) OpenMP-offload compile flags. This is the one part of -fopenmp-target-fast that MFC did not already pass: on AFAR 24.3, -fopenmp-target-fast expands to -O3 -fopenmp-assume-no-nested-parallelism -fopenmp-assume-no-thread-state, and the first two are already in the flags. #1450 dropped -fopenmp-target-fast because it was in the reproducer for #1449.

The flag promises that no kernel changes OpenMP ICVs (the runtime's internal settings, such as thread counts) inside a target region, which lets the compiler drop the per-kernel thread-state bookkeeping. MFC's only runtime-setting call is omp_set_default_device, made on the host before any kernel launches. The GPU macros emit no num_threads, omp_set_*, or nested parallelism.

Results

HPCFund MI210 (gfx90a), AFAR 24.3.0, release build (-O3), no case optimization, ./mfc.sh bench --mem 2, same node as the master baseline. Grind time (lower is better):

case master this PR speedup
5eq_rk3_weno3_hll 2.702 2.259 1.20x
5eq_rk3_weno3_hllc 2.495 2.094 1.19x
5eq_rk3_weno3_lf 2.448 2.078 1.18x
hypo_hll 1.885 1.402 1.35x
ibm 5.449 4.466 1.22x
igr 2.618 2.217 1.18x
viscous_weno5_sgb_acoustic 4.024 3.442 1.17x

Geometric mean: 1.21x. Same-node run-to-run noise in these measurements is about 1-2%. The Frontier (AMD) Bench job on this PR gives an independent MI250X number.

Testing

207 tests on the same build: all viscous, hypoelastic, WENO7, IBM (where #1449's corruption appeared), and chemistry tests, plus samples of HLL, HLLC, and LF. 207/207 passed.

Only amdflang GPU builds are affected; other compilers do not see this flag.

This PR was prepared with Claude Code (AI-assisted).

Acknowledgement

  • I confirm this PR meets the above expectations and reflects my own understanding and real-world context.

No MFC kernel changes OpenMP ICVs on the device, so the per-kernel
thread-state bookkeeping can go. On AFAR 24.3 (MI210) every benchmark
case gets 1.17-1.35x faster (geomean 1.21x); 207 tests pass. This is
the only part of -fopenmp-target-fast (dropped in MFlowCode#1450) not already set.
Copilot AI balanced review requested due to automatic review settings October 3, 2026 03:44

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The isolated compiler-flag change matches MFC’s device runtime usage and has targeted correctness and performance validation.

Review effort: Balanced
Findings: None

What changed in this PR

Adds a verified LLVMFlang OpenMP-offload optimization that removes unnecessary device thread-state bookkeeping.

Changes:

  • Adds -fopenmp-assume-no-thread-state for LLVMFlang GPU builds.
  • Documents the assumption and measured performance benefit.
File Description
cmake/​MFCTargets.cmake Adds the optimization flag to LLVMFlang offload compilation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@codecov

codecov Bot commented Oct 3, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 62.80%. Comparing base (c7feed2) to head (198c5c0).

Additional details and impacted files
@@           Coverage Diff           @@
##           master    #1940   +/-   ##
=======================================
  Coverage   62.80%   62.80%           
=======================================
  Files          86       86           
  Lines       22385    22385           
  Branches     3304     3304           
=======================================
  Hits        14060    14060           
  Misses       6073     6073           
  Partials     2252     2252           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@sbryngelson
sbryngelson merged commit 98c1f4e into MFlowCode:master Oct 3, 2026
86 checks passed
@sbryngelson
sbryngelson deleted the amdflang-no-thread-state branch October 3, 2026 22:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants