nki(matmul_fp32_fp16_fp8): NKI (Trainium) implementation - #288
Open
bowencui123 wants to merge 4 commits into
Open
bowencui123 wants to merge 4 commits into
bowencui123 wants to merge 4 commits into
Conversation
Split out of the consolidated NKI branch cecilia/feature/nki-vector-add (nki-all-operators, PR #259) so each operator can be reviewed on its own. - also carries the operator's `config.yaml` change from the NKI branch Co-Authored-By: Cecilia123li <68335867+Cecilia123li@users.noreply.github.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Q38kGmXvyoeM1qtCbheSL
This was referenced Aug 29, 2026
…k_size_n`, `block_size_k`) Tunables mirror the Triton search space (`BLOCK_SIZE_M`/`BLOCK_SIZE_N`/`BLOCK_SIZE_K`); defaults are the previous constants, so autotune=False is unchanged. candidates are filtered to divisors of M/N/K; candidate timing runs on one core Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012Q38kGmXvyoeM1qtCbheSL
- NKI_MATMUL_NUM_CORES + the two-step reconciliation collapse to one rule: split M blocks across the LNC cores when they divide evenly, else 1 core (that was the net effect before; no knob, no dance). - The fp8/double_row machinery (DOUBLE_ROW kernel branch, strided fp8 PSUM transpose, NKI_MATMUL_DOUBLE_ROW env) was unreachable in the benchmark: the sweep's only fp8 dtype is e4m3fn, which neuronx-cc rejects before TRN3. run() now rejects all fp8 with that message; 324 -> 258 lines. - config.yaml plots addition reverted (unrelated to the NKI backend; the other matmul configs on main don't carry it). Validated on trn2 (runtime-trace framework, case 0, warmup 2 / repeat 10): fp32 default 1.0505 ms (torch 1.0424), fp16 default 0.2861 ms (torch 0.2589), autotune sweep timed 7 candidates for real and replayed the winner (block_size_m/n/k = 512/1024/1024) at 1.0507 ms. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011Wvs1tztZTGQFD78YdZaha
bowencui123
force-pushed
the
bowen/nki/matmul_fp32_fp16_fp8
branch
from
August 30, 2026 19:52
772f306 to
6b9763c
Compare
Collaborator
Author
|
Review pass (Bowen's assignment): Triton/cuTile alignment + scaffold removal + real autotune & default runs on trn2 (post-#309 runtime-trace timing framework, case 0, #288 matmul_fp32_fp16_fp8 — CHANGES PUSHED
🤖 Generated with Claude Code |
Code is AST-identical to the previous commit (verified); the design notes live in the PR review comments and git history. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011Wvs1tztZTGQFD78YdZaha
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
NKI (AWS Trainium) implementation of matmul_fp32_fp16_fp8, split out of the consolidated NKI branch
cecilia/feature/nki-vector-add(nki-all-operators, #259) so each operator can be reviewed independently.Files: M benchmarks/operators/matmul_fp32_fp16_fp8/config.yaml, A benchmarks/operators/matmul_fp32_fp16_fp8/impl_nki.py
Status: imports and exposes run()/get_last_config() on trn2 (nki 0.6.0); not individually re-benchmarked in this split
config.yamlchange from the NKI branchImplementation by @Cecilia123li. Timing/identity infrastructure: #261; Trainium peak/roofline infra: #262.
🤖 Generated with Claude Code
https://claude.ai/code/session_012Q38kGmXvyoeM1qtCbheSL
Autotune (4eb4f1a)
block_size_m,block_size_n,block_size_k(= 128/512/128 PE tiles x tiles-in-block)BLOCK_SIZE_M/BLOCK_SIZE_N/BLOCK_SIZE_Kautotune=Falsekeeps the previous constants (default numbers unchanged). Validation on trn2, case 0 (default run + autotune code path with the candidate timer stubbed — no sweep;--autotuneruns a real sweep):