Skip to content

nki(mean_reduction): NKI (Trainium) implementation - #292

Open
bowencui123 wants to merge 9 commits into
mainfrom
bowen/nki/mean_reduction
Open

bowencui123 wants to merge 9 commits into
mainfrom
bowen/nki/mean_reduction

Conversation

@bowencui123

@bowencui123 bowencui123 commented Aug 29, 2026

Copy link
Copy Markdown
Collaborator

NKI (AWS Trainium) implementation of mean_reduction, split out of the consolidated NKI branch cecilia/feature/nki-vector-add (nki-all-operators, #259) so each operator can be reviewed independently. Supersedes #184 (older per-operator branch: legacy neuronxcc.nki imports; this is the migrated import nki version).

Files: A benchmarks/operators/mean_reduction/impl_nki.py

Status: imports and exposes run()/get_last_config() on trn2 (nki 0.6.0); not individually re-benchmarked in this split

Implementation by @Cecilia123li. Timing/identity infrastructure: #261; Trainium peak/roofline infra: #262.

🤖 Generated with Claude Code

https://claude.ai/code/session_012Q38kGmXvyoeM1qtCbheSL

Autotune (7dc2c2c)

NKI tunables Triton counterpart note
block_size_m (rows per tile, <=128) BLOCK_M/BLOCK_N columns are reduced whole per row

autotune=False keeps the previous constants (default numbers unchanged). Validation on trn2, case 0 (default run + autotune code path with the candidate timer stubbed — no sweep; --autotune runs a real sweep):

# initial run
[mean_reduction] default : verify=OK (1s) 
[mean_reduction] autotune: verify=OK (0s) last_config={'block_size_m': 128} trace_records=1 
STUB_EXIT=0
  Torch (Neuron) baseline FAILED: torch-on-Neuron verification failed: Tensor-likes are not close!
Params      |    Dtype |  Torch(ms) |      NKI(ms) |  Speedup(N)
n=2097152   | fp16     |        nan |       0.0581 |         nan

Split out of the consolidated NKI branch cecilia/feature/nki-vector-add
(nki-all-operators, PR #259) so each operator can be reviewed on its own.
Supersedes PR #184 (older per-operator branch).

Co-Authored-By: Cecilia123li <68335867+Cecilia123li@users.noreply.github.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012Q38kGmXvyoeM1qtCbheSL
bowencui123 and others added 8 commits August 29, 2026 08:37
Tunables mirror the Triton search space (`BLOCK_M`/`BLOCK_N`); defaults are the previous constants,
so autotune=False is unchanged. columns are reduced whole per row

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012Q38kGmXvyoeM1qtCbheSL
Merges NKI backend timing into results/csv/mean_reduction_default.csv, run
against this branch's impl_nki.py on trn2.3xlarge with the LNC2
execution contract (NEURON_LOGICAL_NC_CONFIG=2, NEURON_RT_NUM_CORES=1,
NEURON_CC_FLAGS="--target trn2 --lnc 2").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AQseF7nyesBh8KZAp8g7Cm
…, LNC2)

Merges NKI backend timing into results/csv/mean_reduction_autotune.csv, run with
--autotune against this branch's impl_nki.py on trn2.3xlarge, LNC2
execution contract. All cases pass correctness verification.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AQseF7nyesBh8KZAp8g7Cm
# Conflicts:
#	results/csv/mean_reduction_autotune.csv
#	results/csv/mean_reduction_default.csv
…per column block, fp32 partials), no host pad; XLA torch baseline sums in fp32 for 16-bit inputs; rerun benchmarks

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ScXYNjrrKGgDUVNHxv7HJt
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant