Skip to content

cecilia/feature/nki-batched-matmul & nki-matmul-fp32-fp16-fp8 - #174

Closed
Cecilia123li wants to merge 9 commits into
mainfrom
cecilia/feature/nki-batched-matmul
Closed

Cecilia123li wants to merge 9 commits into
mainfrom
cecilia/feature/nki-batched-matmul

Conversation

@Cecilia123li

Copy link
Copy Markdown
Collaborator

Added NKI kernel implementations for batched matmul and matmul fp32 fp16 fp8

@Cecilia123li
Cecilia123li marked this pull request as draft August 4, 2026 18:36
@Cecilia123li Cecilia123li changed the title Cecilia/feature/nki-batched-matmul & nki-matmul-fp32-fp16-fp8 cecilia/feature/nki-batched-matmul & nki-matmul-fp32-fp16-fp8 Aug 4, 2026
@Cecilia123li
Cecilia123li marked this pull request as ready for review August 4, 2026 18:37
@bowencui123

Copy link
Copy Markdown
Collaborator

Superseded by #269 (bowen/nki/batched_matmul), which carries the newer consolidated version of this operator from cecilia/feature/nki-vector-add (nki-all-operators, #259) — migrated import nki API plus the later fixes — re-based onto main as an operator-only diff (that PR also carried matmul_fp32_fp16_fp8; see #288 for it). Closing this older per-operator PR to keep one reviewable PR per operator; the branch is left in place.

Related: #261 (deterministic NEFF identity for NKI timing), #262 (Trainium peak/roofline infra).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants