Skip to content

MiniMax-H3 / ggml-cuda: cuDNN fused attention on Ampere and newer (B200 max render 15.4 to 11.3 s) - #3

Open
oobabooga wants to merge 9 commits into
perf/h3-gguf-speed-v3from
perf/h3-cudnn-attn
Open

oobabooga wants to merge 9 commits into
perf/h3-gguf-speed-v3from
perf/h3-cudnn-attn

Commits

  1. Commits on Oct 4, 2026

  2. Commits on Oct 8, 2026

  3. Commits on Oct 9, 2026