Repository navigation
MiniMax-H3: flash attention in the video VAE, faster tile stitching, ggml-cuda speed patch - #17
danielhanchen wants to merge 4 commits into
Conversation
…rry ggml-cuda speed patch The H3 video VAE decoder is a 36-layer ViT run per 16x16 latent tile. With only --diffusion-fa (what Studio passes) it used mul_mat + scale + softmax attention, materialising a 32 x L x L f32 score matrix per tile and layer. --diffusion-fa now also covers it; SD_H3_VAE_FLASH_ATTN=0 restores the old path. The text encoder is unchanged (that is what --fa would also switch, which moves the conditioning). scripts/unsloth/ggml-patches/ carries two ggml-cuda commits on top of the pinned submodule (BF16 cuBLAS for large-batch quantized matmul behind GGML_CUDA_QUANT_CUBLAS_MIN_BATCH, a cached device integrated flag, a vectorized row copy); the prebuilt workflow applies them after the submodule checkout.
process_tiles_2d copied and blended each pixel across all planes in the inner loop, striding a whole output plane (960x544 floats for H3) per element. A video tile has frames x channels planes, so every element was a cache and TLB miss: on MiniMax-H3 960x544x124 (105 tiles) this was about 10 s of single-threaded host time per decode. Loop planes outermost and precompute the per-column and per-row blend weights; each element still computes old + new * wy * wx in the same order, so the output is bitwise identical.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Reached Codex review convergence at oobabooga#2. |
Problem
A MiniMax-H3 render (960x544, 124 frames) spends 30 to 140 s after sampling, mostly in the video VAE decode:
--diffusion-fadoes not reach it. It falls back to mul_mat + scale + softmax over 32xLxL f32 score matrices, about 6 s of GPU time per decode on a B200.process_tiles_2dsplits and merges tiles with the plane index in the inner loop, so every element is a cache and TLB miss across 105 tiles: about 10 s of single-threaded host time per decode.cudaGetDevicePropertiesabout 2,800 times per decode, and row copies use a scalar kernel.Changes
src/pipeline/diffusion_engine.cpp:--diffusion-faalso enables flash attention in the H3 video VAE. The text encoder is untouched. Kill switch:SD_H3_VAE_FLASH_ATTN=0. Documented indocs/minimax_h3.md.src/runtime/tiling.cpp: the tile split and merge loops walk planes outermost. Output is bitwise identical.scripts/unsloth/ggml-patches/0001-ggml-cuda-h3-speed.patch, applied by the prebuilt workflow after the ggml submodule checkout (the submodule points at a repository we do not push to; a patch that stops applying fails the build):GGML_CUDA_QUANT_CUBLAS_MIN_BATCH. Off by default, so default output is unchanged.GGML_CUDA_CPY_ROWS=0.Results
MiniMax-H3 UD-Q3_K_XL, 960x544, 124 frames, 4 steps, same seed, each arm a fresh
sd-clion the same Colab VM, built from this branch with the patch applied.--sage-attn--sage-attnAccuracy
SD_H3_VAE_FLASH_ATTN=0reproduces master exactly.Tests
tiling.cppagainst the same build and compares their output bitwise over 9 cases, including circular padding and encode: identical.test-backend-ops: CPY passes, and MUL_MAT passes 1186/1186 with the BF16 path forced on.Not covered