Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
ac39900
Carry the pipelined long-sequence ggml-cuda FlashAttention patch
danielhanchen Oct 1, 2026
ddd8ac4
ggml patch 0002: 128-column long-sequence FlashAttention tile on sm80…
danielhanchen Oct 1, 2026
fbc3a40
docs: MiniMax-H3 long-sequence flash attention switches
danielhanchen Oct 1, 2026
a565a22
tiling: batched tile callback, merge tile planes on several threads
danielhanchen Oct 1, 2026
32861d8
MiniMax-H3: batch video VAE tiles per decoder graph, keep weights res…
danielhanchen Oct 1, 2026
1f1bef0
ggml-patches: 0003 fused cuBLAS bias epilogues and narrow short-row R…
danielhanchen Oct 1, 2026
dbb4106
MiniMax-H3: value-preserving DiT graph rewrites
danielhanchen Oct 1, 2026
49244df
Carry the fused DiT ops ggml patch; document the H3 graph switches
danielhanchen Oct 1, 2026
36963a4
MiniMax-H3: keep the folded MLP path compiling with SD_USE_UPSTREAM_GGML
danielhanchen Oct 1, 2026
69babaf
ggml patch 0002: long-sequence flash attention keeps the stock grid, …
danielhanchen Oct 1, 2026
b92b2a2
MiniMax-H3: video VAE tile batching opt-in (SD_H3_VAE_TILE_BATCH=auto|N)
danielhanchen Oct 1, 2026
2e76e65
Merge branch 'master' into perf/h3-gguf-speed-v2
oobabooga Oct 8, 2026
ff84251
Merge remote-tracking branch 'origin/perf/h3-gguf-speed' into pr18-v2
oobabooga Oct 8, 2026
5ab239d
Build without the fused H3 ops when ggml lacks patch 0004
oobabooga Oct 8, 2026
703f519
Document that the H3 ggml speedups need the carried patches
oobabooga Oct 8, 2026
5f222b8
Merge remote-tracking branch 'origin/perf/h3-gguf-speed' into perf/h3…
oobabooga Oct 8, 2026
d2f58e3
Merge remote-tracking branch 'origin/perf/h3-gguf-speed' into perf/h3…
oobabooga Oct 9, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions cmake/ggml.cmake
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,18 @@ endif()
target_include_directories(${SD_LIB} PRIVATE "${sd_ggml_private_include}")
set_property(TARGET ${SD_LIB} PROPERTY SD_GGML_PRIVATE_INCLUDE_DIR "${sd_ggml_private_include}")

# The fused DiT ops come from scripts/unsloth/ggml-patches/0004; an unpatched ggml keeps the unfused H3 graph.
set(sd_ggml_header "${sd_ggml_private_include}/../include/ggml.h")
if(EXISTS "${sd_ggml_header}")
set_property(DIRECTORY APPEND PROPERTY CMAKE_CONFIGURE_DEPENDS "${sd_ggml_header}")
file(STRINGS "${sd_ggml_header}" sd_ggml_h3_fused_ops REGEX "ggml_rope_pe_permute\\(")
endif()
if(sd_ggml_h3_fused_ops)
target_compile_definitions(${SD_LIB} PUBLIC SD_GGML_H3_FUSED_OPS)
else()
message(STATUS "ggml lacks the fused DiT ops patch: MiniMax-H3 uses the unfused graph")
endif()

if(SD_USE_UPSTREAM_GGML)
target_compile_definitions(${SD_LIB} PUBLIC SD_USE_UPSTREAM_GGML)
message(WARNING "Using upstream GGML: INT8 tensorwise/convrot is disabled and FP8 weights are converted to F16 at load time. Some operators may be unsupported and performance may be lower than with patched GGML.")
Expand Down
28 changes: 28 additions & 0 deletions docs/minimax_h3.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,34 @@ which removes most of its attention cost without touching the text encoder. Set
keep that attention by default (on Vulkan the flash attention decode was slower);
`SD_H3_VAE_FLASH_ATTN=1` turns it on there.

On NVIDIA GPUs from Ampere on, the unmasked DiT and VAE attention runs a pipelined
long-sequence variant of the ggml-cuda flash attention kernel (up to 4x faster at H3
shapes; the arithmetic per output element is unchanged). `GGML_CUDA_FA_LONGSEQ=0` restores the
stock kernel. `GGML_CUDA_FA_LONGSEQ_NCOLS=128` opts into a wider tile that is faster on A100 and L4
but not bit-identical.

The video VAE decodes one 16x16 latent tile per decoder graph by default.
`SD_H3_VAE_TILE_BATCH=auto` puts several tiles into one graph, sized from free device memory (at
most 4 unless `SD_H3_VAE_TILE_BATCH_MAX` raises it), and `SD_H3_VAE_TILE_BATCH=N` forces N; the
batched projections can round differently from the per-tile decode on some GPUs, so it is
opt-in. The decoder weights stay on the device across temporal chunks
(`SD_H3_VAE_KEEP_RESIDENT=0` releases them after every chunk), and the decoder blocks use a
table-based rotary embedding and a fused SwiGLU (`SD_H3_VAE_GRAPH_OPT=0` restores the previous
graph). Each batched tile still goes through its own attention call.

The DiT blocks use fused ggml ops (CPU and CUDA) for the work around the matmuls and attention:
partial RoPE with the attention relayout and the K/V scale and F16 cast, per-segment adaLN
modulation and gated residuals written in place, and the MLP Linear scales folded into those ops
and the swiglu. The result is bit-identical to the unfused graph. `SD_H3_GRAPH_FAST=0` restores
the unfused graph; `SD_H3_FAST_QKV=0`, `SD_H3_FAST_MLP=0`, `SD_H3_FAST_SEGMENTS=0` and
`SD_H3_FAST_VIEWS=0` turn off one part each.

The long-sequence flash attention kernel, the fused cuBLAS epilogues and the fused DiT ops come from
the ggml patches in `scripts/unsloth/ggml-patches`, which the Unsloth prebuilt binaries carry. A
source build gets them by applying the patches to the `ggml` submodule before configuring
(`for p in scripts/unsloth/ggml-patches/*.patch; do git -C ggml apply "../$p"; done`); without
them the build uses the stock ggml kernels and the unfused DiT graph.

## First/last-frame conditioning

Add `--init-img` for I2VA, or both `--init-img` and `--end-img` for FL2VA:
Expand Down
563 changes: 563 additions & 0 deletions scripts/unsloth/ggml-patches/0002-ggml-cuda-fa-longseq.patch

Large diffs are not rendered by default.

Large diffs are not rendered by default.

1,020 changes: 1,020 additions & 0 deletions scripts/unsloth/ggml-patches/0004-ggml-h3-dit-fused-ops.patch

Large diffs are not rendered by default.

67 changes: 58 additions & 9 deletions src/core/ggml_extend.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -207,15 +207,10 @@ ggml_tensor* ggml_ext_gelu_quick(ggml_context* ctx,
return x;
}

ggml_tensor* ggml_ext_linear(ggml_context* ctx,
ggml_tensor* x,
ggml_tensor* w,
ggml_tensor* b,
bool force_prec_f32,
float scale) {
if (scale != 1.f) {
x = ggml_ext_scale(ctx, x, scale);
}
ggml_tensor* ggml_ext_linear_matmul(ggml_context* ctx,
ggml_tensor* x,
ggml_tensor* w,
bool force_prec_f32) {
if (x->ne[2] * x->ne[3] > 1024) {
// workaround: avoid ggml cuda error
int64_t ne2 = x->ne[2];
Expand All @@ -232,6 +227,19 @@ ggml_tensor* ggml_ext_linear(ggml_context* ctx,
ggml_mul_mat_set_prec(x, GGML_PREC_F32);
}
}
return x;
}

ggml_tensor* ggml_ext_linear(ggml_context* ctx,
ggml_tensor* x,
ggml_tensor* w,
ggml_tensor* b,
bool force_prec_f32,
float scale) {
if (scale != 1.f) {
x = ggml_ext_scale(ctx, x, scale);
}
x = ggml_ext_linear_matmul(ctx, x, w, force_prec_f32);
if (scale != 1.f) {
x = ggml_ext_scale(ctx, x, 1.f / scale);
}
Expand Down Expand Up @@ -796,6 +804,47 @@ ggml_tensor* ggml_ext_attention_ext(ggml_context* ctx,
return kqv;
}

ggml_tensor* ggml_ext_attention_prepared(ggml_context* ctx,
ggml_backend_t backend,
ggml_tensor* q,
ggml_tensor* k,
ggml_tensor* v,
int64_t n_head,
int64_t N,
float kv_scale) {
GGML_ASSERT(q->type == GGML_TYPE_F32 && k->type == GGML_TYPE_F16 && v->type == GGML_TYPE_F16);
const int64_t L_q = q->ne[1];
const int64_t d_head = v->ne[0];
if (backend == nullptr || d_head < 64) {
return nullptr;
}

// Same expressions as ggml_ext_attention_ext so the softmax scale rounds identically.
float scale = (1.0f / sqrt((float)d_head));

auto kqv = ggml_flash_attn_ext(ctx, q, k, v, nullptr, scale / kv_scale, 0, 0);
if (!ggml_backend_supports_op(backend, kqv)) {
return nullptr;
}
ggml_flash_attn_ext_set_prec(kqv, GGML_PREC_F32);
if (kv_scale != 1.0f) {
kqv = ggml_ext_scale(ctx, kqv, 1.0f / kv_scale);
}
kqv = ggml_view_4d(ctx,
kqv,
d_head,
n_head,
L_q,
N,
kqv->nb[1],
kqv->nb[2],
kqv->nb[1] * n_head,
0);
kqv = ggml_ext_cont(ctx, kqv);
kqv = ggml_reshape_3d(ctx, kqv, d_head * n_head, L_q, N); // [N, L_q, C]
return kqv;
}

ggml_tensor* ggml_ext_layer_norm(ggml_context* ctx,
ggml_tensor* x,
ggml_tensor* w,
Expand Down
21 changes: 21 additions & 0 deletions src/core/ggml_extend.h
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,12 @@ ggml_tensor* ggml_ext_gelu_quick(ggml_context* ctx,
ggml_tensor* x,
bool inplace = false);

// The matrix multiply of ggml_ext_linear without its input/output scaling and bias.
ggml_tensor* ggml_ext_linear_matmul(ggml_context* ctx,
ggml_tensor* x,
ggml_tensor* w,
bool force_prec_f32);

ggml_tensor* ggml_ext_linear(ggml_context* ctx,
ggml_tensor* x,
ggml_tensor* w,
Expand Down Expand Up @@ -223,6 +229,21 @@ ggml_tensor* ggml_ext_attention_ext(ggml_context* ctx,
float kv_scale = 1.0f,
bool sage_attn = false);

// Flash attention on inputs that are already in the head-major layout the kernel reads:
// q: F32 [d_head, L_q, n_head * N], k/v: F16 [d_head, L_k, n_kv_head * N] with the kv_scale
// already applied. Builds the same flash_attn_ext + output scaling as ggml_ext_attention_ext
// (skip_reshape, unmasked), so the result is identical; returns nullptr when the backend lacks
// flash attention for these inputs (the caller then uses ggml_ext_attention_ext).
// return: [N, L_q, n_head * d_head]
ggml_tensor* ggml_ext_attention_prepared(ggml_context* ctx,
ggml_backend_t backend,
ggml_tensor* q,
ggml_tensor* k,
ggml_tensor* v,
int64_t n_head,
int64_t N,
float kv_scale);

ggml_tensor* ggml_ext_layer_norm(ggml_context* ctx,
ggml_tensor* x,
ggml_tensor* w,
Expand Down
24 changes: 24 additions & 0 deletions src/model/common/ggml_block.hpp
Original file line number Diff line number Diff line change
Expand Up @@ -204,6 +204,30 @@ class Linear : public UnaryBlock {
force_prec_f32 = force_prec_f32_;
}

// For callers that fold this layer's input/output scaling into neighbouring fused ops: the
// weight when forward() is exactly ggml_ext_linear(x, weight, nullptr, prec_f32(),
// effective_scale()) (no bias, weight scale, adapter, INT8 or FP8 path), nullptr otherwise.
ggml_tensor* foldable_weight(GGMLRunnerContext* ctx) {
ggml_tensor* w = params["weight"];
if (bias || has_weight_scale || ctx->weight_adapter || w->type == GGML_TYPE_I8) {
return nullptr;
}
#ifndef SD_USE_UPSTREAM_GGML
if (w->type == GGML_TYPE_F8_E4M3 || w->type == GGML_TYPE_F8_E5M2) {
return nullptr;
}
#endif
return w;
}

float effective_scale(GGMLRunnerContext* ctx) const {
return ctx->linear_scale > 0.f ? ctx->linear_scale : scale;
}

bool prec_f32() const {
return force_prec_f32;
}

ggml_tensor* forward(GGMLRunnerContext* ctx, ggml_tensor* x) override {
ggml_tensor* w = params["weight"];
const float scale = ctx->linear_scale > 0.f ? ctx->linear_scale : this->scale;
Expand Down
Loading
Loading