Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
48 changes: 45 additions & 3 deletions .github/workflows/unsloth-sd-prebuilt.yml
Original file line number Diff line number Diff line change
Expand Up @@ -425,7 +425,8 @@ jobs:
# prefix (it ships as libcublas / libcublas-dev), so it has to go in the
# other list or apt cannot find it. cudart-dev is what carries the headers;
# cudart alone is the runtime and does not compile anything.
sub-packages: '["nvcc", "cudart", "cudart-dev", "thrust"]'
# nvrtc-dev: cudnn-frontend's headers include nvrtc.h (it dlopens libnvrtc itself, so nothing links it).
sub-packages: '["nvcc", "cudart", "cudart-dev", "thrust", "nvrtc-dev"]'
non-cuda-sub-packages: '["libcublas", "libcublas-dev"]'

# This leg rebuilt every object on every run: 3903 s of the 4121 s job on
Expand Down Expand Up @@ -456,6 +457,29 @@ jobs:
# keeps what it compiled.
save: false

- name: cuDNN headers (build time only)
run: |
set -euo pipefail
# ggml-cuda's cuDNN attention path (GGML_CUDA_CUDNN) compiles against cudnn.h, but the
# library itself is opened with dlopen when the first attention op runs and is NOT
# bundled: a host without a loadable libcudnn.so.9 keeps the ggml attention kernels.
# Only the headers package is fetched (80 KB), pinned by version and checksum.
deb=libcudnn9-headers-cuda-12_9.27.0.42-1_amd64.deb
curl -fsSL --retry 3 -o "$RUNNER_TEMP/$deb" \
"https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/$deb"
echo "68243ab04422da565e0bafec00f64ac08cfd82d805c8015c0c503af7ab4dd1d6 $RUNNER_TEMP/$deb" | sha256sum -c -
dpkg -x "$RUNNER_TEMP/$deb" "$RUNNER_TEMP/cudnn-deb"
# The package ships cudnn*_v9.h; the unversioned names normally come from the -dev
# package's alternatives, so link them here.
inc="$RUNNER_TEMP/cudnn-include"
mkdir -p "$inc"
for f in "$RUNNER_TEMP"/cudnn-deb/usr/include/x86_64-linux-gnu/cudnn*_v9.h; do
b="$(basename "$f")"
ln -s "$f" "$inc/${b%_v9.h}.h"
done
test -e "$inc/cudnn.h"
echo "CUDNN_INCLUDE_DIR=$inc" >> "$GITHUB_ENV"

- name: Build sd-cli + sd-server (CUDA)
working-directory: src
run: |
Expand All @@ -474,6 +498,8 @@ jobs:
-DSD_WEBP=OFF -DSD_WEBM=OFF \
-DGGML_NATIVE=OFF \
-DSD_CUDA=ON \
-DGGML_CUDA_CUDNN=ON \
-DGGML_CUDA_CUDNN_INCLUDE_DIR="$CUDNN_INCLUDE_DIR" \
-DCMAKE_CUDA_ARCHITECTURES="$CUDA_ARCHS" \
-DCMAKE_C_COMPILER_LAUNCHER=ccache \
-DCMAKE_CXX_COMPILER_LAUNCHER=ccache \
Expand Down Expand Up @@ -502,6 +528,14 @@ jobs:
patchelf --set-rpath '$ORIGIN' "$BIN/$exe"
done
ldd "$BIN/sd-cli" | sed -n '1,40p'
# cuDNN and NVRTC are opened at run time only; a link-time dependency would stop the binaries
# loading on every host without it.
for exe in sd-cli sd-server; do
if readelf -d "$BIN/$exe" | grep -qE 'NEEDED.*(cudnn|nvrtc)'; then
echo "ERROR: $exe links libcudnn or libnvrtc; they must stay run-time dlopens" >&2
exit 1
fi
done

- name: Package bundle
env:
Expand All @@ -511,8 +545,16 @@ jobs:
LABEL: Linux-Ubuntu-22.04-x86_64-cuda12
COMMIT: ${{ needs.resolve.outputs.commit }}
SOURCE_REPO: ${{ github.repository }}
LICENSE_FILE: ${{ github.workspace }}/src/LICENSE
run: python3 tooling/scripts/unsloth/package_bundle.py
LICENSE_FILE: ${{ runner.temp }}/LICENSE
run: |
set -euo pipefail
# The binaries compile in cudnn-frontend's header-only templates, so its MIT notice ships too.
{
cat src/LICENSE
printf '\n\n---- cudnn-frontend (https://github.com/NVIDIA/cudnn-frontend) ----\n\n'
cat src/build/_deps/ggml_cudnn_frontend-src/LICENSE.txt
} > "$LICENSE_FILE"
python3 tooling/scripts/unsloth/package_bundle.py

- name: Upload bundle
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
Expand Down
22 changes: 22 additions & 0 deletions docs/minimax_h3.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,28 @@ shapes; the arithmetic per output element is unchanged). `GGML_CUDA_FA_LONGSEQ=0
stock kernel. `GGML_CUDA_FA_LONGSEQ_NCOLS=128` opts into a wider tile that is faster on A100 and L4
but not bit-identical.

Builds configured with `-DGGML_CUDA_CUDNN=ON` (the Linux CUDA prebuilt is) can run that unmasked
attention, and the `--sage-attn` attention, through cuDNN's fused attention. By default it replaces
the flash attention kernel on Ada (sm89) and B200-class (sm100) GPUs and the sage kernel on sm100 only,
where it measured faster; on an RTX 3090 (sm86) both existing kernels were faster, on an RTX 6000 Ada
the sage kernel was, and other architectures were not measured.
Only `cudnn.h` is needed at build time (`GGML_CUDA_CUDNN_INCLUDE_DIR`; the header-only
cudnn-frontend v1.26.0 is downloaded by CMake unless `GGML_CUDA_CUDNN_FRONTEND_DIR` points at a
checkout). `libcudnn.so.9` is opened when the first attention op runs and is not shipped: put a cuDNN 9
built for the same CUDA major version as the binary (cuDNN 9 for CUDA 12 for the prebuilt) on the
library path or beside the binary, or point `GGML_CUDA_CUDNN_LIB` at it (cuDNN 9.27 needs nothing
else; older 9.x releases build these kernels with NVRTC and also need that CUDA major's `libnvrtc`).
Without a loadable cuDNN, on Turing and older, for attention calls under about a million scores
(where the extra conversions cost more than cuDNN saves) and for any shape cuDNN declines, the
existing kernels run. The first call of each attention shape builds a cuDNN plan (about 0.4 to 1.3 s
on a B200, logged once). cuDNN runs F16 with F32 accumulation: for the H3 DiT at 960x544x124 it
takes about 8 ms per attention call on a B200 against 28 ms for the sage kernel and 43 ms for the
ggml flash attention kernel, and it is closer to an exact (fp64) reference than either, so frames
and audio differ slightly from the previous kernels. `GGML_CUDA_CUDNN_ATTN=0` restores them and
`GGML_CUDA_CUDNN_ATTN=1` uses cuDNN for the flash attention op on any Ampere or newer GPU;
`GGML_CUDA_CUDNN_SAGE=0` / `=1` does the same for the sage op only, and
`GGML_CUDA_CUDNN_ATTN_BF16=1` runs cuDNN in BF16 (less accurate, same speed).

The video VAE decodes one 16x16 latent tile per decoder graph by default. `SD_H3_VAE_TILE=N`
uses N x N latent tiles instead (20 decodes about 0.7 s faster at 960x544x124 on B200); the tile
seams move, so the frames differ from the default (36 dB PSNR at 20) and it is opt-in.
Expand Down
Loading