Skip to content

Add a Linux ROCm leg to the prebuilt pipeline, so AMD hosts stop landing on the CPU build - #14

Open
LeoBorcherding wants to merge 5 commits into
unslothai:masterfrom
LeoBorcherding:prebuilt-amd-legs
Open

LeoBorcherding wants to merge 5 commits into
unslothai:masterfrom
LeoBorcherding:prebuilt-amd-legs

Conversation

@LeoBorcherding

@LeoBorcherding LeoBorcherding commented Sep 15, 2026 •

Copy link
Copy Markdown

Studio's MiniMax-H3 GGUF path asks this mirror for an sd-cli built for rocm. The latest release has none: macOS, Linux CPU, Linux CUDA, Linux aarch64, Windows CPU. So the installer falls back to leejet's upstream ROCm build, which doesn't carry the H3 fixes in this repo, and when that binary doesn't list a GPU, video.py drops to the CPU build.

That's unslothai/unsloth#8814:

Minimax H3 won't run on AMD card using Linux

7900 XTX on CachyOS, H3 goes installed to load_failed within a second. A Discord report this week on the same card and distro is the other shape of it: H3 Q3_K fills 30 GB of RAM and 36 GB of swap while VRAM sits at 1.5 GB.

What this adds

A build-linux-rocm leg in unsloth-sd-prebuilt.yml, same rule as the CUDA leg: continue-on-error, not in the coverage gate, so it can never hold back the CPU, Apple and Vulkan assets. The Linux and Windows Vulkan legs this PR first added came to master with #22, so they were dropped when master was merged in.

leg runner asset
build-linux-rocm ubuntu-24.04 Linux-Ubuntu-24.04-x86_64-rocm-7.14.0
  • The leg builds against TheRock 7.14.0 wheels like upstream's build.yml, for 20 consumer targets (gfx1010 to gfx1201, including the gfx1150 to gfx1153 APUs), and ships the ROCm userspace it linked against: every library ldd resolves from the wheel tree, rocBLAS / hipBLASLt kernel trees, and the per-target .kpack archives the kpack-split libraries load. The binaries get rpath $ORIGIN/lib, the libraries $ORIGIN, and the leg fails if anything is still unresolved with the wheel tree hidden. Upstream's zip doesn't bundle the runtime, so it only runs on a host with a matching ROCm, which is the #8814 failure.
  • package_bundle.py gains KEEP_LAYOUT=1, which ships the staged tree as is (sd-cli, sd-server, lib/, .kpack/); the other legs still land flat.
  • src/model_manager.cpp: a device whose free-memory report exceeds its total was treated as Vulkan's budget underflow and given 0 bytes. On a ROCm integrated GPU (Strix Halo) ggml reports free memory from /proc/meminfo and total from hipMemGetInfo, so free is above total and every allocation was refused: every render aborted. That rejection now applies to Vulkan only, as upstream did in fix: handle GPU memory reports and LLM encoding failures leejet/stable-diffusion.cpp#2020, and a failed LLM prompt encode returns an error instead of aborting (the same upstream change).

Asset names are what Studio's resolve_release_asset already matches for rocm, so a ROCm host picks this up with no Studio change. Not here: Windows ROCm.

Validation

what result
the leg, run as written on a GitHub-hosted ubuntu-24.04 runner builds, bundles 21 ROCm libraries and 20 kpack archives, ldd gate passes; zip 1802 MiB (GitHub's per-asset limit is 2 GiB)
the bundle on a Strix Halo (gfx1151) host with ROCm 7.2.1 installed Z-Image 1024x1024 and MiniMax-H3 640x384x56 render on the GPU; the only file it maps from /opt/rocm is the optional libhsa-amd-aqlprofile64
the bundle in a clean ubuntu:24.04 container with no ROCm (only /dev/kfd, /dev/dri) both render on the GPU (the #8814 case)
upstream's ROCm zip that Studio falls back to today, same host renders, but loads HIP, hipBLAS, rocBLAS and the HSA runtime from the system ROCm
this tree without the model_manager.cpp change on gfx1151 every render aborts (available 0.00 MB device)
merged with the open MiniMax-H3 stack and #21 merges cleanly; the combined tree builds and renders on HIP, Vulkan and CUDA

Not run: a discrete AMD GPU (only the gfx1151 APU was available).

@oobabooga

Copy link
Copy Markdown
Member

Reached Codex review convergence at oobabooga#5.

@oobabooga

Copy link
Copy Markdown
Member

I pushed one change, 68136c0. On a ROCm integrated GPU (Strix Halo gfx1151) ggml reports free memory from /proc/meminfo (about 120 GB) and total from hipMemGetInfo (64 GB). ModelManager::check_capacity took free > total as Vulkan's budget underflow and gave 0 bytes, so every render aborted (available 0.00 MB device). Since this leg ships gfx1150 to gfx1153 kernels, those hosts would have received a build that cannot render. The commit is the needed part of upstream leejet#2020: the rejection now applies only to Vulkan, and a failed LLM encode returns an error instead of aborting. The rest of that upstream commit conflicts with #19 and isn't needed. I also updated the description: the Vulkan legs it listed came to master with #22.

Ran the build-linux-rocm leg as written on a GitHub-hosted ubuntu-24.04 runner (source = this branch), then tested its zip on a DevLab Strix Halo against upstream's ROCm zip, which is what Studio falls back to today. Renders: Z-Image-Turbo 1024x1024 (8 steps) and MiniMax-H3 640x384x56 (4 steps), Studio flags, in seconds.

Upstream zip (master-813) This PR's bundle
Size 268 MiB 1802 MiB (GitHub's per-asset limit is 2 GiB)
ROCm libraries loaded from the host's /opt/rocm HIP, hipBLAS(Lt), rocBLAS, HSA runtime, comgr, ... only the optional libhsa-amd-aqlprofile64
Z-Image on the host (ROCm 7.2.1 installed) 88.3 s 89.9 s
H3 on the host 351.4 s 321.6 s
Clean ubuntu:24.04 container, no ROCm, only /dev/kfd + /dev/dri needs a matching system ROCm Z-Image 92.4 s, H3 326.6 s, both on the GPU
CI job Status
ROCm leg, as written pass
Bundle on gfx1151 pass
Bundle in a clean container pass
Merged tree, HIP pass, no test-only patches needed

The bundle test's first attempt failed because the artifact download was corrupt; the same zip's SHA-256 matched on the rerun.

Not run: a discrete AMD GPU (only the gfx1151 APU was available).

@oobabooga oobabooga changed the title Add Linux ROCm and Vulkan legs to the prebuilt pipeline, so AMD hosts stop landing on the CPU build Add a Linux ROCm leg to the prebuilt pipeline, so AMD hosts stop landing on the CPU build Oct 9, 2026
@oobabooga

Copy link
Copy Markdown
Member

Reached Codex review convergence at oobabooga#5.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants