Repository navigation
Add a Linux ROCm leg to the prebuilt pipeline, so AMD hosts stop landing on the CPU build - #14
LeoBorcherding wants to merge 5 commits into
Conversation
…g back to upstream or the CPU build
# Conflicts: # .github/workflows/unsloth-sd-prebuilt.yml
|
Reached Codex review convergence at oobabooga#5. |
|
I pushed one change, Ran the
The bundle test's first attempt failed because the artifact download was corrupt; the same zip's SHA-256 matched on the rerun. Not run: a discrete AMD GPU (only the gfx1151 APU was available). |
|
Reached Codex review convergence at oobabooga#5. |
Studio's MiniMax-H3 GGUF path asks this mirror for an sd-cli built for
rocm. The latest release has none: macOS, Linux CPU, Linux CUDA, Linux aarch64, Windows CPU. So the installer falls back to leejet's upstream ROCm build, which doesn't carry the H3 fixes in this repo, and when that binary doesn't list a GPU,video.pydrops to the CPU build.That's unslothai/unsloth#8814:
7900 XTX on CachyOS, H3 goes
installedtoload_failedwithin a second. A Discord report this week on the same card and distro is the other shape of it: H3 Q3_K fills 30 GB of RAM and 36 GB of swap while VRAM sits at 1.5 GB.What this adds
A
build-linux-rocmleg inunsloth-sd-prebuilt.yml, same rule as the CUDA leg:continue-on-error, not in the coverage gate, so it can never hold back the CPU, Apple and Vulkan assets. The Linux and Windows Vulkan legs this PR first added came to master with #22, so they were dropped when master was merged in.build-linux-rocmLinux-Ubuntu-24.04-x86_64-rocm-7.14.0build.yml, for 20 consumer targets (gfx1010 to gfx1201, including the gfx1150 to gfx1153 APUs), and ships the ROCm userspace it linked against: every librarylddresolves from the wheel tree, rocBLAS / hipBLASLt kernel trees, and the per-target.kpackarchives the kpack-split libraries load. The binaries get rpath$ORIGIN/lib, the libraries$ORIGIN, and the leg fails if anything is still unresolved with the wheel tree hidden. Upstream's zip doesn't bundle the runtime, so it only runs on a host with a matching ROCm, which is the #8814 failure.package_bundle.pygainsKEEP_LAYOUT=1, which ships the staged tree as is (sd-cli,sd-server,lib/,.kpack/); the other legs still land flat.src/model_manager.cpp: a device whose free-memory report exceeds its total was treated as Vulkan's budget underflow and given 0 bytes. On a ROCm integrated GPU (Strix Halo) ggml reports free memory from/proc/meminfoand total fromhipMemGetInfo, so free is above total and every allocation was refused: every render aborted. That rejection now applies to Vulkan only, as upstream did in fix: handle GPU memory reports and LLM encoding failures leejet/stable-diffusion.cpp#2020, and a failed LLM prompt encode returns an error instead of aborting (the same upstream change).Asset names are what Studio's
resolve_release_assetalready matches forrocm, so a ROCm host picks this up with no Studio change. Not here: Windows ROCm.Validation
lddgate passes; zip 1802 MiB (GitHub's per-asset limit is 2 GiB)/opt/rocmis the optionallibhsa-amd-aqlprofile64/dev/kfd,/dev/dri)model_manager.cppchange on gfx1151available 0.00 MB device)Not run: a discrete AMD GPU (only the gfx1151 APU was available).