Skip to content

Sync with upstream master-951-f89d9b1 - #24

Merged
oobabooga merged 142 commits into
masterfrom
sync-upstream-20261009
Oct 10, 2026
Merged

oobabooga merged 142 commits into
masterfrom
sync-upstream-20261009

Conversation

@oobabooga

@oobabooga oobabooga commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

Syncs the fork with upstream master-951-f89d9b1 (2026-10-09). The last sync, #15 on 2026-09-21, brought the tree up to master-889-c678dfe, so this adds 62 upstream commits. Among them are fixes Studio users need:

  • the Qwen-Image-2.1 VAE scaling and flow schedule;
  • the Z-Image f16 overflow on CUDA;
  • Metal mmap buffers;
  • the new tile planner;
  • a large ggml update (f583f39 -> d25b121).

Merge this with "Create a merge commit", not squash or rebase. #15 was squashed, so git never recorded it as a merge. Git still thinks the last sync was the 2026-08-07 merge, and a plain git merge re-raised every conflict #15 had already resolved. The first commit here (e739cd3) is git merge -s ours c678dfe. It records #15's sync point without changing a file: the tree is identical to master, and c678dfe..master contains only fork files. The real merge (8a48526) therefore used c678dfe as its base. A squash would throw that away again.

What needed more than a textual merge

Seven files conflicted. Beyond those, these upstream changes overlap in meaning with fork code:

Behaviour that changes with upstream

  • Z-Image output changes on CUDA, ROCm, Vulkan and Metal. Every comparison against pure upstream f89d9b1 that this PR ran is pixel-identical (Z-Image 1024 untiled/tiled and 1008 tiled on CUDA, SD1.5 and Z-Image on Metal), so the change comes from upstream.
  • First OOM retry tile: a decode or encode that runs out of memory now first retries with 256-pixel tiles (6x6 at 1008), not Keep VAE tiles overlapping on the out-of-memory retry (no more seam lines), and retry the encode too #21's half-latent tiles (3x3).
  • H3 overlap with --vae-tiling: the H3 video VAE now uses --vae-tile-overlap (default 0.5) instead of a forced 0.25. Studio passes --vae-tiling only under its low-VRAM offload policies.
  • Encode tile multiplier removed: image VAE encode no longer doubles the tile size.

Verification

Arms are master 4a98feb vs this branch, and pure upstream f89d9b1 where noted.

CUDA, local (RTX 6000 Ada sm89, RTX 3090 sm86, cuDNN on)

  • test-backend-ops, full suite: 16597/16597 on both GPUs.
  • H3, 960x544x124, 4 steps:
    • Default mode: master 98.6 / 99.3 s, this branch 98.9 / 95.2 s, upstream 144.5 / 143.5 s. Max mode: master 73.0 s, this branch 72.5 s.
    • Frames are 109 dB from master (77 of 124 identical) and audio is bit-identical. Repeat runs are identical.
  • Z-Image:
    • Matches upstream exactly at 1024 untiled/tiled and 1008 tiled.
    • At 496 tiled this branch uses 3x3 tiles where upstream uses 2x2. Mean difference from the untiled image across the 240-256 px tile boundary: 2.5-2.7 here vs 4.3-4.6 upstream (columns) and 0.6 vs 1.4-1.8 (rows).
  • VRAM caps on the 3090:
    • Decode retry works (5 GiB free).
    • Encode retry works (img2img, 3.2 GiB free) where pure upstream fails with "failed to encode init image".
  • The retry-loop stop rule, tested by linking the real sd_tiling_seam_safe_tile_size into an emulation of the loop: all 2382 cases (dims 4-400, overlap 0.5/0.25/0.1, scale 8/16) finish within 6 attempts.

AMD Strix Halo gfx1151 (DevLab runners)

  • Builds on HIP, Linux Vulkan and Windows Vulkan. All op tests pass: FLASH_ATTN_EXT, CPY, MUL_MAT, RMS_NORM, ROPE, the fork's fused ops, ADD, MUL, CONV_2D_DW.

  • H3 at 640x384x56:

    Backend Master This branch Output vs master
    HIP 280.8 / 286.9 s 284.3 / 284.1 s differs; see CPU check below
    Linux Vulkan 142.6 / 146.0 s 137.1 / 134.4 s video 106.8 dB
    Windows Vulkan (interleaved) 303.0 / 272.1 s 273.3 / 246.2 s video 106.8 dB, audio bit-identical
  • CPU reference check (2 steps, each arm vs its own CPU-backend render):

    • The CPU renders of master and this branch match: audio bit-identical, video 106.8 dB.
    • HIP: this branch is slightly closer to the CPU than master (video 30.96 vs 30.19 dB, audio 5.50 vs 5.30 dB).
    • Linux Vulkan, three seeds: video equal, audio 6.79 / 7.61 / 9.00 dB vs master's 7.16 / 7.31 / 9.61. The sign varies by seed, the GPU-vs-CPU distance itself is only 7-10 dB, and Windows Vulkan audio is bit-identical. I read this as noise.
  • HIP OOM retries:

    • img2img encode at 3.2 GiB free: retries with 256-pixel tiles.
    • 1536 untiled decode at 10 and 12 GiB free: retries and finishes (11x11 tiles).

Metal (M1, 8 GB)

  • Op tests pass.
  • SD1.5 512 untiled, 512 tiled and 504 tiled match pure upstream pixel for pixel. Timings are equal across master, this branch and upstream (sampling about 42.7 s).
  • Z-Image matches upstream untiled and at 768 tiled. At 496 tiled this branch uses 3x3 tiles where upstream uses 2x2. Mean difference from the untiled image across the 236-260 px tile boundary: 4.39 vs 5.53 (columns) and 1.58 vs 4.76 (rows).

Prebuilt pipeline legs, replayed from this branch's unsloth-sd-prebuilt.yml on GitHub-hosted runners (signing skipped)

  • Every leg builds: macOS arm64 (including the load gate), macOS x86_64, Linux x86_64, Linux aarch64, Linux Vulkan, Windows CPU, Windows Vulkan, and Linux CUDA 12.8.
  • The CUDA leg covers 8 architectures with cuDNN on, passes the NEEDED gate, and produces a 1192.9 MiB zip.
  • That CUDA zip renders H3 on the Ada with cuDNN 9.27 loaded, and renders Z-Image on both GPUs.

After merging

The prebuilt workflow names a release after the highest master-NNN-sha tag reachable in this repository. Upstream's tags after master-813 were never copied here, so releases would still be called master-813-.... Pushing upstream's tag fixes the name:

git fetch https://github.com/leejet/stable-diffusion.cpp tag master-951-f89d9b1
git push origin master-951-f89d9b1

Not covered

  • CUDA runtime on sm80, sm90 and sm120 (it compiles in the leg).
  • Discrete AMD GPUs.
  • One HIP case got worse: Z-Image 1536 with only about 9 GiB free on the Strix Halo iGPU, where master finished and this branch fails in the DiT. Pure upstream f89d9b1 fails the same way at 9 and 8 GiB free. The new ggml reports the iGPU's real free memory, where the old one reported system memory (about 59 GB), so the DiT now switches to segments and runs out partway. This comes from upstream, not from the merge.

stduhpf and others added 30 commits August 11, 2026 22:57
Co-authored-by: leejet <leejet714@gmail.com>
Co-authored-by: leejet <leejet714@gmail.com>
Co-authored-by: leejet <leejet714@gmail.com>
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-10T02:37:27.264392Z d0d47fc Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: af486f1678

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

float attn_scale; // Override flash-attention K/V scaling; 0 keeps the model default
const char* tokenizer; // tokenizer.json path or main=FILE,clip-l=FILE,clip-g=FILE assignments; required for PiD and Lens
bool sage_attn;
int conditioning_cache_size; // Maximum cached conditioning entries per context; 0 disables caching (default: 4)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep public parameter structs ABI-compatible

When an application compiled against the previous public header loads the updated shared library, its sd_ctx_params_t storage ends at sage_attn, but the new sd_ctx_params_init() uses the library's larger sizeof(sd_ctx_params_t) and overwrites beyond the caller's allocation; callers that skip the initializer still have this new field read out of bounds. The appended image_preprocess fields cause the same problem for the image/video generation structs, so these additions need a size/version mechanism or versioned types before they can be exposed safely.

AGENTS.md reference: AGENTS.md:L69-L76

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not changing this. The new fields are upstream's (leejet#2028, leejet#2034, leejet#2059); this PR only merges them. The fork has no shared-library consumer to break: the prebuilt bundles ship statically linked sd-cli and sd-server (no libstable-diffusion), and Studio runs those executables, so nothing loads a newer library against an older header. Adding a size/version scheme to upstream's structs here would also diverge from upstream's API on every later sync.

Comment on lines +614 to +615
const sd_lora_t* loras,
int lora_count);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve the existing conversion entry point

Existing binaries call the exported C symbol convert_with_components with 11 arguments, but this change keeps the same symbol while requiring two more. On an in-place shared-library upgrade, the implementation therefore interprets unspecified register/stack contents as loras and lora_count; a non-null garbage pointer can be dereferenced immediately while logging or applying LoRAs. Preserve the old signature as a wrapper and add a separately named/versioned entry point for the LoRA arguments.

AGENTS.md reference: AGENTS.md:L69-L76

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not changing this. The two LoRA arguments come from upstream (leejet#2079, merging LoRA weights on conversion); this PR only merges them. The fork distributes statically linked sd-cli and sd-server only, so no existing binary resolves convert_with_components from an upgraded shared library. A fork-only wrapper and renamed entry point would diverge from upstream's API on every later sync.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. 🎉

Reviewed commit: 8e412971c2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@oobabooga

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d0d47fc2af

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

// Each batch item has a base pass and, when enabled, a hires pass.
g_preview.job_count = static_cast<int64_t>(std::max(1, params.batch_count)) * (params.hires.enabled ? 2 : 1);
g_preview.job_no = 0;
sd_set_preview_callback(preview_callback, PREVIEW_PROJ, 1, true, false, nullptr);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Decouple SDAPI progress from projection support

When the loaded model lacks a PREVIEW_PROJ mapping, the preview callback is never invoked, so /sdapi/v1/progress remains at step/progress zero for the entire request. This affects the newly supported PixArt family: it uses a 4-channel VAE, but preview_image() only recognizes SD1/SD2/SDXL in its 4-channel projection branch and returns before calling this callback. Track progress independently of latent projection, or add an appropriate PixArt preview path.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not changing this here. This is upstream's server code (the SDAPI preview/progress endpoints from leejet#2093 and leejet#2096, and PixArt from leejet#2047), merged unchanged; the fork does not touch examples/server or preview_image(). Studio does not use /sdapi/v1/progress or PixArt through sd.cpp. A fix belongs upstream, where it would reach this fork on the next sync.

Comment thread src/convert.cpp
Comment on lines +410 to +415
for (int i = 0; i < lora_count; i++) {
ModelManager::LoraSpec spec;
spec.path = loras[i].path != nullptr ? loras[i].path : "";
spec.multiplier = loras[i].multiplier;
spec.is_high_noise = loras[i].is_high_noise;
spec.required = true;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Skip disabled LoRAs when building conversion specs

When the array contains at least one nonzero multiplier, has_active_loras() selects the in-memory path, but this loop also turns every zero-multiplier entry into a required source. Consequently, a caller that keeps a missing or optional LoRA entry disabled with multiplier 0 alongside an active LoRA fails in prepare_lora_sources(), even though the active check explicitly treats that entry as inactive. Omit zero-multiplier entries from the generated specs.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not changing this here. build_lora_specs() and has_active_loras() are upstream's code from leejet#2079 (LoRA merging on conversion), merged unchanged; the fork does not modify src/convert.cpp, and Studio does not run sd.cpp conversion with LoRAs. Diverging from upstream in this file for it would cost a conflict on every later sync; the fix belongs upstream.

@oobabooga

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d0d47fc2af

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

}
return true;
};
if (!load_image("init_image", 0, init_image) ||

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Normalize grayscale inputs before generation

Preserving the source channel count here allows grayscale and grayscale-alpha PNGs to enter generation with 1 or 2 channels. For ordinary 3-channel models, ensure_image_tensor_channels() in src/pipeline/image.cpp only handles >=3 -> 3, so these valid inputs reach its GGML_ABORT and terminate the CLI/server process; 2-channel inputs also abort for 4-channel models. Convert 1/2-channel images to the requested RGB/RGBA layout before passing them to the generation pipeline, or extend that helper to handle both layouts.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not reproduced. On this branch (CPU build of d0d47fc), Z-Image img2img with a grayscale (L) and a grayscale+alpha (LA) PNG completes through both entry points: sd-cli -i, and sd-server POST /sdcpp/v1/img_gen with init_image as base64 (the from_json_str path this comment points at). The log shows the IMG2IMG branch and a successful encode_first_stage for both, with no abort. This loader is also upstream's code (leejet#2021, leejet#2028), merged unchanged.

@oobabooga

Copy link
Copy Markdown
Member Author

@codex review

1 similar comment
@oobabooga

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d0d47fc2af

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +256 to +257
for (int c = 0; c < channels; ++c) {
float fill = rgba[channels == 1 ? 0 : c] / 255.f;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Use the alpha component for grayscale-alpha padding

When a C API caller supplies a two-channel grayscale-alpha image with mode=fit-pad, the second output channel reads rgba[1] (green) rather than rgba[3] (alpha). Consequently, even the default #000000ff creates transparent padding instead of opaque black, and any explicitly requested padding alpha is ignored; map channel 1 to the RGBA alpha component when channels == 2.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not changing this here. src/runtime/image_preprocess.cpp is upstream's (leejet#2028), merged unchanged, and the case needs a C API caller passing a 2-channel image with fit-pad: the fork's sd-cli and sd-server decode inputs to RGB/RGBA before preprocessing, and Studio sends RGB/RGBA. The fix belongs upstream.

@oobabooga

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Swish!

Reviewed commit: d0d47fc2af

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@oobabooga
oobabooga merged commit 50746ec into master Oct 10, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.