Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
142 commits
Select commit Hold shift + click to select a range
b4f1fd6
fix: preserve "token_refiner" token for MiniMax H3 LoRAs (#1864)
stduhpf Aug 11, 2026
487de75
fix: fail with a message when MiniMax-H3 is run in img_gen mode (#1863)
danielhanchen Aug 11, 2026
bcc7e29
feat: support INT8 ConvRot safetensors (#1857)
leejet Aug 11, 2026
06c359f
fix: replace free_compute_buffer with runner_done in vae (#1872)
LostRuins Aug 12, 2026
fabe481
sync: update ggml (#1873)
leejet Aug 12, 2026
de298c2
fix(ci): trigger builds for ggml updates
leejet Aug 12, 2026
6100d83
feat: add taeh3 support (#1874)
stduhpf Aug 19, 2026
58b6cb6
fix: prevent gallocr hash overflow in tiny graph-cut segments (#1880)
fszontagh Aug 19, 2026
1706b32
fix: re-clamp streaming VRAM budget to currently free memory (#1878)
fszontagh Aug 19, 2026
88b044b
fix: mark graph cuts with both a prefix and a suffix (#1883)
wbruna Aug 19, 2026
760717a
fix: make max_order of lms sampler configurable (#1885)
vmobilis Aug 19, 2026
16304cc
fix: guard against missing sampler/scheduler names (#1887)
vmobilis Aug 19, 2026
97d2990
chore: format code
leejet Aug 19, 2026
12ee60d
fix: use sd_get_preview_interval() (#1907)
vmobilis Aug 25, 2026
0a565f2
feat: configurable image / video compression (#1909)
vmobilis Aug 25, 2026
50d6405
feat: support standard Qwen3-VL weights for MiniMax-H3 (#1910)
leejet Aug 25, 2026
be0e344
feat: load scaled FP8 weights without upfront conversion (#1913)
leejet Aug 27, 2026
2c92949
fix: match exact weights in LLM config detection (#1923)
leejet Aug 30, 2026
afd5306
feat: add LTX-2.5 support (#1893)
pwilkin Aug 30, 2026
c797899
fix: correct MiniMax H3 reference audio encoding (#1886)
jk212h20 Aug 30, 2026
dc4000d
fix: correct MiniMax H3 audio Euler steps (#1908)
jk212h20 Aug 30, 2026
2540a4f
feat: use backend-native FP8 matmul when supported (#1916)
leejet Aug 30, 2026
134c821
sync: update ggml
leejet Aug 30, 2026
d9b6e27
feat: additional `--preview-interval` values (#1915)
vmobilis Aug 30, 2026
9029655
feat: support numbering for preview images (#1895)
vmobilis Aug 30, 2026
40e605f
fix: use carrier sampling for MiniMax H3 audio (#1924)
leejet Aug 30, 2026
6b3edaa
feat: generalize temporal tiling across video VAEs (#1926)
leejet Aug 30, 2026
6c57cc3
feat: prefetch streamed layers during compute (#1905)
assouan Sep 6, 2026
462d675
refactor: unify runner lifecycles and weight residency (#1940)
leejet Sep 6, 2026
dbb6112
feat: add verbose logging and log-level selection (#1941)
leejet Sep 6, 2026
80bac2d
feat: enable single-GPU auto-fit with tiered parameter placement (#1942)
leejet Sep 6, 2026
d8fb10c
fix: reuse graph cut plans across CFG passes (#1943)
leejet Sep 6, 2026
31ab2b2
refactor: split ggml extensions and move implementations to cpp files…
leejet Sep 7, 2026
9cdb6b6
fix: preserve K-quantized embedding weights (#1936)
zjn20030811 Sep 7, 2026
d04e895
fix: correct SDXL embeddings loading (#1939)
wbruna Sep 7, 2026
6b47fec
refactor: unify model source and weight lifecycle management (#1956)
leejet Sep 10, 2026
469fc49
docs: reflect GGML_MAX_NAME value change in rpc docs (and in ggml_ext…
stduhpf Sep 10, 2026
14eddb3
refactor: split generation pipeline out of stable-diffusion.cpp (#1957)
leejet Sep 10, 2026
b68d586
fix: enable VAE decode tiling fallback without auto-fit (#1932)
Hmission Sep 10, 2026
e95ab96
fix: preserve BF16 embedding weights for get_rows (#1959)
xledx Sep 11, 2026
3191b23
fix: handle invalid option numbers (#1961)
vmobilis Sep 11, 2026
e06b205
feat: expose the loaded model version name through the public API (#1…
fszontagh Sep 11, 2026
7f986a9
feat: add SenseNova U1.5 support (#1935)
Maphist0 Sep 11, 2026
5ebce93
fix: reuse graph plans when scale parameters change (#1963)
leejet Sep 11, 2026
7f410a3
feat: add linear and attention scale overrides (#1964)
leejet Sep 11, 2026
44dd137
feat: preserve explicit backend assignments during auto-fit (#1967)
leejet Sep 13, 2026
9a97738
fix: guard GPU memory capacity and propagate encoding failures (#1958)
leejet Sep 13, 2026
4a7da26
fix: bound plain-text runs in parse_prompt_attention regex (#1919)
fszontagh Sep 13, 2026
0bd72f0
feat: add Wan2.2 S2V (audio+img-to-video) support (#1925)
noctrex Sep 13, 2026
ca37fad
fix: validate vision projector output dim against LLM hidden size (#1…
fszontagh Sep 13, 2026
5a5400b
fix: resolve MSVC narrowing conversion warnings (#1969)
leejet Sep 13, 2026
42d6c0a
feat: Add generation parameters into video metadata (#1901)
CAHbKA-IV Sep 13, 2026
4964abd
feat: support external Hugging Face tokenizer JSON files (#1973)
leejet Sep 14, 2026
f9ddc0f
refactor: require external Gemma 2 and GPT-OSS tokenizers (#1974)
leejet Sep 14, 2026
07a85c7
feat: support Brownian tree noise in all noise injection samplers (#1…
wbruna Sep 14, 2026
59c23bc
fix: use tokenizer-specific pre-tokenization rules (#1975)
leejet Sep 14, 2026
3161505
fix: remove vision_model. from ununsed tensors (#1983)
leejet Sep 16, 2026
cc515a0
perf: eliminate temporary allocations in Philox rounds (#1982)
leejet Sep 16, 2026
269e726
fix: honor flash attention flag in LLM text encoder attention (#1987)
linxuhao Sep 18, 2026
656a135
refactor: remove obsolete unused tensor filtering (#1984)
leejet Sep 18, 2026
adcac69
perf: pad small attention heads to 64 for MMA Flash Attention (#1992)
leejet Sep 18, 2026
2ea8aff
perf: update ggml for faster direct convolutions (#1993)
leejet Sep 18, 2026
3e037a8
perf: accelerate VAE direct 3D convolutions (#1996)
leejet Sep 19, 2026
9982c9c
fix: propagate CUDA driver dependency to shared library consumers
leejet Sep 19, 2026
d32b4e8
fix: prevent clip_preprocess center crop from exceeding the resized i…
akhenakh Sep 19, 2026
275ab58
perf: reduce CPU overhead in graph execution and sampling (#1997)
leejet Sep 19, 2026
17860c0
perf: parallelize host tensor elementwise and broadcast ops (#1998)
leejet Sep 19, 2026
1330ceb
feat: support building with upstream ggml (#1999)
leejet Sep 19, 2026
137f740
feat: add Qwen Image 2.1 support (#1994)
leejet Sep 20, 2026
008ca5b
feat: restore legacy fp8 handling when building with upstream ggml (#…
wbruna Sep 20, 2026
b8248a8
fix: avoid passing ggml logs as format strings (#2002)
wbruna Sep 20, 2026
15f335d
feat: add LLaDA-Image support (#1968)
fszontagh Sep 20, 2026
187b256
feat: add native CUDA SageAttention support (#2005)
leejet Sep 20, 2026
b56c686
fix: avoid narrowing conversion in SigVQ patch embedding and format code
leejet Sep 20, 2026
c678dfe
docs: update CONTRIBUTING.md
leejet Sep 20, 2026
74988b2
fix: reject video models in image generation (#2017)
leejet Sep 21, 2026
2726dd3
fix: update ggml to prevent permute metadata truncation (#2018)
leejet Sep 21, 2026
78557f8
fix: honor reference image resize opt-out in server requests (#2011)
AzizMuminov Sep 21, 2026
97d932b
fix: restrict VAE tiling retries to allocation failures (#2019)
leejet Sep 21, 2026
6dcb5bb
fix: handle GPU memory reports and LLM encoding failures (#2020)
leejet Sep 21, 2026
e012065
fix: preserve reference image dimensions in server requests (#2007)
mikemikimike Sep 22, 2026
e112ab5
fix: add alpha channel input for Qwen Image 2.1 and relative docs (#2…
CarlGao4 Sep 22, 2026
ac45422
fix: honor reference image resize settings in OpenAI edits (#2025)
leejet Sep 22, 2026
2bb7294
perf: cache MiniMax H3 text conditioning (#1966)
xledx Sep 22, 2026
28b454b
feat: add configurable image input preprocessing (#2028)
leejet Sep 22, 2026
c92d73c
fix: preserve alpha when upscaling RGBA images with ESRGAN (#2029)
leejet Sep 22, 2026
241518b
feat: add latent2rgba preview for Qwen-Image 2.1 (#2032)
stduhpf Sep 23, 2026
e6281b6
feat: add configurable conditioning cache for all models (#2034)
leejet Sep 23, 2026
2dc7f54
feat: add Qwen Image 2.1 prefix KV cache (#2035)
leejet Sep 23, 2026
3674693
fix: add graph cuts for MiniMax-H3 text conditioning (#1900)
assouan Sep 23, 2026
70c1dbc
perf: run one-frame Wan VAE convolutions as 2D convolutions (#2038)
nanguoyu Sep 23, 2026
2a4ebba
ci: automatically close PRs from organization-owned forks
leejet Sep 23, 2026
500ef5f
fix: map mmapped weights through Metal buffers instead of CPU buffers…
nanguoyu Sep 23, 2026
88411ef
refactor: centralize circular RoPE and extend image model support (#2…
leejet Sep 23, 2026
caa111a
feat: optimize cfg special cases with guidance schdeule (#2033)
stduhpf Sep 24, 2026
4dfe8f5
feat: add a stand-alone upscale endpoint to the server (#2026)
nbeerbower Sep 24, 2026
740c7ae
feat: add configurable Qwen cache types and early cache scheduling (#…
leejet Sep 24, 2026
1a2330d
fix: reserve 128 MiB headroom when selecting monolithic execution (#2…
leejet Sep 24, 2026
b167b94
fix: align Qwen Image 2.1 flow schedule with official defaults (#2048)
leejet Sep 24, 2026
0a9340c
fix: scale Qwen Image 2.1 VAE convolutions (#2054)
leejet Sep 25, 2026
5e7b291
fix: keep the server frontend install out of a parent pnpm workspace …
Yi-111-a Sep 25, 2026
510bccf
fix: map Qwen Image 2.1 LoRAs to fused MLP weights (#2057)
leejet Sep 25, 2026
4c3cf75
fix: load safetensors index shards without recursion (#2058)
leejet Sep 25, 2026
39ada08
feat: add PixArt model family support (#2047)
losewayy Sep 25, 2026
19bbbca
refactor: define VAE tile dimensions in image pixels (#2059)
leejet Sep 25, 2026
2f88688
refactor: align PixArt weights with upstream layout (#2061)
leejet Sep 25, 2026
168f7b8
fix: improve GGUF metadata parsing and reader selection (#2062)
losewayy Sep 27, 2026
9947eeb
fix: handle filesystem errors during lora and upscaler cache scan (#2…
Yi-111-a Sep 27, 2026
61a83b8
docs: update hugging face links for MageFlow diffusion and VAE models…
akleine Sep 27, 2026
f8890b9
feat: add Ming-Image Design support (#2063)
leejet Sep 27, 2026
47e83d7
sync: update ggml
leejet Sep 27, 2026
ede32a6
fix: load INT8 convrot LLM embeddings correctly (#2067)
leejet Sep 27, 2026
27f7c43
fix: update ggml to fix INT8 convrot backend fallback (#2069)
leejet Sep 27, 2026
42ab1c1
fix: keep scaled INT8 convrot matmuls on Vulkan (#2070)
leejet Sep 27, 2026
3f8527a
feat: enable hip INT8 tensorwise matmul and convrot (#2071)
leejet Sep 27, 2026
d4bfcd5
ci: bundle HIP runtime libraries with Windows ROCm release (#2090)
harkgill-amd Oct 6, 2026
9db0db6
perf: avoid redundant permute+cont in non-interleaved RoPE (#2102)
daniandtheweb Oct 6, 2026
b6947b8
fix: avoid f16 overflow in Z-Image quantized matmuls on CUDA (#2074)
losewayy Oct 6, 2026
bbdcad4
feat: add Z-Image L2P support (#2075)
aapakhomov Oct 6, 2026
f4b5326
feat: support for merging LoRA weights on model conversion (#2079)
wbruna Oct 6, 2026
c150a6b
fix: correct step count adjustment for custom sigmas list (#2084)
wbruna Oct 6, 2026
019bcfc
fix: correct non-circular tile placement and blending (#2088)
shawn-mengchen-xu Oct 6, 2026
203a882
refactor: shorten log levels and move source locations to the end (#2…
leejet Oct 6, 2026
96f67fe
feat: Server: Add support for previews via sdcpp API (#2093)
stduhpf Oct 6, 2026
fb8ee71
refactor: unify circular and non-circular tiling (#2105)
leejet Oct 6, 2026
1da55eb
feat: enable sdapi latent2RGB previews (#2096)
stduhpf Oct 6, 2026
0963357
fix: keep MiniMax-H3 VAE weights resident across temporal chunks (#2103)
losewayy Oct 6, 2026
fd43b54
refactor: escape newlines in prompt logs and clarify source separator…
leejet Oct 6, 2026
4a89229
docs: document video reference image support and constraints
leejet Oct 6, 2026
8cc5ad8
sync: update frontend
leejet Oct 6, 2026
a1ded76
chore: resolve MSVC warnings in ggml extensions
leejet Oct 6, 2026
a4a9669
fix: avoid scanning cwd during server startup (#2086)
mikemikimike Oct 8, 2026
e16d26a
fix: align SD3 encoder token chunks before conditioning (#2111)
srelus Oct 8, 2026
228c707
perf: use ggml rope apply op on supported backends (#2113)
leejet Oct 8, 2026
7867f6d
fix: bump RPC protocol patch version for rope apply op (#2115)
leejet Oct 9, 2026
f2f168f
fix: preserve singleton output dimension when merging LoRA (#2116)
leejet Oct 9, 2026
f89d9b1
feat: add Iris-3B text-to-image support (#2117)
leejet Oct 9, 2026
e739cd3
Record upstream c678dfe (master-889) as merged, its content already s…
oobabooga Oct 9, 2026
8a48526
Merge upstream master-951-f89d9b1 into the fork
oobabooga Oct 9, 2026
af486f1
Document the tile overlap guard and the encode retry
oobabooga Oct 9, 2026
8e41297
Trim the tiling comments
oobabooga Oct 10, 2026
d0d47fc
Drop upstream's workflow that closes PRs from organization forks
oobabooga Oct 10, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions .github/workflows/build.yml
Original file line number Diff line number Diff line change
Expand Up @@ -564,6 +564,18 @@ jobs:
if: ${{ ( github.event_name == 'push' && github.ref == 'refs/heads/master' ) || github.event.inputs.create_release == 'true' }}
uses: pr-mpt/actions-commit-hash@01d19a83c242e1851c9aa6cf9625092ecd095d09 # v2

- name: Bundle HIP runtime libraries
if: ${{ ( github.event_name == 'push' && github.ref == 'refs/heads/master' ) || github.event.inputs.create_release == 'true' }}
run: |
$ErrorActionPreference = "Stop"
$binPath = (rocm-sdk path --bin).Trim()
if (-not $binPath) { throw "rocm-sdk path --bin returned empty" }
foreach ($dll in @("amdhip64_7.dll", "rocm_kpack.dll", "amd_comgr.dll")) {
$f = Get-ChildItem -Path $binPath -Filter $dll -ErrorAction SilentlyContinue
if (-not $f) { throw "no match for $dll in $binPath" }
Copy-Item $f.FullName -Destination .\build\bin\ -Force
}

- name: Pack artifacts
if: ${{ ( github.event_name == 'push' && github.ref == 'refs/heads/master' ) || github.event.inputs.create_release == 'true' }}
run: |
Expand Down
4 changes: 4 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,10 @@ If you want to update a third-party dependency, please open an issue first inste

## Pull Requests

When contributing from a fork, use a fork under your personal GitHub account and enable **Allow edits from maintainers**. This lets maintainers make follow-up fixes directly on the PR branch.

PRs from organization-owned forks are automatically closed when opened or reopened because GitHub does not support this maintainer-edit option for those forks. Submit the changes from a personal fork instead. See [GitHub's documentation](https://docs.github.com/en/pull-requests/how-tos/work-with-forks/allowing-changes-to-a-pull-request-branch-created-from-a-fork).

Keep each PR focused on one clear change. Large or overly complex PRs are harder to review and may not be merged.

Do not include test code or test scripts in commits or PRs. Keep them local and report verification results in the PR description.
Expand Down
4 changes: 4 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,7 @@ API and command-line option may change frequently.***
- [PiD](./docs/pid.md)
- [LongCat Image](./docs/longcat_image.md)
- [Z-Image](./docs/z_image.md)
- [Z-Image L2P](./docs/z_image_l2p.md)
- [MiniT2I](./docs/minit2i.md)
- [SenseNova U1.5](./docs/sensenova_u1.md)
- [Ovis-Image](./docs/ovis_image.md)
Expand All @@ -64,6 +65,9 @@ API and command-line option may change frequently.***
- [HiDream-O1-Image](./docs/hidream_o1_image.md)
- [Ideogram4](./docs/ideogram4.md)
- [LLaDA-Image](./docs/llada_image.md)
- [Ming-Image Design](./docs/ming_image.md)
- [PixArt](./docs/pixart.md)
- [Iris-3B](./docs/iris.md)
- [Image Edit Models](./docs/edit.md)
- [FLUX.1-Kontext-dev](./docs/kontext.md)
- [Qwen Image Edit series](./docs/qwen_image_edit.md)
Expand Down
Binary file added assets/iris/example.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added assets/qwen/qwen-image-2.1-alpha-in1.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added assets/qwen/qwen-image-2.1-alpha-out1.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added assets/qwen/qwen-image-2.1-alpha-out2.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added assets/wan/Wan2.2_A14B_vace_r2v.mp4
Binary file not shown.
2 changes: 2 additions & 0 deletions docker/Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -40,4 +40,6 @@ RUN printf '#!/bin/sh\nexec /sd.cpp/bin/sd-cli "$@"\n' > /sd-cli && \
printf '#!/bin/sh\nexec /sd.cpp/bin/sd-server "$@"\n' > /sd-server && \
chmod +x /sd-cli /sd-server

WORKDIR /sd.cpp

ENTRYPOINT [ "/sd-cli" ]
2 changes: 2 additions & 0 deletions docker/Dockerfile.cuda
Original file line number Diff line number Diff line change
Expand Up @@ -57,4 +57,6 @@ RUN printf '#!/bin/sh\nexec /sd.cpp/bin/sd-cli "$@"\n' > /sd-cli && \
printf '#!/bin/sh\nexec /sd.cpp/bin/sd-server "$@"\n' > /sd-server && \
chmod +x /sd-cli /sd-server

WORKDIR /sd.cpp

ENTRYPOINT [ "/sd-cli" ]
2 changes: 2 additions & 0 deletions docker/Dockerfile.musa
Original file line number Diff line number Diff line change
Expand Up @@ -42,4 +42,6 @@ RUN printf '#!/bin/sh\nexec /sd.cpp/bin/sd-cli "$@"\n' > /sd-cli && \
printf '#!/bin/sh\nexec /sd.cpp/bin/sd-server "$@"\n' > /sd-server && \
chmod +x /sd-cli /sd-server

WORKDIR /sd.cpp

ENTRYPOINT [ "/sd-cli" ]
2 changes: 2 additions & 0 deletions docker/Dockerfile.sycl
Original file line number Diff line number Diff line change
Expand Up @@ -29,4 +29,6 @@ FROM intel/oneapi-basekit:${SYCL_VERSION}-devel-ubuntu24.04 AS runtime
COPY --from=build /sd.cpp/build/bin/sd-cli /sd-cli
COPY --from=build /sd.cpp/build/bin/sd-server /sd-server

WORKDIR /sd.cpp

ENTRYPOINT [ "/sd-cli" ]
2 changes: 2 additions & 0 deletions docker/Dockerfile.vulkan
Original file line number Diff line number Diff line change
Expand Up @@ -41,4 +41,6 @@ RUN printf '#!/bin/sh\nexec /sd.cpp/bin/sd-cli "$@"\n' > /sd-cli && \
printf '#!/bin/sh\nexec /sd.cpp/bin/sd-server "$@"\n' > /sd-server && \
chmod +x /sd-cli /sd-server

WORKDIR /sd.cpp

ENTRYPOINT [ "/sd-cli" ]
10 changes: 8 additions & 2 deletions docs/backend.md
Original file line number Diff line number Diff line change
Expand Up @@ -156,8 +156,14 @@ the runner's graph-cut capacity checks.

Runtime capacity checks also leave 512 MiB of currently free device memory for
backend scratch buffers and pipelines, including with explicit backend assignments.
They cap stale free-memory reports by the device's total memory minus tracked
resident allocations and reject reports that exceed the device's total memory.
They cap free-memory reports by the device's total memory minus tracked
resident allocations. Vulkan reports exceeding total memory are rejected because
its heap-budget subtraction can underflow. Other backends use the cap instead of
treating such reports as zero free memory. Failed checks log the reported free and
total memory alongside tracked weight and runtime allocations.
With `--mmap`, device-backed mappings count toward these budgets at their full
mapped-file size, once per device buffer even when multiple parameter blocks
share it. Mappings retained in the loader cache continue to count.

Components are considered in `diffusion`, `te`, `vae` order so that repeatedly
used diffusion weights have priority. Each component's weights use the first
Expand Down
10 changes: 10 additions & 0 deletions docs/caching.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,16 @@

Caching methods accelerate diffusion inference by reusing intermediate computations when changes between steps are small.

### Conditioning Cache

Conditioning results are cached per model context using an LRU cache. The default
capacity is **0 (disabled) for `sd-cli`** and **4 entries for `sd-server` and the C
API**. Set `--conditioning-cache-size N` to change the limit; `0` disables caching.
For example, `sd-cli -m model.safetensors -p "a cat" --conditioning-cache-size 4`
enables the cache in the CLI. The C API option is
`sd_ctx_params_t::conditioning_cache_size`, initialized by `sd_ctx_params_init()`.
This cache is independent of the diffusion-step `--cache-mode` options below.

### Cache Modes

| Mode | Target | Description |
Expand Down
3 changes: 3 additions & 0 deletions docs/edit.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,9 @@ Stable-diffusion.spp also supports basic Unet-based editing models like instruct

## Configuring Reference Modes (`--ref-image-args`)

For a one-time input transform before reference presets and model processing,
including cropping, padding, and resizing algorithms, see [Image preprocessing](./image_preprocessing.md).

Different DiT-based editing models require different configurations to process reference images correctly (e.g., whether to use a Vision Language Model (VLM) encoder or pass VAE-encoded images directly to the DiT).

To simplify this, we provide **Presets**. By default, the system automatically selects the best preset based on the model architecture. However, you can override this using the `--ref-image-args` argument.
Expand Down
2 changes: 2 additions & 0 deletions docs/esrgan.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@

You can use ESRGAN—such as the model [RealESRGAN_x4plus_anime_6B.pth](https://github.com/xinntao/Real-ESRGAN/releases/download/v0.2.2.4/RealESRGAN_x4plus_anime_6B.pth)—to upscale the generated images and improve their overall resolution and clarity.

RGBA images, including Qwen Image 2.1 output, keep their alpha channel during model upscaling and hires fix. ESRGAN processes the RGB channels; the alpha channel is resized with bilinear interpolation and recombined with the upscaled image.

- Specify the model path using the `--upscale-model PATH` parameter. example:

```bash
Expand Down
173 changes: 173 additions & 0 deletions docs/image_preprocessing.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,173 @@
# Image preprocessing

Use `--image-preprocess` to transform each image input once, before generation:

```sh
sd-cli ... \
--image-preprocess "target=init,mode=crop-resize,filter=lanczos,antialias=true" \
--image-preprocess "target=mask,filter=nearest-exact" \
--image-preprocess "target=ref,index=0,mode=fit-pad,width=768,height=768,filter=bicubic"
```

CLI and server image loaders decode at the original resolution. The generation
entry point merges input defaults with user rules and prepares one transformed
image per input. The original pipeline then consumes those images, including
its mandatory canvas adaptation, reference resizing, and encoder preprocessing.

```text
native-resolution image
-> input defaults + user overrides
-> one input transform
-> original generation pipeline and model-specific processing
```

These rules do not override internal VAE, CLIP/VLM, ControlNet, or pixel-patch preprocessing.
`--ref-image-args` retains its existing meaning and runs after this input transform.

## Inputs and defaults

| `target` | Input | Default geometry | Indexed? |
| --- | --- | --- | --- |
| `init` | img2img image or video first frame | Center crop to the generation aspect ratio, then resize | No |
| `end` | Video last frame | Center crop, then resize | No |
| `mask` | Inpainting mask | Inherit init geometry; otherwise center crop, then resize | No |
| `control` | Control image | Center crop, then resize | No |
| `ref` | Reference images | Preserve source dimensions | Yes |
| `ip-adapter` | IP-Adapter image | Preserve source dimensions | No |
| `id` | PhotoMaker identity images | Preserve source dimensions | Yes |
| `control-frame` | Control video frames | Center crop, then resize | Yes |

Canvas defaults use the aligned generation dimensions. Reference, IP-Adapter,
and identity inputs use their original dimensions unless overridden. Default
resampling is nearest for images and nearest-exact for masks.

These defaults are shared by CLI, server, and C API. Moving geometry out of
the loaders replaces the previous CLI/server BOX/sRGB resizing, so default
pixels are not guaranteed to match earlier builds.

Reference video and audio preprocessing are outside these image rules.
Preprocessing options apply to `img_gen` and `vid_gen`, not standalone upscale
or ADetailer mode. ADetailer clears the user's rules for its internal crops.

## Rules

Rules are comma-separated `key=value` lists. Repeat the CLI option or separate
rules with semicolons. Every rule requires a `target` and at least one option.
Rule syntax and input compatibility are checked when image/video generation
starts. Unknown keys, invalid values, duplicate keys in a rule, missing images,
and out-of-range indices cause generation to fail with an error log.

Omit `index` to configure every image of that type; otherwise use a zero-based
index. CLI directory inputs follow filename order. Indexed rules override
type-wide rules field by field, regardless of order. At equal specificity,
the last value for a field wins. `auto` selects the input preset.

| `mode` | Input transform |
| --- | --- |
| `auto` | Use the input's default geometry |
| `none` | Keep source dimensions without resizing, cropping, or padding |
| `stretch` | Resize to the target dimensions |
| `crop` | Crop a target-sized rectangle without resizing; fail if the source is too small |
| `crop-resize` | Crop to the target aspect ratio, then resize |
| `fit-pad` | Fit the entire image inside the target dimensions, preserving aspect ratio, then pad |

`width` and `height` must be specified together as positive integers. They
override the input transform's dimensions, not the generation or encoder size.
For a native-size preset, specifying dimensions without a mode selects stretch.
`mode=none` with explicit dimensions different from the source is contradictory
and is rejected.

`anchor=center|top|bottom|left|right` selects crop/padding placement.
`pad_color=#RRGGBB` or `#RRGGBBAA` selects padding, defaulting to opaque black.
A grayscale mask uses the first color component.

`filter=auto|nearest|nearest-exact|bilinear|bicubic|lanczos` selects resampling.
`antialias=auto|true|false` enables antialiasing automatically for filtered
downscaling; explicit true requires bilinear, bicubic, or Lanczos.
Filtered RGBA resizing uses premultiplied alpha.

`canny=true|false` enables edge detection for any supported image target,
defaulting to `false`. It runs once after geometry, before the original
generation pipeline, including with `mode=none`. Grayscale, grayscale-alpha,
RGB, and RGBA inputs are supported; alpha is preserved.

Each input has its own Canny setting. Indexed rules can enable or disable it
for individual references, identity images, or video control frames.

```sh
--image-preprocess "target=init,mode=fit-pad,canny=true"
--image-preprocess "target=ref,index=0,mode=none,canny=true"
--image-preprocess "target=control-frame,index=2,canny=true"
```

Init and mask sources must have the same dimensions. The mask inherits the
init crop, resize, and padding coordinates, while retaining its own filter,
padding value, and Canny setting. Conflicting mask geometry is rejected. An
omitted mask remains absent until the original pipeline creates its default mask.

## Downstream behavior

`mode=none` only skips the input geometry transform. For example:

```sh
--image-preprocess "target=init,mode=none" \
--image-preprocess "target=ref,mode=none"
```

The init image is still adapted to the generation canvas by the original
pipeline. Reference images still follow `--ref-image-args` and model-specific
resizing. CLIP retains its fixed input dimensions and normalization. HiDream-O1
retains its original pixel-reference and visual preprocessing.

Existing sharing between consumers is preserved: for example, Wan img2video
uses the same adapted first frame for VAE conditioning and CLIP. High-resolution
passes reuse the prepared images and apply their original size adaptation;
they do not apply the user's crop a second time.

To disable reference resizing before VAE encoding, use
`--ref-image-args "resize_before_vae=false"` or the server field
`"ref_image_args": "resize_before_vae=false"`. This is separate from
`target=ref,mode=none`, which only skips input geometry. Model constraints
still apply.

## Server requests

Native image/video requests and SDAPI accept `image_preprocess` as a string or
an array of rule strings:

```json
{
"image_preprocess": [
"target=init,mode=fit-pad,filter=bicubic",
"target=mask,filter=nearest-exact",
"target=ref,index=0,mode=none"
]
}
```

OpenAI-compatible requests accept it through
`<sd_cpp_extra_args>{...}</sd_cpp_extra_args>` in the prompt.
Request rules replace server-default rules. Generation metadata records the
user rules; image encodings and channel conventions are unchanged.

## C API

Set `image_preprocess` on the existing image/video generation parameters.
The `generate_image()` and `generate_video()` signatures are unchanged:

```c
sd_img_gen_params_t params;
sd_img_gen_params_init(&params);
/* Set prompt, original-resolution input images, and generation options. */
params.image_preprocess.rules = "target=init,mode=crop-resize,filter=lanczos;"
"target=mask,filter=nearest-exact";
bool ok = generate_image(ctx, &params, &images, &count);
```
Both generation parameter initializers set `image_preprocess.rules` to `NULL`,
selecting input presets. Rule strings are borrowed for the synchronous call.
The library owns temporary transformed pixels; caller images and arrays are
not modified. Add `canny=true` to the desired target's rule in
`image_preprocess.rules` to enable Canny.
The parameter structs have grown; applications and bindings must be rebuilt.
7 changes: 4 additions & 3 deletions docs/int8_convrot.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,17 +84,18 @@ The floating-point output is reconstructed as
Y[r, o] ~= A[r, o] * s_x[r] * s_w[o] + b[o]
```

The packed runtime activation tensor contains the I8 activation rows and their floating-point row scales. Linear layers that share the same input and convrot group size reuse this packed tensor, avoiding repeated rotation and activation quantization within the graph.
The packed runtime activation tensor contains the I8 activation rows and their floating-point row scales. Linear layers with an input scale of `1` that share the same input and convrot group size reuse this packed tensor, avoiding repeated rotation and activation quantization within the graph. For other input scales, activations are scaled before explicit convrot quantization, and the linear output is unscaled before adding bias.

## Backend support

- CPU provides the portable regular Hadamard, activation quantization, INT8 matrix multiplication, and scale restoration implementations.
- NVIDIA CUDA devices with compute capability 7.5 or newer use the native accelerated path. For H256, CUDA fuses the rotation, row-wise maximum reduction, and activation quantization. It uses cuBLAS for I8 x I8 to I32 GEMM and a CUDA kernel for scale restoration and bias addition.
- Vulkan and other GPU backends do not currently have dedicated INT8 convrot kernels. They use the backend scheduler to fall back to CPU, which is expected to be substantially slower than the CUDA path.
- AMD HIP devices in the CDNA, RDNA3 (including RDNA3.5), and RDNA4 families use the same INT8/H256 kernels with hipBLAS for I8 x I8 to I32 GEMM. Other AMD architectures fall back to CPU for INT8 matrix multiplication.
- Vulkan provides native H256 activation quantization and INT8 matrix multiplication when the build and device support accelerated packed INT8 dot products. Unsupported configurations and GPU backends without these kernels use the backend scheduler to fall back to CPU.

LoRA adapters are applied at runtime without modifying the INT8 weights. The INT8 convrot path computes the base linear output, while LoRA, LoHa, LoKr, and raw weight-difference adapters compute their output corrections from the original, unrotated activation and add them to the base output. `--lora-apply-mode auto` selects this path for models containing INT8 tensorwise weights. If `immediately` is requested, sd.cpp falls back to runtime application because merging an adapter would require dequantizing and rotating its weight update, then recalculating the per-row scales and requantizing the result.

The dedicated CUDA convrot activation path currently requires a group size of `256`; other supported group sizes use CPU execution.
The dedicated CUDA and HIP convrot activation paths currently require a group size of `256`; other supported group sizes use CPU execution.

## Example

Expand Down
24 changes: 24 additions & 0 deletions docs/iris.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
# How to Use

Iris-3B generates images directly in pixel space and uses Qwen3-VL-4B-Instruct as the text encoder. No VAE is required.

## Download weights

- Download Iris-3B
- safetensors: https://huggingface.co/speridlabs/iris-3b/tree/main (`model.safetensors` in the root directory)
- Download Qwen3-VL-4B-Instruct
- safetensors: https://huggingface.co/Comfy-Org/Krea-2/tree/main/text_encoders
- gguf: https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct-GGUF/tree/main

## Examples

```
.\bin\Release\sd-cli.exe --diffusion-model ..\models\diffusion_models\iris-3b.safetensors --llm ..\models\text_encoders\Qwen3-VL-4B-Instruct-Q4_K_M.gguf -p "a lovely cat" --cfg-scale 3 -H 1024 -W 1024 --diffusion-fa
```

<img width="256" alt="iris-3b example" src="../assets/iris/example.png" />

## Notes

- Width and height must be multiples of 16. Do not pass `--vae`.
- Captions are limited to 300 tokens including the assistant-turn suffix. Positive and negative prompts use the same template.
Loading