Skip to content

Keep VAE tiles overlapping on the out-of-memory retry (no more seam lines), and retry the encode too - #21

Merged
oobabooga merged 4 commits into
masterfrom
fix/vae-oom-retry-tiles
Oct 9, 2026
Merged

oobabooga merged 4 commits into
masterfrom
fix/vae-oom-retry-tiles

Conversation

@danielhanchen

@danielhanchen danielhanchen commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Problem

When --vae-tiling is not passed and the untiled VAE decode runs out of memory, decode_first_stage retries with spatial tiling at rel_size 0.5, half the latent per axis (backend_fit.cpp, prepare_vae_decode_retry_tiling). For most sizes that is a tile one latent over half of the axis (round(dim / 2)), and sd_tiling_calc_tiles turns it into two tiles overlapping by 2 * tile - dim = 0 or 1 latent. The smootherstep cross-fade then has nothing to blend over, so the tile edge shows as a hard line through the image. With Qwen-Image-2.1 (16x VAE) the edges also carry the VAE's border artifacts, so it is a broken band rather than a thin line.

Examples of the retry geometry on master:

image latent retry tile tiles per axis overlap (latents)
Qwen-Image-2.1 1184x864 74x54 37x27 2, 2 0, 0
Qwen-Image-2.1 1024x672 64x42 32x21 3, 2 16, 0
Z-Image 1008x1008 126x126 63x63 2, 2 0, 0

The same two-tile case is reachable from --vae-tiling too, at the few sizes where the default 32-latent tile is just over half the axis: 57 to 63 latents on decode (Qwen-Image-2.1 992x992 overlaps by 2 latents) and 113 to 127 latents on encode (64-latent tiles).

The encode side had no retry at all: an untiled VAE encode that runs out of memory fails the whole job (vae encode compute failed).

Change

  • sd_tiling_seam_safe_tile_size (runtime/tiling.cpp), applied per axis at the end of VAE::get_tile_sizes: if the tile size the caller asked for would leave tiles overlapping by less than half the target overlap (or under 2 latents), use the nearest tile size that does not, trying smaller tiles first and never more than twice the requested size, so an explicitly small tile stays small. Sizes that already overlap enough are untouched, so the default --vae-tiling geometry only changes at the sizes listed above.
  • The decode retry keeps rel_size 0.5 for its first attempt (now 36x26 instead of 37x27 at 1184x864, 3x3 tiles with about 0.47 overlap). If that still fails, it halves the tiles again, down to 1/8 of the latent, before giving up (an axis already below 1/8 keeps its size). Before, a failed first retry ended the job.
  • encode_to_vae_latents retries a failed encode the same way, on a copy of the tiling params so later calls are unaffected. Encode starts at rel_size 0.25 because image VAE encode tiles are scaled up 2x in get_tile_sizes.

Results

B200, one exclusive card per run, master (1d02858, the build Studio pins; master only adds a CI commit) against this branch, same seed and prompt. The VRAM cap is real: a second process holds the rest of the card so the untiled decode's allocation fails and the retry runs. Seam score = the largest mean absolute difference along any row or column in a 64 px band, against the untiled decode of the same latent from the same binary (0 to 255 levels).

case free VRAM master: tiles, seam this PR: tiles, seam
Qwen-Image-2.1 Q4_K_M 1184x864, decode retry 15.7 GiB 37x27, 2x2, 110.01 36x26, 3x3, 2.85
Z-Image-Turbo Q4_K_M 1008x1008, decode retry 17.7 GiB 63x63, 2x2, 24.85 62x62, 3x3, 8.48
Qwen-Image-2.1 1184x864 img2img, decode retry 15.7 GiB 37x27, 2x2, 105.40 36x26, 3x3, 3.44
Qwen-Image-2.1 992x992, --vae-tiling uncapped 32x32, 2x2, 23.20 30x30, 3x3, 4.48
Qwen-Image-2.1 1184x864, --vae-tiling uncapped 32x32, 3x2, 3.30 identical output
Z-Image-Turbo 1008x1008, --vae-tiling uncapped 32x32, 6x6, 13.10 identical output
  • Three repeats per capped arm gave the same scores. Untiled renders are byte-identical between two master runs and between master and this branch, so the only difference is the VAE tiling.
  • Master's seams sit on its tile boundary (column 588, boundary at 592 px, at 1184x864; row 503, boundary at 504 px, at 1008x1008).
  • The first retry's compute buffer is smaller than master's (1777 vs 1897 MB at 1184x864, 1562 vs 1613 MB at 1008x1008), so no case that fit before stops fitting.
  • Qwen-Image-2.1 2048x2048 txt2img with flash attention and 13.7 GiB free: master's single retry (64x64 tiles) runs out of memory and the job fails. This branch retries again at 32x32 and finishes.
  • Qwen-Image-2.1 2048x2048 img2img with 13.7 GiB free: master fails at vae encode compute failed. This branch retries the encode with 64x64 tiles (3x3), then retries the decode twice, and finishes.

Decode time (decode_first_stage, seconds, including the failed untiled attempt):

case master retry this PR retry this PR, --vae-tiling passed up front untiled, uncapped
Qwen-Image-2.1 1184x864 2.76 / 2.89 / 3.42 3.73 to 4.81 (n=5, median 4.03) 1.46 / 1.47 / 1.47 0.71 to 0.88
Z-Image-Turbo 1008x1008 2.48 / 2.50 / 2.71 2.70 to 5.97 (n=5, median 3.26) 1.05 / 1.05 / 1.31 0.37 to 0.38

The retry is about 1 s slower than master's because it decodes 9 tiles instead of 4. It only runs after an out-of-memory failure.

Before / after

Out-of-memory retry under a VRAM cap. Left: untiled decode. Middle: master's retry (2x2 tiles, no overlap). Right: this branch (3x3 overlapping tiles). Top: Qwen-Image-2.1 Q4_K_M 1184x864. Bottom: Z-Image-Turbo Q4_K_M 1008x1008.

OOM retry before and after

Tests

  • A standalone seam test drives process_tiles_2d with an identity tile function that adds a different constant to each tile, for every axis length from 8 to 256 latents and each geometry used here (retry at 0.5, 0.25 and 0.125 of the latent, encode retry, default 32-latent decode and 64-latent encode tiles). It checks that a constant input comes back unchanged (blend weights sum to 1) and that no step between neighbouring outputs is as large as the per-tile offset. On master it fails for 186 of 249 lengths with the 0.5 retry, 67 with the encode retry, and at 63 latents (default decode) and 127 latents (encode). With this PR it passes for all of them.
  • Builds: CUDA (12.8, sm_100) and CPU sd-cli, no new warnings in the touched files.

Validation on other hardware

hardware result
RTX 3090 (sm86), real VRAM cap (5 GiB free), Z-Image-Turbo 1008x1008 decode retry master 2x2 tiles, seam 11.45 on the tile boundary, 38.4 dB vs the untiled decode; this branch 3x3, seam 3.30, 41.1 dB
RTX 3090, 3.2 GiB free, img2img 1008x1008 master fails (vae encode compute failed); this branch retries the encode and the decode and finishes
Strix Halo gfx1151, ROCm 7.2.1, real hipMalloc failure (6.3 GiB free) master 2x2, seam 9.22 on the boundary, 38.7 dB; this branch 3x3, seam 3.18, 41.2 dB
--vae-tiling Z-Image 496x496 (62 latents) on CUDA, HIP, Vulkan (Linux and Windows), Metal 2x2 -> 3x3; the error along the old tile boundary drops (Linux Vulkan: column band 4.16 -> 2.48, row band 1.74 -> 0.58; Windows Vulkan: 4.41 -> 2.61, 1.78 -> 0.60; Metal: 5.54 -> 4.56, 4.76 -> 1.76); repeated renders are identical
Normal sizes (1008, 1024 tiled), untiled renders, MiniMax-H3 identical to master on CUDA, HIP, Vulkan (Linux and Windows, MSVC) and Metal

Merged with the MiniMax-H3 stack (#17 to #20), the only conflict is src/runtime/tiling.h (both add a declaration; keep both). The combined tree gives the same retry and tiling output as this branch alone.

Not covered

  • Wan and LTX video VAEs were not rendered. Their spatial tiles go through the same get_tile_sizes, so they only change at the thin-overlap sizes.
  • An untiled Qwen-Image-2.1 2048x2048 decode with enough free VRAM aborts in binbcast.cu (GGML_ASSERT(s02 <= UINT32_MAX)) on master and on this branch alike. That is not an out-of-memory failure, so the retry does not run. It is not fixed here.

When an untiled VAE decode fails, the retry tiled at half the latent per
axis. For most sizes that is a tile one latent over half, which
sd_tiling_calc_tiles turns into two tiles overlapping by 0 or 1 latent:
the cross-fade has nothing to blend over and the tile edge shows as a hard
line through the image (Qwen-Image-2.1 1184x864: 110 levels).

get_tile_sizes now picks the nearest tile size whose tiles overlap by at
least half the target (never under 2 latents), preferring smaller tiles so
memory never grows. This also fixes the --vae-tiling path at the few sizes
where it hit the same two-tile case (57..63 latents decode, 113..127
encode).

The retry also halves the tiles again (down to 1/8 of the latent) when the
first tiled attempt still does not fit, and an untiled VAE encode that
fails is now retried with tiles the same way instead of failing the job.
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for security reviews. Please try again later.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1a03bc768d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/runtime/tiling.cpp
auto overlaps_enough = [&](int size) {
int num_tiles;
float overlap_factor;
sd_tiling_calc_tiles(num_tiles, overlap_factor, small_dim, size, tile_overlap_factor, false);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Skip non-circular sizing for circular VAE axes

When --circularx or --circulary is combined with VAE tiling, this helper still selects the replacement size using non-circular geometry. For example, a latent axis of 57 with the default tile size 32 is changed to 28: the previous circular layout visits offsets 0, 19, and 38 and has unit aggregate blend weight, while size 28 makes process_tiles_2d visit 0, 14, 28, 42, and 56 even though the calculation reports four tiles, causing 26 positions to receive up to approximately twice the intended weight. This corrupts the tileable output with a bright band, so circular state needs to be included in the sizing decision or this adjustment should be bypassed for circular axes.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in ee5829c: get_tile_sizes now takes the circular flags and skips sd_tiling_seam_safe_tile_size on a circular axis, so those axes keep the tile size master used. Confirmed with a constant-input process_tiles_2d check: on a circular 57-latent axis the default 32 tile sums to unit weight, while the adjusted 28 tile does not.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-09T19:32:31.506190Z 59c912a New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@oobabooga

Copy link
Copy Markdown
Member

Reached Codex review convergence at oobabooga#7.

@oobabooga

oobabooga commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

After Codex's review on the fork I pushed one change, 59c912a: an explicitly small tile no longer grows (the seam fix's upward search stops at twice the requested size, so --vae-tile-size 8x8 --vae-tile-overlap 0.01 keeps 8x8 instead of 20x20 on a 63-latent axis), and the halving retry leaves an axis already below 1/8 alone. The default and retry geometries need at most 2x across 8-256 latents, and their outputs below are unchanged (1008 / 496 tiled and the 5 GiB decode retry re-rendered identically). It conflicts with the MiniMax-H3 stack (#17 to #20) only in src/runtime/tiling.h, where both add a declaration at the end of the file; keep both. I tested the merged tree (stack + this PR + #14) on every backend, and this PR's retry and tiling outputs there are identical to this branch alone.

Tested ee5829c against master 6321a69 with Z-Image-Turbo Q4_K_M at 8 steps. "Seam" is the largest mean absolute difference along any row or column against the untiled decode of the same latent (0 to 255).

Case Master PR
RTX 3090, real VRAM cap (5 GiB free), 1008x1008 decode retry 2x2 tiles, seam 11.45 on the tile boundary, 38.4 dB 3x3, seam 3.30, 41.1 dB
Strix Halo gfx1151, ROCm, real hipMalloc failure (6.3 GiB free) 2x2, seam 9.22 on the boundary, 38.7 dB 3x3, seam 3.18, 41.2 dB
RTX 3090, 3.2 GiB free, img2img 1008x1008 fails: vae encode compute failed encode retried with 3x3 tiles, then the decode; finishes

--vae-tiling at a thin size, Z-Image 496x496 (62 latents): master tiles 2x2, the PR 3x3. Mean error against the untiled decode in the band around the old tile boundary:

Backend Columns, master → PR Rows, master → PR
CUDA (3090) 4.59 → 2.65 1.78 → 0.60
Vulkan, Linux 4.16 → 2.48 1.74 → 0.58
Vulkan, Windows (MSVC) 4.41 → 2.61 1.78 → 0.60
Metal (M1) 5.54 → 4.56 4.76 → 1.76

Normal sizes (1008, 1024 tiled), untiled renders and MiniMax-H3 renders are identical to master on CUDA, HIP, Vulkan (Linux and Windows) and Metal, and repeated renders are identical.

CI job Status
HIP, tiling + H3 pass
HIP, capped retry pass (at 5.3 GiB free the DiT fails first in both arms; 6.3 GiB isolates the VAE)
Vulkan, Linux pass
Windows, MSVC + Vulkan pass
Windows, with repeats + images pass
Merged tree: HIP, Vulkan, Windows pass

This PR has no GitHub CI checks of its own. It was built here on Linux (CUDA, CPU, HIP, Vulkan), Windows (MSVC) and macOS (Metal). Harness.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 59c912afe3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/core/backend_fit.cpp
Comment on lines +491 to +492
const float first_rel_size = encode ? 0.25f : 0.5f;
const float min_rel_size = 0.125f;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Scale the encode retry floor with its tile factor

For ordinary image VAEs, get_tile_sizes multiplies relative encode tile sizes by encoding_factor = 2.0, but this shared floor stops retrying once rel_size_* reaches 0.125. Consequently encode retries use effective axis fractions of 0.5 and 0.25 and then give up, never trying the advertised one-eighth-sized tile; an encode that still OOMs at one quarter but would fit at one eighth therefore fails unnecessarily. The encode floor must account for the encoding factor, just as first_rel_size already does.

Useful? React with 👍 / 👎.

Comment on lines +2746 to +2748
while (latents.empty() &&
sd::backend_fit::prepare_vae_retry_tiling(retry_tiling_params, false, true)) {
latents = first_stage_model->encode(n_threads, x, retry_tiling_params, circular_x, circular_y);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reconfigure circular padding before tiled encode retries

When image VAE tiling starts disabled and an encode OOM triggers this retry under circular_x or circular_y, configure_image_vae_axes has already enabled circular padding inside the VAE graph. This newly tiled call also passes the circular flags to process_tiles_2d, so every tile wraps both at the tiler level and internally within its own convolution graph. The explicit-tiling path avoids this by disabling the model's circular axis whenever an axis is split; the retry path needs the same reconfiguration or it produces artifacts at internal tile boundaries for circular img2img/control/reference encodes.

Useful? React with 👍 / 👎.

@oobabooga

Copy link
Copy Markdown
Member

Reached Codex review convergence at oobabooga#7.

@oobabooga
oobabooga merged commit 363c91f into master Oct 9, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants