Skip to content

perf: avoid redundant permute+cont in non-interleaved RoPE - #2102

Open
daniandtheweb wants to merge 1 commit into
leejet:masterfrom
daniandtheweb:non-interleaved-rope
Open

daniandtheweb wants to merge 1 commit into
leejet:masterfrom
daniandtheweb:non-interleaved-rope

Conversation

@daniandtheweb

@daniandtheweb daniandtheweb commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR reduces the number of full ggml_cont passes in apply rope for the non-interleaved RoPE layout from 3 copies to 1.

After the single required H/L reorder, the two contiguous halves of the head dim (i, i + d_head/2) are read as views and the rotation is assembled with one ggml_concat. This eliminates 2 of the 3 full copies. The interleaved path is unchanged.

The output is identical to the previous implementation.

As for performance goes, on my RX 7800XT an Anima generation at 1024x1024, cfg 5 goes from 2.21 s/it to 2.15s/it.

This performance optimization possibility was found by a Pi agent running Qwen 3.8 27B locally.
The optimization was then made by me and the AI helped with the review.

Checklist

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant