Materialize the MySQL sign text filter as a derived table - #1008
Merged
Intelli merged 5 commits intoOct 2, 2026
Merged
Conversation
MySQL cannot semi-join the rowid IN (UNION ...) subquery the sign text filter builds, so it runs it as a dependent subquery, re-probing all eight branches for every candidate row and never using the line_N prefix indexes. Wrapping the union in a derived table lets MySQL materialize it once through those indexes. Other databases build the same query as before.
❌ Deploy Preview for coreprotect failed. Why did it fail? →
|
Contributor
|
Thanks - the rare-prefix benchmark demonstrates a useful improvement. Before merging, please address the opposite case: a common prefix combined with a narrowly restricted lookup. The derived table currently materializes text matches across all sign history, while location, time, world, and user restrictions remain outside it. A lookup for one sign location can therefore process a large global match set instead of checking only that location’s candidate rows. Please preserve efficient execution for these restricted lookups, either by carrying the relevant restrictions into the materialized query or using a query shape that lets the optimizer choose appropriately. |
The derived table materialized every sign row matching the text across the whole table before the location, time and user restrictions were applied, so a common prefix at one sign location scanned the full table. MySQL now gets the same inline line predicates DuckDB already uses: restricted lookups filter the rows found through the wid, time or user index, and unrestricted lookups still reach the line prefix indexes through an index merge.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
On MySQL, a sign lookup with a text filter runs a correlated subquery once per candidate row and never uses the
line_Nprefix indexes. On 200,000 sign rows it took 2.2 seconds. Wrapping the union in a derived table brings it to under a millisecond with the same result.The problem
A sign lookup with a text filter (
/co lookup a:sign f:<text>, or theSignAPIprefix filters) goes throughLookupRaw.appendSignMessageFilters(database/LookupRaw.java:1584-1609). For each filter it builds eight selects, one per sign line, and appends:MySQL cannot turn an
IN (... UNION ...)subquery into a semi-join. It runs it as aDEPENDENT SUBQUERY: for every row the outer query considers, it re-evaluates all eight branches with a primary key lookup on that row. Theline_1_prefix_indextoline_8_prefix_indexindexes thatDatabase.java:869creates for exactly this filter are never used.Measured on MySQL 8.4 with CoreProtect's own
co_signDDL, 200,000 rows, three of them matching:rowid IN (UNION ...)DEPENDENT SUBQUERY/DEPENDENT UNION,eq_refonPRIMARYrowid IN (SELECT rowid FROM (UNION ...) signFilterMatches)DERIVED/UNION,rangeonline_1_prefix_index...line_8_prefix_indexBoth returned the same three rows in the same order. MariaDB 11.8 behaves the same way: 2,684 ms for the upstream shape with the same
DEPENDENT SUBQUERYplan, 0.86 ms wrapped, identical rows. The cost of the upstream shape grows with the size of the sign table, since every candidate row pays for eight lookups.The fix
On MySQL only, the union is wrapped in a derived table:
MySQL materializes a derived table once, so each branch runs one indexed range scan and the outer query semi-joins on the result. The bindings are unchanged.
SQLite, DuckDB and ClickHouse build exactly the query they built before. DuckDB already takes a separate path. I only measured MySQL and MariaDB, so the other backends are left as they are.
Behaviour change
None. Same rows, same order.
Risk
Low. The derived table alias is required by MySQL and does not collide with anything in the outer query. The same query runs unchanged on MariaDB.
Testing
Build:
mvn packagepasses.MySQL 8.4 and MariaDB 11.8 (Docker
mysql:8.4,mariadb:11.8), CoreProtect'sco_signtable with 200,000 generated rows: EXPLAIN and timings above, identical result rows for both shapes on both servers.Row parity on SQLite: the 47-step scenario on Paper 26.2 matches upstream row for row, as expected since only the MySQL query changes.