Skip to content

Materialize the MySQL sign text filter as a derived table - #1008

Merged
Intelli merged 5 commits into
PlayPro:masterfrom
tricrotism:for-upstream/mysql-sign-filter
Oct 2, 2026
Merged

Intelli merged 5 commits into
PlayPro:masterfrom
tricrotism:for-upstream/mysql-sign-filter

Conversation

@tricrotism

Copy link
Copy Markdown
Contributor

Summary

On MySQL, a sign lookup with a text filter runs a correlated subquery once per candidate row and never uses the line_N prefix indexes. On 200,000 sign rows it took 2.2 seconds. Wrapping the union in a derived table brings it to under a millisecond with the same result.

The problem

A sign lookup with a text filter (/co lookup a:sign f:<text>, or the SignAPI prefix filters) goes through LookupRaw.appendSignMessageFilters (database/LookupRaw.java:1584-1609). For each filter it builds eight selects, one per sign line, and appends:

AND rowid IN (
  SELECT signFilterRows.rowid FROM co_sign signFilterRows WHERE signFilterRows.face = 0 AND (line_1 LIKE ? AND line_1 LIKE ?)
  UNION ALL SELECT ... line_2 ...
  ... eight branches ...
)

MySQL cannot turn an IN (... UNION ...) subquery into a semi-join. It runs it as a DEPENDENT SUBQUERY: for every row the outer query considers, it re-evaluates all eight branches with a primary key lookup on that row. The line_1_prefix_index to line_8_prefix_index indexes that Database.java:869 creates for exactly this filter are never used.

Measured on MySQL 8.4 with CoreProtect's own co_sign DDL, 200,000 rows, three of them matching:

Query shape EXPLAIN Time
upstream rowid IN (UNION ...) DEPENDENT SUBQUERY / DEPENDENT UNION, eq_ref on PRIMARY 2,188 ms
rowid IN (SELECT rowid FROM (UNION ...) signFilterMatches) DERIVED / UNION, range on line_1_prefix_index ... line_8_prefix_index 0.59 ms

Both returned the same three rows in the same order. MariaDB 11.8 behaves the same way: 2,684 ms for the upstream shape with the same DEPENDENT SUBQUERY plan, 0.86 ms wrapped, identical rows. The cost of the upstream shape grows with the size of the sign table, since every candidate row pays for eight lookups.

The fix

On MySQL only, the union is wrapped in a derived table:

AND rowid IN (SELECT rowid FROM (<same union>) signFilterMatches)

MySQL materializes a derived table once, so each branch runs one indexed range scan and the outer query semi-joins on the result. The bindings are unchanged.

SQLite, DuckDB and ClickHouse build exactly the query they built before. DuckDB already takes a separate path. I only measured MySQL and MariaDB, so the other backends are left as they are.

Behaviour change

None. Same rows, same order.

Risk

Low. The derived table alias is required by MySQL and does not collide with anything in the outer query. The same query runs unchanged on MariaDB.

Testing

Build: mvn package passes.
MySQL 8.4 and MariaDB 11.8 (Docker mysql:8.4, mariadb:11.8), CoreProtect's co_sign table with 200,000 generated rows: EXPLAIN and timings above, identical result rows for both shapes on both servers.

Row parity on SQLite: the 47-step scenario on Paper 26.2 matches upstream row for row, as expected since only the MySQL query changes.

MySQL cannot semi-join the rowid IN (UNION ...) subquery the sign text filter builds, so it runs it as a dependent subquery, re-probing all eight branches for every candidate row and never using the line_N prefix indexes. Wrapping the union in a derived table lets MySQL materialize it once through those indexes. Other databases build the same query as before.
@netlify

netlify Bot commented Sep 23, 2026

Copy link
Copy Markdown

❌ Deploy Preview for coreprotect failed. Why did it fail? →

Name Link
🔨 Latest commit c46a2df
🔍 Latest deploy log https://app.netlify.com/projects/coreprotect/deploys/6ab3eb6ce4beb10008dab91a

@Intelli

Intelli commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Thanks - the rare-prefix benchmark demonstrates a useful improvement. Before merging, please address the opposite case: a common prefix combined with a narrowly restricted lookup.

The derived table currently materializes text matches across all sign history, while location, time, world, and user restrictions remain outside it. A lookup for one sign location can therefore process a large global match set instead of checking only that location’s candidate rows.

Please preserve efficient execution for these restricted lookups, either by carrying the relevant restrictions into the materialized query or using a query shape that lets the optimizer choose appropriately.

The derived table materialized every sign row matching the text across the whole table before the location, time and user restrictions were applied, so a common prefix at one sign location scanned the full table. MySQL now gets the same inline line predicates DuckDB already uses: restricted lookups filter the rows found through the wid, time or user index, and unrestricted lookups still reach the line prefix indexes through an index merge.
@Intelli
Intelli merged commit 2f50109 into PlayPro:master Oct 2, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants