Skip to content

add generated upgrade test harness - #477

Draft
aagbsn wants to merge 22 commits into
mainfrom
add_437_clickhouse_upgrade_test
Draft

aagbsn wants to merge 22 commits into
mainfrom
add_437_clickhouse_upgrade_test

Conversation

@aagbsn

@aagbsn aagbsn commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

This implements a test harness to verify clickhouse upgrades for our database schema. #437

@aagbsn
aagbsn marked this pull request as draft August 11, 2026 05:16
@hellais

hellais commented Aug 13, 2026

Copy link
Copy Markdown
Member

We should also read the changelog for any potential tricky breaking changes and ideally run the target version with the API + data pipeline to make sure that it's able to work without any problems.

Scrolling through the changelog for "backward incompatible change" here are some highlights worth being careful of:

Downgrading after upgrading may cause data loss. Propagate data types serialization versions to nested data types
https://clickhouse.com/docs/resources/changelogs/oss/2026#263-backward-incompatible-change

So we need to be extra careful once we go past 26.3 (maybe we should stop 1 LTS behind until we have done thorough testing).

Especially tricky are ones where they say things like:

Renamed functions searchAny and searchAll to hasAnyTokens and hasAllTokens for better consistency with existing function hasToken
https://clickhouse.com/docs/resources/changelogs/oss/2025#2510

Which means that we need to make sure that we aren't using these functions, otherwise they will crash at runtime.

Disallow truncating replicated databases
https://clickhouse.com/docs/resources/changelogs/oss/2025#backward-incompatible-change-10

Might apply to us in the data pipeline

It might also be worth passing the whole changelog history into an llm and give it our codebase to hunt for potential breakages.

In any case I would suggest for sure upgrading to strictly less than 26.3, since if we mess that upgrade path up, we have no way to rollback.

aagbsn and others added 3 commits August 17, 2026 16:51
Cross-checked sql/001_schema.sql against every ClickHouse fixture in
ooni/backend (ooniapi/services/{oonimeasurements,ooniprobe,oonirun,
testlists}) column by column.

- event_detector_changepoints was accidentally built from backend's
  oonimeasurements column set (`*_current_state` enums) while credited
  to devops's schema.sql, which defines a different, incompatible set
  of columns (last_ts, *_obs_w_sum, *_w_sum, current_mean). Corrected
  to match devops's schema.sql, since that's the source of truth for
  what's actually deployed on oonidata_cluster.
- Added event_detector_cusums (present in devops's schema.sql, missed
  in the original copy) and url_priorities (absent from devops's
  schema.sql entirely, despite being referenced by the ansible
  clickhouse_custom_grants for oonitestlists; ported from backend's
  CollapsingMergeTree definition as ReplicatedCollapsingMergeTree).
- Documented, rather than silently picking a side on, four places
  where devops and backend disagree with each other: fastpath column
  types/indexes, analysis_web_measurement PARTITION BY/ORDER BY,
  jsonl's extra date/source/update_time columns, and faulty_measurements'
  async_insert settings. Followed devops in all four; see the header
  comment in sql/001_schema.sql for specifics.
- Flagged eleven tables that exist in backend's legacy ooniprobe/
  oonirun/testlists fixtures but nowhere in devops's schema.sql
  (test_groups, accounts, session_expunge, counters_test_list,
  counters_asn_test_list, msmt_feedback, fingerprints_dns,
  fingerprints_http, asnmeta, incidents, oonirun) as an open question
  rather than adding them speculatively.
- Fixed harness/scenarios.py's SQL statement loader: it split the
  schema file on bare ";" characters, which broke once the new
  documentation comments above used semicolons as normal punctuation.
  Now strips full-line "--" comments before splitting.

Generated with Claude Sonnet 5 (Claude Code / Cowork).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…truth

Add .github/workflows/clickhouse_upgrade_test.yml with two jobs: the
recommended staged (LTS-hop) upgrade path and the naive direct-jump
path (continue-on-error, since it's diagnostic, not a merge gate).
Each node-upgrade and ON CLUSTER DDL check is its own workflow step
with its own pass/fail checkmark, timing, and log, rather than one
opaque job -- built on a new ci_step.py CLI (setup/upgrade-node/
verify-ddl/report/teardown) where each subcommand is a fresh process
that recovers cluster state via `docker inspect` instead of requiring
shared state between steps.

Refactored harness/scenarios.py so the new granular step functions
(setup_step, upgrade_node_step, verify_ddl_step, step_ok) are what
both ci_step.py and the existing local-use scenario_staged_lts()/
scenario_direct_jump() call -- the CI path and `make test` can no
longer silently diverge. Added harness/compose.py:current_env() and
report.py:render_ci_steps_report() to support this. Added a
docker-compose mem_limit (CI runners are more memory-constrained than
a dev laptop) and a .gitignore for __pycache__/generated results/.

Also fixed sql/001_schema.sql's obs_web table using a real
`SHOW CREATE TABLE` dump run against production: added the missing
probe_id column, all three minmax indexes, and the PARTITION BY
clause derived from bucket_date. Confirmed fastpath already matched
prod exactly.

Generated with Claude Sonnet 5 (Claude Code / Cowork).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Cross-checked citizenlab, citizenlab_flip, jsonl,
analysis_web_measurement, event_detector_changepoints,
event_detector_cusums, and faulty_measurements against `SHOW CREATE
TABLE` run directly on production.

- jsonl, faulty_measurements: exact matches, no changes.
- citizenlab / citizenlab_flip: their ZK paths are swapped relative
  to their table names (citizenlab's data lives under
  .../citizenlab_flip/{shard} and vice versa -- a swap-pair pattern,
  not a bug to tidy up). Also added citizenlab_flip, which had been
  missing from this file entirely.
- analysis_web_measurement: added 4 missing columns
  (top_dns_rule_id/top_tcp_rule_id/top_tls_rule_id/probe_id) and the
  same 3-index minmax trio obs_web has.
- event_detector_changepoints and event_detector_cusums: rewritten
  from scratch. devops' own schema.sql -- not just backend's copies
  of it -- turned out to be stale/inaccurate for both: production
  uses plain, non-replicated ReplacingMergeTree for both tables (no
  ON CLUSTER coordination between replicas), no PARTITION BY on
  either, and entirely different column sets than schema.sql
  describes. Worth flagging to whoever maintains
  scripts/cluster-migration/schema.sql separately from this test.

Every table in sql/001_schema.sql has now been checked against a live
SHOW CREATE TABLE, not just against devops/backend's copies of it.

Generated with Claude Sonnet 5 (Claude Code / Cowork).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
aagbsn added a commit that referenced this pull request Aug 17, 2026
Patch 1 — fix-false-positive-replication-errors.patch (touches harness/validate.py, harness/scenarios.py, harness/report.py):

Fix false-positive replication-error detection that failed real CI run

PR #477's staged-upgrade job failed at "Hop 1/4: upgrade ch2" (run
32041884883) with CANNOT_READ_ALL_DATA logged on ch1/ch3. Root cause,
confirmed from the raw job log: forcibly recreating ch2's container (how
this harness simulates an in-place upgrade) drops the other nodes' live
connections to it, which ClickHouse logs as a NETWORK/CANNOT_READ_ALL_DATA
error regardless of version -- a harmless, self-healing side effect of the
container bounce, not a compatibility problem. The write-then-read-back
probe and row-count convergence checks in that same step had already
confirmed replication was fine.

harness/validate.py now snapshots system.errors counters before each node
bounce and diffs after, classifying new errors as "transient" (expected
from any container recreate: NETWORK/CANNOT_READ_ALL_DATA/REPLICA-session/
etc, non-gating) vs "hard" (CHECKSUM/UNKNOWN_FORMAT/TOO_OLD/CORRUPTED --
only these fail a hop). This also fixes a second, related bug: the old
`last_error_time > now() - INTERVAL 30 MINUTE` window re-flagged errors
from earlier hops in every later step of the same CI job. Verified against
the exact values from the failing run (see harness/scenarios.py's
_hop_ok): the real hop1-ch2 step now evaluates to PASS.

Also hardened harness/scenarios.py:upgrade_node_step() with a try/except
(matching the existing pattern in setup_step()) so a docker/compose-level
failure produces a clean, diagnosable failed-step result instead of an
uncaught traceback -- a no-regrets fix, not the actual root cause here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

Patch 2 — stop-before-26.3-and-audit-findings.patch (touches harness/versions.py, README.md):

Recommend stopping upgrade at 25.8.29.51 for now; record audit findings

Addresses the remaining points from hellais's review on PR #477:

- harness/versions.py now exposes RECOMMENDED_NOW = "25.8.29.51" (last LTS
  strictly before 26.3) separately from LTS_HOPS, since 26.3 ships a
  backward-incompatible nested-data-type serialization change that can
  make downgrading lossy -- per the review comment. LTS_HOPS still walks
  the full ladder to 26.7.3.19 so this harness keeps validating that leg;
  it's just not yet the production recommendation.
- README documents a point-by-point response to the review comment,
  including two new confirmed findings from grepping ooni/backend and
  ooni/data: searchAny/searchAll/hasAnyTokens/hasAllTokens are absent from
  both codebases (clean), and the TRUNCATE-on-replicated-table pattern
  flagged by the reviewer exists in *two* places doing the identical
  citizenlab_flip truncate+insert+exchange sequence (ooni/data's
  oonipipeline AND ooni/backend's analysis service -- not yet confirmed
  which is actually deployed), plus in two backend test fixtures
  (oonirun, ooniprobe) that truncate replicated tables in CI.
- Still open and called out explicitly rather than silently dropped: full
  changelog sweep across every 24.9-26.7 release, running the target
  version against ooni/backend's API and oonipipeline's own test suite,
  and confirming which ClickHouse version introduced the
  truncate-replicated restriction.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
aagbsn and others added 2 commits August 17, 2026 18:09
Patch 1 — fix-false-positive-replication-errors.patch (touches harness/validate.py, harness/scenarios.py, harness/report.py):

Fix false-positive replication-error detection that failed real CI run

PR #477's staged-upgrade job failed at "Hop 1/4: upgrade ch2" (run
32041884883) with CANNOT_READ_ALL_DATA logged on ch1/ch3. Root cause,
confirmed from the raw job log: forcibly recreating ch2's container (how
this harness simulates an in-place upgrade) drops the other nodes' live
connections to it, which ClickHouse logs as a NETWORK/CANNOT_READ_ALL_DATA
error regardless of version -- a harmless, self-healing side effect of the
container bounce, not a compatibility problem. The write-then-read-back
probe and row-count convergence checks in that same step had already
confirmed replication was fine.

harness/validate.py now snapshots system.errors counters before each node
bounce and diffs after, classifying new errors as "transient" (expected
from any container recreate: NETWORK/CANNOT_READ_ALL_DATA/REPLICA-session/
etc, non-gating) vs "hard" (CHECKSUM/UNKNOWN_FORMAT/TOO_OLD/CORRUPTED --
only these fail a hop). This also fixes a second, related bug: the old
`last_error_time > now() - INTERVAL 30 MINUTE` window re-flagged errors
from earlier hops in every later step of the same CI job. Verified against
the exact values from the failing run (see harness/scenarios.py's
_hop_ok): the real hop1-ch2 step now evaluates to PASS.

Also hardened harness/scenarios.py:upgrade_node_step() with a try/except
(matching the existing pattern in setup_step()) so a docker/compose-level
failure produces a clean, diagnosable failed-step result instead of an
uncaught traceback -- a no-regrets fix, not the actual root cause here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Addresses the remaining points from hellais's review on PR #477:

- harness/versions.py now exposes RECOMMENDED_NOW = "25.8.29.51" (last LTS
  strictly before 26.3) separately from LTS_HOPS, since 26.3 ships a
  backward-incompatible nested-data-type serialization change that can
  make downgrading lossy -- per the review comment. LTS_HOPS still walks
  the full ladder to 26.7.3.19 so this harness keeps validating that leg;
  it's just not yet the production recommendation.
- README documents a point-by-point response to the review comment,
  including two new confirmed findings from grepping ooni/backend and
  ooni/data: searchAny/searchAll/hasAnyTokens/hasAllTokens are absent from
  both codebases (clean), and the TRUNCATE-on-replicated-table pattern
  flagged by the reviewer exists in *two* places doing the identical
  citizenlab_flip truncate+insert+exchange sequence (ooni/data's
  oonipipeline AND ooni/backend's analysis service -- not yet confirmed
  which is actually deployed), plus in two backend test fixtures
  (oonirun, ooniprobe) that truncate replicated tables in CI.
- Still open and called out explicitly rather than silently dropped: full
  changelog sweep across every 24.9-26.7 release, running the target
  version against ooni/backend's API and oonipipeline's own test suite,
  and confirming which ClickHouse version introduced the
  truncate-replicated restriction.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@aagbsn
aagbsn force-pushed the add_437_clickhouse_upgrade_test branch from 4dd3e8c to 52df4e9 Compare August 17, 2026 16:10
aagbsn and others added 16 commits August 17, 2026 18:50
Run 32044578317 on PR #477 confirmed the false-positive fix works (the
full 24.8.6.70 -> 25.3.14.14 hop passed clean, including further transient
non-gating blips) and then hit a real, hard failure during
25.3.14.14 -> 25.8.29.51: once ch1+ch2 were on 25.8.29.51, a merge on
either produced a part whose mark file ch3 (still on 25.3.14.14) could not
parse at all (`Code: 79. Unknown mark file extension: '4'.
INCORRECT_FILE_NAME`), permanently stuck its replication queue, and failed
the write-then-read-back probe for the first time in the run.

Tried extensively (8+ attempts across clickhouse.com's combined changelog
pages, version anchors, GitHub's raw CHANGELOG.md, per-release GitHub
pages, and the docs repo's raw markdown) to find the exact changelog entry
for this and came up empty -- every source either 404s/robots-blocks or
truncates to only the most recent 1-2 months regardless of prompting.
Confirmed the exact intermediate version numbers via
endoflife.date's API instead.

harness/versions.py's LTS_HOPS now walks 25.4.13.22, 25.5.11.15,
25.6.13.41, and 25.7.8.71 between the two LTS releases (mirrored in
.github/workflows/clickhouse_upgrade_test.yml as 8 hops instead of 4, kept
in sync by the existing sanity-check assertion) so the next CI run
localizes which specific monthly release introduces the incompatible mark
format instead of only knowing it's somewhere in a 5-month span.
RECOMMENDED_NOW is pulled back to 25.3.14.14 -- the only hop confirmed
clean in real CI -- until that lands. README's TL;DR and PR #477
review-response sections updated to match; job timeout bumped 60->90min
for the added steps.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ings

Run 32047534149 (staged-upgrade, with the bisection hops from the
previous patch) came back conclusive: 24.8.6.70 -> 25.3.14.14 ->
25.4.13.22 -> 25.5.11.15 -> 25.6.13.41 -> 25.7.8.71 all upgrade cleanly,
node by node, zero hard errors -- only the expected transient bounce
noise. The failure reappears exactly and only at 25.7.8.71 -> 25.8.29.51,
same failure family as run 32044578317 but a different specific
manifestation this time (Code: 226 NO_FILE_IN_DATA_PART, missing
columns_substreams.txt, new .cmrk4 mark-file extension). This rules out a
gradual drift across the 25.3-25.8 span: it's one version boundary,
25.8.29.51, changing the on-disk compact-part format in a way no earlier
binary in this range can read.

Ties this to a changelog note found earlier (v25.12: "Enable advanced
shared data for JSON by default... after that change downgrade to
versions before 25.8 will be not possible, because these versions won't
be able to read new data parts with JSON column") -- scoped to JSON and to
downgrading, but naming 25.8 as where the underlying substream-based part
format landed. Our failing table (citizenlab) has no JSON column, so this
is likely that same infrastructure applying more broadly than advertised.
Flagged as corroborating, not confirmed -- still haven't gotten the actual
changelog text to read directly despite repeated attempts.

harness/versions.py's module docstring is rewritten to record this as a
confirmed result rather than an open investigation. RECOMMENDED_NOW stays
at 25.3.14.14 rather than bumping to 25.7.8.71: the newly-confirmed-clean
releases are all non-LTS with ~1 month of support each, so "safe to pass
through" isn't the same claim as "worth resting on". README's "Real CI
findings", "What was and wasn't verified" sections updated with the run 3
results and both run URLs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… hop

Adds continue-on-error to every upgrade-node/verify-ddl step so a hard
failure on one node (e.g. the confirmed 25.7.8.71 -> 25.8.29.51 mark-file
incompatibility) doesn't stop the remaining nodes in that hop from
upgrading too. The job's pass/fail signal moves to `ci_step.py report`,
which now returns 1 if any recorded step failed, since individual steps
no longer do. This lets the next CI run reveal whether upgrading the
lagging node to the same version also clears its stuck replication
queue, or whether the divergence is permanent.
Run 32122682392 (continue-on-error added last patch) completed the full
8-hop staged-upgrade ladder: both the 25.8.29.51 mark-file incompatibility
and the 26.3.17.110 nested-type serialization change (flagged in PR #477
review) turn out to be transient, self-healing mixed-version friction --
the lagging node's stuck replication queue clears as soon as its own
upgrade finishes, in both cases. Neither is a structural block.

Promotes RECOMMENDED_NOW to 26.7.3.19 and adds PRODUCTION_HOPS, a 4-hop
runbook (skipping the monthly bisection releases, which existed only to
localize the incompatibility in CI) with an operational caveat: upgrade
all three nodes back-to-back for the two hops that hit a real
incompatibility, rather than spacing them out like every other hop.

Also fixes a reporting gap in validate.py: NO_FILE_IN_DATA_PART wasn't in
HARD_ERROR_NAME_PATTERNS, so hop6-ch2's report said "no incompatibility
errors" even though it failed from exactly one (caught only via the
separate replication_queue_problems() check, not system.errors).
Extends the harness with a second, additional CI job that verifies the
actual OONI data pipeline -- not just synthetic seed data -- survives the
production upgrade path. Design per direct instruction: download real
data exactly once per run, re-verify integrity + query correctness at
each hop, and walk PRODUCTION_HOPS (the real 4-hop runbook) rather than
every diagnostic bisection waypoint.

- harness/real_data.py: new module. Snapshots real-data tables (row count
  + order-independent cityHash64 checksum) across all 3 nodes, loads real
  OONI measurements once via ooni/data#160's downloader/fastpath
  containers, takes a golden snapshot, then per hop: upgrades all 3 nodes
  (reusing scenarios.upgrade_node_step unmodified), diffs against the
  golden snapshot, and re-runs ooni/data's own pytest suite against the
  real api-oonimeasurements service.
- docker-compose.real-data.yml: new overlay adapting ooni/data#160's
  tests/integration stack onto this project's existing 3-node replicated
  cluster instead of a single throwaway node. Pinned to the
  add_end_to_end_tests branch since #160 is still unmerged -- update the
  ref once it merges.
- sql/001_schema.sql: adds fingerprints_dns, fingerprints_http,
  obs_web_ctrl, obs_http_middlebox, obs_openvpn -- the tables
  oonipipeline's observations workflow and fastpath actually write, with
  per-table provenance notes and manual signedness fixes where
  oonipipeline's own DDL generator is known-wrong (Int32/Int8 vs the
  already-verified obs_web table's UInt32/UInt8 convention).
- harness/compose.py: multi-compose-file support (files= param throughout)
  plus inspect_exit_code()/run_oneoff() for polling one-shot containers
  and running a fresh `verify` pytest pass per hop.
- harness/scenarios.py: extracted apply_schema() from
  load_schema_and_seed() so the real-data scenario can reuse the schema
  load without the synthetic seed rows; generalized step_ok() to cover
  every non-upgrade-node step shape via a single top-level "ok" fallback.
- harness/report.py: render_ci_step() now handles the real-data scenario's
  step shapes (setup-real-data, load-real-data, golden-snapshot,
  verify-e2e, real-data-hop).
- ci_step.py: 6 new subcommands (setup-real-data, load-real-data,
  golden-snapshot, verify-e2e, real-data-hop, teardown-real-data) wiring
  harness/real_data.py into discrete, individually pass/fail-able CI
  steps, matching the existing setup/upgrade-node/verify-ddl pattern.
- .github/workflows/clickhouse_upgrade_test.yml: new real-data-upgrade
  job (checkout ooni/data at add_end_to_end_tests, sanity-check
  PRODUCTION_HOPS alignment, setup/load/snapshot/4-hop-loop/report/
  teardown). Deliberately not run on pull_request -- depends on an
  unmerged external branch and real network downloads -- only on push to
  main and workflow_dispatch (scenario: real-data or all).
- README.md: documents the new job's design, the load-once/verify-per-hop
  rationale, the external-branch caveat, and marks PR #477 review point 4
  ("run the target version against the real API + data pipeline") as
  addressed.

Verified offline (no Docker daemon available in this sandbox, consistent
with every prior patch in this project): all Python modules compile, all
YAML/compose files parse and interpolate, SQL is structurally sound
(balanced parens/backticks, unique table names), and every new step shape
round-trips correctly through step_ok()/render_ci_step(). Real validation
deferred to an actual GitHub Actions run.
a820098 extracted apply_schema() out of load_schema_and_seed() so the
real-data scenario could reuse the schema load without the synthetic
seed rows. That extraction left load_schema_and_seed()'s seed-loading
loop still calling entry.execute(...), but entry was only ever assigned
inside apply_schema()'s own scope now -- broke setup for both
staged-upgrade and direct-jump-upgrade (NameError: name 'entry' is not
defined), confirmed in CI run 32132374150.
Wires harness/real_data.py's real-data scenario (and the ci_step.py
subcommands added in a820098) into an actual CI job -- these existed
but weren't invoked anywhere yet. Adds the real-data-upgrade job
(checkout ooni/data, sanity-check PRODUCTION_HOPS, setup/load/
snapshot/4-hop-loop/report/teardown) plus the real-data/all
workflow_dispatch scenario options. See a820098's commit message and
README.md's "Real-data end-to-end scenario" section for the design.
…y.py)

Existing scenarios only compare data-at-rest between checkpoints; this
adds a continuous read/write canary that runs for the entire
PRODUCTION_HOPS rollout and directly proves the three claims that matter
for a live-cluster go/no-go: no full-cluster downtime, no blocked
writes, and no lost or corrupted data, not just "data matched at the
start and end."
decrease the number of upgrade steps after identifying which releases
caused issues, add a more aggressive upgrade experiment, target the most
recent LTS release.
ensure that content is not corrupted, not just row counts, after each
upgrade hop
@aagbsn

aagbsn commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

ClickHouse 24.8 → 26.8: query audit + staged rollout checklist

Query audit. Scanned ooni/backend (705a7da, incl. legacy api/ and fastpath) and ooni/data (9ca7326). Ran the oonimeasurements and oonipipeline test suites, ~90 extracted queries, all DDL and the web-analysis query against real 24.8.6.70 and 26.8.9.10 servers loaded with identical data. No query breaks or returns different results. One default change has to be pinned (async inserts), plus a few cosmetic diffs listed at the bottom.

Rollback. Tested by writing OONI-shaped tables (fastpath, obs_web) on the new version, then starting the previous binary on the same data dir:

hop rollback, defaults rollback, with pins
24.8 → 25.3 ✅ n/a
25.3 → 25.8 ✅ ✅
25.8 → 26.3 ❌ 25.8 detaches every part as broken, tables come up empty ✅
26.3 → 26.8 ✅ ✅
25.8 → 26.8 ❌ same as above ✅

So the point of no return is removing the merge_tree pins, not installing 26.3. With the pins, the cluster can go all the way to 26.8 and still roll back to 25.8. Removing them is a separate, deliberate step at the end.

⚠️ Older servers refuse to start if the merge_tree config has a setting they don't know. Each pin has to ship with the binary that introduces it, and come out again if you roll back.


Phase 0: prep (still on 24.8, no node changes)

  • Add async_insert: 0 to the write and default profiles (ansible/group_vars/clickhouse/vars.yml). 26.3 turns async inserts on by default, and fastpath sends one INSERT per measurement: measured 3 ms → 61 ms per insert. 24.8 already knows this setting, so it's safe to deploy now.
  • Make the merge_tree config in the clickhouse role depend on the installed version, so it can carry the pins below.
  • Rerun the staged-upgrade CI with the pins and check whether the hop 2/3 trailing-node errors go away.
  • Add a ClickHouse version matrix to backend CI. Its tests pin 25.2 / latest / 22.8, never the production version (Add end to end tests data#160 has the pattern).
  • Take a fresh backup before each phase.

Phase 1: 24.8.6.70 → 25.3.14.14 (reversible, no pins needed)

  • Rolling upgrade, one node at a time
  • Soak: fastpath lag, API errors, system.replication_queue, system.detached_parts = 0

Phase 2: 25.3 → 25.8.29.51 (reversible)

  • Ship merge_tree.write_marks_for_substreams_in_compact_parts: 0 with the 25.8 binary. The changelog says servers older than 25.5 can't read the new compact parts; our test read them fine, but the pin is cheap insurance.
  • Rolling upgrade, all 3 nodes back to back
  • Soak. Keep it short: 25.8 LTS support ended 2026-08-29.
  • Rollback = 25.3 binary and remove the pin

Phase 3: 25.8 → 26.3.17.110 (kept reversible by the pins)

  • Ship these with the 26.3 binary, keeping the Phase 2 pin:
    serialization_info_version: basic, string_serialization_version: single_stream, propagate_types_serialization_versions_to_nested_types: 0
  • Check that async_insert is actually 0 for the fastpath user (SELECT getSetting('async_insert'))
  • Rolling upgrade, back to back
  • Soak
  • Rollback = 25.8 binary and remove these 3 pins

Phase 4: 26.3 → 26.8.9.10 (still reversible to 25.8)

  • Rolling upgrade, keeping all pins
  • Set asynchronous_metrics_key_values_mode: both. 26.8 turns the per-CPU/disk/network metrics on /metrics into labelled series, which breaks dashboards that use the old names.
  • Soak

Phase 5: commit (point of no return)

  • Decide explicitly; take a fresh backup
  • Remove the merge_tree pins. New parts get the new formats, so downgrading below 26.x is no longer possible.
  • Revisit async_insert for fastpath (batch writes, or keep it at 0)

Code follow-ups (non-blocking)

  • oonimeasurements/routers/data/list_analysis.py:134 and list_observations.py:251: add a unique tiebreaker to ORDER BY. Rows with tied sort keys come back in a different order on 26.8, so page contents differ.
  • Informational: Float32 parsed from text is exactly rounded from 26.7 (1-ULP differences, text-loaded fixtures only; clickhouse-driver inserts are unaffected). db_stats.bytes / row_count values change. Some default-valued Float columns come back as 0.0 instead of 0 through clickhouse-driver.
  • Existing bugs, unrelated to the upgrade: fastpath obs_openvpn DDL uses Uint8, which fails on every version; oonipipeline checkdb reports false diffs (Datetime64 casing, probe_id type).

@aagbsn

aagbsn commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

We should also read the changelog for any potential tricky breaking changes and ideally run the target version with the API + data pipeline to make sure that it's able to work without any problems.

Scrolling through the changelog for "backward incompatible change" here are some highlights worth being careful of:

Downgrading after upgrading may cause data loss. Propagate data types serialization versions to nested data types
https://clickhouse.com/docs/resources/changelogs/oss/2026#263-backward-incompatible-change

So we need to be extra careful once we go past 26.3 (maybe we should stop 1 LTS behind until we have done thorough testing).

Especially tricky are ones where they say things like:

Renamed functions searchAny and searchAll to hasAnyTokens and hasAllTokens for better consistency with existing function hasToken
https://clickhouse.com/docs/resources/changelogs/oss/2025#2510

Which means that we need to make sure that we aren't using these functions, otherwise they will crash at runtime.

Disallow truncating replicated databases
https://clickhouse.com/docs/resources/changelogs/oss/2025#backward-incompatible-change-10

Might apply to us in the data pipeline

It might also be worth passing the whole changelog history into an llm and give it our codebase to hunt for potential breakages.

It looks like there shouldn't be any breaking changes with regard to these

In any case I would suggest for sure upgrading to strictly less than 26.3, since if we mess that upgrade path up, we have no way to rollback.

This is apparently possible by pinning and disabling specific features. Each phase of the upgrade can be completed and let to 'soak' before proceeding in case of any bugs that are surfaced that we didn't find before hand. However, sitting at a LTS release prior to 26.3 keeps us out of the support window.

There is also a decision to be made with regard tohow inserts are performed:

fastpath and async inserts. fastpath (10 workers in prod) sends one INSERT per measurement and waits for each one. From 26.3, every INSERT is async by default (async_insert=1, wait_for_async_insert=1): the server holds the row in a buffer until a flush timer fires (adaptive, up to 200 ms), then writes one part for everything buffered. That suits many concurrent clients, but not a client that waits on every row.

Measured with fastpath's real row shape on a ReplicatedReplacingMergeTree table, 10 concurrent workers. This was a single-node sandbox with embedded Keeper, so production latencies will be higher; the ratios are what matters.

setup throughput p50 insert parts created
24.8 today 302 inserts/s 30 ms 800 per 800 rows
26.8 defaults 48 inserts/s 204 ms 110 per 800 rows
26.8 async_insert=0 250 inserts/s 36 ms 800 per 800 rows
26.8 async, wait=1, async_insert_busy_timeout_max_ms=20 216 inserts/s 46 ms 137 per 800 rows
26.8, fastpath sending 25 rows per INSERT, async_insert=0 ~6,500 rows/s 34 ms 120 per 3,000 rows

On 26.8 defaults, fastpath's write capacity drops about 6×. When the workers fall behind, the 500-slot queue fills and fastpath's HTTP handler blocks. The ooniprobe service's submit then times out, and measurements go to the S3 fallback (MISSED_MSMNTS).

What the options mean for reliability:

  • async_insert=0: today's behaviour.
  • async_insert=1, wait_for_async_insert=1 (the 26.3 default): same guarantees as a sync insert. The client is acknowledged only after the part is written and committed to Keeper, and errors come back to the client. In our test, a row that violated a constraint was rejected with an error, exactly as a sync insert would be. The only cost is latency. If a flush fails, every insert in that batch gets the error and can retry.

@aagbsn

aagbsn commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

ClickHouse upgrade 24.8 → 26.8 LTS: executive summary

Bottom line: the upgrade is safe to go ahead as a staged rolling upgrade (24.8 → 25.3 → 25.8 → 26.3 → 26.8), provided the config pins below go out with each new version. No application code changes are needed in ooni/data or ooni/backend. The risk is almost entirely in operations: replication between old and new versions during the rollout, and keeping the option to roll back.

Scope: every one of the 7,348 ClickHouse changelog entries from 24.9 through 26.8.12 was checked against our code and config. 79 touch something we use, and the important ones were reproduced on real servers. Each version step was also rehearsed on a 3-node cluster set up like production.

Concerns raised in review: both are clear

  • TRUNCATE: the 25.3 change only affects a database type we don't use. TRUNCATE TABLE event_detector_cusums SYNC worked at every step, run from either the upgraded node or the old one.
  • Renamed/removed functions (searchAny/searchAll etc.): none are used. Those two functions never existed in our current version, and none of the 7 functions removed since 24.8 appear in our code.

Most important steps (in order)

# Step Why Status
1 Never skip a version step. Mixing 24.8 and 26.8 knocks the old node out of the coordination service (Keeper) the cluster depends on. Each planned step on its own is clean. Verified
2 Ship a MergeTree compatibility pin with each new version: 1 at 25.8, 4 at 26.3, 5 at 26.8. Without the 25.8 pin, old replicas stop replicating new data mid-rollout. We had thought of this pin as optional insurance; it is required. The pins also keep rollback possible. Verified
3 At 26.8, pin the Keeper snapshot format (write_snapshot_version: 6). A 26.8 Keeper writes snapshots that no older version can read, so rollback is impossible without this pin. Newly found. Verified
4 At 26.8, add the skip-index pin (packed_skip_index_max_bytes: 0). Otherwise older replicas can't use the index behind the API's measurement lookup, and those lookups become full scans during the rollout or after a rollback. Newly found. Verified
5 Keep async_insert: 0 pinned before 26.3. Otherwise fastpath write capacity drops about 6×, and measurements fall back to S3. Already planned
6 Spend at least 1 hour on 26.3 before moving to 26.8. This gives insert deduplication time to switch to 26.8's new format. Required by ClickHouse
7 Check that data1-3 and notebook CPUs support AVX2 before 26.8. The standard 26.8 build won't start without AVX2. Not yet checked
8 Raise systemd's stop timeout to 180 s. 26.8 waits up to 120 s to shut down cleanly, but systemd kills it at 90 s on every restart. Config change
9 Take a backup before each step and pause Airflow during each node restart. Standard safety. Process
10 Remove all pins only as a deliberate final step. After that, rolling back is no longer possible. Decision point

Test in staging before the 26.8 step

  • Memory limits: 26.8 accounts memory more strictly, and it lowers its own limit when Airflow or JupyterHub are using RAM, so heavy queries could fail where they succeed today.
  • Write and merge load: the big analysis insert now runs multi-threaded, and column statistics are now built automatically during merges.
  • Notebook users: a few query results change, for example 64-bit integers are no longer quoted in JSON and NOT binds more loosely. They should be told.

Found along the way (unrelated to the upgrade)

  • The replication bandwidth caps in our ansible vars are never applied.
  • The per-user query quotas are never enforced.

Both come from gaps in the ansible role we use, and are worth fixing separately.

Full report with evidence, config snippets and the list of all 79 relevant changes: clickhouse-changelog-review-24.8-to-26.8.md.

@aagbsn

aagbsn commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

ClickHouse 24.8.6.70 → 26.8 LTS: changelog review for ooni/data and ooni/backend

Summary

I went through every ClickHouse changelog entry from 24.9 to 26.8, plus the 26.8 patch releases up to 26.8.12.53. That is 7,348 entries across all categories, not only the "Backward Incompatible" sections. I checked each one against ooni/data, ooni/backend and the ooni/devops ClickHouse config. 79 entries touch something OONI uses. I tested the ones that matter against the real 24.8.6.70, 25.3.14.14, 25.8.29.51, 26.3.17.110 and 26.8.9.10 binaries, and ran each planned hop of the rolling upgrade on a 3-node cluster set up like production (embedded Keeper on every node).

  • TRUNCATE: nothing to change. The 25.3 change only blocks TRUNCATE DATABASE on databases that use the Replicated database engine. ooni is an Atomic database, and OONI only runs TRUNCATE TABLE, which works on 26.8 (tested, including on ReplicatedReplacingMergeTree tables).
  • Renamed or removed functions: none are used. searchAny/searchAll were only added in 25.7, so code running against 24.8 can't be calling them. None of the 7 functions removed between 24.8.6.70 and 26.8.9.10 (list below) appear in the code.
  • No query or API code changes are required in ooni/data or ooni/backend.
  • Config and rollout changes are required. Two of them are new since the clickhouse version is out of support period #437 plan:
    1. The 25.8 write_marks_for_substreams_in_compact_parts: 0 pin is required, not just insurance. Without it, 25.3 replicas stop replicating parts written by 25.8 during the rolling upgrade (reproduced).
    2. A 26.8 Keeper writes snapshots that no older version can load, so rolling back from 26.8 is impossible unless write_snapshot_version is pinned (reproduced; the pin fixes it).
    3. Add packed_skip_index_max_bytes: 0 to the 26.8 merge_tree pins. Without it, 26.3 replicas, and any rollback, can't use the skip indexes on parts that 26.8 wrote, including fastpath's measurement_uid_idx (reproduced).
    4. Don't skip hops. A Keeper cluster with both 24.8 and 26.8 nodes loses the old node (reproduced). Each planned hop on its own is clean.

1. hellais' two questions

TRUNCATE. These are all the TRUNCATE statements in the three repos:

Where Statement Status
oonipipeline/tasks/updaters/citizenlab_test_lists_updater.py:130 TRUNCATE TABLE citizenlab_flip works on 26.8; removed anyway by ooni/data#190
oonipipeline/cli/commands.py:329 TRUNCATE TABLE event_detector_cusums SYNC (only with --truncate-cusums) works on 26.8
backend/analysis/analysis/citizenlab_test_lists_updater.py:106 TRUNCATE TABLE citizenlab_flip (legacy updater) works on 26.8
test fixtures (oonipipeline/tests/conftest.py, ooniapi/services/*/tests/conftest.py) TRUNCATE TABLE <t> works on 26.8

The changelog entry (#76651, 25.3) is about the Replicated database engine. ooni is created with CREATE DATABASE ooni ON CLUSTER oonidata_cluster (Runbooks.md:1082), which has no ENGINE clause and so defaults to Atomic. The tables are ReplicatedMergeTree tables inside that Atomic database, and that combination is unaffected. Results on a real server:

Statement 24.8.6.70 26.8.9.10
TRUNCATE TABLE on ReplicatedReplacingMergeTree in an Atomic DB (with and without SYNC) ok ok
TRUNCATE TABLE on plain ReplacingMergeTree ok ok
TRUNCATE TABLE on a table inside a Replicated database ok ok
TRUNCATE DATABASE on an Atomic database ok ok
TRUNCATE DATABASE on a Replicated database ok Code 48 NOT_IMPLEMENTED

In production the tables Airflow truncates (event_detector_cusums, citizenlab_flip) are ReplicatedReplacingMergeTree (devops scripts/cluster-migration/schema.sql). I ran the exact statement from cli/commands.py:329, TRUNCATE TABLE event_detector_cusums SYNC without ON CLUSTER, on one node of a 2-replica pair. The data was removed on both replicas, later inserts replicated normally, and the replication queue had no errors. The result was the same with both replicas on 24.8, both on 26.8, and one on 26.3 with the other on 26.8 (the last-hop state).

Renamed or removed functions, settings, engines and types. I checked these three ways:

  • Diffing the real servers. I dumped system.functions, system.settings, system.merge_tree_settings, system.server_settings, system.table_engines and system.data_type_families from running 24.8.6.70 and 26.8.9.10 servers and diffed them.
    • Removed functions: dateTimeToSnowflake, dateTime64ToSnowflake, snowflakeToDateTime, snowflakeToDateTime64, detectProgrammingLanguage, kql_array_sort_asc, kql_array_sort_desc. None are used.
    • Removed engines and types: LiveView, ExternalDistributed, Object, TIME. None are used. The Object hits in the code are JavaScript and Python identifiers, not ClickHouse column types.
    • No query-level or MergeTree settings were removed. 8 server settings were removed (page_cache_*, cgroup_memory_watcher_*, format_alter_operations_with_parentheses); none of them are in our config.
    • Settings that are obsolete in 26.8 and still used by OONI: none. The two bandwidth settings in clickhouse_config are still valid server settings, but they are never rendered into the config at all (see §7).
  • Grepping every identifier the changelog calls renamed, removed, deprecated, forbidden or rejected. That is every backticked name in those entries plus the Backward Incompatible sections. Only 32 appear anywhere in the repos, all generic names (MergeTree, DateTime, toDate, max_execution_time and the like). I read each matching entry, and none change behaviour OONI relies on.
  • Grepping directly for searchAny, searchAll, hasAnyTokens, hasAllTokens, LIVE VIEW, Object('json'), allow_experimental_object_type, generateSeries, deltaSumTimestamp, and ORDER BY tuple() on Replacing/Collapsing/Summing tables (forbidden since 25.12). No hits.

2. How the review was done

  • Scope: every entry in the ClickHouse changelogs for 24.9 through 26.8, in all categories (bug fixes, improvements, performance, new and experimental features, build, backward incompatible), plus the 26.8.2 through 26.8.12.53 patch-release notes. That is 7,348 entries after de-duplication.
  • Reading: the entries were split into 9 slices, and each slice was read in full, one entry at a time. Every entry that could plausibly touch OONI was checked against ooni/data @07e13bd, ooni/backend @705a7da and ooni/devops @114ccbb, plus the ansible role idealista.clickhouse_role 3.5.1 that renders our server config. Entries for features OONI doesn't use (Iceberg, Kafka, S3 disks, vector and text search, the Web UI, AI functions and so on) were dropped.
  • Behaviour changes: tested with clickhouse local or a real server on both versions where possible. Setting defaults were diffed between real 24.8.6.70 and 26.8.9.10 servers.
  • Mixed-version cluster harness: 3 nodes, each running clickhouse-server with embedded Keeper, oonidata_cluster with 1 shard and 3 replicas, the same macros as production, and an Atomic ooni database. For each planned hop:
    • Nodes are upgraded one at a time on the same data directories.
    • The Keeper leader is moved onto the new node and onto an old node.
    • From both new and old nodes, a workload runs: CREATE ... ON CLUSTER ReplicatedReplacingMergeTree, INSERT, ALTER DELETE with mutations_sync, OPTIMIZE FINAL, EXCHANGE TABLES ON CLUSTER, REPLACE PARTITION FROM a temporary table, TRUNCATE ON CLUSTER SYNC, and DROP ... SYNC.
    • Each node is then checked for replication-queue errors, detached parts, read-only replicas, row-count agreement and Keeper health.
  • Keeper rollback: Keeper snapshots written by each version were loaded by the previous versions.

3. What the mixed-version tests found

Hop (rolling, one node at a time) MergeTree pins on upgraded nodes Result
24.8.6.70 → 25.3.14.14 none clean
25.3.14.14 → 25.8.29.51 none ❌ the 25.3 replicas can't execute merges or mutations on parts written by 25.8 (Code: 79 Unknown mark file extension: '4'), the replication queue backs up, and row counts diverge (1980 on 25.3 vs 2969 on 25.8) until the last node is upgraded
25.3.14.14 → 25.8.29.51 write_marks_for_substreams_in_compact_parts: 0 clean
25.8.29.51 → 26.3.17.110 the above plus serialization_info_version: basic, string_serialization_version: single_stream, propagate_types_serialization_versions_to_nested_types: 0 clean
26.3.17.110 → 26.8.9.10 the same 4 pins replication clean, but 26.3 replicas can't use skip indexes on parts written by 26.8 (see 4c)
26.3.17.110 → 26.8.9.10 4 pins plus packed_skip_index_max_bytes: 0 clean, and skip indexes work on both versions
24.8.6.70 ↔ 26.8.9.10 in one Keeper cluster (a skipped hop) none ❌ once a 26.8 Keeper is leader, the 24.8 follower loops on Failed to parse request from log entry: Operation 503 is unknown and stops serving
same, with Keeper feature_flags off on the 26.8 nodes none clean

Keeper snapshot rollback:

Snapshot written by Loaded by Result
25.3 24.8 ok
25.8 25.3 ok
26.3 25.8 ok
26.8 (default) 26.3, 25.8, 24.8 ❌ Keeper doesn't start: Unsupported snapshot version 8, "Manual intervention is necessary"
26.8 with write_snapshot_version: 7 26.3, 25.8, 24.8 ❌ same error, version 7
26.8 with write_snapshot_version: 6 26.3, 25.8, 24.8 ok, all znodes present

Keeper feature flags that each version turns on by default (from ftfl):

Version On by default
24.8 filtered_list, multi_read
25.3 same as 24.8 (knows remove_recursive but leaves it off)
25.8 + check_not_exists, create_if_not_exists, remove_recursive, multi_watches
26.3 + check_stat, persistent_watches, try_remove, list_with_stat_and_data
26.8 + get_children_recursive, max_request_size

On each planned hop, neither the server workload nor a keeper-client battery (create, set, rm, rmr, ls, find_big_family, sync) caused any parse errors on the older Keeper. Only mixing 24.8 with 26.8 breaks.

Several runs show transient Code: 999 Session expired errors right after the test forced a Keeper leader change. They appeared in the 24.8 → 25.3 hop too and are not a version issue, but expect a few in production while each node restarts.

4. Required changes (config and rollout)

4a. The 25.8 compact-part pin is required. In #437 Phase 2, change "cheap insurance" to required. Ship it with the 25.8 binary and keep it until Phase 5.

4b. Pin Keeper's snapshot format while rollback is still possible. Ship this with the 26.8 binary only. 26.3 and older reject the setting (Code 115 Unknown setting 'write_snapshot_version'), so it must also be removed when rolling back. The role renders clickhouse_keeper.coordination_settings from vars, so it can go there or in a config.d file:

<clickhouse>
  <keeper_server>
    <coordination_settings>
      <write_snapshot_version>6</write_snapshot_version>
    </coordination_settings>
  </keeper_server>
</clickhouse>

Remove it in Phase 5 (point of no return). Also back up /var/lib/clickhouse/coordination before each hop.

4c. Add a fifth MergeTree pin with 26.8. packed_skip_index_max_bytes now defaults to 1M, so 26.8 writes each part's skip indexes into a single skp_idx.packed file, and versions before 26.6 ignore it. Tested: with a part written by 26.8 on defaults, a measurement_uid-style lookup reads 1954/1954 granules on 26.3 against 1/1954 on 26.8. With the pin it reads 1/1954 on both. That lookup is SELECT ... FROM fastpath WHERE measurement_uid = :uid in oonimeasurements/routers/v1/measurements.py:199, which depends on measurement_uid_idx. So during the 26.3/26.8 window, and after any rollback, it would become a full scan. The pin has to be server-level: set as a table SETTINGS, it stops the table from attaching on older versions (UNKNOWN_SETTING).

<merge_tree>
  <write_marks_for_substreams_in_compact_parts>0</write_marks_for_substreams_in_compact_parts>
  <serialization_info_version>basic</serialization_info_version>
  <string_serialization_version>single_stream</string_serialization_version>
  <propagate_types_serialization_versions_to_nested_types>0</propagate_types_serialization_versions_to_nested_types>
  <packed_skip_index_max_bytes>0</packed_skip_index_max_bytes>   <!-- new, 26.8 only -->
</merge_tree>

4d. Don't skip hops. Keep 24.8 → 25.3 → 25.8 → 26.3 → 26.8. The changelog itself says upgrades spanning more than a year are unsupported (#71385). If a hop ever has to be skipped, turn the new flags off on the upgraded nodes until the last node is done. Older versions reject flag names they don't know (Code 36 Invalid feature flag), so this override also has to ship only with the newer binary:

<keeper_server><feature_flags>
  <check_not_exists>0</check_not_exists><create_if_not_exists>0</create_if_not_exists>
  <remove_recursive>0</remove_recursive><multi_watches>0</multi_watches>
  <check_stat>0</check_stat><persistent_watches>0</persistent_watches><try_remove>0</try_remove>
  <list_with_stat_and_data>0</list_with_stat_and_data>
  <get_children_recursive>0</get_children_recursive><max_request_size>0</max_request_size>
</feature_flags></keeper_server>

4e. Soak on 26.3 for at least an hour before going to 26.8. 26.8 only understands the new unified insert-deduplication hash (#107886). The migration path is to run with compatible_double_hashes, which is 26.3's default, for longer than the deduplication window: replicated_deduplication_window_seconds defaults to 3600 on 26.3. Don't set insert_deduplication_version in any config: 26.8 refuses to start with anything other than new_unified_hash.

4f. Already in the plan (unchanged): async_insert: 0 in the default and write profiles (#97590); the four existing format pins; and asynchronous_metrics_key_values_mode: both on 26.8.

4g. Give systemd enough time to stop the server. In 26.8 shutdown_wait_unfinished goes from 5 s to 120 s (#110838). The role installs its own /etc/systemd/system/clickhouse-server.service, which has no TimeoutStopSec, so systemd's 90 s default will SIGKILL the server partway through a clean shutdown on every rolling restart. Add a drop-in with TimeoutStopSec=180, or set shutdown_wait_unfinished to 60. Pause the Airflow DAGs before restarting a node either way.

5. Decide or test in staging

# Change Why it matters for OONI Suggested action
1 Default x86 build needs AVX2 (26.6, #105019) A host without AVX2 won't start 26.8 at all. The repos don't record the CPU models of data1-3 or notebook. grep -m1 -o avx2 /proc/cpuinfo on each host before Phase 4; use the amd64compat build if it's missing
2 Unknown top-level config keys now stop the server (26.8, #100332) The role's rendered config.xml and users.xml start cleanly on 26.8 (tested). Hand-placed config.d/users.d files on the hosts aren't in the repo. Start 26.8 against a copy of each host's /etc/clickhouse-server before Phase 4
3 Memory limits: external-library allocations are now counted (25.8); the global limit now shrinks when other processes use RAM (26.6, memory_worker_dynamic_hard_limit); GROUP BY, ORDER BY and hash joins spill to disk at 50% of the limit (24.12, 26.5) API users have 1 GB limits. Airflow runs on data1 and JupyterHub on notebook, next to ClickHouse. Queries that fit today may fail, or may now spill to the ClickHouse tmp directory instead of failing. Replay the top memory users from system.query_log in staging; watch tmp disk. memory_worker_dynamic_hard_limit: 0 restores 24.8 behaviour.
4 max_insert_threads defaults to the number of cores (26.8) The big INSERT INTO analysis_web_measurement SELECT ... (web_analysis.py:506) writes in parallel: more parts and more memory Compare parts and peak memory in staging; if worse, set max_insert_threads=1 on that query
5 Automatic column statistics built during merges (auto_statistics_types, 26.4) Applies to existing tables too; adds merge CPU and changes the planner. Rollback to 24.8 was fine in testing. Measure merge load in staging; pin auto_statistics_types to empty (server-level) if it's too costly
6 MV insert deduplication on by default (deduplicate_blocks_in_dependent_materialized_views, 26.2) Only matters if counters_test_list / counters_asn_test_list / tls_consistency_matview exist in production with Replicated targets. The repo DDL uses non-replicated targets. SHOW CREATE them in production; pin to 0 if they're live and replicated
7 Fix for a mutation CHECKSUM_DOESNT_MATCH with serialization_info_version='basic' (26.8, #113588) That's our 26.3 rollback pin. The fix isn't backported to any 26.3 patch (checked up to 26.3.34.136). It didn't reproduce in the harness, whose ALTER DELETE ran on compact parts across restarts. Keep the 26.3 phase short; watch system.replication_queue during Phase 3
8 Web terminal on by default (enable_webterminal, 26.6) On notebook, nginx forwards /click to 8123 as the passwordless default user Set enable_webterminal: false everywhere
9 Notebook and HTTP users will see different results: 64-bit ints no longer quoted in JSON (25.8); NOT precedence (26.3: NOT (NULL) IS NULL is 1 on 24.8 and 0 on 26.8); date_time_input_format=best_effort (26.5); FINAL skips partition pruning (26.3/26.5) No production code path is affected: everything uses the native protocol and parameterised queries Tell JupyterHub users; optionally pin output_format_json_quote_64bit_integers=1 on notebook
10 Newer patch releases exist: 25.8.33.6, 26.3.34.136, 26.8.12.53 26.8.12.53 fixes wrong GROUP BY/DISTINCT/window results for keys like d + INTERVAL 1 MONTH (#121496; doesn't match any OONI query) and a Keeper session drop under backpressure (#121559). 26.8.10.6 adds TLS hostname verification (we don't use TLS between nodes). Consider 26.8.12.53 as the final target after re-running the harness; I haven't tested that binary

6. Behaviour changes that need no action but are worth knowing

  • Insert deduplication: repeated INSERT ... SELECT is no longer block-deduplicated (26.1), and the window for plain inserts drops from 1 week to 1 hour (25.10). A re-run of the analysis INSERT will add rows until merge. analysis_web_measurement is a ReplacingMergeTree, and the pipeline already OPTIMIZEs partitions after reprocessing (cli/commands.py:196), so nothing changes in practice. fastpath's retried single-row inserts are still deduplicated (tested).
  • Joins: the default join algorithm is now direct,parallel_hash,hash,ie_join (24.12), and ANY/OUTER joins can be rewritten into cheaper forms (25.9 to 26.5). The API's ANY LEFT JOIN test_groups USING (test_name) and the pipeline's grace_hash join gave identical results on 24.8 and 26.8. Avoid serving API traffic from 26.1 to 26.4, which had an ANY JOIN bug with today() filters; the planned hops don't land there.
  • Query settings now on by default: the query condition cache, compile_expressions, optimize_syntax_fuse_functions (sum/count/avg results and column names unchanged), read_in_order_use_virtual_row, join runtime filters and use_skip_indexes_on_data_read. These are performance changes; fastpath-shaped lookups read the same marks.
  • Server-wide scheduling and time limits: CPU concurrency control is on by default (2× cores, fair scheduler), so big pipeline jobs no longer starve fastpath inserts and API queries. max_execution_time is now enforced reliably, including on INSERT (26.7), so API queries that used to overrun their 30 s limit may now fail with a timeout error.
  • Mutations and deletes: mutations go through the new analyzer (26.6); the pipeline's ALTER DELETE gave identical results. ALTER DELETE is faster when a whole part matches (25.5), which helps event_detector_changepoints. Several lightweight DELETE and count() fixes landed.
  • Values: min/max/argMax now always skip NaN (26.4), DateTime64 range is extended to 0000-9999 instead of clamping at 1900-2299 (26.7), and Float parsing is exact (26.7, already known). No OONI query result changes in testing.
  • Keeper: durability and quorum fixes, and a fix for possible data loss when a replicated INSERT times out during Keeper recovery (26.8.3). Keeper no longer repairs a broken snapshot or log by itself (26.1), so add "move the bad snapshot aside" to Runbooks.md.
  • Ops: databases always load asynchronously (25.2), so gate each rolling step on system.replicas (no readonly replicas, queue drained) rather than on the port being open. The sampling profiler is on by default (25.10), so system.trace_log grows a little.

The full list of 79 entries, with severity and PR link, is in the appendix.

7. Pre-existing problems found along the way (not caused by the upgrade)

  • max_replicated_sends/fetches_network_bandwidth_for_server are not applied. They're set in clickhouse_config (vars.yml:47-48, "1GB/s 50% utilization cap"), but idealista.clickhouse_role 3.5.1's config.xml.j2 never renders them. They're the only two clickhouse_config keys it drops. Replication traffic is currently uncapped.
  • The SQL quotas are created empty and attached to no user. The role's DDL/QUOTA.j2 only reads item.settings and item.to, and DDL/USER.j2 ignores item.quota. So the oonimeasurements and oonitestlists limits in vars.yml (12000 queries / 600 s and so on) aren't enforced today.
  • 24.8 bug with DETACH/ATTACH on ReplacingMergeTree tables without a version column (fixed in 25.4). After DETACH PARTITION and ATTACH PARTITION on such a table with 10 or more parts, FINAL can keep the wrong row: in the reproduction, 24.8 returned 2 instead of 12, and 26.8 returned 12. Affected tables include obs_web, obs_web_ctrl, analysis_web_measurement, citizenlab and faulty_measurements. Until production is past 25.4, avoid DETACH/ATTACH round-trips on them. The runbook's ATTACH PARTITION ... FROM obs_web_bak was not affected in the same test.

8. Recommended changes by repository

  • ooni/data: nothing required. Optional: put max_insert_threads=1 on the analysis INSERT if staging (§5 row 4) shows it's worse. Keep the ClickHouse version matrix from Add end to end tests data#160, with 26.8 in it.
  • ooni/backend: nothing required. Add a ClickHouse 26.8 CI job, since backend tests currently pin 25.2, latest and 22.8, never the production version. Watch for API TIMEOUT_EXCEEDED errors after Phase 4 (§6).
  • ooni/devops (updates to the clickhouse version is out of support period #437 checklist):
    • Phase 2: the write_marks_for_substreams_in_compact_parts pin is required (4a)
    • Phase 3: soak on 26.3 for at least 1 h before Phase 4 (4e); keep the phase short (§5 row 7)
    • Phase 4, with the 26.8 binary: add packed_skip_index_max_bytes: 0 to the merge_tree pins (4c); add Keeper write_snapshot_version: 6 (4b); add enable_webterminal: false (§5 row 8)
    • Before Phase 4: AVX2 check on all four hosts; test-start 26.8 against each host's real config tree (§5 rows 1-2)
    • systemd drop-in TimeoutStopSec=180 (4g)
    • Rollback steps for Phase 4: use the 26.3 binary and remove packed_skip_index_max_bytes and write_snapshot_version, both of which 26.3 rejects. Back up coordination/ before each hop.
    • Phase 5: remove all the pins, including write_snapshot_version
    • Fix the two role gaps in §7, or render those settings through config.d

9. Limits of this review

  • The harness is small. It's local and single-machine, with small tables and no production load or network latency. It shows compatibility, not performance; §5 lists what needs a staging run with production-shaped data.
  • Patch notes reviewed only for 26.8. Only the 26.8.x patch lines were reviewed; later 25.3.x, 25.8.x and 26.3.x patches are backports of fixes that already appear in the minor-release changelogs. 26.8.12.53 wasn't run.
  • Some findings come from the reviewers' own tests. Each ACTION/CHECK item comes from an entry that was read in full and checked against the code. I re-ran the mixed-version, Keeper, snapshot, skip-index and TRUNCATE results myself. The memory, deduplication, JSON and parsing findings come from each reviewer's own clickhouse local comparisons of the two versions.

Appendix: all 79 entries that touch OONI
Release PR Severity Change
25.8 #84171 ACTION write_marks_for_substreams_in_compact_parts enabled by default; servers older than 25.5 cannot read new Compact parts. Reproduced in the 3-node harness: 25.3 replicas stall with "Unknown mark file extension" when 25.8 writes parts.
26.3 #94859 ACTION propagate_types_serialization_versions_to_nested_types enabled by default; parts written by 26.3+ with nested String types unreadable by older versions (downgrade unsafe).
26.3 #97590 ACTION async_insert turned on by default (compatibility < 26.2 restores false).
26.8 #111816 ACTION MergeTree setting packed_skip_index_max_bytes now defaults to 1M, so skip indexes are written into one skp_idx.packed archive per part. Versions before 26.6 treat these indexes as not materialized. compute_exact_num_defaults_for_sparse_columns is also on by default now.
26.8 #111715 ACTION Keeper default write_snapshot_version raised to 8 (the setting is new in 26.8 and hot-reloadable).
24.11 #71385 CHECK ReplicatedMergeTree no longer creates a missing metadata_version Keeper node; the changelog says ClickHouse does not support upgrades that span more than a year and asks users to upgrade gradually.
24.12 #71406 CHECK Automatic spill of GROUP BY / ORDER BY to disk based on memory usage (max_bytes_ratio_before_external_group_by / _sort, default 0.5 in 26.8; the settings do not exist in 24.8).
25.8 #74079 CHECK output_format_json_quote_64bit_integers default changed 1 -> 0: (U)Int64 values are no longer emitted as quoted strings in JSON* output formats.
25.8 #84082 CHECK All allocations by external libraries are now counted by the memory tracker; reported query memory can rise and queries can newly fail with MEMORY_LIMIT_EXCEEDED. Related: E00886 (25.11) now tracks temporary hash-join result allocations.
26.1 #93886 CHECK Keeper CHECK_STAT and TRY_REMOVE extensions are now enabled by default. A 26.8 Keeper also advertises remove_recursive, multi_watches, persistent_watches, list_with_stat_and_data, get_children_recursive, check_not_exists and create_if_not_exists by default (the 4lw ftfl output). Only a problem if a hop is skipped. Staged hops tested clean; a direct 24.8/26.8 mix takes the old Keeper node down (reproduced).
26.2 #95970 CHECK Deduplication turned ON for all inserts by default, including async inserts and dependent materialized views (deduplicate_blocks_in_dependent_materialized_views 0 -> 1; new deduplicate_insert setting).
26.4 #101275 CHECK MergeTree setting auto_statistics_types now defaults to non-empty, so column statistics are built automatically for every suitable column (26.8.9.10 default: 'basic, uniq_v2'; materialize_statistics_on_insert=1 in 26.8.9.10).
26.6 #107886 CHECK Server setting insert_deduplication_version default changes to new_unified_hash: deduplication is per whole INSERT, not per part/partition; in 26.8 it is a deprecated guard and only new_unified_hash is supported (server refuses to start with other values).
26.6 #105019 CHECK Default x86 build now targets x86-64-v3 (AVX2); CPUs without AVX2 need the amd64compat build.
26.6 #104964 CHECK MemoryWorker now dynamically lowers the server hard memory limit to (RSS + MemAvailable) * max_server_memory_usage_to_ram_ratio (memory_worker_dynamic_hard_limit=1). 26.7 also adds memory_worker_rss_speculative_reserve_ratio=1.0, which raises MEMORY_LIMIT_EXCEEDED earlier (E04120).
26.6 #106255 CHECK The /webterminal HTTP endpoint (an interactive clickhouse-client over WebSocket) is now production and enabled by default (enable_webterminal=1).
26.8 #109006 CHECK max_insert_threads default goes from 0/1 to auto (the number of CPU cores). This parallelizes the write side of INSERT SELECT, and in some cases of plain INSERT. It can change the number of parts and the order of inserted rows.
26.8 #110838 CHECK The server now has an introspection port, and the shutdown_wait_unfinished server setting goes from 5 to 120 seconds.
26.8 #100332 CHECK Server refuses to start with UNKNOWN_ELEMENT_IN_CONFIG on unknown config options (disable with <skip_check_for_incorrect_settings>).
26.8 #113588 CHECK Fixed Compact-part mutations giving different results with serialization_info_version='basic' depending on whether serialization info was in memory or reloaded, which caused CHECKSUM_DOESNT_MATCH between ReplicatedMergeTree replicas.
24.10 #61473 INFO use_concurrency_control handling fixed, so concurrency control is actually enforced; in 26.8 concurrent_threads_soft_limit_ratio_to_cores also defaults to 2 (was 0 = unlimited in 24.8).
24.10 #70322 INFO The compatibility setting in the default profile now also sets default MergeTree settings at server startup.
24.11 #70977 INFO Replacing merge algorithm optimized for non-intersecting parts.
24.12 #70788 INFO join_algorithm default is now 'direct,parallel_hash,hash,ie_join' (was 'default'), so parallel_hash is used when applicable; with query_plan_join_swap_table=auto (E00017) the build side may also be swapped.
25.2 #75302 INFO format_alter_operations_with_parentheses now defaults to true, so ALTER command lists are formatted as (cmd), (cmd). This breaks replication only with servers older than 24.3.
25.2 #74772 INFO async_load_databases is now always on: the server accepts connections before all tables are loaded.
25.2 #75870 INFO The 24.12 switch of the default join to parallel_hash is now in settings history, so compatibility < 24.12 restores the plain hash join.
25.3 #69236 INFO New query condition cache (use_query_condition_cache, on by default in 26.8). It is extended to ORDER BY ... LIMIT (TopK) queries in 26.8 (E03034) and restored after lightweight DELETE (E03126).
25.4 #77976 INFO Attached parts of MergeTree tables are now attached in block order, which matters for ReplacingMergeTree.
25.4 #79080 INFO The query condition cache (use_query_condition_cache) is now enabled by default.
25.4 #78041 INFO parallel_distributed_insert_select now also applies to INSERT SELECT from ReplicatedMergeTree, and its default went 0 -> 2 (E01743, 25.7).
25.5 #79907 INFO compile_expressions (JIT compilation of expressions) is now enabled by default.
25.5 #79307 INFO ALTER ... DELETE mutations now create an empty part, without rewriting, for parts where every row matches.
25.5 #79838 INFO Sending crash reports is now enabled by default.
25.7 #83488 INFO Keeper now enables the create_if_not_exists, check_not_exists and remove_recursive feature flags by default, so servers can send new request types.
25.8 #84747 INFO concurrent_threads_scheduler default changed from round_robin to fair_round_robin (named max_min_fair in 26.8). The server-wide concurrency limit is also active by default in 26.8 (concurrent_threads_soft_limit_ratio_to_cores 0 -> 2).
25.10 #87414 INFO replicated_deduplication_window_seconds default lowered from 604800 (1 week) to 3600 (1 hour); E01246 (25.9) raises replicated_deduplication_window from 1000 to 10000 blocks.
25.10 #88515 INFO Keeper internal replication switches to async mode by default; upgrading from Keeper older than 23.9 needs an intermediate step or async_replication=0.
25.10 #88209 INFO Global sampling profiler enabled by default (global_profiler_real/cpu_time_period_ns 0 -> 10s for all server threads).
25.11 #89403 INFO ANY LEFT/RIGHT JOIN may be rewritten to ALL INNER JOIN (and, per E01275/25.9, to SEMI/ANTI JOIN) when the filter makes it equivalent.
25.11 #71334 INFO DDL ON CLUSTER queries are now executed on each host with the original query user's context (access checks on every replica) instead of the DDL worker's context.
25.12 #91830 INFO INSERT ... SELECT block deduplication reworked (insert_select_deduplicate, which in 26.8 is replaced by deduplicate_insert_select='enable_when_possible'): INSERT SELECT is deduplicated only when the SELECT is 'stable' (ORDER BY ALL, single stream) or an insert_deduplication_token is given.
26.1 #92951 INFO INSERT SELECT deduplication was reworked. In 26.8 insert_select_deduplicate is obsolete. deduplicate_insert_select = enable_when_possible deduplicates INSERT SELECT into Replicated* tables only when the SELECT is 'stable' (ORDER BY ALL, single stream) or when insert_deduplication_token is set.
26.1 #90677 INFO use_variant_as_common_type is enabled by default. if/multiIf/CASE, UNION and array literals with incompatible types now return Variant(...) instead of throwing NO_COMMON_TYPE.
26.1 #93407 INFO use_skip_indexes_on_data_read is now on by default: skip indexes are applied while data is read (streaming) rather than only during index analysis.
26.1 #94168 INFO When Keeper finds a broken snapshot or inconsistent changelogs, it now throws instead of aborting or cleaning up files automatically, so manual intervention is required.
26.2 #89314 INFO enable_join_runtime_filters is now on by default. Related correctness fixes: E06323 (LEFT ANTI JOIN with multiple keys), E06359 (ANY->SEMI/ANTI conversion with unmatched rows), E05888 (join reordering).
26.2 #95300 INFO concurrent_threads_scheduler default is now max_min_fair (was fair_round_robin).
26.3 #97680 INFO NOT precedence now follows the SQL standard: it binds looser than IS NULL/BETWEEN/LIKE/arithmetic.
26.4 #101494 INFO Fixed a large INSERT performance regression caused by deduplicate_insert='enable' (the default since 26.2): hashing is deferred to the sink and done in batches.
26.4 #102036 INFO Fixed wrong row ordering for queries with ORDER BY that use the grace_hash join algorithm (also E06233: fixed grace_hash wrong results for non-equi joins).
26.4 #100448 INFO min/max/argMin/argMax now always skip NaN (previously the result depended on where the NaN was).
26.4 #102900 INFO minmax_count_projection and trivial count() are no longer disabled for good after a lightweight DELETE once the masked parts have been merged away (same as E05926).
26.4 #100408 INFO Fixed a spurious TOO_MANY_ROWS for SELECT count() with max_rows_to_read when parts have non-aligned granule boundaries.
26.4 #102961 INFO Plain INSERTs without MVs now request max_insert_threads concurrency slots, not max_threads, which avoids CC slot starvation under high INSERT throughput.
26.5 #89334 INFO date_time_input_format and cast_string_to_date_time_mode defaults changed from basic to best_effort.
26.5 #104705 INFO New setting defer_partition_pruning_after_final (default 1) keeps the 26.3 behaviour of skipping partition pruning under FINAL. Setting it to 0 restores pre-26.3 pruning.
26.5 #103880 INFO Fixed the ANY OUTER JOIN to INNER/SEMI/ANTI conversion optimizers mis-handling filters with non-deterministic functions (now, today, rand), which could silently drop rows.
26.5 #104285 INFO max_bytes_ratio_before_external_join is now 0.5 by default (and E05644/26.4 added auto-spill of hash/parallel_hash joins to grace hash). Hash joins spill to disk once the right side exceeds half of the available memory.
26.5 #103148 INFO Fixed old clients failing with UNEXPECTED_PACKET_FROM_SERVER when inserting into a newer server via remote() or Distributed (the server sent Progress unconditionally at the end of the insert).
26.5 #105029 INFO Fixed silent MV data loss after EXCHANGE TABLES or CREATE OR REPLACE of a materialized view's source table. This was a regression from E05408 (26.5), which itself fixed view dependencies being lost on RENAME/EXCHANGE.
26.5 #104427 INFO DROP USER/ROLE/QUOTA/SETTINGS PROFILE now also removes references to the dropped entity from other access entities on disk, so dangling IDs are no longer resurrected at restart (they used to cause ACCESS_ENTITY_NOT_FOUND).
26.6 #98884 INFO Mutations (ALTER DELETE/UPDATE, lightweight DELETE) are now analyzed with the new analyzer.
26.6 #102928 INFO INSERT pipelines no longer reserve CPU slots up to max_threads at start (concurrent_threads_lazy_allocation=true), which fixes SELECTs slowing down during concurrent INSERTs under concurrency control.
26.6 #104939 INFO ALTER TABLE ... REPLACE PARTITION ... FROM an empty source partition is now rejected with BAD_ARGUMENTS unless allow_replace_partition_from_empty_source=1.
26.6 #108330 INFO use_skip_indexes_on_data_read (default true since 26.1) can now be reverted with compatibility, as an escape hatch for a performance regression where the on-data-read path defeats minmax/set/bloom_filter mark-range pruning.
26.6 #107481 INFO SELECT ... FINAL and OPTIMIZE FINAL no longer return duplicates after ATTACH PARTITION ... FROM (or CLONE AS / MOVE PARTITION) moves parts from a plain MergeTree into a Replacing/Summing/AggregatingMergeTree: the adopted part's merge level is reset to 0.
26.7 #107907 INFO DateTime64 range extended from [1900, 2299] to [0000, 9999]. Values that used to be clamped are now kept exactly, and saturation now goes to the new limits.
26.7 #108361 INFO Insert deduplication always uses the unified hash; the legacy modes are removed (the server refuses to start only if insert_deduplication_version is explicitly set to a legacy value).
26.7 #110726 INFO merge_selector_enable_heuristic_to_lower_max_parts_to_merge_at_once is now enabled by default: the merge selector lowers max parts per merge based on how full the partition is.
26.7 #109792 INFO max_execution_time with timeout_overflow_mode='throw' could sometimes never cancel a query; now it is enforced reliably. INSERT into MergeTree also now honours cancellation and max_execution_time while writing many parts (E04137).
26.7 #109086 INFO precise_float_parsing is now on by default and also applies to input formats and numeric literals (known item).
26.8 #106215 INFO read_in_order_use_virtual_row is now on by default. ORDER BY LIMIT n over many parts reads only the parts that can contribute.
26.8 #110320 INFO The native protocol sends String columns with a separate size stream, but only when both peers are at protocol revision 54489 or higher.
26.8 #113654 INFO Keeper durability fixes: Raft state file crash-safety (possible two leaders); also E03664 (snapshot dir fsync), E03728 (CORRUPTED_DATA after crash during log truncation), E03680 (follower dropping out of quorum after large snapshot), E03679 (empty Set refused over soft memory limit), E03755 (client held for session_timeout during leader change).
26.8 #115334 INFO SET = DEFAULT now respects settings constraints; previously 'SET readonly = DEFAULT' could leave readonly mode.
26.8.3.105 #111196 INFO Fixed potential data loss when an INSERT into ReplicatedMergeTree was cancelled or timed out while ClickHouse was recovering an uncertain Keeper commit.
26.8.10.6 #116724 INFO Outbound TLS connections (secure native/Distributed, secure Keeper, remoteSecure, s3/url) now verify the certificate hostname. openSSL.client.extendedVerification=false restores the old behaviour.
26.8.12.53 #120182 INFO Fixed wrong GROUP BY/DISTINCT/LIMIT BY/window results when the key is plus or minus of a partition column with a non-integer constant (e.g. d + INTERVAL 1 MONTH), because the allow_*_partitions_independently optimizations, now on by default, wrongly treated such a key as injective.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants