Skip to content

feat: nightly load test - #574

Open
kollegian wants to merge 5 commits into
mainfrom
feat/nightly-load-regression
Open

kollegian wants to merge 5 commits into
mainfrom
feat/nightly-load-regression

Conversation

@kollegian

Copy link
Copy Markdown

No description provided.

@cursor

cursor Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

PR Summary

Medium Risk
Changes nightly CI topology (two chains, new env, genesis-funded keys) and crypto/key wiring for load tests; production controller paths are mostly test harness and keygen utilities.

Overview
Nightly benchmark now runs two isolated chains per schedule: saturation (mock_balances, configurable EVM transfer profile) and steady-state (vanilla SEID_IMAGE, nightly_steady_state, 30 minutes, real balances). Each subtest gets its own chain id (<base>-saturation / <base>-steady-state), workload label for metrics/alerts, and requires SEID_IMAGE in addition to the existing mock image env.

EVM funding for steady-state: adds keygen.DeriveEVM() (hex private key, 0x EVM address, sei bech32 “cast” of the same 20 bytes) and reuses a shared bech32Address helper from cosmos Derive. The steady-state run genesis-funds the cast address and mounts the hex root key into the sei-load Job for profile dispersal.

sei-load harness: load is sent only to the first RPC follower; receipts and the inclusion gate use the second. Profile rendering adds __RECEIPT_ENDPOINT__. The Job template stamps sei.io/seiload-workload, optionally mounts a root-key Secret (fsGroup + 0400), and validates workload as a K8s label value. Job create/wait is centralized in runJob; mnemonic secrets use generic createKeySecret.

Reviewed by Cursor Bugbot for commit 0937db2. Bugbot is set up for automated code reviews on this repo. Configure here.

@seidroid seidroid Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This PR adds TestNightlyLoadRegression, which funds a throwaway EVM root key at genesis, runs a 200 TPS sei-load Job against a fresh 4-validator chain, and gates included TPS, block-interval p50/p90 and the validators' CPU and memory against a recorded baseline; it also pulls the profile-fill, Job-run, block-page and key-Secret helpers out for sharing. Nothing blocks, so this is an approve with two non-blocking notes; I dropped codex's retry-pod finding because the Job template sets backoffLimit: 0 with restartPolicy: Never, so the Job has only one seiload pod and seiloadRunTimes cannot pick a failed earlier attempt.

Non-blocking

1 finding on the changed lines, as inline comments.

  • The gate logic that decides nightly pass/fail (Measure's window-coverage and edge handling, nearest-rank percentile, Check's direction and max(pct, abs) allowance, and parseBaseline's rejection rules) has no unit tests, and all of it is pure functions. A sign or allowance bug would only show up as a spurious or missed REGRESSION in the nightly, so a small table test in the package would catch it at PR time instead.
1 nit, not posted on the code
  • test/integration/loadregression_test.go:158 — time.Sleep here ignores ctx, so a SIGTERM during the wait isn't handled until the sleep ends. The wait is at most about 30s because finished >= to has already been checked.

seidroid review · decision approve · session 2aab5e4ee3f949a5afe79adca1434e1e · turn resp_claude_fd6d583295e5e45ce233f683c52f645f · item b7ce020afa5550ada791e035cb83e9f5

Findings: 0 blocking | 2 non-blocking | 1 posted inline

// change to one is reviewed with the others. TestNightlyLoadRegression in
// test/integration runs it.
//
// To change the workload, edit profile.json (sei-load's profile schema) and

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion — The package doc says a mismatch between profile.json and baseline.json's config block "fails this package's tests", but test/integration/loadregression has no _test.go files. Today a mismatch only shows up when the nightly starts and fails the Comparable check as UNEVALUABLE, not at PR time. Either add the unit test that calls RunConfig/RecordedBaseline/Comparable, or change the doc.

@seidroid seidroid Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

review found nothing blocking.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want higher recall? High effort reviews run extra passes and find more bugs. A team admin can switch effort levels in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 7544b84. Configure here.

Comment thread test/integration/loadregression_test.go Outdated

@bdchatham bdchatham left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for putting this together, the fail-closed handling and the split between UNEVALUABLE and REGRESSION are really thoughtful. Most of my comments come down to one direction, where the harness runs the load and reports whether the run was valid, and alerts become the validation layer for performance.

On the PR body, could you link a companion platform PR in the description? I'm thinking new rules for committed TPS, block interval p50/p90 and validator CPU/memory against a 7-day median, plus scoping the existing benchmark alerts to workload!="load-regression".

Comment thread internal/keygen/evm.go
return evmIdentityFromKey(priv)
}

func evmIdentityFromKey(priv *btcec.PrivateKey) (EVMIdentity, error) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we add a fixed-key vector test for evmIdentityFromKey, the way hd_test.go pins Derive? A wrong address here would only show up as an unfunded root on a live chain.

Comment thread internal/keygen/evm.go Outdated
h.Write(uncompressed[1:])
addr := h.Sum(nil)[12:]

converted, err := bech32.ConvertBits(addr, 8, 5, true)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The ConvertBits and bech32 Encode steps mirror keygen.go:77, so I'd pull them into one helper and keep the address format defined in a single place.

// image (included TPS 0.003%, p90 1.6%, CPU 1.2%, memory 1.8%). p50 is the
// exception: it moved 7.1% between those runs, so it keeps a 10% allowance.
// Revisit them once the baseline holds a week of nightlies.
package loadregression

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the first subpackage under test/integration, and our shared harness code lives in harness/ today, so I'd lean towards harness/loadregression or folding it into the test package.

// recorded with and what the tracker needs to be trustworthy.
func RenderProfile(chainID, sendEVM, receiptEVM string) string {
return strings.ReplaceAll(bench.FillProfile(profileTmpl, chainID, []string{sendEVM}),
"__RECEIPT_ENDPOINT__", receiptEVM)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adding __RECEIPT_ENDPOINT__ to bench.FillProfile would keep every profile placeholder in one replacer, rather than a second ReplaceAll layered on top.

Comment thread test/integration/loadregression_test.go Outdated
hc := &http.Client{Timeout: 10 * time.Second}
// rpc-0 takes the sends; rpc-1 takes none and is where receipts, heads and
// the measured blocks are read.
sendNode, receiptNode := ch.rpcNodes[0], ch.rpcNodes[1]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ch.rpcNodes[0] and [1] would panic if RPCNodes ever drops below 2, so a len check that fails as UNEVALUABLE keeps that failure readable.

Comment thread test/integration/loadregression_test.go Outdated
}
t.Logf("load-regression result: %s", result)

v := loadregression.Check(baseline, m.Metrics)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd like the test to fail only when the run itself isn't valid data (halt or lag, a short window, missing coverage, a high revert ratio), and let alerts own the regression call.

Comment thread test/integration/loadregression/gate.go Outdated
@@ -0,0 +1,260 @@
package loadregression

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With alerts owning the comparison, I think the baseline, threshold and Check logic plus resources.go can go, since PrometheusRules can compare each night against the trailing nightlies without a hand-maintained baseline file.

Comment thread go.mod Outdated
github.com/pelletier/go-toml/v2 v2.2.4 // indirect
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 // indirect
github.com/prometheus/client_golang v1.23.2 // indirect
github.com/prometheus/client_golang v1.23.2

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Once resources.go is gone, client_golang and prometheus/common can drop back to indirect.

Comment thread test/integration/loadregression_test.go Outdated
t.Fatalf("UNEVALUABLE: %v", err)
}

promURL := mustEnv(t, "PROMETHEUS_URL")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dropping PROMETHEUS_URL also means the harness needs no Prometheus access, which is one less piece of CronJob wiring to land in platform.

Comment thread test/integration/loadregression_test.go Outdated
Image: s.seiloadImage,
DurationMinutes: s.durationMin,
ProfileCM: profileCM,
Workload: "load-regression",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's keep Workload: "load-regression", since it's the label the new alerts scope on and the one the existing benchmark alerts will need to exclude.

@bdchatham

Copy link
Copy Markdown
Collaborator

@kollegian one more thought after sitting with this a bit longer. With the regression call moving to alerts, this suite and TestNightlyBenchmark end up differing mostly in configuration, since both already run 4 validators and 2 RPC followers with the same EVM tuning through provision and runSeiload. So I'd lean towards folding this work into the benchmark rather than landing a second suite.

The two runs still answer different questions, so I'd keep both workloads. The benchmark measures saturation throughput, and this run measures what a fixed 200 TPS load costs the chain, which is where CPU, memory and block interval become comparable from night to night.

A shape I think would work:

  1. TestNightlyBenchmark becomes a table of workloads run as subtests (saturation and steady-state), each on its own chain with its own Workload label.
  2. The root-key funding moves into runSeiload as an option, so any workload can run on a real-balance chain.
  3. Receipts turn on for both workloads, read from the follower that takes no sends, which also closes the gap where the benchmark only measures mempool admission today.
  4. The test keeps the validity checks, and the PrometheusRules judge each workload by its label, with the existing < 500 TPS alert scoped to workload="saturation".

One thing I haven't checked is whether the AMM and ERC20 scenarios run on the mock_balances image. If they do, both workloads could share one image, and if not the steady-state workload stays on vanilla. The helper extractions in this PR already lay most of the groundwork, so I think this fits naturally as the next revision here.

@bdchatham bdchatham left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kollegian

Copy link
Copy Markdown
Author

@kollegian
kollegian requested a review from bdchatham October 1, 2026 20:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants