Find flaky tests before CI does.
Most flaky tests pass on a fast, idle laptop and fail on a slow, shared CI runner: a goroutine or thread is scheduled late, a fixed sleep or timeout expires first, or two pieces of work race. shakedown recreates those conditions locally. It runs your test suite many times in parallel, each run in a fresh process, while keeping most CPU cores busy. Then it tells you which tests failed only some of the time.
$ shakedown go -runs 20 ./internal/...
shakedown go: 20 runs, parallel 4, load 14, output /tmp/shakedown-1234
[ 1/20] run-002 pass 4.1s
[ 2/20] run-001 FAIL 4.3s TestEmailReloadFencesPreviousAccount
...
shakedown go: 20 runs in 31s (parallel 4, load 14)
19 passed, 1 failed, 0 timed out (812 tests seen)
FLAKY: passed in some runs, failed in others
1/20 5% example.com/app/internal/service TestEmailReloadFencesPreviousAccount
first failure (run-001): /tmp/shakedown-1234/failures/...TestEmailReloadFencesPreviousAccount.log
go install github.com/dubee/shakedown@latest
shakedown <adapter> [flags] [runner arguments]
| Adapter | Runs | Per-test results from |
|---|---|---|
go |
go test -json -count=1 (default ./...) |
go test -json events |
pytest |
pytest --junitxml=... |
JUnit XML |
jest |
npx jest --json --outputFile=... |
Jest JSON report |
junit |
any command; put {junit} (a file) or {junitdir} (a directory) where it should write JUnit XML |
JUnit XML |
cmd |
any command | exit status only |
Examples:
shakedown go -runs 50 ./internal/...
shakedown go -race -shuffle -runs 5 ./...
shakedown pytest -runs 30 -run test_checkout tests/
shakedown jest -runs 20 -cmd "yarn jest" src/
shakedown junit -runs 10 -- mvn -q test -Dsurefire.reportsDirectory={junitdir}
shakedown junit -runs 10 -- cargo nextest run --profile ci # with JUnit output configured to {junit}
shakedown cmd -runs 25 -- make integration-test
Flags go before the runner's arguments. Use -- when the runner's arguments start with a dash.
| Flag | Default | Meaning |
|---|---|---|
-runs |
10 | Total runs. |
-parallel |
4 | Runs executing at the same time. |
-load |
auto | CPU cores to keep busy. Auto is all cores but 4, or 0 with -race. 0 turns load off. |
-timeout |
10m | Kill a run that takes longer and report it as hung. |
-race |
off | Race detector (go). |
-shuffle |
off | Random test order each run (go, pytest with pytest-randomly, jest). |
-run |
Only matching tests (go -run, pytest -k, jest -t). |
|
-cmd |
Replace the runner command, e.g. "poetry run pytest". |
|
-out |
temp dir | Where run logs and failure output go. |
-json |
Also write the full summary, including every run, as JSON. | |
-dir |
cwd | Working directory for each run. |
-failfast |
off | Stop starting new runs after the first failure. |
-quiet |
off | No line per finished run. |
Exit status: 0 all runs passed, 1 at least one failure, 2 bad arguments or setup error.
- Fresh process per run. Every run is a new process, so no state carries over between
runs the way it can with in-process repeats like
go test -count=N. - CPU load. shakedown starts busy-loop threads, one per core of load, so the tests get only the cores that are left, like on a busy CI runner.
- Process isolation. Each run gets its own process group. When a run exits but processes it started are still running, shakedown reports it ("left processes running") and kills them. Leftover processes such as browsers or servers are a common source of flakes in later tests. (Not available on Windows.)
- Hang detection. A run over
-timeoutis killed; tests that started but never finished are reported as hung. - Grouping. Results are grouped by test across runs, split into flaky (failed sometimes) and always failing, with the first failure's output saved to a file.
Each run gets SHAKEDOWN_RUN (1, 2, ...) and SHAKEDOWN_RUN_DIR in its environment. Runs
share the working directory, so tests that write fixed files, ports, or databases may
collide when run in parallel; use these variables to separate them, or lower -parallel.
- Start with
-runs 20under load. A test that fails even once is timing-sensitive and will eventually fail in CI. - Run
-race -shuffleseparately. It finds data races and tests that depend on order or on state another test left behind. Tests with wall-clock budgets may fail under-raceonly because it slows code 2 to 10 times. - Once you have a fix, hammer just that test:
shakedown go -runs 100 -run '^TestFoo$' ./pkg.