Repository navigation
Flaky test_flight_animation_export_gif: intermittent bus error in VTK off-screen rendering #1078
Description
Activity
@thc1006 which branch and commit were you working at?
Thanks for picking this up, @Gui-FernandesBR.
I was on my PR #1054, branch
enh/reproducible-montecarlo-seeding, at commitda2ba5cf. The useful part is that it is not specific to my branch: the same crash happens on develop. The Tests run on developd31822d3failed the same way.Both crash in the
Pytest (macos-latest, 3.14)job, duringpytest tests/integration:Fatal Python error: Bus error File ".../tests/integration/test_plots.py", line 79 in test_flight_animation_export_gif ... Bus error: 10 pytest tests/integration --cov=rocketpy --cov-append- develop
d31822d3: run 29697106388, thePytest (macos-latest, 3.14)job - my ENH: reproducible Monte Carlo via per-simulation-index seeding #1054
da2ba5cf: run 29728954730, the same job
A few things point to a flake rather than a real failure:
- The other matrix jobs in each run show
The operation was canceled, not a test failure. That is fail-fast cancelling the siblings once the macOS job dies (it leaves orphanpytestandXvfbprocesses behind), so a run can look like every job failed when only macos-3.14 actually crashed. - It is intermittent on the same code: an earlier commit of mine,
e5ba9404, passedPytest (macos-latest, 3.14), andda2ba5cfcrashed it. - Same rendering stack in the passing and crashing runs: vtk 9.6.2, pyvista 0.48.4, matplotlib 3.11.1, pillow 12.3.0, so it is not a version bump.
- My change only touches Monte Carlo seeding, nothing in the plotting path.
So it reads as a native crash in the VTK off-screen GIF export on the macOS runner. Happy to send a small PR marking the animation tests with a rerun (e.g.
pytest-rerunfailures) if that is the direction you would prefer.- develop
Picking this up (thanks @thc1006). One wrinkle worth flagging before choosing the approach:
- rerun alone won't catch it. A
Bus erroris a native SIGBUS that kills the interpreter, sopytest-rerunfailureshas nothing left to rerun — the process is already gone. - fork isolation is out on this matrix.
pytest-forkedwould contain the crash, but it needsos.fork, and the test matrix includeswindows-latest.
So I'd go at the root cause instead. The crash is VTK's off-screen OpenGL on the headless Linux runner (the
setup-headless-display-actiongives a display, but the GL path is still the fragile bit). Forcing Mesa software rendering removes exactly this class of intermittent bus error and keeps the tests running (no coverage loss, no new deps, no per-test markers):env: LIBGL_ALWAYS_SOFTWARE: "1" GALLIUM_DRIVER: llvmpipe
It's a no-op on the macOS/Windows legs (no Mesa), so it's matrix-safe. I'll open a small PR adding that to
test_pytest.yamlandtest-pytest-slow.yaml. If it turns out to still flake after that, the clean fallback is a rerun layer for residual soft failures — but software rendering should take out the hard crash. Happy to go a different direction (separate CI step, skip) if a maintainer prefers.- rerun alone won't catch it. A
I have been hitting this repeatedly over the last two days, so I went and counted rather than guessing. Two things came out that change the picture, @wuisabel-gif, and one of them cuts against what I wrote when I opened this.
It is not Linux, and it is not macOS. It is 3.14.
Failing
Pytestjobs across the last 40 completedTestsruns:macos-latest, 3.14 9 ubuntu-latest, 3.14 5 everything else 0Both 3.10 jobs are clean on every platform, and Windows is clean on both versions. The ubuntu 3.14 crash is the same stack as the macOS one, frame for frame:
File ".../pyvista/plotting/plotter.py", line 2239 in render File ".../pyvista/plotting/plotter.py", line 5690 in write_frame File ".../rocketpy/plots/flight_plots.py", line 1079 in _run_animation File ".../tests/integration/test_plots.py", line 79 in test_flight_animation_export_gifSo the headless GL setup on the Linux runner is a real suspect, but it cannot be the whole story: the same crash happens on macOS, which does not go through that path, and neither platform crashes on 3.10 with the same rendering stack. Whatever this is, Python 3.14 is the variable both sides share. Forcing Mesa would presumably fix the Linux half and leave the macOS half where it is.
I should also correct the title I gave this issue. It is not always a bus error:
exit 139 (SIGSEGV) 8 exit 138 (SIGBUS) 2 exit 1 4Same test either way, so the first two are one fault showing up as two signals rather than two problems. The four ordinary failures are unrelated to this and I have not looked at them.
The cancellation is separate, and it is a one-line fix.
test_pytest.yamlsets nofail-fast, so the default applies and the other five jobs are cancelled the moment this one dies. A contributor sees six red jobs and has to open each to find out that five of them never finished. My #1098 has now been through three runs: every one of them died here, and every one took five healthy jobs down with it.fail-fast: falseon that matrix would not fix the crash, but it would stop one flaky test from deciding the result of the whole matrix, and it is worth having regardless of how the VTK side is resolved.Happy to send that as a small PR now, separate from whatever you land for the crash itself, if that suits.
Sent the cancellation half as #1100, separate from the crash so it does not get in your way, @wuisabel-gif.
It is
fail-fast: falseon the two multi-job matrices and nothing else.linters.ymlanddocs.ymlalso have matrices but both are single entry, so there is no sibling to cancel and the setting would do nothing there.The cost is runner time, and I put the numbers in the PR rather than waving at it: a full successful matrix is about 52 minutes across the six jobs, and a failing run currently stops short of that. If the extra minutes are not worth it, the narrower version is to set it on the main matrix only.
Nothing in it touches the VTK crash, which is still the actual bug.
#1100 is in, so the exit-code half of this is closed: the retry now matches 135, 138 and 139 rather than 138 alone, which was letting eight of the ten crashes I counted straight through.
The crash itself is still yours, @wuisabel-gif, and #1084 seems to have done it: #1054 was rebased onto it and its macOS 3.14 job has gone green twice since, on a leg that had been crashing all week.
Reacted by Gui and Isabel WuThe two mitigation commits are both present in current
develop(cb6106a717207dd8fc2dfe1446d80ff75022f21b):f40f18e3: isolates the GIF export and retries native exits 135, 138, and 139;3656d2a4: setsfail-fast: falseon both multi-job test matrices.
As a post-merge check, both full test runs for #1144 completed successfully on Python 3.14 for Ubuntu and macOS (
31739388805and31740477205), along with the remaining matrix jobs. That does not prove a native flake can never recur, but the failure and cancellation behaviors tracked here now have dedicated handling ondevelop.Unless another occurrence has been observed after these commits, this issue looks ready to close and reopen with a new run/job link if the isolated retry is exhausted.
Reacted by Gui- linked a pull request that will close this issueMNT: do not let one Python version cancel the other in the slow matrix #1100
on Aug 14, 2026 - linked a pull request that will close this issueBUG: fly the parachute the simulation sampled #1098
on Aug 14, 2026
What happens
tests/integration/test_plots.py::test_flight_animation_export_gif, and the other VTK/PyVista off-screen animation tests near it, sometimes crash the whole pytest process:It is intermittent. The same commit and the same environment can pass one run and crash the next. When it does crash, the interpreter dies rather than a test failing cleanly, so the coverage upload is skipped and the whole matrix goes red.
Why it looks like a flake rather than a code regression
Possible directions
pytest-rerunfailures(@pytest.mark.flaky(reruns=2)), so an occasional OpenGL hiccup does not take down the whole matrix.I ran into this while working on #1054, which is a seeding change that does not touch plotting. Happy to send a small PR marking the animation tests with a rerun if that is the approach you would prefer.