Skip to content

Move NS8 module tests from DigitalOcean to QEMU #8192

Description

@stephdl

Every NS8 module runs its automated tests on DigitalOcean droplets. We want to stop using DigitalOcean and run the same tests on a QEMU virtual machine inside the GitHub runner. The shared workflow test-module-qemu.yml in ns8-github-actions does this, and it is a drop-in replacement for the DigitalOcean one: same legs, same cost, same variables passed to the tests.

The move is also a chance to check which modules really test the upgrade. Each run has two legs: install tests the new image on a clean node, and update installs the last stable release first. What each leg checks is decided by the Robot Framework suite of the module, through the ${SCENARIO} variable. A suite that ignores it gets no upgrade test: the update leg is dropped and both distros run install. Upgrade tests have to be written by hand, module by module.

Proposed solution

For each repository, open one pull request following what was done in ns8-crowdsec, ns8-sogo and ns8-webserver:

  • edit the existing .github/workflows/test-module.yml, no new file: point uses: to NethServer/ns8-github-actions/.github/workflows/test-module-qemu.yml@v1, drop the secrets: block, add the permissions: block and the run-name: line
  • if the repository calls test-on-digitalocean-infra.yml directly, replace the module-info job and the scenario matrix with that single uses:
  • delete the local test-module.sh if there is one: the shared script is used instead

No other input is needed. The documentation is in docs/test-module-qemu.md.

Separately, when the upgrade test is missing, add the IF '${SCENARIO}' == 'update' branch in tests/*.robot, as done in ns8-webserver and ns8-piler. It can be the same pull request or a later one.

Prerequisite: NethServer/ns8-github-actions#67, which makes the QEMU workflow a drop-in replacement.

Runner cost

Each leg (one distro and one scenario) takes one GitHub runner for the whole test. A run has one leg per distro, so 2 runners with Rocky Linux 9 and Debian 13, the same as DigitalOcean today. When the suite reads ${SCENARIO}, one distro runs install and the other runs update, and the distro of each scenario changes with the run number. When it does not, both distros run install, and the run shows a notice saying the update leg was dropped.

Checklist

Each repository has two boxes: the first one for the switch to QEMU, the second one for the upgrade test in its Robot Framework suite.

Already on QEMU:

Using the shared DigitalOcean test-module.yml, only uses: changes:

Calling DigitalOcean directly, the jobs are replaced and the local test-module.sh is deleted:

Special cases:

  • NethServer/ns8-core: tests the core itself, not a module, through its own tests.yml. Work in progress in ci(tests): run the core tests on QEMU ns8-core#1328, which needs feat(test-on-qemu): test ns8-core itself ns8-github-actions#62
  • nethesis/ns8-nethvoice: on hold, nothing changed for now. Its workflow runs two DigitalOcean chains: the shared test-module.yml, and a FIAS end-to-end test that calls test-on-digitalocean-infra.yml directly with RUN_FIAS_E2E, its own scenario matrix and an optional fias_image, so four droplets per run. The QEMU workflow cannot pass RUN_FIAS_E2E to the suite, and the DigitalOcean runs have all been cancelled since August 2026, so there is no green baseline to compare with. To decide with the maintainer: whether FIAS stays in every run or only on demand, how to pass it, and why the current runs are cancelled.
    • upgrade test
  • NethServer/ns8-grafana (ci: run module tests on QEMU ns8-grafana#114): its workflow called test-on-ubuntu-runner.yml, which no longer exists, and was disabled
    • upgrade test
  • NethServer/ns8-hermes-agent: has its own DigitalOcean workflows and does not use the shared ones
    • upgrade test

When all repositories are done:

  • disable the DigitalOcean workflows in ns8-github-actions (test-on-digitalocean-infra.yml and test-module.yml)
  • remove the do_token secrets, at the very end

Alternative solutions

Running two virtual machines on one runner, one for install and one for update, does not fit: a standard runner has 16 GB of memory for a public repository, and each virtual machine needs 8 GB. Running every scenario on every distro on each run doubles the cost, and no module needs it today: a module that does can call the workflow once per scenario.

Additional context

Modules published in NethForge (collabora, dependencytrack, dokuwiki, lamp, mariadb, matrix, n8n, passbolt, postgresql, prometheus, sogo, webserver, wordpress, grafana, hermes-agent) are not found by add-module <name> on a new node, where NethForge is disabled. An upgrade test for them enables it first with cluster/alter-repository, as done in NethServer/ns8-webserver#228.

Modules that ship with the core need nothing special: the QEMU workflow gives the image under test to the core installer on the install leg only. This was checked on ns8-traefik and ns8-metrics.

The QEMU virtual machine has no public address, unlike a droplet. A test that needs a real Let's Encrypt certificate or a connection from the Internet cannot run there. No module suite seems to need it today; ns8-core excludes its letsencrypt tests.

The repositories without automated tests are not listed: ns8-docs, ns8-images, ns8-nethforge, ns8-repomd, ns8-rocky-iso, ns8-terraform-infra, ns8-ui-lib, ns8-user-manager. ns8-terraform-infra holds the DigitalOcean setup and stays as it is for now. Other private ns8- repositories under nethesis may need the same work later.

See also

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

Relationships

None yet

Development

No branches or pull requests

Issue actions