A new module can hang forever on Debian 13. It happens when the module is added right after another module was removed from the same node.
The removed module frees its Linux user ID (UID). The new module gets the same UID. The node agent then enables "linger" for the new user, but systemd never starts the user manager (user@<uid>.service). The module agent never starts, so the add-module task never completes.
Steps to reproduce
- Take a node running Debian 13 (systemd 257).
- Add a rootless module, for example
api-cli run add-module --data '{"image":"ghcr.io/nethserver/imapsync:1.3.3","node":3}'. It gets imapsync2 with UID 1004.
- Remove it:
api-cli run remove-module --data '{"module_id":"imapsync2","preserve_data":false}'.
- Add it again at once. The image is already in cache, so this is fast.
Expected behavior
The second add-module completes. imapsync3 runs.
Actual behavior
imapsync3 gets UID 1004 again. add-module never completes. user@1004.service stays inactive.
12:04:19 userdel: delete user 'imapsync2'
12:04:24 useradd: new user: name=imapsync3, UID=1004
12:04:26 agent@node: loginctl enable-linger imapsync3
12:04:26 systemd: Finished user-runtime-dir@1004.service
12:04:29 systemd: Stopped user-runtime-dir@1004.service
Starting the user manager by hand unblocks it: systemctl start user@1004.service. The agent starts and create-module completes.
It failed on the first try, on a real node. The same thing happens in the core test suite on Debian 13: the openldap suite removes openldap1, then samba1 gets its UID and hangs.
Rocky Linux 9 (systemd 252) does not show it in the same test suite.
Components
- NethServer core 3.22.0
- Debian GNU/Linux 13 (trixie), systemd 257.13-1~deb13u1
Possible fix
In core/imageroot/var/lib/nethserver/node/actions/add-module/50update, start user@<uid>.service explicitly after loginctl enable-linger. It does nothing when the service already runs.
A new module can hang forever on Debian 13. It happens when the module is added right after another module was removed from the same node.
The removed module frees its Linux user ID (UID). The new module gets the same UID. The node agent then enables "linger" for the new user, but systemd never starts the user manager (
user@<uid>.service). The module agent never starts, so theadd-moduletask never completes.Steps to reproduce
api-cli run add-module --data '{"image":"ghcr.io/nethserver/imapsync:1.3.3","node":3}'. It getsimapsync2with UID 1004.api-cli run remove-module --data '{"module_id":"imapsync2","preserve_data":false}'.Expected behavior
The second
add-modulecompletes.imapsync3runs.Actual behavior
imapsync3gets UID 1004 again.add-modulenever completes.user@1004.servicestays inactive.Starting the user manager by hand unblocks it:
systemctl start user@1004.service. The agent starts andcreate-modulecompletes.It failed on the first try, on a real node. The same thing happens in the core test suite on Debian 13: the openldap suite removes
openldap1, thensamba1gets its UID and hangs.Rocky Linux 9 (systemd 252) does not show it in the same test suite.
Components
Possible fix
In
core/imageroot/var/lib/nethserver/node/actions/add-module/50update, startuser@<uid>.serviceexplicitly afterloginctl enable-linger. It does nothing when the service already runs.