Skip to content

linux: stop the guest sleeping the host, and close the cwd fd - #113

Merged
fwsGonzo merged 2 commits into
masterfrom
fuzz-clock-nanosleep-and-cwd-fd-leak
Aug 21, 2026
Merged

fwsGonzo merged 2 commits into
masterfrom
fuzz-clock-nanosleep-and-cwd-fd-leak

Conversation

@perbu

@perbu perbu commented Aug 2, 2026

Copy link
Copy Markdown
Collaborator

Two bugs found by fuzz/syscall_fuzz.cpp on a 10h ARM64 campaign. Both are in shared code, not ARM64-specific.

clock_nanosleep let the guest park the VMM thread. The guest's timespec and flags went straight to the host clock_nanosleep(), so a large tv_sec — or TIMER_ABSTIME with a far-future deadline — blocked indefinitely. SYS_nanosleep was already a no-op to prevent exactly this, but modern glibc routes nanosleep() through clock_nanosleep, so that no-op covered almost nothing. Now validates like Linux (EINVAL on bad clockid, unknown flags, out-of-range tv_nsec, negative tv_sec) and returns without sleeping. The old error branch was dead code: clock_nanosleep(3) returns errno positively and never sets errno, so failures were reported to the guest as success.

set_current_working_directory() leaked a directory fd per call. The fd is kept outside m_fds, so neither reset_to() nor ~FileDescriptors() reached it, and re-setting dropped the previous one. An embedder re-applying its policy per warm fork leaks one fd per request. Each leaked fd also pins ~24 KiB of unreclaimable kernel slab on a 16 KiB-page host.

The fd leak is what wrecked the fuzzing campaign: workers reached ~100k fds in minutes, SUnreclaim hit 4.4 GB and MemAvailable fell to 83 MB, after which children died during VM setup and were tallied as 93 artifact-less "crashes". Killing the workers returned 4.2 GB instantly.

Verification

  • The three artifacts that used to hang for 60+ seconds now replay in <1 ms.
  • Worker fd count flat at 13 over 260k executions; SUnreclaim back to ~235 MB.
  • A 5h43m re-run after both fixes: 183M executions, 0 crashes, 0 timeouts, 0 OOMs (same fuzzer and corpus that produced 93 crashes before).
  • All 6 ARM64 unit suites pass.

Adds a regression test to tests/unit/syscalls.cpp. Note that file is only registered for AMD64 in tests/unit/CMakeLists.txt, so it runs in AMD64 CI; on ARM64 I verified the fix with a standalone guest instead.

perbu and others added 2 commits August 2, 2026 21:58
Both found by fuzz/syscall_fuzz.cpp on a 10h ARM64 campaign.

clock_nanosleep passed a guest-controlled timespec and flags straight to
the host clock_nanosleep(), so a large tv_sec -- or TIMER_ABSTIME with a
far-future deadline -- blocked the VMM thread indefinitely. SYS_nanosleep
was already a no-op to prevent exactly this, but modern glibc routes
nanosleep() through clock_nanosleep, so that no-op covered almost nothing.
Validate as Linux does and return without sleeping. The old error branch
was also dead: clock_nanosleep(3) returns errno positively and never sets
errno, so every failure was reported to the guest as success.

set_current_working_directory() opens the directory but keeps the fd
outside m_fds, where neither reset_to() nor ~FileDescriptors() reached it,
and re-calling the setter dropped the previous one. An embedder that
re-applies its policy per warm fork leaked an fd per request; each leaked
fd also pins ~24 KiB of unreclaimable kernel slab on a 16 KiB-page host,
which is enough to exhaust a machine overnight.
# Conflicts:
#	lib/tinykvm/linux/fds.cpp
#	tests/unit/syscalls.cpp
@fwsGonzo

Copy link
Copy Markdown
Member

Merging with two adjustments made while resolving the conflict against master:

  • Dropped the fds.cpp hunk. The cwd-fd leak was fixed on master in the meantime by replace_working_directory_fd(), which covers the destructor, reset_to() and re-setting the policy. That was the conflict; the file now matches master. The analysis in the description still stands, it just landed by another route.
  • Added threads().suspend_and_yield() to the success path. SYS_nanosleep gained it on master for the cooperative-scheduling reason (a returns-immediately sleep turns any poll-with-backoff loop into a livelock; Go's sysmon is exactly that). Since — as this PR points out — glibc routes nanosleep() through clock_nanosleep, this is where that yield actually matters, and without it the fix here would have regressed the Go case. It is a no-op when no other thread is runnable.

The clock_nanosleep validation and the regression test are unchanged. Verified locally: full unit suite passes, test_syscalls included and no hang. (test_elf fails, but identically on pristine master — a missing rust.elf fixture, unrelated.)

@fwsGonzo
fwsGonzo merged commit 2b5d772 into master Aug 21, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants