Skip to content

Speed up port exclusions and avoid redundant port copies - #954

Merged
bee-san merged 8 commits into
masterfrom
perf/port-exclusion-filter
Oct 2, 2026
Merged

bee-san merged 8 commits into
masterfrom
perf/port-exclusion-filter

Conversation

@bee-san

@bee-san bee-san commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Preparing a scan currently clones the ordered port list, linearly searches every excluded port for each candidate, then allocates a second vector. This becomes quadratic for large exclusion lists. Reuse the ordered vector, filter small lists in place, and index larger lists with a bitmap covering the complete u16 port space.

Closes #948.

Behavior and implementation

  • No exclusions: return the existing vector immediately.
  • Fewer than 64 candidate ports: retain with linear membership checks, avoiding bitmap setup.
  • Larger lists: build an 8 KiB bitmap and retain in place, reducing preparation from ports × exclusions to ports + exclusions.
  • Keep the bitmap helper out of line: release assembly shows the async poll frame drops from 8,424 bytes to 344 bytes, avoiding bitmap stack setup on every wake.
  • Preserve generated serial/random order, duplicate ports, 0/65535 boundaries, scanner intervals and all probe scheduling/result semantics.

Simplification: one private helper replaces the existing filter/collect expression. No dependencies, public APIs, runtime, socket, timeout or retry changes. A measured slowdown in an exploratory short/dense case was fixed before the final comparison.

Matched production benchmark

Compared merged Tokio master b7119f7f3eb1601e81126ef3572699ed7b3b7c0b with implementation 270561d553688acbd99f61c460cdfa7c76c1423e. Linux 7.2.6-1-cachyos, Intel Core Ultra 7 165U, Rust 1.98.1, release/LTO, identical lockfile and benchmark harness. For each of twelve cases: two independent 20-sample Criterion runs per build, A/B then B/A, pinned CPU 0, 500 ms warmup/2 s target measurement, no concurrent builds. Both repetitions are retained because CPU frequency is not fixed.

This calls the production Scanner::run_with_status with zero target addresses: it measures port preparation plus empty scan setup, not full network scan latency. It opens no sockets, performs no DNS and runs no scripts. Runtime and Scanner construction are outside the timer.

Input Master A / B Optimized A / B
65,535 ports, no exclusions 166.11 / 115.69 µs 3.65 / 2.33 µs
65,535 ports, 4 exclusions 175.72 / 181.34 µs 29.77 / 30.23 µs
65,535 ports, 4,096 exclusions 8.811 / 8.470 ms 59.85 / 63.21 µs
64 ports, 4,096 exclusions 6.23 / 10.09 µs 2.48 / 3.04 µs

All twelve cases improved in both repetitions, including the one-port and 16-port cases. Default scans save only a fraction of a millisecond; socket I/O/timeouts usually dominate total scan time. The large exclusion case is 134–147× faster in this preparation measurement.

Reproduction, every case, raw samples, confidence intervals and binary hashes.

Validation

  • cargo test --locked: 74 passed, zero failed; one pre-existing ignored doctest. One focused regression compares an independent membership oracle against empty lists, endpoints, duplicates, all ports and frozen random order.
  • cargo clippy --locked --all-targets -- --deny warnings: passed, including the benchmark.
  • cargo fmt --check: passed.
  • cargo doc --locked --workspace --all-features --no-deps --document-private-items: passed.
  • cargo bench --locked --bench benchmark_helpers --no-run: passed for baseline and candidate; the saved executables completed all 48 final case/build/repetition runs.
  • Network-free Python harness assertions: exclusion arguments, attempted socket counts and expected listeners passed for Linux/macOS/Windows. python3 -m py_compile .github/scan-benchmark/scan_bench.py: passed.
  • git diff --check and complete branch self-review: passed.

The CI-only real-scan harness adds a sweep excluding 1,024 ports, reusing its existing listeners. It checks excluded results and counts only attempted sockets for throughput. Full hosted tests and matched release TCP/UDP comparisons passed as detailed below.

Hosted validation

Linux x64/ARM64 and Windows scan benchmarks and all four test jobs passed for implementation 270561d. macOS diagnostic scans caught No buffer space available (ENOBUFS, error 55) on both master and candidate, so those failed datasets cannot support speed claims. A prior macOS run found every listener but failed the existing performance threshold; inspecting release assembly then motivated keeping the bitmap out of the recurring async poll frame.

The workflow now provisions the disposable macOS runner's TCP accounting budget once, from its default 1/32 to 1/8 of physical memory when needed, applying identical conditions before either build is timed. It prints the actual before/after limit and allocation counters. Every workload, sample count, listener assertion and acceptance threshold is retained. Immediate untimed diagnostics record TCP states, memory accounting and errors on any failed pair. The interface and default budget are documented in Apple's XNU TCP initialization and memory accounting interface. This changes CI configuration only; it does not change user machines or scanner socket behavior.

Final head cd6a23387f898a771577aebabbe5db906c2cde98:

  • Test matrix: all four jobs completed successfully (Linux x64/ARM64, macOS ARM64, Windows x64).
  • Matched scan benchmarks: all four jobs completed successfully; 792 measured invocations across both builds, nine interleaved repetitions per scenario and 27 for one open port. All expected TCP/UDP listeners found, no result mismatches, no subprocess failures and no diagnostic scans triggered. Existing throughput/wall-time/RSS gates passed on every platform.
  • macOS reported an initial TCP budget of 234,881,024 bytes and configured 939,524,096 bytes, before timing either build. The exclusion-sweep median wall time was 1.359 s for master versus 1.106 s for the candidate on this runner. This individual hosted result does not imply a comparable speedup for all scans: default macOS scans were 9.3% slower and UDP was 0.4% slower, both within the unchanged gate. Linux x64's exclusion sweep was essentially unchanged (0.795 versus 0.796 s); Linux ARM64 was 0.439 versus 0.430 s.
  • Python syntax checks also include macos_tcp_memory.py; network-free interface checks verified the 48-byte ABI, TCP subsystem lookup, 64-bit limit values and raise-only configuration.

Simplification/self-review covered the final complete branch diff: the temporary standalone diagnostic script was removed in favor of a failure capture in the existing harness. The small macOS helper is limited to the documented XNU memory-accounting interface; no runtime dependency or probe behavior changed. git diff --check passes and the worktree is clean.

Merged-revision verification

Merged as 67c08a33100362c54b56fde784b0fd234f70fa11. Its tree matches the tested head exactly (a96119553a1a5fd75c14d9a0be60baf987e86849). Local master was fast-forwarded; the original working checkout was preserved.

  • Merged master tests: all four platform jobs completed successfully.
  • Merged master build pipeline: all eight Linux/macOS/Windows platform builds and the Debian package completed successfully. Tag-only publication was skipped as expected.
  • Downloaded the pipeline's x86_64-linux-rustscan.tar.gz artifact and verified its SHA-256 against the packaged checksum: 03c64d11de0b1603feae154cbbdc76478244a49a0dcb3c35268ecdf057a188bb. Extracted binary --version and --help both exited 0 (rustscan 2.4.1); these checks perform no scans.

@bee-san
bee-san merged commit 67c08a3 into master Oct 2, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Optimize port exclusion filtering and avoid copying the port list

1 participant