Skip to content

feat: outage measurement, TUI path trace, and history sparkline - #59

Merged
servak merged 7 commits into
mainfrom
feature/outage-mtr-history
Sep 24, 2026
Merged

servak merged 7 commits into
mainfrom
feature/outage-mtr-history

Conversation

@servak

@servak servak commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Summary

1. Outage measurement (for failover / maintenance tests)

  • An outage starts after 3 consecutive lost probes and ends at the next successful probe. It is recorded with start/end, duration, and lost probe count; the start is the send time of the first lost probe. Shorter runs count as ordinary packet loss. The threshold is fixed, not configurable.
  • Out-of-order results: concurrent TCP/HTTP probes can report a timeout after a later probe has already succeeded. Results are therefore buffered for timeout + interval and applied in send-time order. The buffer is flushed when the event channel closes, so final reports are complete.
  • Where it shows up:
    • Host list: the LastFailTime cell also shows the outage duration, e.g. 16:44:20 (DOWN 5.20s) in red while down and 16:44:15 (0.50s) in yellow after recovery.
    • TUI detail panel: outage count, total downtime, longest outage, and the most recent outages.
    • A cross-target timeline printed on TUI exit and in batch table output.
    • New json/csv fields: outages[], outage_count, total_downtime_ms, longest_outage_ms.

2. MTR-style path trace in the TUI

  • New internal/trace package. Each round it sends ICMP echo with TTL 1..30 and matches replies by echo ID/seq, including the datagram quoted in Time Exceeded / Dest Unreachable messages (handles IPv4 IHL options and IPv6).
  • Uses a raw socket when available and falls back to unprivileged datagram ICMP. On macOS the fallback also receives Time Exceeded, so it works without root (verified).
  • p toggles a live path panel that follows the selected host. Any target type works: for tcp://, https://, dns://, … the host part is traced.
    • The trace starts asynchronously and runs at most one round per second, so router ICMP rate limiting doesn't show up as fake loss.
    • Probes within a round are spread evenly over the interval (like mtr). While the destination is unknown, it probes at most 8 TTLs past the farthest hop that has replied. Sending TTL 1..30 back to back made the home router treat the traffic as an ICMP flood and drop ICMP for every host behind it, which also broke mping running on other machines.
    • Hops still waiting for their first reply show waiting. The panel also notes when loss appears only at intermediate hops, which is usually ICMP rate limiting rather than real loss.
  • There is no standalone mping mtr subcommand; the trace is only available in the TUI.

3. History column and percentiles

  • The host list shows the last 20 probes as a sparkline, and × marks a failure. Bar height is how many milliseconds an RTT exceeds the host's median (+5/10/20/50/100/200/500ms). A host's normal jitter therefore stays flat, and +100ms or more is drawn in the warning color.
  • RTT p50/p95/p99 in the detail panel and in json/csv.

Fixes

  • Data race between probing and the UI: SortBy/GetMetrics returned live pointers that the UI read without the lock. They now return deep-copied snapshots, and panels are refreshed from the latest snapshot on every update.
  • Selection: the selection now stays on the same host when re-sorting moves it to another row, and returns to that host after a filter that hid it is cleared.
  • Descending sort: it uses a strict ordering, so equal rows no longer swap places on every refresh.

Cleanup

Test plan

  • go test -race ./..., golangci-lint run (0 issues)
  • Unit tests:
    • Outage detection: 3-probe threshold, short loss ignored, ongoing outages, reordering, late results.
    • Exports, sparkline, percentiles, LastFailTime cell.
    • Selection across re-sort and filter.
    • ICMP reply parsing (v4/v6, IHL options, rejects), target → host extraction, rate-limit detection.
  • Network integration tests for path tracing (-tags integration)
  • E2E: stopped and restarted a local server during batch → 1.600s outage = 8 lost × 200ms
  • TUI in tmux: History column, outage detail, path view following the selection
  • Path view without root on macOS (11 hops, reverse DNS)
  • Paced probing: 8.8.8.8 and 1.1.1.1 traced in parallel reach the destination in 15/15 rounds, with 0/200 loss to the gateway meanwhile (the burst version never reached either destination and caused shared loss on the LAN)
  • Cross-compiled for every goreleaser target (linux/windows/darwin × amd64/arm64/386)
  • Raw socket on macOS (setuid root): ICMP monitoring and the TUI path view in the same process with no cross-talk (29/29 echo replies). The ICMP prober ignores the tracer's Time Exceeded messages by echo ID.
  • Raw socket (root) run on Linux
  • IPv6 path trace on a v6-enabled network (no v6 route here; parsing is covered by unit tests)

🤖 Generated with Claude Code

@servak servak changed the title feat: outage measurement, MTR path trace, and history sparkline feat: outage measurement, TUI path trace, and history sparkline Sep 23, 2026
Base automatically changed from feature/batch-automation to main September 23, 2026 07:46
@servak
servak force-pushed the feature/outage-mtr-history branch from 901f432 to d792af3 Compare September 23, 2026 08:41
- Record an outage when 3 or more consecutive probes are lost, with its
  start, end, duration and lost probe count. It starts at the first lost
  probe and ends at the next success; shorter runs are ordinary loss.
- Buffer results for timeout + interval and apply them in send order, so
  a timeout reported after a later success does not split an outage.
  The buffer is flushed when the event channel closes.
- Return deep-copied snapshots from SortBy/GetMetrics so the UI no
  longer reads live metrics without the lock.
- Add RTT percentiles (nearest rank over the history).
- Use a strict ordering for descending sorts so equal rows keep their
  order across refreshes.
- Replace SetBeepEnabled with Options.DisableBeep for batch runs.
- Host details show RTT p50/p95/p99 and the recent outages with their
  total and longest downtime.
- Table reports end with a cross-target outage timeline; json/csv gain
  outages[], outage_count, total_downtime_ms, longest_outage_ms and the
  RTT percentiles.
- FormatLastFailCell adds the related outage duration to the last
  failure time (red while down, yellow after recovery).
- FormatSparkline draws recent probes with bar height by how many ms an
  RTT exceeds the median, so normal jitter stays flat.
- Insert a History sparkline column after Loss and show the outage
  duration in the LastFailTime cell.
- Record the selected host on every selection change and restore it by
  name, so re-sorting keeps the same host selected. A filter that hides
  it shows the first row meanwhile without losing the selection.
- Refresh the detail panel from the latest snapshot on every update.
Each round sends ICMP echo requests with TTL 1..30 and keeps per-hop
loss and RTT statistics. Replies are matched by echo ID/seq, including
the datagram quoted in Time Exceeded / Destination Unreachable messages
(IPv4 with IHL options, IPv6). A raw socket is preferred, with a
fallback to unprivileged datagram ICMP, which still receives Time
Exceeded on macOS. Hop addresses are reverse-resolved in the background.

Network integration tests run with -tags integration.
Press p to show an MTR-style trace next to the host list. It follows
the selection and traces the host part of any target (tcp://, https://,
dns://, ...). The trace starts off the UI goroutine and runs at most one
round per second so router ICMP rate limiting does not show up as loss.
Hops still waiting for a reply show 'waiting', and loss only at
intermediate hops is marked as likely rate limiting.
@servak
servak force-pushed the feature/outage-mtr-history branch from 8963f2a to ef2f0f5 Compare September 23, 2026 22:05
Each round sent TTL 1..30 back to back, and kept doing so when the
destination did not answer. Home routers took the burst for an ICMP
flood and dropped ICMP for every host behind them for a while, so
opening the path view on one machine made mping fail on the others.

- Spread a round's probes evenly over the interval, like mtr.
- While the destination is unknown, probe at most 8 TTLs past the
  farthest hop that has replied instead of up to 30.
@servak
servak force-pushed the feature/outage-mtr-history branch from a652c70 to e5346bc Compare September 24, 2026 04:35
@servak
servak merged commit 72d90eb into main Sep 24, 2026
4 checks passed
@servak
servak deleted the feature/outage-mtr-history branch September 24, 2026 05:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant