Skip to content

Repository files navigation

HomeLab

A segmented home network and its observability stack, managed as code.

CI Digest drift Last commit Open issues Dependabot License: MIT

pfSense Proxmox VE TrueNAS Ubuntu Server Docker Compose WireGuard

Prometheus Alertmanager snmp_exporter Grafana Loki Alloy

SOPS age Suricata Zeek Wazuh Velociraptor Caddy step-ca gitleaks

OpenTofu Packer Ansible Home Assistant Jellyfin Immich

Start here · Architecture · Network · Observability · Security · Runbooks · Decisions · Roadmap


Six VLANs and the untagged switch-management LAN behind a pfSense firewall, default-deny between every segment, with a Prometheus/Loki/Grafana stack watching all of it. Every config in this repository is the config that runs, validated on every push.

It started as a place to practise security work and turned into the network the house actually depends on, which changed the requirements considerably — a broken experiment is a learning opportunity, a broken DHCP server is a domestic incident.

Arriving without the context? docs/runbooks/successor-handover.md is the front door: what this estate is, what to check on day one and in what order, what fails soonest if nobody touches anything, where the secrets are and what is needed to decrypt them, and what can be switched off. This repository is the operator-facing documentation — the wiki on oracle is the household's, and ADR-0011 is why they are different documents for different readers.

Highlights

  • Network segmented by trust, not by function. Six VLANs; IoT, media and guest segments are terminal outward — nothing on them initiates anywhere else, and each carries a tripwire that logs anything which gets past that. Inbound is a separate question, answered one host at a time: since 2026-09-16 more-trusted segments reach smaug on CasaBonita on named ports, so the televisions can have a media server without the segment ceasing to be terminal in the direction that matters (ADR-0016), and since 2026-09-28 Home Assistant reaches the Hue bridge on Skids and nothing else there (ADR-0035). Default deny holds everywhere except the trusted workstation segment and the switch LAN, both of which are listed rather than counted. Why
  • Full observability pipeline for a mixed estate. Grafana Alloy agents push metrics and logs from Linux hosts; snmp_exporter polls the four devices that can't run an agent (firewall, switch, UPS, iLO). One agent config, deployed identically everywhere. How
  • Dashboards and alerting as code. 7 provisioned dashboards, 144 panels, and 157 alert rules — 138 metric-based in Prometheus, 19 log-based in Loki — sharing one Alertmanager routing tree. No dashboard exists only in a database.
  • Secrets encrypted in-repo with SOPS + age. Per-device credentials, decrypted at deploy time into gitignored paths, with git log showing which credential rotated and when — but never to what. Why
  • CI that actually validates the infrastructure. docker compose config, promtool, amtool, alloy fmt, a real Loki boot to parse the LogQL rules, dashboard-JSON and datasource checks, every dashboard's PromQL parsed, plus gitleaks over the full history.
  • CI that validates the documentation too. Eleven assertions cross-check this prose against the configs it describes — counted claims (rules, dashboards, panels, Alloy agents, VLANs, ADRs, runbooks), the SNMP inventory and the compute table against docs/network.md, the host/stack and ports tables against compose.yaml, ADR numbering, firewall posture against docs/firewall-claims.yaml, guest rows against each other, this file's outstanding-purchase count against the roadmap's buy table, every critical alert's runbook_url against the runbook and heading it names, and a ban on image versions in prose — Dependabot edits only compose.yaml, so a version written anywhere else is stale from the next bump. A document that disagrees with the repository fails the build. That opening count is now one of the claims, read from the check registry rather than kept by hand: it said six while ten ran, and the list beside it had been overtaken by four — drift in the sentence advertising that drift gets caught.
  • Supply chain pinned by digest. Every image carries both a tag and a sha256: digest, so a moved tag cannot change what deploys. CI enforces it; make pin-digests re-resolves them from the registry. Every docker run in the Makefile, the scripts, the workflow and the runbooks resolves its image from compose.yaml too, so an image that is not pinned there cannot be run at all.
  • Documented decisions and runbooks. 79 ADRs covering what was chosen and what was rejected — including the costs accepted knowingly; 44 runbooks for the operations that are easy to get wrong at 1am, one of which is the handover page a successor reads first. Every critical alert links to one.

Architecture

The home network: the ISP gateway in bridge mode, the pfSense firewall morpheus, the core switch neo and the 9U rack across the top; below them one panel per VLAN in its patch-cable colour, holding every addressed host — Winterfell's monitoring, wiki and household hosts, Hicks' workstations by role, CasaBonita's NAS and screens, ImaginationLAN's hypervisor and its ten guests, Skids' IoT by class, and the guest network.

Click for the full-resolution SVG. Drawn from docs/network.md and docs/hardware.md; how to edit it is in docs/diagrams/.

Each panel's footer says what that segment reaches, and that is the short version. Default deny holds for every segment except Hicks and the switch LAN, both of which reach further than any picture of exceptions suggests. network.md's Reaches column is the segment-level summary. The host-scoped passes underneath it are in each segment's notes in the same file, and the two together are the current state. Two examples: Saruman reaching prometheus on 9090/3100, and smaug over NFS. ADR-0013 holds the method and the reasoning, and describes the ruleset as it stood on 2026-09-01; the Hicks interface was narrowed the day after. A count was the wrong instrument and this README carried the wrong count for months. Segment colour matches the patch cable in the rack. A dashed border means the segment is terminal outward: nothing on it initiates a connection to another internal segment, though named inbound passes may still reach it. Data flow and the maintained Mermaid topology are in docs/architecture.md.

Stack

Layer Tool Where Role
Firewall / routing pfSense on FreeBSD 16 morpheus VLANs, Kea DHCP, Unbound, Suricata, NUT, default-deny, the one WireGuard rdr
Virtualisation Proxmox VE Saruman Lab hypervisor; iLO on shiva
Storage TrueNAS 25.10 smaug 2× 18 TB ZFS mirror erebor, the SMB share, and the Docker the media stack runs under
Media Jellyfin, Audiobookshelf, Navidrome smaug Quick Sync transcoding on the NAS, audiobooks with synced progress and music over Subsonic (#140, #141); the one stack deployed from TrueNAS rather than by make deploy
Household services Caddy, step-ca, Home Assistant, AdGuard Home, Immich, Paperless-ngx, Vaultwarden and more trinity The sensitive tier, behind its own CA, and the house's DNS filter
Wiki Wiki.js and Postgres oracle The household's documentation, and the off-host backup copies
Metrics Prometheus prometheus 30-day retention capped at 12 GiB, remote-write receiver
Logs Loki prometheus Single-binary, filesystem storage
Collection Grafana Alloy every Linux host but smaug, which is scraped through node_exporter instead node + cAdvisor metrics, Docker/journal/syslog/auth logs
Network polling snmp_exporter prometheus pfSense, switch, UPS, iLO
Alerting Alertmanager prometheus Severity routing, inhibition
Visualisation Grafana prometheus 7 provisioned dashboards
Lab observability Prometheus, Loki, Grafana alexander The lab's own Prometheus; only liveness crosses to the estate's, never telemetry
Security tooling Wazuh, Velociraptor odin SIEM and endpoint forensics for the lab domain
Network sensor Zeek fenrir East-west traffic on the lab bridge, from a tc mirror
Lab provisioning Packer, OpenTofu, Ansible phoenix VM templates, guests cloned from them, and the ad.matrix.elysium domain's configuration
Secrets SOPS + age in the repo Encrypted in-repo, decrypted at deploy time
CI GitHub Actions GitHub Lint, config validation, secret scanning, digest pinning, close keywords in prose

Repository layout

.
├── stacks/
│   ├── observability/        # the estate's stack on prometheus — nine services
│   │   ├── prometheus/       #   config, file_sd targets, 138 alert rules
│   │   ├── alertmanager/     #   routing and inhibition
│   │   ├── loki/             #   single-binary config + 19 LogQL rules
│   │   ├── alloy/            #   the agent config directory, shipped to every host
│   │   ├── snmp-exporter/    #   generator.yaml is the source of truth
│   │   └── grafana/          #   provisioning + 7 dashboards
│   ├── sensitive/            # the household's tier on trinity — its own CA,
│   │                         #   leaves over ACME (ADR-0034, ADR-0037)
│   ├── media/                # Jellyfin, Audiobookshelf and Navidrome on smaug,
│   │                         #   under TrueNAS's own Docker (ADR-0040)
│   ├── wiki/                 # Wiki.js on oracle (ADR-0011, ADR-0015)
│   ├── lab/                  # the lab's own stack on alexander — six services,
│   │                         #   never remote-writes to VLAN 99 (ADR-0020)
│   ├── soc/                  # Wazuh and Velociraptor on odin (ADR-0030)
│   ├── sensor/               # Zeek on fenrir, on a mirror of the lab bridge (ADR-0068)
│   └── scratch/              # a DISPOSABLE copy of soc on diabolos (ADR-0071)
├── packer/  tofu/  ansible/  # lab templates, guests and domain, run from phoenix
├── secrets/                  # SOPS-encrypted; see secrets/README.md
├── scripts/                  # bootstrap, render, validate, check_docs… — see its README
├── systemd/                  # the timers that back up, verify and converge — see its README
├── .github/                  # CI, Dependabot, the ruleset on main — see OVERVIEW.md
├── SECURITY.md               # disclosure policy and known exposure
├── docs/
│   ├── architecture.md  network.md  hardware.md
│   ├── observability.md  security.md  roadmap.md  changelog.md
│   ├── diagrams/             # the network diagram (SVG) and its predecessors
│   ├── adr/                  # 79 ADRs — architecture decision records
│   └── runbooks/             # 44 runbooks; successor-handover.md is the front door
└── Makefile                  # make help

Every directory with more in it than its name says has a page of its own:

  • stacks/: each stack's README covers its host, its services and what it deliberately leaves out.
  • scripts/: every script by purpose, with the make target that runs it.
  • systemd/: every timer, its host and its schedule, and how a job that stops running pages.
  • .github/: the workflows, the ruleset on main, Dependabot and the templates.
  • docs/diagrams/: the network diagram, what it is drawn from, and how to keep it current.
  • packer/, tofu/, ansible/ and secrets/.

Quick start

Requires Docker with the compose plugin, plus sops, age and openssl.

git clone https://github.com/Gerrrt/HomeLab.git && cd HomeLab

make secrets-init     # generate an age keypair, create the encrypted secrets file
make secrets-edit     # fill in real values
make certs ARGS=--ca  # create the lab CA
make certs ARGS="--host grafana.matrix.elysium --ip 10.0.99.20 --dns grafana"    # Grafana's leaf
make validate         # everything CI runs
make up               # render config and start the stack

The two certs steps are not optional: Grafana serves https from that leaf and Prometheus verifies it with the CA, so make up renders nothing until they exist. Details in docs/runbooks/generate-certificates.md.

Grafana on :3000 over https, Prometheus on :9090. Alertmanager binds to 127.0.0.1 and is reached through Grafana (#70). Grafana's certificate is signed by the lab's own CA, so a browser warns and curl needs -k until you trust certificates/ca.pem — step 4 of that runbook. Full procedure, verification steps and troubleshooting in docs/runbooks/deploy-stack.md.

That is the first deploy. After it, the monitoring host deploys itself: a timer runs scripts/converge.sh hourly, which fetches main, refuses it unless the tip carries GitHub's signature, fast-forwards and runs the same make up — recording what it deployed and refusing to overwrite anything edited on the host (#99, ADR-0021, docs/runbooks/converge-the-host.md). A host is installed with HOMELAB_CONVERGE_APPLY=0, which makes the timer fetch, verify and report what it would deploy without applying it, and that runbook has the step that lets it act — so on a host still in that mode, a merge reaches the stack only when someone runs make up there.

$ make help
  up               Render config and start the stack
  converge         Fetch main, verify it, fast-forward and deploy (ARGS=--dry-run)
  down             Stop the stack (volumes are preserved)
  reload           Hot-reload Prometheus, Alertmanager and snmp-exporter (no restart)
  secrets-init     Generate an age keypair and create the encrypted secrets file
  secrets-edit     Edit the encrypted secrets in $EDITOR
  validate         Run every check CI runs
  check-docs       Verify the documents agree with the configs
  backup           Quiesce the stack, archive its volumes to ./backups/, verify, copy to oracle
  restore          Restore the stack's volumes from a backup set (ARGS="--from <stamp>")
  install-timers   Install and enable the systemd timers on this host (needs sudo)
  secrets-verify-backup  Check a backup age key decrypts the secrets (KEY=/path/to/keys.txt)
  ...

The timers are what stop backup, backup-firewall, snmp-verify and now deployment itself being things someone has to remember, and the alert rules that come with them fire on a job having stopped being run rather than only on one that failed (#77). One job deliberately has no timer: secrets-verify-backup needs a human to mount removable media, so it gets a ninety-day deadline and an alert instead. See docs/runbooks/schedule-maintenance.md.

Dashboards

Rendered from the running stack by make screenshots on 2026-10-04, over a 24-hour window. Every dashboard but Logs and Security is here; docs/images/README.md explains why those two are deliberately left out, and what in these renders is a known fault rather than the steady state.

Host Overview dashboard: CPU, memory, load, storage and network for every host running an Alloy agent, with a table of firing host alerts across the top.

Docker Containers dashboard: per-container CPU, memory, network and filesystem writes from cAdvisor, alongside restart counts, CPU throttling and a container inventory.

Network & Firewall dashboard: pfSense pf state table and packet filter drops, MokerLink switch interface throughput and link status, and HPE iLO chassis power draw and hardware health.

UPS & Power dashboard: APC power source, battery charge, output load, runtime, voltage and time on battery, under the banner recording when the pack was fitted and proven.

Observability Stack dashboard: every scrape target with its staleness, then Prometheus, Loki, Alertmanager and the Alloy agents — the collection path watching itself.

What runs it

The entire observability stack runs on a 2012 MacBook Pro with Ubuntu Server on it. Four SNMP devices at a 60-second interval, Alloy agents, and 30 days of metrics, on hardware that was otherwise going to landfill. The rest of the estate is as unglamorous:

Host Hardware Job
morpheus HP ProDesk 600 G4 Mini, a second NIC on an M.2 adapter pfSense firewall
neo MokerLink 26-port managed switch Core switching
Saruman HPE ProLiant DL360 Gen9, iLO 4 as shiva Proxmox VE, and the lab's guests
smaug Lenovo ThinkServer TS150 TrueNAS and the media stack
trinity HP ProDesk 600 G4 DM The household's tier
prometheus Apple MacBook Pro (2012) The observability stack
oracle Dell Inspiron 15 The wiki and the off-host copies
mjolnir APC Smart-UPS X 1500 Power for the rack and everything on its PDU

All of it but the two laptops, trinity and the NAS sits in a 9U open-frame rack. Details in docs/hardware.md.

Security posture

Segmentation rationale, threat model, secrets handling, and an explicit account of what this repository deliberately does not publish (full MAC addresses, owner-linked device names, camera placement) are in docs/security.md.

Historical credential exposure in this repository's git history is documented there too, along with the runbooks to remediate it — including the parts not yet done. SECURITY.md carries the disclosure policy and a summary of what is known.

Container images are pinned by tag and digest. A tag is a mutable pointer; a digest is the content hash, so a moved tag cannot change what gets deployed. CI enforces it, and make pin-digests re-resolves them.

compose.yaml is the only place an image may be named, including images no service runs — the tar that takes backups and the scanner CI runs are both profile-gated entries there. CI parses every docker run, pull and create in the repository and requires each to resolve its image through scripts/image-for.sh, because the pin that caused this rule was not a wrong one but a missing one, and no amount of grepping finds those.

Roadmap

Open work is tracked in Issues, grouped into milestones; docs/roadmap.md is the shape of it — what is outstanding, what gates it, and why it is in that order. What happened, and what it found, is docs/changelog.md, dated and never rewritten.

Every purchase still outstanding is in one place, the roadmap's Everything still to buy: 0 items now, one later, and a rule that nothing joins them without a decision. This sentence used to carry the list itself, name three purchases coupled to the UPS work and omit the tier's host entirely, which is how one ProDesk came to be bought for two jobs. It names no current item for the same reason: the roadmap moves faster than a front page is reread.

License

MIT

About

Cybersecurity HomeLab documentation. Here are my notes, setups, and configurations for infrastructure, applications, and networking.

Topics

Resources

Security policy

Stars

0 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages