Skip to content

network — Design notes

Implementation and reference detail for the network module (OPNsense firewall, zones, DNS, per-module rules, reverse proxy). For the catalog entry see README.md; for installation see INSTALL.md; for test coverage see TEST.md.

This module is the canonical source of aliases.json; the deployed copy in /home/tappaas/config/ is seeded from here. The canonical zones.json lives with network-manager (../tappaas-cicd/manager/network-manager/zones.json), which transforms it into the live per-installation config/zones.json at install time.

Capabilities

Capability Entry point Purpose
network:proxy services/proxy/*.sh → caddy-manager Caddy reverse-proxy registration per consumer module
network:rules services/rules/*.sh → rules-manager Per-module ingress/egress firewall rules compiled from module.json

A consumer module opts in by adding the capability to its dependsOn, e.g.:

"dependsOn": ["cluster:vm", "network:proxy", "network:rules"]

Firewall bootstrap (prebuilt image)

The firewall is installed during foundation bootstrap, before tappaas-cicd exists. config-firewall.sh imports a preconfigured OPNsense image (built by GitHub Actions, published as a Release — see .github/workflows/build-opnsense-image.yml) that boots straight to a working firewall at 10.0.0.1 — LAN, DNS/DHCP, static hosts, hostname, API enabled — with no GUI clicking and no installer (this replaced the earlier dvd-ISO + manual console-installer procedure).

Deployment-unique credentials, over one-shot SSH: OPNsense has no API to create/rotate API keys, so the unique API key + root password are applied by replacing /conf/config.xml (rendered from firewall-config.xml.template) over the image's well-known bootstrap SSH, then rebooting. The deployed config has SSH off, so the bootstrap credentials die with the push; they are only ever reachable on the isolated bootstrap LAN. API credentials for cicd land in ~/.opnsense-credentials.txt.

The disk is an expandable UFS disk (32G) — the earlier nano image's fixed raw layout could not be grown and caused disk-full update failures.

Files

File Role
config-firewall.sh Bootstrap orchestrator — run on a PVE node; creates the VM from the prebuilt image and pushes the unique config.
firewall-config.xml.template Parameterized OPNsense config.xml (placeholders @APIKEYS@, @ROOT_PW_HASH@). Other values are TAPPaaS conventions (mgmt 10.0.0.0/24, internal/mgmt.internal, static hosts). No private keys/certs are committed; OPNsense regenerates the GUI cert on first boot.
migrate-firewall-to-network.sh Supervised migration of a legacy firewall module install to the network name (ADR-007 P8).
update.sh Module update: OPNsense software update → conditional reboot → DNS health check → zone-manager → proxy/trunk reconfiguration.

Bootstrap vs cicd

Stage Configures Tooling
Bootstrap VM + expandable UFS disk; LAN 10.0.0.1; DHCP; Unbound+Dnsmasq DNS; mgmt static hosts; hostname; API key enabled config-firewall.sh + seeded config.xml (no cicd)
Later VLANs/zones, per-module reverse proxy, firewall rules, ACME certs opnsense-controller (zone-manager/caddy-manager/rules-manager) in tappaas-cicd

Maintenance note: firewall-config.xml.template must track the OPNsense version in network.json — OPNsense's config.xml format can change across releases.

Naming note (ADR-007 P8): the module was renamed firewall → network; the OPNsense host is still firewall.mgmt.internal (the host rename is a deferred supervised migration). update.sh resolves config/network.json first and falls back to the legacy config/firewall.json.

Per-module firewall rules

The network:rules capability lets each module declare its inbound and outbound network contract in its own JSON. Rules are compiled deterministically into OPNsense and reconciled on every update.

Schema fields (see module-fields.json)

Field Purpose
ports[] Network ports the module exposes for inbound traffic. Source of truth for ingress validation.
ingress[] Inbound traffic permitted to those ports (from, ports, protocol, description).
egress[] Outbound exceptions beyond the source zone's access-to (to, ports, protocol, description).
aliases{} Module-local OPNsense aliases referenced by alias:<name>. Overrides identically named entries in aliases.json.

The from / to fields accept:

  • A zone name (must exist in zones.json)
  • The literal "internet" (resolves to WAN-side any)
  • Another module name — resolved to an OPNsense host alias tappaas_module_<vmname> populated with the peer's FQDN (<vmname>.<zone0>.internal). OPNsense's Unbound resolves the FQDN against dnsmasq, so DHCP-driven IP changes flow through without rule rewrites.
  • "alias:<name>" — references a module-local alias or a global entry in aliases.json.

Worked example

{
  "vmname": "litellm",
  "zone0": "srv_work",
  "dependsOn": ["cluster:vm", "network:proxy", "network:rules"],
  "ports": [
    { "port": 4000, "protocol": "TCP", "description": "LiteLLM API" }
  ],
  "ingress": [
    { "from": "srv_work", "ports": [4000], "description": "Intra-zone consumers" },
    { "from": "dmz",      "ports": [4000], "description": "Reverse proxy" }
  ],
  "egress": [
    { "to": "alias:llm_providers", "ports": [443],   "description": "Upstream LLM APIs" },
    { "to": "vllm",                "ports": [11434], "description": "Local inference fallback" }
  ],
  "aliases": {
    "llm_providers": {
      "type": "host",
      "addresses": ["api.anthropic.com", "api.openai.com"],
      "description": "Approved upstream providers"
    }
  }
}

Rule identity

Every compiled rule carries a canonical description used for idempotent upsert and orphan detection. There are two prefixes:

tappaas-module:<vmname>:<direction>:<peer>:<port>[/<protocol>]      # manual rules
tappaas-svcdep:<consumer>:<service>:<provider>:<port>[/<protocol>]  # auto-pinholes

Examples: - tappaas-module:litellm:ingress:srv_work:4000 - tappaas-module:litellm:egress:vllm:11434 - tappaas-module:hassosova:egress:iot-home:5353/UDP - tappaas-svcdep:hassosova:mqtt:mosquitto:1883 ← auto-pinhole

The owner-module is always position 1; that's the consumer for tappaas-svcdep and the rule-owning module for tappaas-module. list-rules, verify-rules, reconcile, and remove-rules recognise both prefixes.

Auto-pinholes from service dependencies

Manual ingress / egress entries cover bespoke firewall policy. For the common case — "this module just needs to talk to the service it depends on" — the network:rules service can synthesise the pinhole automatically.

A service provider opts in by dropping a pinhole.json in its service directory declaring the ports the service answers on:

// <provider-module>/services/<service>/pinhole.json
{
  "ports": [
    { "port": 4000, "protocol": "TCP", "description": "LiteLLM API" },
    { "port": 4001, "protocol": "TCP", "description": "Prometheus scrape" }
  ]
}

A consumer module triggers the auto-pinhole simply by declaring the service — in dependsOn or in integratesWith — and declaring network:rules in its own dependsOn:

{
  "vmname": "translation-agent",
  "zone0": "srv_work",
  "dependsOn": [
    "cluster:vm",
    "network:rules",           // opt into per-module firewall
    "litellm:llm-proxy"        // provider:service with a pinhole.json
  ]
}

When the consumer's install runs, rules-manager walks both dependsOn and integratesWith, finds each provider's pinhole.json, and emits one auto-pinhole per declared port — only when:

  1. The two modules are in different zones (intra-zone traffic flows freely).
  2. The consumer's zone is not already in the provider zone's access-to (zone-level rule covers it).
  3. The consumer's zone is in the provider zone's pinhole-allowed-from (policy gate).

If condition 3 fails, the auto-pinhole is skipped with a warning — the operator chose this trade-off on the original ticket: a missing zone policy should not block an install, just be loud about what wasn't done.

Auto-pinholes are owned by the consumer: they're created when the consumer is installed, recomputed on reconcile (so a changed declaration re-applies), and removed on the consumer's teardown — regardless of the provider's state.

Hardness and reachability are separate properties (#632). dependsOn and integratesWith differ in LIFECYCLE — whether a missing provider blocks the install — not in whether traffic may flow. A module can legitimately be optional-but-reachable, which is exactly what integratesWith exists to express, so both lists earn a pinhole and both are subject to the same three conditions above. Walking only the hard list meant a cross-zone optional integration got no path; worse, where a module had previously declared the coordinate as a hard dependency, the rule already in OPNsense became one nothing declared, and the next reconcile --apply pruned it as extra.

The provider is resolved the way the installer resolves it — an environment's own <provider>-<env>.json wins over the shared <provider>.json — so a consumer in an environment reaches the provider deployed alongside it. Looking the provider up by its bare name silently emitted no rule for an environment-deployed provider, even for a hard dependency install-module.sh had already validated.

Sequence bands

Band Range Source Purpose
0 0–99 OPNsense auto Anti-lockout
1 100–999 zone-manager Infrastructure (DHCP, NTP, ICMP); caddy-reach — the DMZ gateway /32 on tcp/80+443, seq 990/991 (see below)
2 1000–9999 zone-manager Foundation deny defaults
3 10000–19999 rules-manager ingress Per-module pinholes
4 20000–29999 rules-manager egress Per-module egress exceptions
5 30000–39999 zone-manager Zone-level rules: gateway + access-to allows, then rfc1918 block + internet
6 40000–49999 Manual Logging-only
7 50000–59999 Manual Administrator overrides

Within bands 3 and 4, each module receives a deterministic 100-slot range based on a stable hash of its vmname. Slot collisions are detected at compile time.

Rules use quick (first match wins; lower sequence = higher priority). Band 5 sits above the module bands so a zone's rfc1918 catch-all block (which isolates a zone from unlisted internal networks) is evaluated after per-module pinholes — otherwise it would shadow them and silently break cross-zone module connectivity. Within band 5 each zone gets a deterministic 100-slot range (stable hash of the zone name; cross-zone slot collisions are harmless since each zone's rules bind to its own interface). The intra-slot offsets are fixed — base+0 gateway DNS, base+1..+85 one pass per access-to zone, base+86..+88 gateway NTP, DHCP and ICMP, base+90..+92 rfc1918 block, base+99 internet — so adding a zone to access-to never renumbers or collides with the block/internet rules.

What a zone may ask of the firewall (#399, #384)

A zone reaches its own gateway for the services the firewall runs for it — DNS (tcp/udp 53), NTP (udp 123), DHCP (udp 67) and ICMP — and nothing else. The gateway rule used to pass every port, which put the admin GUI (8443) and SSH (22) in reach of every zone, guest and IoT included, behind a password only. Every zone this module writes rules for also gets two blocks at seq 980/981 (band 1, ahead of every pass): tcp 8443 and 22 to (self), every address the firewall holds. That closes the other route in — an access-to edge grants another zone's whole subnet, and that zone's gateway is this firewall. Manual zones get no rules and keep their reach: mgmt (LAN, also covered by OPNsense's anti-lockout), the NetBird overlay and the admin VPN. Caddy is reached by its own rule, below.

Anti-spoofing (#159), seq 970 — the first rule of every zone. A packet arriving on a zone's interface must carry a source in that zone's subnet; Zone X block spoofed source drops everything else (from ! <subnet>), before any pass can match a forged address. OPNsense's own DHCP rules come earlier, so a client's 0.0.0.0 discover still works. Fragments need no rule: pf's scrub in all fragment reassemble (the OPNsense default) reassembles every packet before filtering and drops overlapping fragments, so no fragment reaches a rule. Both are always on — TAPPaaS takes the decision, not a field.

Why caddy-reach is in band 1 (#366, ADR-021 D3)

Split-horizon DNS resolves every published name to the DMZ gateway, where os-caddy listens. A client zone has no access-to: dmz — and under ADR-021 D3b must not, since that would grant the whole DMZ subnet on every port, the firewall's own GUI and SSH included (#618) — so the zone's own band-5 rfc1918 block at base+90 would drop that traffic. The pass therefore sits at seq 990/991, low in band 1, where first-match-quick evaluates it before the block. Band 1 is load-bearing here, not cosmetic: move this rule into band 5 and every published name stops resolving usefully from every client zone.

It is L3 reachability only — a /32 on two ports, never the DMZ subnet. Authorization stays where it can be expressed per service: Caddy's proxyAllowedZones access list and the Authentik identity gate. zone_manager._configure_caddy_reachability emits it for:

Candidate Why
any zone with internet (or all) it can already reach that Caddy the long way round, via the WAN hairpin, so this opens no new destination
an Overlay zone with a non-empty access-to (D3a) overlay peers are not a routed segment and carry no internet, yet resolve the same answer. admin is the live case; its bridge names the interface its peers arrive on (e.g. wireguard). A dormant overlay (empty access-to) stays out
the dmz zone itself (D3b) it was skipped as "reaches the gateway locally", but that was the unrestricted Zone dmz -> gateway rule which #399 narrows to DNS/NTP/DHCP/ICMP — after which a DMZ workload would lose tcp/443 to Caddy

A fully-isolated zone (empty access-to, no internet) gets no rule and cannot reach a published name — the correct outcome, by the same mechanism that denies it the internet.

Validation

Compile-time, before any OPNsense API call:

  1. Schema — fields/types match module-fields.json.
  2. Zone existence — every zone-named peer exists in zones.json.
  3. Policy — every ingress.from is in destination zone's pinhole-allowed-from.
  4. Port consistency — every ingress.ports value is in module.ports[].
  5. Module existence — every module-named peer has a corresponding <peer>.json on disk.

Egress to a zone not in the source zone's access-to is permitted but logged as a warning (exceptions are intentional).

firewallType: "NONE" fallback

When network.json declares firewallType: "NONE", the service scripts skip OPNsense entirely and rules-manager prints the rules in a human-readable form for manual entry into the operator's firewall. Exit code 0.

CLI: rules-manager

rules-manager add-rules <module>          # compile + apply
rules-manager reconcile <module>          # diff + apply + prune orphans
rules-manager remove-rules <module>       # delete all rules/aliases for module
rules-manager verify-rules <module> [--deep]
rules-manager list-rules [--module <n>] [--orphans]
rules-manager create-alias <name> --type host --addresses ip1,ip2
rules-manager remove-alias <name>

# Common flags:
#   --firewall-type opnsense|NONE
#   --zones-file <path>          --aliases-file <path>
#   --modules-dir <path>         --check-mode
#   --output text|json           --no-ssl-verify
#   --credential-file <path>     --debug

The CLI is normally invoked by the services/rules/*.sh capability scripts on the consumer module's lifecycle hooks, but can be run directly for debugging.

Physical network: switches & WiFi (ADR-008)

Beyond OPNsense (L3), zones.json is also reconciled onto Proxmox bridges/trunks, physical switches, and WiFi APs so a zone's VLAN is carried everywhere it needs to be. These providers live in scripts/ and run on tappaas-cicd (symlinked into ~/bin):

Script Role
network-manager reconcile orchestrator — runs every provider in order (opnsense → proxmox → switch → ap)
switch-controller physical switches (controllers / switches / ports → trunk + access VLANs)
ap-controller WiFi APs (SSID → VLAN via the vendor controller)
proxmox-controller node bridge-vids + per-VM trunks
setup-switches.sh interactive switch registration (bootstrap step)
setup-wlan-secrets.sh set WiFi SSID names (in zones.json) + passphrases (0600 secrets file)

Each provider follows a 5-verb contract (interrogate → update-desired → delta → apply → confirm) over two files in ~/config/: switch-configuration-actual.json (reality) and …-desired.json (regenerated from zones.json). Vendor automation is plugin-based (scripts/plugins/<vendor>.sh; unifi.sh shipped, manual.sh is the by-hand fallback).

Full command reference, the inventory model, and how to add a switch brand: scripts/README.md.

Test network on a dedicated physical port

test-network.sh stands up a throwaway, isolated test network served on a spare physical NIC of the node running the firewall VM — separate from the VLAN trunk that carries the production zones. Plug a switch or AP into the spare port and you get an isolated, internet-connected sandbox without touching zones.json or the trunk.

test-network.sh                       # interactive: pick a vacant port, default 172.17.3.1/24
test-network.sh --port enp3s0         # non-interactive port choice
test-network.sh --subnet 172.17.9.1/24 --bridge testbr2
test-network.sh --status              # show current state
test-network.sh --delete              # tear down in reverse order
test-network.sh --check-mode ...      # dry run, no changes

What it does, in order (and reverses on --delete):

  1. Finds the node hosting the firewall VM (pvesh), discovers vacant physical ports, and prompts for one.
  2. Creates a Linux bridge on that node and enslaves the port, persisted in /etc/network/interfaces (backed up first) and applied with ifreload.
  3. Attaches the bridge to the firewall VM as a new virtio NIC (qm set --netN), which appears in OPNsense as vtnetN.
  4. Drives the OPNsense side via test-network-manager (an opnsense-controller entry point): assigns the interface with a static gateway IP, enables DHCP, and installs the routing/firewall policy.

Routing policy (asymmetric, per the issue):

From → To Action Notes
test → internet allow OPNsense automatic outbound NAT covers 172.16/12
test → internal (RFC1918, incl. mgmt) block isolation
mgmt → test allow return traffic is stateful, so test→mgmt stays blocked

All OPNsense artefacts (interface, DHCP range, rules) carry a test-net description so teardown removes exactly what setup created.

Note: test-network-manager is a console-script entry point in the opnsense-controller package. Rebuild that package (e.g. via update-tappaas / nixos-rebuild) so the wrapper lands in PATH; test-network.sh falls back to python3 -m opnsense_controller.test_network_cli when the wrapper is not yet present.

See docs/test-network-setup.md for the full setup/teardown runbook, options, verification, and troubleshooting.

Troubleshooting: Unbound / DNSBL

Symptom: DNS resolution fails on the mgmt network (clients cannot resolve via 10.0.0.1), typically during zone-manager --execute or right after a firewall update. Unbound logs show ModuleNotFoundError: No module named 'dns'.

Root cause (Python version mismatch): OPNsense 25.x compiles Unbound against Python 3.11 but ships dnspython only as py313-dnspython (py311-dnspython does not exist in the OPNsense repos). The unbound.inc PHP plugin unconditionally adds python iterator to the module-config and includes the DNSBL python script, so any operation that regenerates the Unbound config (zone-manager API calls, an OPNsense update) produces a config Unbound cannot load. The prebuilt image works only because it ships a pre-generated /var/unbound/unbound.conf — the failure appears on the first regeneration. This is why update.sh runs the OPNsense update + reboot before zone-manager, followed by a DNS health check.

Second root cause (working-directory path): the DNSBL module is referenced by a relative path, and only one of the two Unbound start paths sets the working directory:

Method Script Has cd /var/unbound/ Result
pluginctl -c unbound_start /usr/local/opnsense/scripts/unbound/start.sh Yes Works
service unbound restart /usr/local/etc/rc.d/unbound No Fails

Restarting via the rc.d script runs unbound-checkconf from the wrong directory: pythonmod: can't open file unbound-dnsbl/dnsbl_module.py for reading. The OPNsense API and configd use pluginctl, so API-driven restarts work.

Recovery (on the firewall, csh):

pluginctl -c unbound_start

This uses start.sh, which handles the working directory correctly. Equivalent verbose form: cd /var/unbound && /usr/local/sbin/unbound-checkconf /var/unbound/unbound.conf && service unbound restart. To diagnose, service unbound status + unbound-checkconf /var/unbound/unbound.conf — running-but-checkconf-fails means the process holds an old working config and any restart will fail.

Permanent-fix options (none applied yet — the update.sh ordering is the mitigation):

  1. Patch unbound.inc to drop the Python module from the module-config (disables DNSBL integration; overwritten by OPNsense updates).
  2. Wait for OPNsense to fix the mismatch (ship py311-dnspython or rebuild Unbound against Python 3.13).
  3. Patch /usr/local/etc/rc.d/unbound to cd /var/unbound before its checkconf calls (fixes the path issue, not the version mismatch).
  4. Disable DNSBL entirely (removes DNS-level ad/malware blocking).

Resilience & ingress without a public IP

Public ingress with no public IP. The default ingress path terminates at Caddy on the OPNsense firewall on a routable WAN IP. A site behind CGNAT / dynamic IP / no-inbound instead publishes via a satellite (ADR-010): an operator-owned VPS runs a blind TLS-passthrough reverse-proxy and forwards :443 over a WireGuard tunnel to Caddy at home — no commercial tunnel SaaS, and TLS keys stay on-site. The satellite is optional and non-invasive (a site with a public IP needs none); see the satellite foundation module.

Internet-outage survival. OPNsense runs a caching recursive resolver (Unbound), so the local TAPPaaS ecosystem keeps resolving and functioning when the internet link drops (§"Troubleshooting: Unbound / DNSBL"). Local-network redundancy via B.A.T.M.A.N. was considered but is not implemented.

Firewall failover. OPNsense can run active-active (CARP) for firewall HA, but that requires an Ethernet-connection handover and is out of scope for a standard TAPPaaS install; single-firewall is the supported topology. Node/service HA and disk redundancy are the cluster module's concern — see cluster/DESIGN.md §"High Availability".