Skip to content

tappaas-cicd — Design notes

The TAPPaaS "mothership" control plane. This document describes the internal component contract and the three-level dispatch that drives every manager and controller uniformly (ADR-007 P10 deliverable). For the catalog entry see README.md; for installation see INSTALL.md; for test coverage see TEST.md. How the mothership pulls its updates and the repository patterns behind it (basic / community / downstream / private, and the developer workflows) are a separate design note: DESIGN-GIT.md.

For the full rationale see docs/design/ADR-007-implementation.md (packages P4 — layout — and P10 — the template + dispatch).

Overview — internal layout

src/foundation/tappaas-cicd/
├── install.sh / update.sh / test.sh   # top-level entry scripts (drive the dispatchers)
├── bootstrap.sh                       # first-boot: clone repo + nixos-rebuild the VM
├── pre-update.sh                      # pre-update pass of the tappaas-cicd module
├── scripts/refresh-control-plane.sh   # self-refresh: pull + relink ~/bin + build components
├── scripts/tappaas-self-prepare.sh    # ExecStartPre 2: the refresh, before the sweep (ADR-017 D3)
├── scripts/run-migrations.sh          # config migrations, between the refresh and the rebuild (ADR-025)
├── migrations/                        # NNNN-<slug>.sh — one-time config rewrites (empty in this release)
├── scripts/tappaas-self-rebuild.sh    # ExecStartPre 3 (root): nixos-rebuild of the mothership
├── scripts/update-tappaas-schedule.sh # renders update-tappaas.timer from site.json (D2)
├── scripts/notify-update-failure.sh   # mails the site owner when a run fails (#651)
├── scripts/site-notify.sh             # the mail itself: site.json email via a node's mailer (ADR-007e)
├── tappaas-cicd.nix / flake.nix       # the cicd VM's NixOS configuration
├── manager/                           # domain-object lifecycle (CONFIG state)
│   ├── install.sh / update.sh / test.sh   # dispatcher: loop child components
│   └── <name>-manager/               # one dir per manager component
├── controller/                        # infrastructure control (RUNTIME state)
│   ├── install.sh / update.sh / test.sh   # dispatcher: loop child components
│   └── <name>-controller/            # one dir per controller component
├── lib/                               # shared libraries, sourced (never copied per component)
└── scripts/                           # foundation bring-up + unit tests (scripts/test/)
  • manager/ — owns config state: JSON config files, schemas, and their validation. A manager may call controllers to realize that config. Managers ship a validate.sh. Examples: identity-manager, site-manager, environment-manager, module-manager, network-manager, health-manager.
  • controller/ — owns runtime state: APIs, network devices, VMs. A controller does not ship a validate.sh. Examples: opnsense-controller, proxmox-controller, switch-controller, ap-controller, identity-controller.
  • lib/ — shared code sourced by components (e.g. common-install-routines.sh). Shared logic lives here once; it is never copied into individual components.
  • top-level entry scripts — install.sh / update.sh / test.sh at the cicd root.

Component contract

Each manager/controller lives in its own subdirectory and exposes a fixed set of files:

<component>/
├── <name>.{ts,py,sh}   # main entry: domain verbs (CRUD for a manager, reconcile for a controller)
├── install.sh          # idempotent: build artifact, place bin/ symlink, one-time setup
├── update.sh           # idempotent: rebuild, re-link, migrate on-disk state if schema changed
├── test.sh             # self-contained tests; exit non-zero on failure
├── validate.sh         # (managers only) schema/reference validation for the domain
└── README.md           # what it owns; manager-vs-controller; which controllers it calls
  • install.sh, update.sh, and test.sh are mandatory and must be idempotent in every component.
  • validate.sh is present for managers and absent for controllers.
  • The preferred entry-point language order is TypeScript → Python → Bash (see "Preferred language" below).

Runtime-state access rule (the F12 carve-out)

Controllers own runtime state — but managers sometimes need to read it (health gates, module list, drift tables). The rule (ADR-007 post-implementation refactor, F12 decision B):

  • A manager MAY read cluster runtime state, but ONLY through the shared lib/ts/src/cluster.ts helpers (one audited choke-point) — never via its own scattered ssh/pvesh/qm calls.
  • Every runtime write goes through a controller. No exceptions for new code. (Grandfathered: network-manager scp's zones.json to the nodes — that distributes a config artifact, it does not mutate a device; and health-manager's update-os verb delegates to update-os.sh, to be revisited when that script is ported.)
  • health-manager is a read-mostly orchestrator under this rule — it stays a manager.

Config root resolution

Every component resolves the config root the same way: $TAPPAAS_CONFIG, else $CONFIG_DIR, else /home/tappaas/config — TAPPAAS_CONFIG is the TAPPaaS-specific override and wins; CONFIG_DIR is the long-standing variable the bash layer exports. A component must never invent a different precedence (managers used to disagree, which made two managers read different roots under the same environment).

Three-level dispatch

The control plane is driven by three trivial levels — adding a component never changes anything above it:

  1. Top level — tappaas-cicd/{install,update,test}.sh call manager/<verb>.sh and controller/<verb>.sh.
  2. Dispatcher level — manager/<verb>.sh and controller/<verb>.sh each loop their child component directories and run the matching verb script. There is no shared runner; each dispatcher is a few lines:
for d in "${here}"/*/; do
    [ "$(basename "${d}")" = TEMPLATE ] && continue   # defensive: scaffold/work dirs stay inert
    [ -x "${d}install.sh" ] || continue
    "${d}install.sh" "$@"
done
  1. Component level — each component's own install.sh / update.sh / test.sh does the actual work.

Adding a component = drop a directory containing the standard verb scripts (copy the nearest real component and edit — see "Scaffolding a new component"). The dispatcher picks it up automatically; nothing above it is edited.

Top-level wiring is additive (option A). The cicd VM keeps its own nixos-rebuild/test for the VM itself and drives the manager/controller dispatchers — the dispatch is added alongside the existing VM lifecycle, it does not replace it.

Compiled components

A component that ships a built/packaged artifact must rebuild the package and refresh its bin/ entry-point symlinks in install.sh/update.sh — not merely copy source. This is what makes a code change get picked up on update.

  • Python (e.g. opnsense-controller, update-tappaas) — nix build / pip install -e of the component's pyproject.toml, then relink its entry points (the former whole-VM pre-update.sh behaviour, now per-component).
  • TypeScript (future) — npm/pnpm install && build, then link the bin.
  • Bash — nothing to compile; just symlink the entry script onto PATH.

The build step must be idempotent: a no-op when its inputs are unchanged.

Testing: fast vs deep slices

Every component test.sh runs a fast, non-disruptive slice by default and a deep slice only when TAPPAAS_TEST_DEEP=1:

  • Fast (default) — schema/CLI/validation + mocked-logic unit tests. No live services, no Authentik mutation, no cluster/VM ops. Quick and safe anytime.
  • Deep (TAPPAAS_TEST_DEEP=1) — adds the disruptive/slow tests: live Authentik reconcile, cluster ops, VM provisioning. These mutate real state, so they are scoped (e.g. zztest- names) and self-cleaning.

The cicd module gate honours the split:

  • module-manager module test tappaas-cicd (fast, ~seconds) runs the quick checks plus Test 11, a lightweight per-component smoke using the already-built bins (validate schemas, CLI loads, config reads) — confirms basic functionality without touching anything.
  • module-manager module test tappaas-cicd --deep runs the VM/variant suites and the manager/ + controller/ dispatchers with TAPPAAS_TEST_DEEP=1, so every component's full suite (offline unit + live tiers) runs.

When adding a component: gate its disruptive tests behind [[ "${TAPPAAS_TEST_DEEP:-0}" == "1" ]], and add a one-line smoke to Test 11.

A third axis: runtime vs source-tree

Fast/deep is about cost. --runtime-only (TAPPAAS_TEST_RUNTIME_ONLY=1) is about subject, and it cuts across both. A check is runtime if it asks something about the running system (a bin loads, a timer is active, a node answers SSH) and source-tree if it asks something about the repository (an offline unit suite over temp fixtures, a generated-doc drift check, a grep guard over a tracked file).

update-module.sh's Step 2 pre-update test is a gate on a mutation, so it passes --runtime-only: a source-tree defect must never abort the update of a healthy module. A stale generated README did exactly that on three consecutive nightly sweeps (2026-09-07..09) — the mothership stopped updating itself and the fleet's shared manager binaries went unrebuilt (#595). Every other caller — an operator, the deep sweep, CI — runs everything.

Keep new checks on the right side of that line, and put source-tree checks inside the guarded region rather than among the runtime ones.

Self-update: the control plane refreshes before the sweep, not inside it

scripts/refresh-control-plane.sh is the mothership updating itself: pull the tracked repositories, relink scripts/*.sh + lib/*.sh into ~/bin, and rebuild every compiled component through the manager/ + controller/ dispatchers.

It runs as the unit's second ExecStartPre step (tappaas-self-prepare.sh), ahead of every module and of the mothership's own nixos-rebuild (ADR-017 D3); pre-update.sh still calls it so module modify tappaas-cicd stands alone. That placement is the point. The work used to live inline in pre-update.sh, which update-module.sh runs at Step 3 — after the Step 2 pre-update test. So the git pull sat behind a test of the code it would replace: a commit that broke a fast-mode check wedged the mothership, because the pull that would carry the fix could no longer run (#595).

Two properties follow, and both matter:

  • Idempotent, so pre-update.sh still calls it (a standalone module modify tappaas-cicd must be correct on its own) and the second run is a no-op.
  • Exit code 10 means STALE: the pull and relink succeeded, but a component failed to build, so its bin is the previous one. Not fatal — the fleet update proceeds on working binaries, and Test 11 is what surfaces a broken build — but never silent either. It reaches the sweep summary and last-update-result.json as control_plane, and makes ok false. A stale control plane is the root cause that otherwise presents as N unrelated-looking module failures.

Config migrations run in that same step, before the rebuild

scripts/run-migrations.sh applies the pending migrations/NNNN-*.sh (ADR-025) from inside tappaas-self-prepare.sh, immediately after the refresh above and before the hand-over. That slot is the whole design:

  • the migrations it runs are the ones the refresh just pulled;
  • it is before tappaas-self-rebuild.sh, so a migration may change config the new system generation will read — updateSchedule is the first one that must (ADR-017 D7);
  • it is before the sweep's first module, cluster included, which pre-update.sh cannot be: that hook runs inside the sweep, after the modules it depends on.

The cost is a contract rather than a hazard: a migration runs before the rebuild and possibly after a refresh that returned STALE, so it may use only bash, jq, coreutils and the files under config/ — no manager binary, nothing from the not-yet-switched generation.

A failure stops the unit: no rebuild, no sweep, no module updated, and the #651 notice names the stage migrate. The site keeps running the old code with a config/ that is untouched or restorable from config/.migrations/backup/. The ledger (config/.migrations/applied) is written after each success, so a migration interrupted mid-apply is simply re-run next sweep — which is safe because every migration is idempotent.

migrations/README.md is the contract for writing one.

Preferred language

In order: TypeScript → Python → Bash. Pick the highest applicable tier for a new component: prefer TypeScript; use Python where an existing package/ecosystem fit makes it cheaper (e.g. extending opnsense-controller); reserve Bash for thin glue, install-time scripts, and small wrappers. Existing Bash/Python components are not rewritten by this rule — they are migrated opportunistically.

Scaffolding a new component

Copy the nearest real component and edit in place (the scaffold TEMPLATE dirs were retired — they had drifted from what real components look like):

# a new TypeScript manager — copy a small real one and rename
cp -r manager/site-manager manager/my-manager

# a new bash controller — copy a small real one and rename
cp -r controller/ap-controller controller/my-controller

Then rename the entry point, rewrite src/ (or the entry script), fill in the verb scripts, and (for a manager) the validate.sh. The dispatcher will run it on the next install/update/test with no changes above the component directory.