Skip to content

module-manager

The module lifecycle manager: install, update, delete, test, reconcile, and snapshot TAPPaaS modules, with scope/source classification lint and environment-aware deployment. It owns the per-module JSON config in config/ and drives the Proxmox cluster (over SSH) to provision and maintain the module's VM.

What it owns

  • Per-module config JSON in config/ (<module>.json, or <module>-<environment>.json for non-default environments), plus .orig backups used for a 3-way merge of operator edits against release updates.
  • The classification of each module: its scope, derived from its stack (ADR-022e v1.1 — the foundation stack is site-scoped: installed in mgmt, source: official unless --allow-fork, --force to delete; every other stack is environment-scoped), and its source (official | community | private | local), validated against module-fields.json. The old tier field is retired (#676): read only when a config has no stack, for one release.
  • The "kind":"module" tag stamped onto every deployed config at install time — the authoritative marker list/show use to tell a deployed module apart from the co-located state files (zones.json, site.json, …). Configs from before the tag fall back to a heuristic (any of dependsOn/provides/ moduleSource, or location before migration 0006); provider-only modules (e.g. templates, no vmid/vmname) are kept.

Standardized verbs (ADR-007) — module-manager

The module-manager CLI presents the standardized verbs on entity module (the verb-alignment front door). It is a thin orchestrator: the CONFIG-layer verbs (list/show/validate) run directly over config/*.json; the LIFECYCLE verbs delegate to the bash scripts below (which stay live until a later retire phase).

Verb Maps to Notes
module list — enumerate deployed modules (--json for the cascade)
module show <m> — one deployed config in full (--json)
module resolve <m> src/resolve.ts desired state: the config plus the schema defaults it does not declare (--json)
module drift <m> src/converge.ts that desired state vs the live guest, per service. --service cluster:vm --json prints the record a converge applies
module validate [<m>] scope/source lint all modules, or one; --allow-fork
module validate --blueprint [<m>\|<dir>]... the ADR-027 blueprint check on module source directories every registered repository's modules, the named ones, or --repo PATH (what the release boundary runs)
module add <m> install-module.sh create + provision; --instance NAME names the instance (ADR-026 D6.4)
module adopt <address> adopt-module.sh a machine that already runs becomes a module (ADR-026 D8.1); nothing on it is changed
module update <m> update-module.sh release update (snapshot + test + 3-way merge) — what the sweep runs (#655)
module modify <m> update-module.sh change the config, then converge: --set field=value (ADR-020) or --unset field (#648). Bare, it is the deprecated spelling of update. --set instance=<new> renames the instance instead (#566); --lockdown runs the module's own lockdown.sh (a satellite → the off-site vault)
module delete <m> delete-module.sh --archive (default) / --remove; --decommission (a machine)
module reconcile <m> src/inspect.ts read-only drift report (default); --apply → leaf converge (src/reconcile.ts)
module test <m> test-module.sh --deep, --runtime-only, --vmid, --zone0
module snapshot-vm <m> snapshot-vm.sh special VM op (not CRUD)

Common options: --config-dir <dir>, --json (list/show/resolve/validate), -h. The leading module entity keyword is optional (it is the only entity).

show vs resolve — show prints the deployed config verbatim; resolve prints what that config means once the module-fields.json defaults for fields it does not declare are filled in. That resolved value is what a converge actually uses, so the two differ exactly where a field is undeclared: show omits cputype, resolve reports host (marked default). There is one resolver behind it, shared by the drift report and the apply path, so the reported desired value and the applied one cannot diverge (ADR-020 D1).

Not to be confused with resolve-module.sh, which answers a different question — where a module's source directory is. list --resolution is that one's reporting front door.

Changing a field: modify --set (ADR-020)

module-manager module modify nextcloud --set cores=8 --set memory=16384
module-manager module modify nextcloud --set zone0=iot --allow-disruption   # subnet change reboots

One verb, one algorithm. Bare, modify is the release update update-tappaas already runs. With --set it writes the field into the deployed config first and then runs that same algorithm — snapshot, 3-way merge, converge, test, updateTime. There is no second apply path to keep in step with the first.

What it refuses, and when. A change is refused up front only when the schema alone can say so:

Class Example What happens
in-place cores, memory, cputype, vmtag applied live, no downtime
grow-only diskSize a grow applies; a shrink is refused at apply time
in-place-reboot zone0, bridge0 needs a guest reboot → deferred unless authorized
migrate node relocates the guest → deferred unless authorized
manual storage reported; moving a disk stays an operator action
immutable / recreate vmid, bios, image* rejected before anything is written

A --set naming an immutable field is rejected whole — if any field in one command cannot be applied, none of them are written, so config and cluster never move apart. A field none of the module's services use is also rejected: writing it would change the config and nothing else.

What a module claims about itself: version and status (#248)

A community user reads these two fields to decide whether a module is worth installing, so they mean something definite. Undercommit, overperform: a module does not claim to be more than it is.

version What it says
0.1.0 first working version — the author has run it, nobody else
0.x.y still iterating; pre-stable, and SemVer says so
1.0.0 an install other than the author's is confirmed working; breaking changes are versioned from here
status What it says Reached when
Development incomplete or untested; not ready for use —
Testing feature complete, author-tested in a live install; wider validation wanted the author has run test-module.sh against a live install
Production validated beyond its author at least one independent install is confirmed working
Deprecated no longer maintained —

The two move together: the step that takes status to Production is the step that earns 1.0.0. archived and external are not maturities — they are states a deployed config carries (ADR-022g moves them to management), and a released module naming one fails the lint.

module-manager validate (and every install) checks this: an unknown status, or a deployment state in a released module, fails. A version that is not SemVer, a missing claim, and the two contradictions — 1.0.0+ while Development, Production below 1.0.0 — warn, because which side is wrong is the author's to say.

Keys beginning _ (_README, _note, _comment) are documentation for whoever reads the file next; no tool reads them, and the schema check passes over them.

validate also reports a field nothing will read (#683). module-fields.json records which provider:service consumes each field, so a module can declare one and wire no consumer: ports is owned by network:rules, and a module that depends on network:proxy but not network:rules gets no ingress validation and no firewall rules from it. It is a warning, not an error — the field may be honest documentation of what a proxy-published module listens on internally — and it fires only where the author can act. Three cases are deliberately silent: a consumer wired through integratesWith counts as wired; a field whose usedBy includes general belongs to no one service; and a self-consumed field, where the coordinate's service is one this module provides, is correct as it stands (network declares ports and provides rules, and must not depend on itself).

This used to be a warning from the three-way merge instead, which meant an operator saw it on every update of an affected module, in a sweep log, about a module-authoring decision they could neither act on nor silence. The merge still classifies the field the same way and still keeps it at top level; it just no longer says so out loud (TAPPAAS_DEBUG=1 if you want the per-merge note).

validate --blueprint — is the module made of what a module is made of? (ADR-027)

Plain validate lints deployed configs. --blueprint reads a module's source directory and checks the set ADR-027 defines — a missing MUST is an error, a missing SHOULD a warning, the same for every module whatever its stack or source:

Where MUST (error) SHOULD (warning)
module <module>.json, install.sh, update.sh, README.md and INSTALL.md (ADR-013); fields.json when the module declares a field no schema defines; services/<s>/ for every service in provides test.sh
services/<s>/ install-service.sh, update-service.sh, delete-service.sh, README.md; fields.json when a field names <module>:<s> in usedBy test-service.sh

A services/<s>/ for a service the module does not declare in provides is a warning: it is dead, or the declaration is missing — and a consumer's dependency on an undeclared service is refused at install. Test fixtures (stack: test) are not modules and are skipped; 00-Template is checked like any module, so the template cannot drift from the rule. A module whose fields have their own schemas/<module>-fields.json (the satellite) is not asked for a fields.json.

module-manager validate --blueprint                       # every module of every registered repository
module-manager validate --blueprint ~/TAPPaaS/src/apps/myapp
module-manager validate --blueprint --repo ~/TAPPaaS      # one repository's catalogue — what the release boundary runs

It exits 1 on any error. It never runs at install time: the person installing a module is not the person who can fix it, so an incomplete module still installs. It runs at every release boundary instead (tappaas-train, once a fortnight, on the commit about to be promoted), together with the check that every service README's generated field section matches its fields.json: a missing MUST stops the boundary, and the result is recorded in release-train.json (RELEASE-TRAIN.md).

Renaming an instance: modify --set instance=<new> (#566)

module-manager module modify hassanova --set instance=hass-delbuschy --dry-run   # what would move
module-manager module modify hassanova --set instance=hass-delbuschy

The instance name is its config file's name (ADR-026 D6.1); the module it is an instance of is named by moduleSource (D6.2). So a rename moves config/<old>.json, its merge baseline .orig and its .meta.json, and repoints anything that named the instance as its Host (node). Nothing on the guest changes: it keeps its own name (vmname), and with it its DNS record, firewall alias and proxy upstream — renaming the guest is a separate, disruptive change this verb does not make. The backup job follows the vmid and is unaffected. --dry-run lists what would move and changes nothing; it is the only verb that has one, because every other change converges the module rather than rewriting names.

It runs alone (no other --set in the same call) and refuses a name that is not a usable instance name or is already taken. A machine is refused: it is named after the host it is (ADR-026 D8), so rename the host and re-register it with module-manager module adopt <address> --instance <new>.

A name off the <module>-<environment> convention is not a fault — the convention is the default when no name is given, and such an instance updates like any other (the release source is the module's own JSON, D6.3). This verb is for when the operator wants the convention back.

A one-way change of management: modify --lockdown (ADR-010 §8.4.4)

module-manager module modify satellite --lockdown

Runs the module's own lockdown.sh <instance>, from its module directory, and nothing else — no --set in the same call, no converge afterwards (a locked-down machine no longer admits the mothership). For a satellite it makes the unmanaged off-site vault: a read-only login on the Site's PBS, the pull, self-patching, and the mothership's key removed last. A module without a lockdown.sh has nothing to lock down and is refused.

Removing a stale field: modify --unset (#648)

module-manager module modify nextcloud --unset legacyField

For the one field nothing else can remove: present in the deployed config, absent from the release source and from <module>.json.orig, and undeclared — the 3-way merge keeps it by rule 2b and warns field '<f>' is not in the schema — kept at top level on every single update. That rule exists to protect operator-added fields, so the merge cannot tell this case apart; saying --unset is how an operator does.

Two refusals, both because the removal would not mean what it looks like:

The field is… Why it is refused
declared in module-fields.json every reader expects it — change it with --set field=value
present in <module>.json.orig the release still defines it, so the next merge re-adopts it — remove it in the source

--unset is gated before any --set in the same command is written, so one modify still applies the whole change or none of it.

--force vs rebootOk — three levers that no longer collide

Some changes need downtime. Whether we are allowed to cause it is a separate question from whether the change needs it, and it has exactly two answers:

  • module modify <m> --force — an operator, now.
  • rebootOk: true on the module, honoured only inside the unattended sweep, and only because the site already accepts downtime in that window (automaticReboot). Default false: silence never authorizes a reboot.

--force is neither. It proceeds past a refusal — a fatally failed pre-update test, an archived/external module — and never reboots anything; site-manager update --force forwards exactly that to every module. Downtime for a whole run is site-manager update --allow-disruption, which still honours rebootOk (ADR-020 v0.10 D8). update-tappaas --force is deprecated (ADR-017 D5).

When a disruptive change is not authorized the converge applies everything else, prints a machine-parseable line, and still exits 0 — not applying a change is not a failure:

⚠ DEFERRED: nextcloud net0 needs a disruptive change (reboot/offline migrate) that is not authorized
  Apply in a maintenance window:  module-manager module update nextcloud --allow-disruption

update-tappaas collects those and ends the sweep with one summary of what is still pending.

reconcile vs drift — both compare declared state with reality, for different readers. reconcile <m> is the operator's three-way report (Released[git] / Desired[~/config] / Actual), field by field, plus the dependency-service section. drift <m> is the two-way record a CONVERGE acts on: which apply unit each change belongs to, what class it is, which hook takes it, and what side effects it drags along. drift --service <p:s> --json is literally the input to update-service.sh --apply-drift, so what you read is what would be applied — there is one differ behind both (ADR-020 D7).

reconcile vs modify — reconcile --apply re-applies the existing config (idempotent converge: each dependency's update-service.sh + the module's own update.sh/install.sh, all run from the module directory), with no snapshot, no tests, no 3-way merge, and no updateTime bump. modify (update-module.sh) changes the config via a release update, then performs the same apply by delegating to reconcile --apply, wrapped in snapshot + pre/post tests + rollback. reconcile is the leaf the site/environment reconcile --deep cascade walks down to.

Service contract — services/<svc>/update-service.sh is the converge for an already-installed module and every service must ship one (enforced by test.sh). install-service.sh is create-only prerequisites; where a service has no create-only work it simply execs update-service.sh. There is no fallback from one to the other: install-service.sh has create semantics (cluster:vm's calls Create-TAPPaaS-VM.sh, which refuses an existing VMID), which is why reconcile previously failed on every VM-backed module.

What reconcile <m> (no --apply) reports — a read-only drift report in two parts:

  1. Config fields — three-way Released[git] / Desired[~/config] / Actual[running VM]; for a module with no vmid the Actual column is N/A and the diff degrades to Released-vs-Desired.
  2. Dependency-service state — for each dependsOn entry, that provider's read-only services/<service>/test-service.sh <module> (the same verifier module test runs): declared firewall rules, NAT rules, discovery relays. For a policy-only module (no VM) this is the whole module, so without it a clean field diff said nothing. --no-services skips it.

Detected drift exits 0 — this is a report, and list --diff plus the --deep cascade propagate the rc. A check that could not run (missing or non-executable test-service.sh) exits 1: unknown state is not clean. A provider that ships no test-service.sh is reported as NOT checked, never as passing.

The service checks cost one child process — usually one firewall API round-trip — per dependency, so they are on for a single reconcile <m> and off for the fleet/cascade paths: list --diff needs --services to include them (and says so when it does not), and the environment reconcile preview passes --no-services.

module-manager module list
module-manager module show nextcloud --json
module-manager module validate --allow-fork
module-manager validate --blueprint --repo ~/TAPPaaS
module-manager module add nextcloud --environment acme
module-manager module reconcile nextcloud                # report: fields + dependency services
module-manager module reconcile nextcloud --no-services   # report: fields only
module-manager module reconcile nextcloud --apply        # converge to the current config
module-manager module list --diff --services             # fleet rollup, services included

Underlying scripts

All bash, linked onto PATH by install.sh. These remain the source of truth (the manager verbs orchestrate them) until a later retire phase.

install-module.sh — install a module

install-module.sh <module-name> [--environment <name>] [--instance <name>]
                  [--allow-fork] [--reinstall] [--<field> <value>]...
  • --environment <name> — target environment (sets the VM name and zone; default env → <module>, otherwise <module>-<env>). --variant <name> is a deprecated alias.
  • --instance <name> — name the instance (ADR-026 D6.4): the config is config/<name>.json instead of the default <module>[-<env>].json, and a VM, if the module deploys one, is named after it. The module stays identified by the config's .moduleSource, so tools ask which module with module_of, never by parsing the name — three cluster nodes are three instances of one module. The name must be a DNS label and not one config/ already uses (site, zones, peer configs…).
  • --allow-fork — permit a site-scoped module (the foundation stack) from a non-official source.
  • --reinstall — delete then install: the only way to replace a deployed config (recovers a failed partial install too). There is no --force: an already deployed module is taken forward with module update <m> (#453).
  • --no-rollback — keep a failed install in place for inspection. By default an add that fails removes what that run created — the deployment if it created the VM, otherwise the config it wrote — so nothing half-installed is left behind for the next update to trip over (#584).
  • --<field> <value> — override any module JSON field.
install-module.sh nextcloud
install-module.sh nextcloud --environment acme
install-module.sh pvehost --instance tappaas2      # one of several instances of a module

update-module.sh — update a module

update-module.sh [options] <module-name>
  • --environment <name> — resolve the installed config name (deprecated alias --variant).
  • --force — proceed although something says no: a pre-update test that failed fatally (exit 2, #635), or a module whose status is archived/external. It never reboots and never overwrites the deployed config (ADR-020 v0.10 D8).
  • --allow-disruption — authorize downtime for this module now: a reboot or an offline migrate. Without it such a change is deferred, not applied.
  • --no-snapshot — skip the pre-update snapshot / rollback.
  • --debug, --silent.

It snapshots the VM, tests, updates, and rolls back on a fatal failure. The Step 2 pre-update test runs --runtime-only: it gates a mutation, so it asks whether the module is healthy, not whether the source tree is correct (#595).

The gate follows the suite's own grading (#635). test.sh exits 2 for a fatal failure and 1 for failed assertions; the gate honours both. Exit 2 aborts the update (unless --force); exit 1 warns and the update proceeds, and those failures become the baseline for Step 6: the post-update test fails the module only on checks that were not already failing, and otherwise prints one machine-readable TEST-WARN: line that update-tappaas collects into the sweep summary and last-update-result.json (test_warnings). A module whose update did not break anything is no longer failed by five unrelated DNS assertions it inherited.

delete-module.sh — delete a module

delete-module.sh <module-name> [--archive|--remove] [--vmid <id>] [--decommission]
                 [--environment <name>] [--yes|-y] [--force]
  • --archive (default) keeps the config; --remove deletes it.
  • --vmid <id> — target a specific VMID.
  • --environment <name> (alias --variant).
  • --yes / -y — skip the confirmation prompt.
  • --force — skip dependency checks; required for site-scoped modules (the foundation stack).
  • A machine (kind: machine) is only unregistered: the config goes, the machine keeps running untouched, --vmid is refused, and the module's own delete.sh is not run (ADR-026 D8.1).
  • --decommission — a machine only: run its module's delete.sh first, which takes the Site's side of it down (a satellite: the OPNsense peer, tunnel server and edge rules), then unregister. A failed decommission keeps the instance registered. The machine itself is still never touched (ADR-010 §8.4.5). Refused for a module that is not a machine.

adopt-module.sh — a running machine becomes a module

adopt-module.sh <address> [--instance NAME] [--zone ZONE] [--wait SECONDS]

<address> is an IP address or a DNS name. It reaches the machine as root with the mothership's key; if that fails, it prints the one command that authorises the key and waits (--wait, default 300s, one try every 20s). It then reads the hostname and /etc/os-release, picks the module for that OS (debian → debianhost, and a Debian machine running Proxmox VE → pvehost, #665; any other OS is refused), names the instance after the hostname, takes zone0 from the active zone whose subnet holds the address, and calls install-module.sh <module> --instance … --address … --zone0 … --os …. Adopting the same machine again changes nothing.

Refused, with nothing written: a key that never works, an OS with no machine module, a Proxmox VE host that is not in this Site's cluster (site.json hardware.nodes — join it with site-manager node add), an instance name another machine has (--instance), an address another instance already has, and an address in no active zone (--zone). It never changes how the machine is reached: key-only SSH is a separate step.

test-module.sh — run a module's tests

test-module.sh [--deep] [--runtime-only] [--vmid <id>] [--zone0 <zone>] <module-name>

--runtime-only skips a module's source-tree checks (offline unit suites, generated-doc drift, grep guards over the repo) and runs only what says something about the live system. update-module.sh passes it for the Step 2 pre-update test, so a source-tree defect cannot abort the update of a healthy module (#595). --deep runs everything regardless.

test-module.sh openwebui
test-module.sh --deep litellm
test-module.sh --runtime-only tappaas-cicd   # what the pre-update gate runs

snapshot-vm.sh — manage a module's VM snapshots

snapshot-vm.sh <module-name> [--list | --cleanup <N> | --restore <N>]

No action = create a snapshot. --list lists; --cleanup <N> keeps the newest N tappaas-* snapshots (hand-made ones are never pruned) and exits non-zero if a delete fails; --restore <N> restores N steps back (1 = most recent).

copy-update-json.sh — copy/normalize a module JSON into config

copy-update-json.sh <module-name> [--variant <name>] [--environment <name>]
                    [--default-environment <name>] [--vmname <v>] [--vmid <v>]
                    [--zone0 <v>] [--proxyDomain <v>] [--<field> <value>]...

Applies environment defaults (vmname suffix, auto-incremented vmid, zone from the environment), validates fields against the schema, and writes canonical config-block form.

module-format.sh — convert JSON form

module-format.sh <to-flat|to-config> <file.json> [--in-place]

validate-module-tier-source.sh — scope/source lint

validate-module-tier-source.sh [--allow-fork] [--quiet] <module.json>

A site-scoped module (stack: foundation) requires source:official (override with --allow-fork); an invalid source is rejected; source:community warns. A module without a stack is warned about and installs environment-scoped. A tier is retired (#676): its presence is warned about, and one that contradicts the stack is named as ignored. The file keeps its historical name. Used standalone and at install time.