The control plane
Antifailure works without a control plane. af up builds an environment on the
machine it runs on, and nothing calls home.
The control plane is what a team adds when one person’s laptop stops being the right place for the answer: environments that outlive a CI job, a reviewer who wants to open one, scheduling across a queue, quotas, and history.
Everything degrades to local
Section titled “Everything degrades to local”AF-CP-001 The control plane at https://cp.example.com could not be reached. Next: Antifailure works without it. Unset control_plane.url to run fully locally.AF-CPL-003 The control plane could not be reached: dial tcp: i/o timeout Next: Environments keep working without it; events are buffered and sent when it returns.Environments keep running and teardown still works, because teardown reads the local journal rather than the control plane.
Running one
Section titled “Running one”Which tag to run
Section titled “Which tag to run”main-fa6c8aa is a published image, and the tag names the commit it was built
from. Every release publishes main-<short sha>, so a newer one is
usually available; list what exists with
curl -s "https://ghcr.io/token?scope=repository:antifailure/control-plane:pull&service=ghcr.io" \ | sed -n 's/.*"token":"\([^"]*\)".*/\1/p' \ | xargs -I{} curl -s -H "Authorization: Bearer {}" \ https://ghcr.io/v2/antifailure/control-plane/tags/listPin a main-<sha> tag, not :latest or a version tag. A sha tag names the
commit it was built from, which is what lets tools/claimcheck fail the build
if the pinned image cannot run the steps below. :latest moves, and the
:v0.1.1 image was published from a different commit than the v0.1.1 git tag,
so neither names anything checkable.
Four steps, and the order is not optional. The first two stand the server up. The second two give it a tenant and an owner, and skipping them is the mistake that makes a fresh control plane look broken: a tenant normally begins when somebody installs the GitHub App, so before that every sign-in lands with no organization, no page in the console can be reached, and nothing explains why.
# 1. Prepare the database. Applies the schema, creates the application role,# and grants it the membership that makes the schema visible to it.docker run --rm \ -e AF_MIGRATION_DATABASE_URL=postgres://owner:...@db:5432/antifailure \ -e AF_DATABASE_URL=postgres://af_app:...@db:5432/antifailure \ ghcr.io/antifailure/control-plane:main-fa6c8aa node bootstrap.mjs
# 2. Serve. Note what is absent: no migration credential, and no AF_MIGRATE.docker run \ -e AF_DATABASE_URL=postgres://af_app:...@db:5432/antifailure \ -e AF_GITHUB_CLIENT_ID=... \ -e AF_GITHUB_CLIENT_SECRET=... \ -e AF_GITHUB_REDIRECT_URI=https://cp.example.com/auth/github/callback \ -p 8080:8080 ghcr.io/antifailure/control-plane:main-fa6c8aa# 3. Create the first organization. It creates no account and grants nobody# anything, so it is not a way in on its own.node apps/api/src/backup-cli.ts create-org \ --url postgres://owner:...@db:5432/antifailure \ --org acme --name "Acme" --github-login acme
# 4. Sign in at https://cp.example.com so your account exists, then make# yourself the owner. It writes an audit entry saying a break-glass was used.node apps/api/src/backup-cli.ts break-glass \ --url postgres://owner:...@db:5432/antifailure \ --org acme --github-login you --role owner --reason "first owner"Both run inside the image, from its working directory, so reach them with
docker exec on the running container or docker run --rm ... node apps/api/src/backup-cli.ts. Installed on a host they are on the path as
af-control-plane-backup, which is how the operations
page writes break-glass.
Step 3 prints what it did:
organization acme (its new uuid)name Acmegithub acmeaudit entry 1
The organization exists and has no members, which grants nobody anything.Sign in through GitHub so your account exists, then:
af-control-plane-backup break-glass --url <admin> --org acme \ --github-login <your login> --role owner --reason "first owner"Steps 3 and 4 take the connection string step 1 used, not the one step 2 serves with. The application role is subject to the row level policies and can neither create an organization nor read across tenants, which is the point of it.
Step 3 is only for the case where no GitHub App is installed yet. Naming
--github-login the account you will later install the App on makes that
installation adopt this organization instead of creating a second one beside it,
because the App derives an organization’s slug from the account’s login.
--dry-run on either reports what would change and writes nothing. Both are
idempotent: running create-org again reports the organization that is already
there and leaves its name and GitHub login exactly as they are, in case an
installation has adopted it since.
Every variable it reads is in the configuration reference, including retention and the schema maintenance that keeps the events table partitioned.
On Kubernetes, use the chart in deploy/helm/antifailure-control-plane, which
runs step 1 as a Job before the Deployment rolls. It installs on any conformant
cluster and is developed against kind.
The Job runs step 1 only, so run step 3 once by hand against the pod:
kubectl exec "$(kubectl get deploy -l app.kubernetes.io/name=antifailure-control-plane -o name)" \ -- node apps/api/src/backup-cli.ts create-org \ --url postgres://owner:...@db:5432/antifailure \ --org acme --name "Acme" --github-login acmeIts values.yaml names every setting on the reference page that an installation
is meant to choose, with the argument for each one written where you set it.
tools/wirecheck compares the reference page against both supported installation
routes, the Terraform module and this chart, and fails the build when either
cannot deliver a variable and no row in tools/docs/wiring-exemptions.tsv gives
a reason. Eight variables have such a row for the chart.
That check was added because the chart could not set 23 of them, including
AF_SITE_ORIGIN, and nothing said so. A missing setting does not present as a
missing setting: the operator portal answers as though the installation has no
customers, the enterprise contact form on a marketing site tells the visitor to
check their network connection, and analytics records nothing. helm install
prints which of these are off in the release it just created.
extraEnv puts anything the chart does not name into the serving container. It
is the escape hatch for a variable added faster than this chart learns it, and
it is deliberately not counted as delivery by the check above, so a new setting
still earns a named value.
Two database roles, on purpose
Section titled “Two database roles, on purpose”The application connects as an unprivileged role that cannot run DDL. A role
that can ALTER TABLE can drop the policies that isolate tenants, so the role
serving requests is deliberately not that role.
antifailure_app is not an account
Section titled “antifailure_app is not an account”Migration 0001_init.sql creates antifailure_app as NOLOGIN. It is a GROUP
role that holds the grants. Nobody can connect as it. The application
connects as a separate login role that is a member of it and owns nothing:
-- Run by the bootstrap step above. Shown here for anyone doing it by hand.CREATE ROLE af_app LOGIN PASSWORD '...' NOSUPERUSER NOCREATEDB NOCREATEROLE NOBYPASSRLS;GRANT antifailure_app TO af_app; -- after the migrations, not beforeGRANT CONNECT ON DATABASE antifailure TO af_app;The grant has to come after the migrations, because that is what creates
antifailure_app.
If you skip it, the schema migrates, the server starts, /health returns 200,
and every query fails with:
ERROR: relation "organizations" does not existA role with no USAGE on the schema is told the relation is not there rather
than that it lacks permission. Check the membership directly:
psql -c "SELECT pg_has_role('af_app', 'antifailure_app', 'MEMBER')" # expects tWhat the unprivileged role cannot do
Section titled “What the unprivileged role cannot do”Verified against a real Postgres rather than asserted:
| Attempt | Result |
|---|---|
ALTER TABLE users DISABLE ROW LEVEL SECURITY |
refused, must be owner of table users |
DROP POLICY self_or_shared_org ON users |
refused, must be owner of relation users |
CREATE TABLE ... |
refused, permission denied for schema public |
UPDATE or DELETE on audit_entries |
refused, permission denied |
ALTER ROLE af_app BYPASSRLS |
refused, needs CREATEROLE |
SELECT with no tenant set |
returns nothing, rather than everything |
Tenant isolation is row level security in Postgres rather than a WHERE clause
in the application. The suite runs every query as a second tenant and asserts it
sees none of the first’s rows, on every table, and fails if a new table appears
that nobody classified.
Connecting an engine
Section titled “Connecting an engine”An engine token is what a self-hosted engine presents. It belongs to the organization rather than to whoever made it, so it keeps working after they leave, and it carries no identity: it can send events and read an environment back, and it cannot reach a key, a member, or another token.
A job in GitHub Actions needs none of this. Give the workflow
permissions: id-token: write and point it at this control plane, which the
pull request the App opens does for you and the AF_CONTROL_PLANE repository
variable does for a file copied by hand, and the engine trades the identity GitHub signs for that job for a short-lived
credential of its own, at POST /v1/engine/token. Nothing is stored in the
repository and nothing has to be rotated. The rest of this section is for an
engine running somewhere GitHub will not vouch for it: a developer’s machine, a
self-hosted runner outside Actions, or another CI system.
Mint one from a terminal.
af login --control-plane https://cp.example.com --scope tokens.manageaf token create ciThe scope has to be asked for by name because minting produces a credential, and
the words appear on the screen where the login is approved. A token from a plain
af login cannot mint one, so a terminal credential is not a credential factory,
and neither is an engine token: only a person who is an owner or an admin right
now can mint.
af token create prints the token once and the export lines to put it in:
export AF_CONTROL_PLANE_URL=https://cp.example.comexport AF_CONTROL_PLANE_TOKEN=aft_...Only the hash is stored, so nothing can show it again. Losing it means minting
another and revoking the old one with af token rm <prefix>, which takes effect
on the next request rather than at the end of a cache window. af token list
shows every token with when each was last used, revoked ones included, because
the question it is usually asked is whether the token that stopped working is
the one you revoked.
AF-CPL-001 No control plane token is configured. Next: Create an engine token in the control plane, then set AF_CONTROL_PLANE_TOKEN. Everything except this command works without one.AF-CP-002 The control plane rejected this engine's token. Next: Create a new engine token in the control plane and set AF_CONTROL_PLANE_TOKEN to it. The old one was revoked, expired, or belongs to a different control plane.Reading an environment
Section titled “Reading an environment”AF-CPL-002 The control plane has no environment called env-pr-41. Next: Check the identifier with 'af env list', or confirm the engine that created it was sending events to this control plane.The second half is usually the answer: an engine with no token, or one pointed at a different control plane, creates environments the control plane never hears about.
Health checks
Section titled “Health checks”/health returns {"ok":true} and does not touch the database. It answers
“is this process running”, which makes it a correct liveness probe and a wrong
readiness probe: a replica that has lost its database still returns 200 and
would still be sent traffic.
So the Helm chart and the Terraform both use /health for liveness only, and a
TCP check for readiness. If you write your own probes, do the same.
The audit log
Section titled “The audit log”Append only, enforced by the grants rather than by the code: the application
role has INSERT and SELECT on it and nothing else, and UPDATE, DELETE
and TRUNCATE are explicitly revoked. Entries are hash chained, so removing one
from the middle leaves a break that anybody can detect.
Related: configuration, GitHub.