Skip to content

Type to search pages.

View .md

The control plane

Antifailure works without a control plane. af up builds an environment on the machine it runs on, and nothing calls home.

The control plane is what a team adds when one person’s laptop stops being the right place for the answer: environments that outlive a CI job, a reviewer who wants to open one, scheduling across a queue, quotas, and history.

AF-CP-001 The control plane at https://cp.example.com could not be reached.
Next: Antifailure works without it. Unset control_plane.url to run fully
locally.
AF-CPL-003 The control plane could not be reached: dial tcp: i/o timeout
Next: Environments keep working without it; events are buffered and sent when
it returns.

Environments keep running and teardown still works, because teardown reads the local journal rather than the control plane.

main-fa6c8aa is a published image, and the tag names the commit it was built from. Every release publishes main-<short sha>, so a newer one is usually available; list what exists with

Terminal window
curl -s "https://ghcr.io/token?scope=repository:antifailure/control-plane:pull&service=ghcr.io" \
| sed -n 's/.*"token":"\([^"]*\)".*/\1/p' \
| xargs -I{} curl -s -H "Authorization: Bearer {}" \
https://ghcr.io/v2/antifailure/control-plane/tags/list

Pin a main-<sha> tag, not :latest or a version tag. A sha tag names the commit it was built from, which is what lets tools/claimcheck fail the build if the pinned image cannot run the steps below. :latest moves, and the :v0.1.1 image was published from a different commit than the v0.1.1 git tag, so neither names anything checkable.

Four steps, and the order is not optional. The first two stand the server up. The second two give it a tenant and an owner, and skipping them is the mistake that makes a fresh control plane look broken: a tenant normally begins when somebody installs the GitHub App, so before that every sign-in lands with no organization, no page in the console can be reached, and nothing explains why.

Terminal window
# 1. Prepare the database. Applies the schema, creates the application role,
# and grants it the membership that makes the schema visible to it.
docker run --rm \
-e AF_MIGRATION_DATABASE_URL=postgres://owner:...@db:5432/antifailure \
-e AF_DATABASE_URL=postgres://af_app:...@db:5432/antifailure \
ghcr.io/antifailure/control-plane:main-fa6c8aa node bootstrap.mjs
# 2. Serve. Note what is absent: no migration credential, and no AF_MIGRATE.
docker run \
-e AF_DATABASE_URL=postgres://af_app:...@db:5432/antifailure \
-e AF_GITHUB_CLIENT_ID=... \
-e AF_GITHUB_CLIENT_SECRET=... \
-e AF_GITHUB_REDIRECT_URI=https://cp.example.com/auth/github/callback \
-p 8080:8080 ghcr.io/antifailure/control-plane:main-fa6c8aa
Terminal window
# 3. Create the first organization. It creates no account and grants nobody
# anything, so it is not a way in on its own.
node apps/api/src/backup-cli.ts create-org \
--url postgres://owner:...@db:5432/antifailure \
--org acme --name "Acme" --github-login acme
# 4. Sign in at https://cp.example.com so your account exists, then make
# yourself the owner. It writes an audit entry saying a break-glass was used.
node apps/api/src/backup-cli.ts break-glass \
--url postgres://owner:...@db:5432/antifailure \
--org acme --github-login you --role owner --reason "first owner"

Both run inside the image, from its working directory, so reach them with docker exec on the running container or docker run --rm ... node apps/api/src/backup-cli.ts. Installed on a host they are on the path as af-control-plane-backup, which is how the operations page writes break-glass.

Step 3 prints what it did:

organization acme (its new uuid)
name Acme
github acme
audit entry 1
The organization exists and has no members, which grants nobody anything.
Sign in through GitHub so your account exists, then:
af-control-plane-backup break-glass --url <admin> --org acme \
--github-login <your login> --role owner --reason "first owner"

Steps 3 and 4 take the connection string step 1 used, not the one step 2 serves with. The application role is subject to the row level policies and can neither create an organization nor read across tenants, which is the point of it.

Step 3 is only for the case where no GitHub App is installed yet. Naming --github-login the account you will later install the App on makes that installation adopt this organization instead of creating a second one beside it, because the App derives an organization’s slug from the account’s login.

--dry-run on either reports what would change and writes nothing. Both are idempotent: running create-org again reports the organization that is already there and leaves its name and GitHub login exactly as they are, in case an installation has adopted it since.

Every variable it reads is in the configuration reference, including retention and the schema maintenance that keeps the events table partitioned.

On Kubernetes, use the chart in deploy/helm/antifailure-control-plane, which runs step 1 as a Job before the Deployment rolls. It installs on any conformant cluster and is developed against kind.

The Job runs step 1 only, so run step 3 once by hand against the pod:

Terminal window
kubectl exec "$(kubectl get deploy -l app.kubernetes.io/name=antifailure-control-plane -o name)" \
-- node apps/api/src/backup-cli.ts create-org \
--url postgres://owner:...@db:5432/antifailure \
--org acme --name "Acme" --github-login acme

Its values.yaml names every setting on the reference page that an installation is meant to choose, with the argument for each one written where you set it. tools/wirecheck compares the reference page against both supported installation routes, the Terraform module and this chart, and fails the build when either cannot deliver a variable and no row in tools/docs/wiring-exemptions.tsv gives a reason. Eight variables have such a row for the chart.

That check was added because the chart could not set 23 of them, including AF_SITE_ORIGIN, and nothing said so. A missing setting does not present as a missing setting: the operator portal answers as though the installation has no customers, the enterprise contact form on a marketing site tells the visitor to check their network connection, and analytics records nothing. helm install prints which of these are off in the release it just created.

extraEnv puts anything the chart does not name into the serving container. It is the escape hatch for a variable added faster than this chart learns it, and it is deliberately not counted as delivery by the check above, so a new setting still earns a named value.

The application connects as an unprivileged role that cannot run DDL. A role that can ALTER TABLE can drop the policies that isolate tenants, so the role serving requests is deliberately not that role.

Migration 0001_init.sql creates antifailure_app as NOLOGIN. It is a GROUP role that holds the grants. Nobody can connect as it. The application connects as a separate login role that is a member of it and owns nothing:

-- Run by the bootstrap step above. Shown here for anyone doing it by hand.
CREATE ROLE af_app LOGIN PASSWORD '...' NOSUPERUSER NOCREATEDB NOCREATEROLE NOBYPASSRLS;
GRANT antifailure_app TO af_app; -- after the migrations, not before
GRANT CONNECT ON DATABASE antifailure TO af_app;

The grant has to come after the migrations, because that is what creates antifailure_app.

If you skip it, the schema migrates, the server starts, /health returns 200, and every query fails with:

ERROR: relation "organizations" does not exist

A role with no USAGE on the schema is told the relation is not there rather than that it lacks permission. Check the membership directly:

Terminal window
psql -c "SELECT pg_has_role('af_app', 'antifailure_app', 'MEMBER')" # expects t

Verified against a real Postgres rather than asserted:

Attempt Result
ALTER TABLE users DISABLE ROW LEVEL SECURITY refused, must be owner of table users
DROP POLICY self_or_shared_org ON users refused, must be owner of relation users
CREATE TABLE ... refused, permission denied for schema public
UPDATE or DELETE on audit_entries refused, permission denied
ALTER ROLE af_app BYPASSRLS refused, needs CREATEROLE
SELECT with no tenant set returns nothing, rather than everything

Tenant isolation is row level security in Postgres rather than a WHERE clause in the application. The suite runs every query as a second tenant and asserts it sees none of the first’s rows, on every table, and fails if a new table appears that nobody classified.

An engine token is what a self-hosted engine presents. It belongs to the organization rather than to whoever made it, so it keeps working after they leave, and it carries no identity: it can send events and read an environment back, and it cannot reach a key, a member, or another token.

A job in GitHub Actions needs none of this. Give the workflow permissions: id-token: write and point it at this control plane, which the pull request the App opens does for you and the AF_CONTROL_PLANE repository variable does for a file copied by hand, and the engine trades the identity GitHub signs for that job for a short-lived credential of its own, at POST /v1/engine/token. Nothing is stored in the repository and nothing has to be rotated. The rest of this section is for an engine running somewhere GitHub will not vouch for it: a developer’s machine, a self-hosted runner outside Actions, or another CI system.

Mint one from a terminal.

Terminal window
af login --control-plane https://cp.example.com --scope tokens.manage
af token create ci

The scope has to be asked for by name because minting produces a credential, and the words appear on the screen where the login is approved. A token from a plain af login cannot mint one, so a terminal credential is not a credential factory, and neither is an engine token: only a person who is an owner or an admin right now can mint.

af token create prints the token once and the export lines to put it in:

Terminal window
export AF_CONTROL_PLANE_URL=https://cp.example.com
export AF_CONTROL_PLANE_TOKEN=aft_...

Only the hash is stored, so nothing can show it again. Losing it means minting another and revoking the old one with af token rm <prefix>, which takes effect on the next request rather than at the end of a cache window. af token list shows every token with when each was last used, revoked ones included, because the question it is usually asked is whether the token that stopped working is the one you revoked.

AF-CPL-001 No control plane token is configured.
Next: Create an engine token in the control plane, then set
AF_CONTROL_PLANE_TOKEN. Everything except this command works without one.
AF-CP-002 The control plane rejected this engine's token.
Next: Create a new engine token in the control plane and set
AF_CONTROL_PLANE_TOKEN to it. The old one was revoked, expired, or belongs to
a different control plane.
AF-CPL-002 The control plane has no environment called env-pr-41.
Next: Check the identifier with 'af env list', or confirm the engine that
created it was sending events to this control plane.

The second half is usually the answer: an engine with no token, or one pointed at a different control plane, creates environments the control plane never hears about.

/health returns {"ok":true} and does not touch the database. It answers “is this process running”, which makes it a correct liveness probe and a wrong readiness probe: a replica that has lost its database still returns 200 and would still be sent traffic.

So the Helm chart and the Terraform both use /health for liveness only, and a TCP check for readiness. If you write your own probes, do the same.

Append only, enforced by the grants rather than by the code: the application role has INSERT and SELECT on it and nothing else, and UPDATE, DELETE and TRUNCATE are explicitly revoked. Entries are hash chained, so removing one from the middle leaves a break that anybody can detect.

Related: configuration, GitHub.