Skip to content

Type to search pages.

View .md

Control plane configuration

The control plane reads its configuration from the environment and refuses to start without what it needs, naming the variable that is missing. A process that starts with a missing secret and fails on the first request that needs it is a process that fails in production rather than at deploy time.

Every variable on this page can be set by the deploy paths this project ships, and that is checked rather than asserted: tools/wirecheck fails a build when a variable documented here has no env block in the Terraform module and no row in tools/docs/wiring-exemptions.tsv saying why it cannot have one. It was written because the two checks that already covered this ground both proved a variable was DOCUMENTED, which a variable nothing could deliver satisfies perfectly. Standing up production has the order for the four features whose credential Terraform must not hold.

Variable What it is
AF_DATABASE_URL The connection string the application uses. This is the unprivileged role, not the owner: it cannot run DDL, because a role that can ALTER TABLE can drop the policies that isolate tenants.
AF_GITHUB_CLIENT_ID The OAuth App’s client identifier.
AF_GITHUB_CLIENT_SECRET The OAuth App’s client secret.
AF_GITHUB_REDIRECT_URI Where GitHub returns the browser after sign in. Must match what the App is configured with exactly.
Variable Default What it does
AF_PORT 8080 The port to listen on.
AF_POOL_MAX 10 Connections in the application pool.
AF_APP_BASE_URL unset The public origin, used to build absolute links.
AF_ADMIN_DATABASE_URL unset The connection string the operator portal uses, and the only credential on this instance that can read across tenants. A second credential rather than a second setting on the first: its role holds BYPASSRLS, which is how an operator reads across tenants, and the application’s role must never hold it, because a different credential is something the application cannot be granted its way into where a privilege is something it can. Optional rather than required, and the process says which at startup: unset means this installation has no operator portal, which is the right default for a single team, and /admin then refuses every request naming this variable rather than answering an empty list that reads like a platform with no customers on it.
AF_ADMIN_POOL_MAX 4 Connections in the operator pool. Small on purpose: it serves a handful of operators rather than customer traffic.
AF_SIGNIN_ALLOWLIST unset GitHub logins, comma or whitespace separated, that may sign in. Unset means any GitHub account may sign in, which is the default and is what Antifailure’s own hosted plane runs: it is the right answer both for an installation whose network already decides who reaches it, and for a product people sign themselves up to. Set but empty means nobody, not everybody: a deployment that lost this value should close, not open. The mode is printed at startup. To close sign-ups on a self-hosted installation, name the logins here, or set it to an empty string to admit nobody at all.
AF_SELF_SERVE_SIGNUP unset Set to 1 to give somebody who signs in with no organization one of their own, on the free plan, owned by them. Unset, that person lands in no organization and waits for a GitHub App installation or an invitation, which is what happened before this existed. Off by default, and the direction is the argument rather than an opinion about convenience: what it grants is a tenant with real quotas and real compute against them, so on an installation where AF_SIGNIN_ALLOWLIST is unset it grants that to anybody who can reach the address, and forgetting the variable has to close the door rather than open it. The organization is named after the GitHub account and carries its login, so installing the App on that account later adopts the same organization rather than creating a second one beside it. The mode is printed at startup next to the allowlist’s, because the two are one sentence: who may sign in, and whether there is anything on the other side of the door. Any value other than 1, 0, true, false or unset stops the process.
AF_INSECURE_COOKIES unset Set to 1 to drop the Secure attribute from cookies. For local development over plain HTTP and nothing else.
AF_TRUSTED_PROXY_HOPS 1 How many proxies every request passes through before it reaches the process, which decides which entry of X-Forwarded-For is believed. Every proxy appends the peer address it saw to the end of that header and leaves whatever the caller sent in front, so the entries a deployment can trust are the last ones, one per proxy, and the client is the entry this many places from the end. 1 is the Azure Container Apps ingress alone, which is what the Terraform module builds, and one ingress controller, which is what the Helm chart assumes; Microsoft documents that only the rightmost entry is provided by Container Apps and everything else must be treated as the caller’s. Set 2 when a Front Door, an Application Gateway or a WAF that also appends to the header sits in front of the ingress. It must be the number of proxies that every request passes through: a proxy some requests can skip is not a trusted hop, because a caller who reaches the inner one directly gets to write the entry this count attributes to the outer one. The address chosen here keys the sign-in, magic link, device code and OAuth callback rate limits and is what the sign-in audit trail records, so a count that is too high hands every caller their own limit and lets them write their own audit entry; a count that is too low limits everybody behind the same outer proxy together. When the header is absent, as it is for a direct connection or the local twin, or when the chosen entry is not an address, the request is limited in one shared bucket rather than exempted, and nothing is recorded as its address. A value that is not a whole number from 1 to 16 stops the process at startup.
AF_MIGRATE unset Set to 1 to apply migrations at startup. Requires AF_MIGRATION_DATABASE_URL.
AF_MIGRATION_DATABASE_URL unset A connection string for a role that may run DDL.
AF_VERSION dev The build’s version, reported by /readyz. Stamped into the image at build time; setting it by hand only makes the endpoint lie.
AF_COMMIT unknown The commit the build came from, reported by /readyz. Stamped the same way.
AF_GITHUB_APP_ID unset The numeric App ID from the GitHub App’s settings page. Needed together with the private key and the webhook secret; setting some and not others stops the process at startup rather than producing a half-working App.
AF_GITHUB_APP_PRIVATE_KEY unset The PEM GitHub generated when the App’s private key was created, or that PEM base64 encoded. Literal \n sequences are turned back into newlines, because most ways of getting a multi-line value into a container flatten it, and the resulting key fails with a message about DECODER routines that sends you somewhere else entirely.
AF_GITHUB_APP_WEBHOOK_SECRET unset The webhook secret set on the App. Every delivery is verified against it before its body is parsed. Unset means /webhooks/github answers 503 rather than accepting unsigned deliveries.
AF_GITHUB_APP_INSTALL_URL unset The public https://github.com/apps/<slug>/installations/new address. When it is set, a person who signs in without an organization gets an Install the GitHub App action. When it is unset they are told the address has not been configured and are still offered Check my GitHub membership, which never depended on it. Either way the startup log says which. Any other origin or path, or a value that is not a URL, stops the process at startup.
AF_SIGNUP_URL unset Where somebody AF_SIGNIN_ALLOWLIST refuses is sent instead. A refused sign-in renders a page rather than a JSON body, and when this is set that page carries one link to it. Unset is the self-hosted default and means the page offers no link: an operator running an allowlist has their own way of being asked, and pointing their users at somebody else’s contact page would be wrong. Never rendered at all when the allowlist is unset, because then nobody is refused. Must be an absolute http or https address, or the process stops at startup.
AF_SITE_ORIGIN unset Every browser origin allowed to post to the routes a page on the marketing site calls: POST /v1/leads, POST /v1/applications and POST /v1/site/events. One whole origin such as https://example.com, or several separated by commas, such as https://example.com,https://www.example.com. A site served on both an apex and a www hostname needs both, because the browser sends the hostname the visitor is standing on and the comparison is exact. These are the only routes on the server that answer a cross-origin browser, and this is the only variable that widens them. Unset means no other origin may post, so a contact form on a separate marketing host cannot submit and reports a network error; the routes still answer curl and a page on this origin. Never a wildcard: there is no value meaning “any origin”. A value carrying a path, a query or a fragment stops the process, because a browser sends only scheme, host and port and such a value could never match, which would allow nobody while looking configured. An empty entry, from a stray comma, stops it too.
AF_LEAD_NOTIFY_EMAIL unset Where an enterprise lead is announced. Unset means leads are recorded and nobody is mailed, which the startup log says and which the form itself tells the person who filled it in. Setting it without a mailer, meaning AF_RESEND_API_KEY and AF_MAIL_FROM, is called out at startup as its own state: that deployment believes it is announcing leads and cannot. Read the queue in either case with af-control-plane-backup leads.
AF_GITHUB_API_BASE https://api.github.com Where the GitHub API lives. For GitHub Enterprise Server, and for tests.
AF_MODEL_PRICES unset What a model costs, as model=input/output in US dollars per million tokens, comma separated: claude-sonnet-5=2/10,gpt-4.1=2/8. Adds to the built-in defaults rather than replacing them. A model with no price is refused rather than charged nothing, because a request that spends money and adds nothing to the total is a spend cap that does not cap spending. A malformed entry stops the process at startup rather than being skipped, since a skipped entry is a model silently falling back to another price.
AF_PROVIDER_KEY_SECRET unset 32 bytes of base64, the secret that seals customers’ Anthropic and OpenAI keys, and the sealing key for version v1. Generate one with openssl rand -base64 32. Unset means keys cannot be stored at all: saving one is refused rather than written in the clear. It must not live in the same place as the database, or a database dump carries both halves. Anything other than 32 bytes stops the process at startup rather than failing later on the one action the feature exists for. Anything that is not canonical base64 stops it too, because Buffer decoding drops characters it does not recognise and a truncated paste would otherwise decode to a short key. On its own it is the whole configuration and no other variable here is needed.
AF_PROVIDER_KEY_SECRETS unset More sealing keys, as v2=<32 bytes of base64>, comma separated, in the same identifier=key grammar as AF_LICENSE_PUBLIC_KEYS and for the same reason: something holding exactly one key cannot rotate without invalidating everything in the field. Merged with AF_PROVIDER_KEY_SECRET rather than replacing it, so a rotation adds one new value and never has to read the old one back out of a vault to compose a combined string. Every key named here can OPEN a stored credential; which one new credentials are sealed under is AF_PROVIDER_KEY_VERSION. Two different keys under one version stops the process, because rows filed under that version were sealed with one of them and there is no safe choice between them. A version is up to 32 characters of lower case letters, digits, dot, dash or underscore. The start-up log prints the versions held, which is the only way to confirm a new revision picked a new key up without decrypting somebody’s credential.
AF_PROVIDER_KEY_VERSION the single configured version Which sealing key version new provider keys are sealed under. Optional while exactly one key is configured, which is every installation that has not rotated. With several configured it is required: the process stops at startup naming the versions it holds, rather than guessing which of somebody else’s keys to seal their credential with. A version nothing is configured for stops it as well. Rotating is: add the new key, set this to it, deploy, re-seal with af-control-plane-backup reseal, then remove the old key. See rotating secrets.
AF_RESEAL_DATABASE_URL unset The connection string af-control-plane-backup reseal uses when --url is absent, which is how the hosted reseal job supplies it: a container app job’s command is not run through a shell, so it could not be an environment reference in the argument list, and a connection string spelled out there would be a database password visible in the revision template. Read only by that command. It must be a role row level security does not apply to, because re-sealing rewrites every tenant’s rows and a tool that re-sealed one tenant’s and reported success would be the worst outcome available.
AF_STRIPE_SECRET_KEY unset The Stripe API key, server side only. Needed together with the webhook secret and AF_STRIPE_PRICE_TEAM, which are the three billing needs to be on; setting some and not others leaves billing off and prints the missing names at startup, because an operator who sets two of three believes billing works and the one they miss is usually the webhook secret, which fails only when a real customer pays. AF_STRIPE_PRICE_ENTERPRISE is not one of the three, and the reason is on its own row.
AF_STRIPE_WEBHOOK_SECRET unset The signing secret for the endpoint registered at Stripe. Every delivery is verified against it, timestamp included, before its body is parsed. Unset means /webhooks/stripe answers 503 rather than accepting unsigned deliveries.
AF_STRIPE_PRICE_TEAM unset The Stripe price the team plan is sold at. A subscription for a price that is not named here is recorded and does not change the plan: somebody who bought through a link nobody configured has paid, and entitling them to the free plan would take away capacity they just bought.
AF_STRIPE_PRICE_ENTERPRISE unset The Stripe price the enterprise plan is sold at, and optional. Unset is a supported state and the expected one wherever Enterprise is agreed with a person rather than bought from a page: billing stays on, Team is still sold, and subscriptions.checkout for enterprise is refused before any call is made to Stripe, with a sentence saying the plan is agreed with a person and where to ask rather than one that reads like an outage. It was required once, so a deployment with a Team price and no Enterprise price was reported as half configured and took no money at all, including for Team. A plan with no price is a plan this installation does not sell self-serve, which is a decision rather than a mistake.
AF_STRIPE_API_BASE https://api.stripe.com Where the Stripe API lives. For tests, which point it at the engine’s own Stripe mock pack, and for nothing else.
AF_HOSTED_REQUIRED_PLAN unset Set to enterprise on a hosted control plane that is sold only to enterprise organizations. Authentication, sign-out and the exits remain reachable; browser procedures, CLI provider operations, model proxy requests and engine ingestion are refused until Stripe grants the enterprise plan. The exits are billing, exporting the organization’s data, deleting the organization, closing an account, and listing and revoking sessions: a plan gate may restrict what the product does for a customer and may never restrict their ability to leave, to retrieve what is theirs, or to secure their account. Any other value stops the process. Setting this while billing is off also stops the process, because otherwise no customer could satisfy the gate, and so does setting it to a plan that has no Stripe price, which is the same contradiction reached the other way: billing can be on while the gated plan itself is not sold self-serve. Leave it unset when self-hosting.
AF_OPERATOR_SETS_PLAN unset Set to 1 on an installation where whoever runs the control plane also decides each organization’s plan. Unset, billing.set is refused and the plan can only come from a signed Stripe delivery, which is the right answer anywhere the people signing in are not the operator: the first person into an organization becomes its owner, an owner holds billing.manage, and on a plane that takes no payment that would be a signed-in stranger granting themselves the largest plan. It is off by default rather than on because the dangerous configuration is the one where nothing has been configured yet, and a flag that has to be remembered would be forgotten by exactly that operator. Set it when you run the control plane for yourself; you can already write the column with psql, and this is the same act with an audit entry. Setting it together with any Stripe variable or with AF_HOSTED_REQUIRED_PLAN stops the process, because a plan that can be granted by hand is not a plan anybody has to buy. Any value other than 1, 0, true, false or unset stops the process.
AF_CONSOLE_DIR /app/console-out Where the console’s build is. The published image carries it at the default and nothing needs setting. Point it elsewhere only if you build console/ yourself. A directory that is not there is not fatal: the API serves normally, the start-up log says the console is missing, and every page answers with that sentence rather than a blank 404 that reads like a routing bug.

These are read by the enterprise entry point, the one in ghcr.io/antifailure/control-plane-enterprise, and by nothing in the community image, which ignores them. The hosted control plane runs the enterprise image. A deployment running the community image sets none of them.

Each was measured against the entry point with the variable present, absent and wrong, rather than read off the code, and the table says what the process did.

Variable Default What it does
AF_EE_SSO_KEY unset, and required 32 bytes of base64 that single sign-on seals every stored client secret and service provider key under, with the organization id bound as additional data. Without it the process exits before it listens, whatever the licence says. Generate one with openssl rand -base64 32 and never change it, because a new key cannot open anything the old one sealed. The Terraform module generates it into Key Vault, so no person ever holds it.
AF_LICENSE_KEY unset The licence. Unset or empty is the one state that is not a refusal: every enterprise route is mounted and answers 402 naming the feature and the licence state. A key that does not parse, or one signed by a key this installation does not trust, stops the process at start-up with exit status 2, because that is a deployment mistake rather than a commercial state. An expired licence starts, keeps working through its grace period, then answers 402 with every enterprise setting kept.
AF_ORG unset The organization the licence was issued to. Required whenever AF_LICENSE_KEY is set, because a licence with nothing to compare against stops the process. A licence issued to a different organization starts and answers 402 as wrong_org.
AF_LICENSE_PUBLIC_KEYS unset The keys a licence may be signed by, as kid=base64,kid=base64. Public keys, not secrets. Required whenever AF_LICENSE_KEY is set, because no build carries a stamped key, so without one no licence can be verified and the process stops.

AF_ENTERPRISE_BASE_URL is where single sign-on and SCIM publish themselves. It defaults to AF_APP_BASE_URL, which is the right answer wherever one origin serves the console and the API, as the hosted control plane does, and the process stops at start-up when neither is set. AF_LICENSE_REVOKED takes a comma separated list of licence identifiers this installation refuses as revoked; nothing publishes such a list, so it is set by hand when one is needed. The enterprise edition also reads AF_PROVIDER_KEY_SECRET, the key in the table above, and seals each organization’s audit stream collector credential under it rather than under a key of its own, because it already reaches every deployment that stores provider keys. Unset, the audit stream routes still answer, saving a destination is refused with 503 naming the variable, the start-up log says no organization can choose its own destination, and an installation sink set in the environment is unaffected. A value that is not 32 bytes of base64 stops the process at start-up with exit status 2.

The audit stream’s variables are on the audit stream page.

The process says what it decided on every start: the extensions it mounted, what the licence permits right now, and whether the audit log is being forwarded. Read those lines after a deploy rather than assuming.

Variable Where it is set What it is
AF_ADMIN_BOOTSTRAP_PASSWORD In the shell that runs the command The password for af-control-plane-backup bootstrap-operator and set-operator-password. The serving process never reads it. It is an environment variable or standard input and deliberately not a command line argument, because an argument is visible in ps to every user on the machine, lands in the shell history file, and on a CI runner is printed by any step that echoes its own invocation. At least twelve characters, and a value that begins or ends with whitespace is refused, since that is almost always a newline a heredoc added and would be part of the password invisibly forever.
Variable Where it is set What it is
AF_CONTROL_PLANE As a repository variable in GitHub, on the customer’s repository The address the workflow the App commits reports back to. It is a variable of the REPOSITORY, never of this process, and the control plane does not read it from its own environment at any point. The committed file carries the control plane’s own address as the variable’s default, so a customer of the hosted control plane sets nothing; the variable exists so a repository can point its runs at a self hosted control plane whose address the App did not know when it wrote the file. It is listed here for the same reason as the token below, which is that this is the page somebody setting up their own installation reads, and a variable named after the control plane is easy to mistake for one the control plane consumes.
AF_CONTROL_PLANE_TOKEN On the engine, or in a CI job An engine token, which the control plane issues and verifies but never reads from its own environment. Somebody running their own control plane creates one by posting to /v1/tokens, then sets it where af runs so the CLI can reach a hosted control plane. It is listed here because this is the page somebody setting up a self-hosted installation reads, and a token the control plane mints is easy to mistake for a variable the control plane consumes. Setting it on the control plane process does nothing at all.

A job in GitHub Actions should set none of that. It asks GitHub for a workflow identity and exchanges it at /v1/auth/github-oidc for a token that expires in fifteen minutes, so there is no secret to paste and none to rotate. The repository has to be claimed once first, and the GitHub guide says why that step is what grants access rather than the signature.

Everything above this section is read by the control plane process itself.

Two gates, and they are not the same one.

AF_SIGNIN_ALLOWLIST decides who may complete a GitHub sign-in at all. An account not on it is refused during the OAuth callback, before any row is written, so a refused person leaves no account behind.

Membership decides what a signed-in person can see, and it is derived from GitHub rather than granted here: an account is a member of an organization only where a GitHub App installation exists for that organization. That installation row is written by /webhooks/github when somebody installs the App, so a control plane with no App configured has no installations, and everybody who signs in lands with no tenant. Somebody can therefore sign in successfully and have no tenant at all, which is what happens to any account added to the allowlist before it is invited anywhere.

There is a third setting and it decides what a signed-in person with no organization finds. AF_SELF_SERVE_SIGNUP=1 gives them one, on the free plan, owned by them, named after their GitHub account. Without it they wait for an installation or an invitation. It is off by default because it hands out a tenant with real quotas, so on an installation with no allowlist it hands one to anybody who can reach the address; forgetting the variable has to close the door rather than open it.

The organization it creates carries the person’s GitHub login, and that is what makes the two paths one path. slugFor derives the same slug from the same login on both sides, so installing the App on that account afterwards adopts the organization the signup made rather than creating a second one beside it. Environments, audit chain and plan survive the step.

All three are needed. The allowlist is a closed door, self serve signup is what makes an open one lead somewhere, and the installation check is what makes both safe.

Closing sign-ups on a self-hosted installation

Section titled “Closing sign-ups on a self-hosted installation”

Sign-in is open by default, which is right for an instance reached only from inside a network and wrong for one on a public address that should admit named people. Two ways to close it, and they are different:

# Only these GitHub accounts.
AF_SIGNIN_ALLOWLIST=ada,grace
# Nobody at all. Note that this is the variable SET to an empty string, which
# is not the same as leaving it unset.
AF_SIGNIN_ALLOWLIST=

The process prints which mode it is in on every start, in one of three sentences, so this is never something to infer from a deployment template.

Under Helm, config.signinAllowlist is the same three states: null for anybody, a list for those accounts, and [] for nobody. In Terraform, signin_allowlist is null, a list, or [], and Terraform will not produce a plan without a value at all, so opening the door stays a decision somebody wrote down.

Closing sign-ups does not by itself stop somebody who is already a member. Their sessions continue until they expire or are revoked, which the operator portal and the Sessions page can do.

When sign-ups are open and AF_GITHUB_APP_INSTALL_URL is set, a new customer can complete the whole path without an operator: sign in with GitHub, install the App on an organization, then choose Check my GitHub membership. The second OAuth exchange reads the installation GitHub just created and grants the membership. The first GitHub administrator to claim an empty organization becomes its owner under the rule below.

When it is unset, that path still exists but nobody can start it from the console. The screen says the address has not been configured and offers Check my GitHub membership on its own, which is the right action for somebody who already belongs to a connected organization and the wrong one for somebody who does not. Unset is a supported state rather than a half configuration, because a self-hosted control plane may grant membership its own way and have no App to point at. It is not the right state for a plane with open sign-ups, and the startup log names it either way so an operator can tell which they have.

It is the App’s public installation page, and only a human with owner access to the GitHub organization that owns the App can produce it.

  1. Open the App’s settings under the owning organization, at Settings, then Developer settings, then GitHub Apps.
  2. If no App exists yet, create one. It needs the same App ID, private key and webhook secret that AF_GITHUB_APP_ID, AF_GITHUB_APP_PRIVATE_KEY and AF_GITHUB_APP_WEBHOOK_SECRET already document, so create it once and take all four values in the same sitting.
  3. Set the App to Any account under Install App, not just the owning account. An App only its owner can install is an App no customer can install.
  4. Read the slug out of the App’s public page URL, https://github.com/apps/<slug>. It is derived from the App name and is not always what you would guess.
  5. The value is that address with /installations/new on the end, and nothing else. No query string and no fragment: both are refused at startup.

On an enterprise-only hosted deployment that owner lands on Plan. Checkout is the only path that can grant the required plan; billing.set is refused, so an owner cannot turn a free organization into an enterprise one without Stripe. That refusal does not depend on Stripe being configured. billing.set is refused on every installation that has not set AF_OPERATOR_SETS_PLAN=1, including one where billing has not been set up yet, because that is the installation on which an owner granting themselves the largest plan would otherwise succeed. The signed subscription webhook changes the plan. Refresh from Stripe asks Stripe for every subscription belonging to that customer and repairs the same state when a webhook never arrives, including the case where no local subscription row exists yet.

It also clears a checkout that cannot be paid. If Subscribe is refused because Stripe has no record of the checkout this organization already opened, Refresh from Stripe asks Stripe for that checkout. When Stripe still has no record of it, the stale checkout is cleared and the next Subscribe opens a new one. Nothing is charged by either step.

The role comes from GitHub, read at sign-in with an installation token: an organization owner on GitHub becomes an admin here, and everybody else becomes a member.

An owner on GitHub deliberately does not become an owner here. That role also holds billing.manage, and who pays is this application’s decision rather than GitHub’s. Promote somebody with the role control on the Members page; a role set that way is marked manual and is never overwritten by a later sign-in.

With one exception, and it is the first sign-in. An organization is created by the installation webhook, before anybody has signed in, so every organization passes once through a state where it has no members at all. The first person to sign in becomes its owner rather than its admin, provided GitHub confirms they administer the organization. Without that, no organization created this way would ever have an owner, and nothing would hold billing.manage. The promotion is marked manual, so a later sync does not take it back, and it is recorded in the audit log as member.bootstrapped.

Two cases where nothing changes rather than something being guessed. Sometimes GitHub will not say what somebody’s role is: no App configured, a rate limit, an outage. An existing membership then keeps the role it already had, because a transient failure must not demote the only administrator out of their own organization. A first sign-in during the same failure gets member, because guessing upward would hand out administrative rights on a timeout, and that applies to the first member of an empty organization as well: GitHub has to say admin for anybody to become an owner. If the App is permanently broken and that leaves an organization with nobody who can act, the way back is break-glass, which is an operator holding the database credential rather than a guess made by a web request.

Sign-in can only ever speak for the person signing in. Sync from GitHub on the Members page reconciles everybody at once, and it is the only thing that takes access away: somebody removed from the GitHub organization keeps their role until it runs, because a person who has been removed has no reason to come back and sign in. It needs members.manage, it refuses an empty member list from GitHub rather than removing every owner, and it records what it changed in the audit log.

Everything on this page is reachable by whoever the role table says can reach it. The console hides what a role cannot do; the server refuses it, and the refusal is what the permission matrix tests, one route against each of the four roles.

Inviting somebody who is not in your GitHub organization

Section titled “Inviting somebody who is not in your GitHub organization”

Membership follows the GitHub App installation, which is right for engineers and useless for the two cases every company has: a finance person who needs the billing page and no repository access, and a contractor who is not in the GitHub organization at all. Invitations on the Members page sends a link.

The link carries a token that exists only in the link. What is stored is its hash, the same way a session is stored, so a leaked backup is a list of hashes rather than a list of ways into your organization. Two consequences worth knowing before you use it:

  • The link is shown to you as well as sent. A control plane with no AF_MAIL_FROM cannot send anything, and an invitation that only existed as an email would silently do nothing there. Copy it and send it however you like.
  • Sending it again produces a NEW link and the old one stops working. The original cannot be resent because it is not stored. That is also the better behaviour: an invitation forwarded to the wrong person is invalidated by asking for a fresh one.

A link expires after fourteen days. An invitation stays good after the person who sent it has left, because it was authorised when it was sent, and the record keeps their name as it was at the time. Accepting adds the account that is signed in, which is not necessarily the address the invitation was sent to: the token is the proof and the address is a label.

Signed in now, under Settings, lists every live session in the organization with who it belongs to, where it came from and when it was last used, and marks the one you are reading it in. It never shows a token or a hash of one.

Signing a session out takes effect on that session’s next request. Removing somebody from the organization signs them out in the same transaction, so there is no window in which a person who is no longer a member still has a working session. Both need sessions.manage; removal needs members.manage.

A session that is not used for twelve hours stops working, and no session lives longer than thirty days however active. The list shows when each one expires so that a session which is about to go on its own can be left alone.

Download a copy, under Settings, produces one JSON file holding people, invitations, repositories, masking rules, egress policy, environments, runs, verdicts, runtimes, credentials by name, billing history and the audit log. It needs data.export.

Every reference in it is the name you already use: a repository is owner/name, a person is their login, an environment is its env id. There is not one internal identifier in the file. Inside it, files holds text keyed by path, and those are the parts you can put straight back: masking.yaml is a masking file the engine reads as it is, and egress.yaml is the egress: block from antifailure.yaml.

What it deliberately does not contain is listed in the file itself, under notIncluded, with the reason for each. Engine token values and provider key material are the important two: an export carrying either would be a way into your CI.

organization.delete is held by an owner and nobody else. It is not a delete statement, and the order is the point:

Step What happens
Stop what is running Every environment is marked torn down, every queued or running run is cancelled, and the organization is suspended so nothing new can be started.
End the subscription Cancelled at Stripe at the end of the period you have paid for. Nothing is refunded and nothing is taken away early.
Wait Nothing else happens until that period ends. Everything still reads, and the deletion can still be called off.
Revoke credentials Engine tokens, provider keys, sessions, and the GitHub App installation, which is removed at GitHub rather than only marked here.
Produce the export The same document as Download a copy, taken before anything is removed, because afterwards there is nothing left to build one from.
Delete The organization and every row belonging to it, including the audit log.

Two things follow from that order and both matter.

A deletion that is interrupted picks up where it stopped. Each step records that it happened in the same transaction as the change it describes, so a process that dies between two steps leaves a record saying exactly which happened. The control plane retries on its own, and Continue now does the next step immediately.

The download link is shown once, when you ask for the deletion. After the organization is gone there is no membership left to authorise a download, so the link is the authorisation. Keep it. It works for seven days, and Destroy the copy removes the held document early if you would rather we did not keep one.

Your database is not touched by any of this, because none of it is here: no snapshot, no masked branch and no captured request body ever reaches this control plane.

Every role can close their own account, including viewer. It erases your name, address, GitHub identity and avatar, removes your memberships, and signs you out everywhere. Signing in again afterwards creates a new account.

It is called closing rather than deleting because the row is not removed. The audit log references it, and that reference is deliberately one the database refuses to break: an audit log whose subject can erase themselves from it is not an audit log. The entries keep the name you had at the time, because the log is a hash chain and rewriting an entry breaks it, and they go when the organization does.

The only refusal is the last owner of an organization. An organization with no owner cannot grant anybody the permission to become one, so make somebody else an owner first, or delete the organization.

Two endpoints, answering two different questions. Point the right thing at the right one.

Endpoint Answers Touches the database
GET /health Is the process alive? No
GET /readyz Can it serve a request? Yes, one trivial query

/health is a static literal, and it stays one. A liveness probe restarts the container when it fails, so wiring it to the database turns a slow Postgres into a restart loop that makes the outage worse.

/readyz takes a connection from the pool the application serves with and asks the database a question. It answers 200 with the build, or 503 with the reason:

{ "ready": true, "version": "v1.0.0", "commit": "31ce3f7" }
{ "ready": false, "version": "v1.0.0", "commit": "31ce3f7",
"reason": "password authentication failed for user \"af_app\"" }

Use /readyz for a deploy gate, and check the commit as well as the status. The first deploy of this application to Azure answered /health with 200 for thirteen minutes while every endpoint that touched a table returned 500: the schema had never applied, because the managed Postgres refused CREATE EXTENSION pgcrypto. A gate watching /health would have called that deploy a success. Checking the commit catches the other half, a rollout that silently did not happen and left the previous build serving.

GitHub is the front door and it needs a route to github.com. A preview environment has none by design, and an isolated network has none at all, so there is a second way in: a link sent to an address that already belongs to a member of an organization.

It is off unless all three variables below are set. Setting one or two of them stops the process at startup and says which are missing, because two of three is a link that goes nowhere or mail that cannot be sent, and both of those fail at the moment somebody is locked out rather than at deploy time.

There is no sign-up on this path. An address receives a link only once somebody has invited it into an organization, the link works once, and it expires in fifteen minutes.

Variable Default What it does
AF_RESEND_API_KEY unset The Resend key the link is sent with. An HTTP mail API rather than SMTP on purpose: it is a request the egress sidecar can capture, which is what lets a preview environment read its own sign-in mail instead of delivering it to somebody.
AF_MAIL_FROM unset The From address. Resend refuses a domain it has not verified, which is a configuration error worth failing loudly on.
AF_PUBLIC_URL unset Where the link points: the origin a browser reaches this deployment on. Wrong here means a link that lands somewhere nobody is serving.
AF_ENV_URL injected Set by Antifailure inside a preview environment: the address of the environment’s first web service, which is the application a person opens. AF_PUBLIC_URL is preferred where a deployment sets one, and this is the fallback, because the address a preview answers on is allocated at run time and no value written in a manifest can be right. Ignored outside a preview, where nothing sets it.
AF_RESEND_BASE_URL https://api.resend.com Where the mail API is. Set it to point at a local capture during development.
AF_PRODUCT_NAME Antifailure The name in the subject line, for a white-labelled deployment.

The events table is partitioned by month. Partitions are created ahead of the writes, because a range-partitioned table with no partition for an incoming row does not slow down, it fails.

Keeping ahead is DDL, so it runs as the migration role and not as the application role. The connection is opened for each pass and closed after it, rather than held idle between them.

Variable Default What it does
AF_MAINTENANCE_DATABASE_URL falls back to AF_MIGRATION_DATABASE_URL The role that creates and drops partitions. When neither is set, this process logs a warning at startup and does not keep the partitions ahead. Something else must.
AF_EVENT_RETENTION_MONTHS unset Drop event partitions entirely older than this many whole months. Unset keeps everything forever, which is the default because retention is an operator’s decision. A value that is not a whole number of months at least 1 stops the process at startup rather than silently keeping everything.
AF_EVENT_ARCHIVE_DIR unset Write a month out as newline delimited JSON here before dropping it.
AF_FAILURE_RETENTION_DAYS 30 How long a group in control_plane_failures survives past its LAST occurrence, not its first: a failure first seen in March and last seen this morning is the most interesting row on the page, and sweeping by its age would delete exactly the long running failure an operator is trying to date. Applied only when this maintenance pass can run, because the application role is granted no DELETE on that table on purpose. A value that is not a whole number of days at least 1 stops the process at startup.

The store of the control plane’s own failures

Section titled “The store of the control plane’s own failures”

Both error handlers write what they caught to standard output and to a grouped table, so an installation with no log aggregation can still answer “what is failing right now” from the operator portal. A row is a fingerprint over the declared route, the method, the error class and the driver code, with a count, so the table’s size is set by the code and not by traffic. It holds at most 500 groups and never a message, a stack, a payload or an organization. See operations for what it can and cannot answer.

Variable Default What it does
AF_FAILURE_STORE on off, 0 or false records nothing. The Logs page then says nothing is being recorded, rather than showing an empty list that reads as a healthy day. The default is on because the table is bounded by the code, the writes are one statement per distinct group per ten seconds rather than one per failure, and a healthy installation writes nothing at all.
  1. Creates the current month and the three after it. This happens unconditionally and first. Nothing below is allowed to prevent it.
  2. Archives each month that retention has condemned, if AF_EVENT_ARCHIVE_DIR is set. The file is written under a temporary name and renamed when it is complete, so a file appearing in the directory always means a whole one.
  3. Drops those months, but only if every archive finished. A failed write costs a retention run rather than the events, because a month deleted with no copy anywhere cannot be undone.
  4. Prunes the default partition by age, a bounded number of rows per pass.

A pass runs at startup and then once a day. A pass that throws is logged and the schedule continues: the failure that matters is running out of partitions, and giving up after one transient error is how that happens quietly.

Nothing needs to be done by hand. Events whose month does not exist land in the default partition rather than failing, and the next pass moves them into the month it creates for them. It detaches the default partition, creates the month, moves the rows through the parent so that Postgres decides where each one goes, and reattaches, all in one transaction.

Ingestion depends on a unique constraint to make retries safe:

INSERT INTO events (...) VALUES (...)
ON CONFLICT (org_id, idempotency_key, occurred_at) DO NOTHING

An engine that sent a batch and lost the response cannot know which half landed, so it sends the batch again and the database drops the copy.

Postgres will not enforce a unique constraint that omits the partition key, so the partition column is necessarily part of that key. received_at is assigned here, by the clock, and would differ between an attempt and its retry: the conflict would never fire and every retry would duplicate. occurred_at is assigned by the sender when the event happened and is resent unchanged, so it does not vary between attempts and costs nothing by being in the key.

The usual objection to partitioning on a value a client supplies is a skewed clock inventing partitions forever. Ingestion already rejects occurredAt more than a day in the future or more than a year in the past, so the live range is bounded before a row reaches the table.

The cost, stated plainly: the idempotency key is now (org_id, idempotency_key, occurred_at) rather than (org_id, idempotency_key). A sender that reuses an identifier under a new timestamp gets two rows where it used to get one. No sender does that by accident, since the identifier and the timestamp are minted together and resent together, but it is a real difference and not a free one.

Each line is one event, as JSON, with timestamps as RFC 3339 text rather than in a driver’s own format, because the file is read by something that is not this process.

Terminal window
# how many events, and over what span
wc -l events_2026_03.jsonl
head -1 events_2026_03.jsonl | jq -r .occurred_at
# everything one environment did
jq -c 'select(.env_id == "env-1234")' events_2026_03.jsonl

An owner opens Administration → Website to edit the public site. The page picker lists the routes in the built site, including documentation. A draft shows in the preview before publication; publishing updates the public content and requests a static refresh. Unchanged text, images and design settings use the version in the site’s source. HTML, CSS and JavaScript blocks run in an isolated iframe rather than in the surrounding page.

The Pages view lists built routes and individual Writing articles. Owners can create a page at a new path or an article under /blog, then edit its title, introduction, summary, rich body, date and topics. A draft URL is available for preview before it exists publicly. Publishing rebuilds its HTML, Markdown version and sitemap entry; new articles also enter the Writing index and RSS feed. The editor links newly authored pages from the site’s Pages index so visitors and crawlers can reach them. Existing pages retain source defaults until an owner changes a field.

The Ask AI panel is optional. AF_CMS_ANTHROPIC_API_KEY is the Anthropic API key used only by the control-plane process for owner-requested edit suggestions. Leave it unset to use the manual editor without AI. The key is never included in the website build, preview messages or published content.

Variable Default What it does
AF_CMS_ANTHROPIC_API_KEY unset Optional server-side Anthropic credential for the owner-only website assistant. Keep it in a secret store; the website build does not read it.

Hosted staging and production read it from the existing Key Vault secret named cms-anthropic-api-key; Terraform stores the secret’s address, not its value. For a Helm installation, set websiteAI.existingSecret to the name of a Kubernetes Secret holding the AF_CMS_ANTHROPIC_API_KEY key. The assistant receives only the selected page fields and recent chat turns. It may suggest copy, styles and section changes, but cannot save or publish them. An owner reviews the proposal, applies it to the draft and publishes separately. Each owner has a daily allowance of 40 requests and 180,000 tokens.

Off unless a surrogate secret is configured, and said out loud at startup either way. There is no fallback to a constant key: a constant key is a surrogate anybody can recompute, which is an organization identifier with extra steps.

Variable Default What it does
AF_ANALYTICS_SURROGATE_SECRET unset 64 hex characters, which is 32 bytes. The key organization surrogates are computed under. Unset records nothing at all, and the dashboard says so rather than showing an empty chart. A value of any other length stops the process at startup rather than on the first event. Generate one with openssl rand -hex 32.
AF_ANALYTICS_OPERATOR_ORG unset The slug of the organization that operates this control plane. Its owners and admins may read the analytics dashboard; nobody else may, whatever permissions they hold in their own organization. Unset means nobody, and the route says which variable to set.
AF_ANALYTICS_RETENTION_DAYS unset Delete raw analytics events older than this many days. The daily aggregates computed from them are kept, because a count of page views by channel has nothing in it that identifies anybody. Unset keeps the raw events forever, which is the default because retention is an operator’s decision.
AF_SITE_ORIGIN unset Every origin the marketing site is served from, comma separated, for the endpoints a browser calls cross origin. Unset refuses every beacon rather than reflecting whatever Origin arrives, which is what a permissive default would do.
AF_POSTHOG_REGION unset us or eu, and nothing else. Mounts the PostHog proxy at /ph, so the marketing site sends its product analytics to this control plane and this control plane forwards it, and a reader’s browser opens no connection to a posthog.com host. Unset mounts nothing, so a site configured to send analytics here is answered 404 rather than quietly reaching a vendor the operator did not choose. The two values select a pair of fixed upstream hosts: there is no setting of any kind that makes this forward to a host outside that pair, which is what stops it being an open forwarder. A PostHog project API key does not carry its region, so read it off the cloud rather than guessing: post the key to https://us.i.posthog.com/flags/?v=2 and to the eu host beside it, and the one that answers 200 rather than authentication_failed is the region to set. AF_SITE_ORIGIN still governs which origins may call it.
AF_POSTHOG_PROJECT_KEY unset The PostHog project API key this process reports its OWN hosted usage under: which hosted MCP tool was called, how it ended, how long it took, and the model, token counts and latency of a model call the control plane brokered. Public by design, like every PostHog project key: it can only write events into one project and reads nothing back. Unset sends nothing, which is the right default for a self-hosted installation, because that installation’s usage is its own. Needs AF_POSTHOG_REGION as well, since a key with no region has nowhere to go and defaulting to a cloud would pick a continent on the operator’s behalf. What is never sent: a tool’s arguments or results, a prompt, a completion, or an organization identifier. The organization is a pseudonym, the same domain separated HMAC AF_ANALYTICS_SURROGATE_SECRET computes for this control plane’s own analytics, so with that unset nothing is reported at all.

Mounted only when AF_POSTHOG_REGION is set. A reader’s browser then connects to this control plane rather than to a posthog.com host: the site is a static export with no server of its own, so the forwarding has to happen on the one process this product already runs on its own hostname.

It is transport and it is not a boundary. It changes the destination the browser connects to, not who receives the data. PostHog, Inc. receives every event, every autocaptured interaction and every session recording either way. What it buys is that a content blocker’s vendor list does not match the request, so the measurement is not silently half missing; that the recorder bundle, the largest and most blockable request the library makes, arrives rather than failing while ingestion looks healthy; and that the reader’s address is dropped in passing. It does not buy the sentence “no third party sees this”, and the privacy page says so.

Separate from the proxy, off by default, and configured by AF_POSTHOG_PROJECT_KEY. The proxy forwards a browser’s requests; this sends events from the control plane about the hosted service it runs.

Two producers, and nothing else has one:

  • Hosted MCP tool calls. The tool name, which is a closed set of the eight tools the surface registers, the outcome (ok, error or refused), and the duration. Never the arguments and never the results: a tool call carries project identifiers, hostnames, table names, SQL and error text, all of it the customer’s, and inspect_recorded_egress alone would ship their outbound destinations to a vendor.
  • Brokered model calls, in PostHog’s own $ai_generation shape: $ai_model, $ai_provider, $ai_input_tokens, $ai_output_tokens, $ai_latency and $ai_trace_id, plus the cost when the provider reported usage to compute one. Never the prompt and never the completion. PostHog’s schema makes $ai_input and $ai_output_choices optional, so omitting them is the supported shape. Only where the control plane itself brokers and bills the call.

Nothing of this kind exists in the engine and nothing of this kind may be added to it. af mcp runs on a customer’s own machine, inside their network. engine/internal/telemetry already carries the engine’s events, it requires a redactor before any sink may write, and it exports to the CUSTOMER’S control plane. A path from there to our analytics vendor would be an outbound flow nobody agreed to, out of a product sold on the promise that production data stays in the customer’s boundary. Engine side numbers travel the route that exists; only a hosted control plane forwards anything onward about its own service.

A failure here never reaches a caller. Both producers sit on load bearing paths, one being a customer’s agent and the other being the proxy that spends their money, so every send is fire and forget, swallows its own errors, and is flushed at shutdown rather than awaited on a request.

It is same site, not same origin. The site is served on an apex and a www hostname, this control plane answers on a third, and those are three different origins sharing one registrable domain. Every forwarded route therefore answers a CORS preflight and echoes exactly one allowed origin from AF_SITE_ORIGIN.

Three separate things keep it from becoming a general forwarder, and none of them replaces the others:

  • The upstream host comes from a closed set of two regions. No request, header or setting can name a different one.
  • The paths that reach PostHog are an allowlist. A path under /ph that is not on it is not a route at all, so it is answered 404 rather than forwarded.
  • A redirect from the upstream is refused rather than followed, so PostHog cannot steer this process at another server.

Nothing of the browser’s is passed upstream except content-type: no cookie, no authorization, and not the visitor’s address. That last one is deliberate and it has a cost. PostHog geolocates from the source address, and behind this every event arrives from one container, so the $geoip_* properties describe the deployment rather than the reader. Forwarding the address would send every visitor’s IP to a third party, which is the disclosure this proxy exists to avoid, and it is the one direction that cannot be undone afterwards.

The analytics stream is a closed schema. An event whose name is not in the catalog is refused and counted, and so is a payload field the catalog does not declare. There is no free-text field of any kind, so a repository name, a branch, a query string or a page URL cannot reach the store even by mistake.

The organization is recorded as a keyed hash rather than as an identifier. The store can count organizations and follow one through a funnel, and it cannot name one without the key.

The application role holds INSERT on the stream and no SELECT. Only the rollup, which runs as the schema owner, ever reads it, and only daily aggregates come back out. A read attempted by the application raises 42501 rather than returning nothing, which is the difference between a mistake somebody sees and one somebody ships.

Three questions need to follow one subject across days or across events, and a daily count cannot. The rollup computes them into tables of counts:

Question How it is computed What bounds it
How many distinct organizations or sessions were active over a window A working set of one row per subject per event per day, then a distinct count over 1, 7 and 28 days The working set is kept for 98 days
How many completed a declared sequence of steps, in order and inside a window The raw stream at full precision, once per subject, stored as how far each got The widest funnel window plus the rollup lookback
Of the organizations first seen in a week, how many came back each week after The working set against the first-seen date in the facts table 12 cohort weeks

The funnels are declared in the catalog rather than built in the interface, and that is deliberate. A funnel builder would need the application to be able to run an arbitrary query against rows that carry a subject surrogate, which is the capability the grants above exist to withhold. The application holds no SELECT on the working set at all: it reads counts, and the tables it can read contain no identifier of any kind.

A retention cell over fewer than 5 organizations is shown as a count rather than as a percentage. A rate over three subjects moves by a third when one of them opens a laptop.

The site sends one event per page a reader lands on, one when the sign-up screen is reached, and one when somebody asks to be contacted. It sets no cookie, loads no third-party script, and keeps its session identifier in sessionStorage, so it dies with the tab and two visits a day apart cannot be joined. A session also ends after thirty minutes idle and after twenty four hours whatever happens, so a tab left open over a weekend is several sessions rather than one identifier held for three days.

Events are queued and sent in batches of up to twenty, flushed every three seconds and on the way out of the page through sendBeacon. A request that fails with a server error or a network failure is retried with a capped, jittered backoff; one refused with a 4xx is not, because a refusal does not become true on the third attempt. A retry cannot double count: the event identifier and timestamp are stamped once when the event happens and resent unchanged, so the second copy collides on the primary key and is recorded as a duplicate.

It turns itself off for a reader who has set Global Privacy Control or Do Not Track, for a browser whose user agent names a crawler, and for a driven browser. The user agent is read in the page and never sent, so the crawler filter only sees crawlers that execute JavaScript: these counts are a floor and a shape rather than an audited total, and the dashboard says so beside them.

A switch on the privacy page turns measurement off for that browser, and opening any page with ?af-analytics=off does the same thing without a click, which is what makes excluding a colleague a link rather than an install. Either is undone by the switch or by ?af-analytics=on. That is the only value the beacon keeps beyond the tab, it is a single flag, and it is never sent anywhere.

The switch takes effect on the page it is pressed on rather than on the next one, and it discards whatever is queued and unsent, because the queue holds events for up to three seconds and sending them because they were captured a moment before the reader objected is the disclosure the switch was pressed to prevent. It reports which of four things is deciding: this reader asked, the browser asked through Global Privacy Control or Do Not Track, the browser looks automated, or this build has no endpoint configured. Only the first is the switch’s to change, and where it is not the switch is not offered.

The referrer, the URL and the query string are turned into a bounded channel, a page shape and a campaign identifier in the browser, so the raw values never cross the network at all. That is a stronger claim than discarding them on arrival, and it is why the normalization lives in the page rather than in a server reading a Referer header.

The endpoint is unauthenticated, because a shared secret in a static page is a secret everybody has. Its counts are therefore a floor and a shape rather than an audited total, which the dashboard says beside them.