# Command reference

URL: https://antifailure.dev/docs/reference/cli

Every command and every flag, generated from the command tree itself.

---

Generated from the command tree, so it cannot fall behind the binary: a flag
added, renamed, or removed changes this page in the same commit, and the build
fails if it does not.

## Global flags

These work on every command.

| Flag | Default | What it does |
| --- | --- | --- |
| `-C`, `--directory` | - | Run as if started in this directory. |
| `--no-color` | `false` | Do not emit colour, regardless of the terminal. |
| `-o`, `--output` | `text` | Output format: text or json. |
| `-q`, `--quiet` | `false` | Print only what was asked for. |
| `-v`, `--verbose` | `false` | Print the underlying cause of an error. |

## Flags on `af` itself

These work on `af` on its own rather than on a command under it.

| Flag | Default | What it does |
| --- | --- | --- |
| `--short` | `false` | With --version, print only the version number. |
| `--version` | `false` | Print the version, commit, and edition. |

## How output adapts

Text output is stable for the same input. There are no timestamps and no
durations in it, so a snapshot test, a diff and two CI logs of the same run
compare cleanly. Timestamps live in `--output json`, where a machine
wants them.

What does vary is layout, and only where there is a terminal to lay anything
out on. Colour, width and the live status line under a long run are decided
once, from the output stream, when the command starts.

| Variable | What it does |
| --- | --- |
| `NO_COLOR` | Any non-empty value turns colour off. It wins over everything, including `AF_FORCE_COLOR`. |
| `AF_FORCE_COLOR` | Any non-empty value turns colour on for a stream that is not a terminal, for a CI system that renders escape codes. |
| `AF_WIDTH` | Lay output out at this many columns rather than measuring the terminal. Clamped to between 40 and 200. |

A stream that is not a terminal, a pipe, a file, or a CI log, is laid out at 80
columns and carries no escape sequences. That is what keeps the output of a
piped run identical from one machine to the next. `TERM=dumb` is
treated the same way.

## Commands

### `af change`

Read the diff and say which checks will exercise what it touched.

What this change touches, and which checks cover it.

Every changed path is classified by a rule that names it, and every check is
reported as selected or not, together with whether the manifest configures it
at all. A check that is selected and unavailable is the line worth reading:
something changed and nothing is going to look at it.

Two things it will not do. It never says a change is safe or risky; it says
which checks exercise which files, and what it cannot see. And a path no rule
recognises selects every check rather than none, because the cost of the two
mistakes is not the same.

In a GitHub Actions job it writes one output per check, so a later step can
skip work this change does not need.

This is the one command that does not need antifailure.yaml. Without one it
still says what the diff touches, and reports every check as unavailable
because nothing is configured to run it.

```
af change [flags]
```

```
# Against the base branch this job names.
af change
# Against a ref you choose, or a diff you already have.
af change --base origin/main
af change --diff pr.patch
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--base` | - | Ref to measure against, defaulting to this job's base branch. |
| `--branch` | - | Branch to read the manifest for, defaulting to the checked out one. |
| `--diff` | - | Read a unified diff from this file instead of asking git. |
| `--head` | - | Ref to measure, defaulting to HEAD. |
| `-w`, `--write` | - | Write the report section here as markdown. |

### `af chaos`

Break this environment on purpose and prove the recovery.

Injects the faults the manifest's chaos block declares into the running
environment, one at a time, and reads what the system did about each one.

The faults are real. A process is killed with SIGKILL, a container is stopped,
a container is frozen, a container is detached from the network, a data
directory is made read only. Nothing is simulated, and nothing is aimed
anywhere but at the containers this environment created: a target is resolved
from the labels the runtime stamped at create time, the ownership is proved
again from the daemon at the instant of the act, and the egress sidecar is
refused whatever a fault asks for, because a fault that can stop the thing
deciding where the environment may connect is a way out rather than an outage.

Around a fault aimed at the database, the durability proof runs. Concurrent
writers commit into a schema of the engine's own while the fault lands, and
afterwards every commit the client was told was committed must still be there
and nothing may be there that no client ever wrote. That needs a record the
database cannot provide, because the claim is about what the database SAID,
and the write ahead log is then read for the evidence that it actually
replayed: the position recovery started from, against the one the control file
named before the crash, and the position it reached, against the last flush a
writer saw.

Anything that could not be established is reported as unverified rather than as
a pass. A fault that was applied and changed nothing is refused, because every
assertion after it would be measuring a system that never broke.

```
af chaos [flags]
```

```
# Inject the manifest's faults and prove what the recovery did.
af chaos

# Against a branch other than the checked out one.
af chaos --branch fix-the-outbox

# The whole result, including the acknowledged commit ledger.
af chaos -o json
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to break, defaulting to the checked out one. |

### `af ci`

Bring an environment up, run everything, write a report, tear it down.

The whole check in one command, for a pull request.

The agents drive the workflows, the invariants are asked of the data, the
migrations are rehearsed against a throwaway branch of the golden, and what the
environment reached for is summarised. Every finding is ranked by the manifest's
policy block, which decides what fails the check and what is only reported.

Load runs when the manifest enables it. A missing traffic source produces a
read-only smoke from the literal safe routes, not a production benchmark.

Teardown happens whatever the outcome, including a failure and including an
interrupt, because an environment that outlives its pull request is the leak
this product exists to prevent. It happens before the report is written, so a
teardown that left something behind is in the report rather than after it.

Only a real finding exits non zero. A blocked run says what was missing and
exits zero, so an incomplete environment is not indistinguishable from a broken
change.

```
af ci [flags]
```

```
# What CI runs: up, migrate, test, load, gate, report, down.
af ci
# --report is Markdown for a person, --report-json is the same run
# for a program.
af ci --report report.md --report-json report.json --keep
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--baseline` | - | Compare queries and plans against a report saved on the base branch. |
| `--branch` | - | Branch to check, defaulting to the checked out one. |
| `--docs` | - | Where documentation links point. |
| `--keep` | `false` | Leave the environment up, for debugging a failure. |
| `--load` | `false` | Generate load even when the manifest's load block is off. |
| `--report` | - | Write the report here as well as to the terminal. |
| `--report-json` | - | Write the same report here as JSON, for a program to read. |
| `--runner` | - | Path to the runner's entry point. |
| `--save-baseline` | - | Save this run's queries and plans, to compare a later branch against. |
| `--timeout` | `30m0s` | Give up after this long. |

### `af doctor`

Check that this machine can run Antifailure, and say how to fix what cannot.

Every check names what to do about a failure. A diagnostic that tells you
something is wrong and stops is worse than no diagnostic, because it costs the
same attention and yields nothing.

```
af doctor
```

```
af doctor
af doctor -o json
```

### `af down`

Remove the environment and everything it created.

Replay the journal in reverse and delete what this environment recorded
creating. Every resource is journaled before it is made, so what teardown
removes is what was actually created rather than what a list somebody
maintained remembers to look for.

Teardown never stops at the first failure. A provider that is unreachable must
not strand the other resources, so each is attempted, and two things survive
the run: anything that could not be removed, and anything this build has no way
to delete, which is left recorded rather than forgotten. Both are named
individually in the output, and exit code 10 means resources are still pending.

So the answer to what this is about to remove is not a sentence here. It is
'af status' for what is running and the pending list this prints for whatever
it could not reach.

```
af down [flags]
```

```
af down
af down --branch feature/checkout
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to tear down, defaulting to the checked out one. |

### `af env`

See and clean up the environments on this machine.

Reads the daemon rather than a registry, because the daemon is the thing that
actually has them. A registry can be wrong; a container either exists or it
does not, and a list that disagrees with reality is worse than no list.

```
af env
```

```
af env list
```

Subcommands:

- [`af env extend`](#af-env-extend) Keep an environment past its lifetime, up to its maximum.
- [`af env list`](#af-env-list) List the environments this machine is holding.
- [`af env prune`](#af-env-prune) List the environments older than a cutoff, and remove them with --yes.
- [`af env pull`](#af-env-pull) Read an environment's record from the control plane.
- [`af env reap`](#af-env-reap) List the environments whose lifetime has ended, and remove them with --yes.

### `af env extend`

Keep an environment past its lifetime, up to its maximum.

Moves an environment's expiry, so a sweep does not take one you are still
using.

There is a bound, and it is the point. No extension may take an environment
past runtime.max_ttl measured from when it was CREATED, not from now, so
extending repeatedly cannot walk the limit forward. Asking for more than the
maximum grants the maximum and says so rather than failing, because being given
less time than you asked for silently is how you come back to an environment
that is gone.

```
af env extend <environment> [flags]
```

```
# The ceiling is measured from when the environment was created, so
# extending twice does not buy twice the time.
af env extend af-orders-feature-checkout-05ca6c --for 2h
af env extend af-orders-feature-checkout-05ca6c --for 2h --reason 'debugging the failing checkout'
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--for` | `4h0m0s` | How long from now the environment should live. |
| `--reason` | - | Why, recorded with the extension. |

### `af env list`

List the environments this machine is holding.

```
af env list
```

```
af env list
af env list -o json
```

### `af env prune`

List the environments older than a cutoff, and remove them with --yes.

An environment nobody tore down holds a database branch, a network, and a
container per service, and the machine that accumulates a dozen of them is a
machine somebody reboots to fix.

Run bare, it removes nothing. It lists every environment on this machine that
is older than the cutoff, whichever repository created it, and stops with the
command that would remove them. Removal needs --yes, and what --yes removes is
exactly what the bare run listed. --dry-run means the same as running bare and
is kept so that a script which passes it keeps working.

The cutoff is --older-than, a day when not given, and the plan prints it, so
the default is never something a reader has to remember. The scope is the
whole machine on purpose: this is the command for a laptop that is full, and
the daemon does not record which repository made what, so a cutoff from here
reaches every project's environments. For a sweep that reads each
environment's own lifetime instead, see af env reap.

--orphaned narrows it to environments that hold networks with nothing attached
and nothing running, which is what a run killed before its teardown leaves.
Each such network still holds one of the thirty or so address ranges Docker's
default pools can hand out, and when they run out no environment can be
created at all. With --orphaned the cutoff is an hour unless --older-than says
otherwise, measured from the environment's newest resource, so one that is
being brought up right now is never taken. Networks without the Antifailure
label are never considered.

```
af env prune [flags]
```

```
# Lists what is older than a day, on this machine, and removes nothing.
af env prune
# Removes exactly what that listed.
af env prune --yes
# Everything on this machine, whatever its age: look, then remove.
af env prune --older-than 0s
af env prune --older-than 0s --yes
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--dry-run` | `false` | List what would be removed and stop, which is also what running bare does. |
| `--older-than` | `24h0m0s` | Only consider environments older than this. |
| `--orphaned` | `false` | Only environments holding networks with nothing attached and nothing running. |
| `--yes` | `false` | Remove what the plan lists. Without it nothing is removed. |

### `af env pull`

Read an environment's record from the control plane.

Reads what the control plane holds for one environment: its branch, its state,
its preview URL, and the golden version it was built from.

This never changes anything locally. The control plane is a record of what
happened, not a source of configuration: what an environment does comes from
the manifest in the repository, on the machine the environment is on. A control
plane that could change what an environment runs would be a control plane that
could change what it masks.

Needs a credential. Run af login, or set AF_CONTROL_PLANE_TOKEN to an engine
token, which is what a build machine with nobody sitting at it uses.

```
af env pull <environment> [flags]
```

```
af env pull af-orders-feature-checkout-05ca6c
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--control-plane` | - | The control plane to read from (default: AF_CONTROL_PLANE_URL, or the hosted instance). |

### `af env reap`

List the environments whose lifetime has ended, and remove them with --yes.

Finds every environment on this machine that has passed the lifetime it was
created with, and nothing else. Run bare, it lists them and removes nothing;
--yes removes them, and a scheduled job passes --yes. --dry-run means the same
as running bare.

The lifetime is read off each environment's own resources, stamped there when
it was created from that repository's runtime.ttl. It is never taken from the
manifest this command was run with, so a repository with a two hour lifetime
cannot remove another project's week long environment on the same machine.

Three things are never removed. An environment whose resources state no
lifetime, which is everything created before this feature existed: use
'af env prune --older-than' for those, where a person names the cutoff. An
environment something is running against, which is deferred to the next sweep
rather than pulled out from under a command. And anything that is not an
environment, such as the shared sidecar image.

An environment you are still using can be kept with 'af env extend'.

```
af env reap [flags]
```

```
# Only environments past the lifetime they were created with. The
# bare run lists them and removes nothing; a scheduled job passes --yes.
af env reap
af env reap --yes
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--dry-run` | `false` | List what would be removed and stop, which is also what running bare does. |
| `--yes` | `false` | Remove what the plan lists. Without it nothing is removed. |

### `af eval`

Run saved agent incidents as regression cases.

```
af eval
```

```
af eval run suite.json
```

Subcommands:

- [`af eval run`](#af-eval-run) Run every named scenario and retain each verdict.

### `af eval run`

Run every named scenario and retain each verdict.

```
af eval run <suite.json> [flags]
```

```
af eval run suite.json --candidate HEAD
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--candidate` | `HEAD` | Candidate Git revision for every saved case. |

### `af explain`

Show the effective configuration, with every default filled in.

The most common configuration bug is a default nobody knew about. This prints
the resolved value of every setting, so "why is it blocking that host" has a
one line answer.

```
af explain
```

```
af explain
af explain -o json
```

### `af explore`

Send agents at a goal with no declared workflow.

An exploration is a goal without a script. The agent reads each page through
the accessibility tree, chooses somewhere to go, and writes down every place
the application cost it effort. It answers the question a workflow cannot ask:
nothing broke, so why would somebody give up here.

It cannot fail your build. Nobody declared what should happen on the pages it
wanders onto, so a finding is an observation and never a red mark. Only a run
that could not start is reported as blocked.

Every choice comes from the goal's seed, so the same seed takes the same path
and every finding arrives with the command that replays it.

The manifest's goal is the default and the flags below override it for one run,
without writing anything: explore as a different persona, from a different
page, in a different window, for a different budget. A viewport of phone is
390x844 with a mobile user agent and a touch screen, tablet is 768x1024,
desktop is 1440x900, and WIDTHxHEIGHT is any size between 320 and 3840 a side.
A budget is a step count such as 8 or a duration such as 5m. A persona the
manifest does not declare is refused, and the refusal names the ones it does.
The report and the artifacts record the persona, the start path and the
viewport that were actually used.

```
af explore [flags]
```

```
# Agents go at a goal with no workflow written for it.
af explore
# The same goal as the owner, from the billing page, on a phone, in eight steps.
af explore --only upgrade-a-plan --persona owner --start /settings/billing --viewport phone --budget 8
af explore --emit-workflow checkout.yaml
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to run against, defaulting to the checked out one. |
| `--budget` | - | Most this run may spend: a step count such as 8, or a duration such as 5m. |
| `--emit-workflow` | `false` | Print the workflow block that replays what was explored, instead of the report. |
| `--focus` | - | A sentence about what to attend to; its words decide which controls are pressed first. |
| `--headed` | `false` | Show the browser rather than running it hidden. |
| `--only` | - | Explore just these goals, by name. |
| `--persona` | - | Explore as this declared persona rather than the goal's. |
| `--runner` | - | Path to the runner's entry point. |
| `--seed` | - | Replay with this seed rather than the one the manifest declares. |
| `--start` | - | Begin at this path rather than the goal's start_path, such as /settings/billing. |
| `--viewport` | - | Window to explore in: phone (390x844, mobile), tablet (768x1024), desktop (1440x900), or WIDTHxHEIGHT. |

### `af fidelity`

What this environment reproduces, component by component, and what it does not.

An inventory of the copy against the thing it is a copy of.

Every line comes from something the engine already knew: the runtime says what
is running, the database provider says which golden the branch came from and
whether its attestation still checks out, the branch says how much it holds and
whether the personas exist in it, and the manifest says which third party hosts
the policy names and what answers for each.

There is a headline number and it is defined on the page it prints: how many of
the measured components are production's own thing rather than a substitution,
a refusal or an absence. What could not be measured is excluded from it and
named, never counted as either a pass or a failure, because a percentage that
quietly absorbs an unknown is worth less than no percentage at all.

The per dimension verdict is the part to read. A change to billing cares about
the third party hosts and not about traffic; a migration cares about the data
and about neither. One averaged number hides whichever of those is yours.

Set fidelity.require in the manifest to fail this command when a dimension is
not fully reproduced.

```
af fidelity [flags]
```

```
# An inventory of the copy against the thing it is a copy of.
af fidelity
af fidelity -o json
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to inventory, defaulting to the checked out one. |

### `af github`

Connect this repository's pull requests to Antifailure.

The pull request integration runs in the repository's own GitHub Actions,
because an environment needs Docker, Postgres and a browser beside the code
under test, and the masked data stays inside the customer's own runner.

One workflow file makes that happen, and these commands write it.

```
af github
```

```
af github init
```

Subcommands:

- [`af github init`](#af-github-init) Write the workflow that checks every pull request.

### `af github init`

Write the workflow that checks every pull request.

Writes .github/workflows/antifailure.yml, the same file af init writes when
the checkout has a github.com remote, into a repository that already has a
manifest. The file calls a reusable workflow in the antifailure repository, so
it is short and rarely needs to change.

It is safe to run twice. A file identical to the template is left as it is and
said to be. A file that differs is left alone unless --force replaces it,
because a workflow somebody edited is theirs.

The manifest gains a github block when it has none, naming the three settings
the file depends on: the mode, whether a comment is left, and the fork policy.

Every secret the check can use is optional and is printed here by name, never
by value. The one repository variable a hosted control plane needs is printed
the same way.

```
af github init [flags]
```

```
# Writes .github/workflows/antifailure.yml and names the optional secrets.
af github init
# Replace a workflow file somebody edited with the template.
af github init --force
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--force` | `false` | Replace a workflow file that differs from the template. |

### `af golden`

Manage the masked copies branches are made from.

A refresh copies production, masks it, reads it back to check the masking, and
publishes it only if that check passes.

A golden that fails verification is never published, so it cannot be branched,
so no environment can ever hold it. That is enforced by the provider rather
than by remembering to check.

```
af golden
```

```
af golden list
```

Subcommands:

- [`af golden gc`](#af-golden-gc) List the goldens past the retention count, and remove them with --yes.
- [`af golden list`](#af-golden-list) List the goldens that exist.
- [`af golden pull`](#af-golden-pull) Bring a published golden onto this machine.
- [`af golden refresh`](#af-golden-refresh) Copy production, mask it, verify it, and publish it.
- [`af golden verify`](#af-golden-verify) Re-check a published golden.

### `af golden gc`

List the goldens past the retention count, and remove them with --yes.

How many to keep comes from database.golden.retain in the manifest, so that
every machine and every runner collects the same way. --keep overrides it for
one run.

Run bare, it lists which versions it would remove and which it would keep, and
removes nothing. --yes removes what the bare run listed. A golden is shared by
every branch of this project on the machine, so the list is worth a look
before it goes.

Two versions are never removed. One is any version an environment is still
branched from: taking away the copy something is running on breaks the
environment rather than tidying it, and that refusal comes from the provider,
which is the only thing that knows. The other is the newest verified golden,
whatever the count says, because a project with nothing left to branch cannot
bring an environment up at all, which is worse than the disk it saved.

```
af golden gc [flags]
```

```
# Lists which versions would go and which stay, and removes nothing.
af golden gc
af golden gc --yes
af golden gc --keep 3 --yes
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch context to use, defaulting to the checked out one. |
| `--keep` | `0` | How many of the newest goldens to keep, overriding database.golden.retain. |
| `--yes` | `false` | Remove what the plan lists. Without it nothing is removed. |

### `af golden list`

List the goldens that exist.

```
af golden list [flags]
```

```
af golden list
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch context to use, defaulting to the checked out one. |

### `af golden pull`

Bring a published golden onto this machine.

One machine holds the production credential and refreshes; every other machine
pulls what it published and never reads production at all. That is what
database.golden.storage and storage_url are for.

With no version, the newest complete one is taken. A version is complete when
its attestation is in the store: the dump is written first and the attestation
second, so a version with only a dump is a publish that did not finish, and it
is invisible here rather than offered.

A pulled golden is NOT trusted because it came from the store. The verification
scan runs again, here, against the database that actually arrived. A pull that
skipped it would make the store a way to get an unverified database branched.

```
af golden pull [version] [flags]
```

```
af golden pull
af golden pull gv_20260830044013_74234e98
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch context to use, defaulting to the checked out one. |

### `af golden refresh`

Copy production, mask it, verify it, and publish it.

```
af golden refresh [flags]
```

```
af golden refresh
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch context to use, defaulting to the checked out one. |

### `af golden verify`

Re-check a published golden.

Branches the golden, reads it back with the detectors, and removes the branch
whether or not the check passed.

Worth doing because a golden published under one set of rules is not verified
under another, and because a golden that arrives by import was never checked
here at all.

```
af golden verify <version> [flags]
```

```
af golden verify gv_20260830044013_74234e98
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch context to use, defaulting to the checked out one. |

### `af inbox`

Read the mail and messages the environment sent.

Every message a captured provider was asked to send is recorded here instead of
being delivered. Nobody receives anything, and the workflow that was waiting on
it can carry on.

The link and code are extracted for you, because an agent following a magic
link should not have to parse HTML to find it.

```
af inbox
```

```
af inbox list
```

Subcommands:

- [`af inbox get`](#af-inbox-get) Show one message in full.
- [`af inbox list`](#af-inbox-list) List what the environment sent.
- [`af inbox wait`](#af-inbox-wait) Block until a matching message arrives.

### `af inbox get`

Show one message in full.

```
af inbox get <number> [flags]
```

```
af inbox get 1
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to read, defaulting to the checked out one. |

### `af inbox list`

List what the environment sent.

```
af inbox list [flags]
```

```
af inbox list
af inbox list --to ada@example.com --limit 5
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to read, defaulting to the checked out one. |
| `--limit` | `50` | How many messages to show. |
| `--to` | - | Only messages addressed to this recipient. |

### `af inbox wait`

Block until a matching message arrives.

Waits for a message, checking what already arrived first.

That order matters. The message has usually been sent before anybody starts
waiting for it, and a wait that only looks forward is how a test passes on a
slow machine and fails on a fast one.

```
af inbox wait [flags]
```

```
# Blocks until the message arrives, or the timeout runs out.
af inbox wait --to ada@example.com
af inbox wait --subject 'Verify your email' --timeout 60s
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to read, defaulting to the checked out one. |
| `--subject` | - | Wait for a subject containing this text. |
| `--timeout` | `1m0s` | How long to wait. |
| `--to` | - | Wait for a message addressed to this recipient. |

### `af incident`

Inspect captured agent evidence and save an immutable replay scenario.

```
af incident
```

```
af incident list
af incident inspect billing-failure
```

Subcommands:

- [`af incident import`](#af-incident-import) Import an SDK capture into this project's local evidence store.
- [`af incident inspect`](#af-incident-inspect) Read retained incident content and missing dependencies.
- [`af incident list`](#af-incident-list) List incidents without hiding malformed records.
- [`af incident save`](#af-incident-save) Freeze an incident, a verified golden and distinct failure/fix assertions.

### `af incident import`

Import an SDK capture into this project's local evidence store.

```
af incident import <capture.json>
```

```
af incident import capture.json
```

### `af incident inspect`

Read retained incident content and missing dependencies.

```
af incident inspect <id>
```

```
af incident inspect billing-failure
```

### `af incident list`

List incidents without hiding malformed records.

```
af incident list
```

```
af incident list
```

### `af incident save`

Freeze an incident, a verified golden and distinct failure/fix assertions.

```
af incident save <id> [flags]
```

```
af incident save billing-failure --scenario billing --pointer /recommendation --original '"charge"' --expected '"review"' --table subscriptions
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--endpoint` | `/af-replay` | Explicitly enabled application replay endpoint. |
| `--expected` | - | JSON value the fix must produce. |
| `--golden` | - | Pin a verified golden; defaults to the capture reference. |
| `--original` | - | JSON value that identifies the original failure. |
| `--owner` | `local` | Owner of the regression case. |
| `--pointer` | - | JSON pointer into the agent outcome. |
| `--scenario` | - | Name the immutable scenario. |
| `--table` | - | Relevant database tables to compare. |

### `af init`

Read the repository and write antifailure.yaml.

Detection reads the repository and proposes a manifest: the services it found,
the port each listens on, the migration command, and, most usefully, a network
policy derived from the SDKs you depend on.

It never runs anything from the repository. Everything it reports names the
file it came from, so you can check the reasoning rather than trust it.

Anything detection is not sure about becomes a question rather than a silent
guess, because a manifest you have to audit is worth less than one you can
read.

A service is identified by the directory it is built and run from, not by its
name, because every source spells the name differently: a Dockerfile and a
language analyzer use the directory, a compose file uses its own key, a
Procfile uses the process name, and a package manifest uses the package. One
application described by several of those is one service, and the name it keeps
comes from the source that identifies an application best, a package manifest
ahead of a compose key ahead of a Procfile process ahead of the directory.
Where one source declares two services in a directory, which is what a compose
file with a web and an admin container on one build context is, they stay two.

A Dockerfile in a subdirectory is built either from that directory, which is
what 'docker build <dir>' does, or from the repository root, which is what a
monorepo image reaching a lockfile at the top of the tree needs. Its COPY lines
say which: a path that exists beside the Dockerfile and not at the root means
the directory, and one that exists only at the root means the root. Where they
do not settle it, this is a question rather than a default, because building
from the wrong one either fails on a missing path or, with COPY . ., succeeds
and produces an image assembled from the wrong directory.

--answer settles a question, and also overrides a value detection read with
confidence, such as a port an EXPOSE line named. An id naming nothing is
refused with the ids that would have worked rather than dropped in silence.

```
af init [flags]
```

```
af init
af init --non-interactive --answer database.present=yes
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--answer` | - | Answer a question, or override a detected value, as id=value. Repeatable. |
| `--force` | `false` | Replace an existing antifailure.yaml with a fresh detection; nothing is merged and its edits are lost. |
| `--non-interactive` | `false` | Do not ask questions; accept every default and report what was assumed. |

### `af insights`

What Postgres can tell you about this change before anybody clicks anything.

A branch is a real database with production's shape in it, which makes some
questions answerable without running the application at all.

The migrations are rehearsed against a throwaway branch and every statement is
timed, so a migration that takes four seconds on an empty test database and
ninety on production row counts is visible before the deploy window rather than
during it. The plans on that branch are compared before and after, which is how
a sequential scan appearing where an index scan was gets found. And the queries
this environment ran are compared against a report saved on the base branch.

Where the migrations take something away, the previous release is built and run
against the migrated branch as well, because a rolling deploy leaves both
releases talking to the same database for the length of the window and nothing
else here checks that. It exits non zero only when a workflow passes without
the migrations and fails with them.

It says what it could not measure, and it names any check the manifest turned
off. A report that silently omits a check reads exactly like a check that found
nothing.

```
af insights [flags]
```

```
# Rehearses the migration against a branch of the golden.
af insights
# Save a report on the base branch, compare against it on this one.
af insights --save baseline.json
af insights --baseline baseline.json
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--against` | - | Which commit the previous release is, overriding the manifest. |
| `--baseline` | - | Compare against a report saved earlier. |
| `--branch` | - | Branch to read, defaulting to the checked out one. |
| `--limit` | `20` | How many queries to show. |
| `--no-rehearsal` | `false` | Skip the migration rehearsal, which is the only check that makes a second branch. |
| `--runner` | - | Path to the runner's entry point. |
| `--save` | - | Save this report to compare against later. |

### `af invariants`

Ask the data the questions the manifest declares.

An invariant is a read only statement that must return no rows, asked of the
branch, so that a flow which appears to succeed while corrupting data is caught
by the data rather than by the screen.

They are asked automatically after the workflows in 'af test' and 'af ci'. This
runs them on their own, which is what you want while writing one, or after a
migration, or when a run failed and you want to know whether the data is the
reason.

Every statement runs inside a transaction opened READ ONLY, so a write is
refused by Postgres rather than trusted not to happen, and each one has its own
timeout. Rows returned means the invariant is violated, and the rows are the
evidence: they are printed, because a check that tells you something is wrong
without telling you which rows has told you to go and do the work yourself.

```
af invariants [flags]
```

```
af invariants
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to ask, defaulting to the checked out one. |

### `af license`

Show the license status of this installation.

This is the community edition. It has no license and needs none.

Everything the engine does is here and stays here: masked environments, sealed
egress, captured mail, agents, load, insights, and teardown. None of it expires
and none of it phones home.

A license adds the enterprise edition, which is a separate binary built from
the ee directory of the same repository: single sign on, SCIM, custom roles,
SIEM streaming, organization wide policy enforcement, customer owned runtime
clusters, enterprise secret managers, and billing.

```
af license
```

```
af license status
```

Subcommands:

- [`af license install`](#af-license-install) Install an enterprise license key.
- [`af license remove`](#af-license-remove) Remove the installed license key.
- [`af license status`](#af-license-status) What this installation is licensed for.

### `af license install`

Install an enterprise license key.

```
af license install <key>
```

```
af license install AF-LICENSE-KEY
```

### `af license remove`

Remove the installed license key.

```
af license remove
```

```
af license remove
```

### `af license status`

What this installation is licensed for.

```
af license status
```

```
af license status
```

### `af load`

Send traffic shaped like production's at the environment.

A weighted mix rather than one endpoint at a fixed rate. Hammering one endpoint
proves that endpoint is fast, which nobody doubted; what breaks under real
traffic is the mix, and the page nobody thinks about that is nine percent of
requests.

Every route is treated as unsafe until the manifest names it safe. A generator
that finds POST /checkout in an access log and exercises it four hundred times
is a generator that charges four hundred cards.

```
af load
```

```
af load smoke
```

Subcommands:

- [`af load compare`](#af-load-compare) Run the same traffic against the base branch too, and report what moved.
- [`af load run`](#af-load-run) Run the full load profile.
- [`af load scenario`](#af-load-scenario) Run the declared journeys against the environment.
- [`af load smoke`](#af-load-smoke) Send a short burst, to check the environment answers under any load at all.
- [`af load sql`](#af-load-sql) Run a concurrent SQL workload against the branch's database.

### `af load compare`

Run the same traffic against the base branch too, and report what moved.

Brings a second environment up from the base revision, branches the same golden
for both so they answer queries over identical rows, sends both the same
weighted mix in the same order under the same seed, and reports every route and
every run wide number that moved.

This is the base branch comparison. It is a different question from the one
'af load run' answers: that measures one build against what production serves,
using the per route p95 in your traffic source, and it is the right question
when you want to know whether a route is slower than the fleet. This one
measures this build against the last one, which is the right question when you
want to know whether your change made it slower.

Each side is first sent a short warm-up that is thrown away, which takes the
first request of every route out of the numbers. Then each side is sent the mix
in rounds, interleaved so that neither side always goes first, with the same
seed for both sides in each round.

Each route is judged ROUND AGAINST ROUND. Every round is a small comparison of
its own, and the change is measured from how those comparisons agreed, with an
interval as wide as the host's own noise between rounds. The intervals hold at
ninety percent for every route together. A limit inside a route's interval is
neither a pass nor a fail, and the report prints the smallest change that route
could have shown on this host, so on a noisy machine the answer is "too close
to say" rather than a regression that is not there. More rounds or a longer
duration narrows it.

What it still cannot control is printed with every report rather than left
implied. The rounds are sequential, because two environments sending traffic
at once on one host would contend with each other and measure that instead.
A difference is a difference, and a threshold under load.comparison.thresholds
is what turns one into a verdict.

With --sql it compares the concurrent SQL workload instead: clients running
whole transactions against each build's own database rather than requests
against its application. Same mix, built once on this build so that neither
side reads its own pg_stat_statements, same client count, same think time and
the same per round seed. The unit of comparison becomes the transaction and
the statement inside it, and each one reports p50, p95 and p99 on both sides.
Throughput becomes committed transactions a second, judged against the same
load.comparison.thresholds.throughput_drop. It needs a load.sql block and
refuses without one.

With --image and --baseline-image it compares two builds of the DATABASE rather
than two builds of the application. Each defaults to the manifest's
database.image, so naming one varies that side alone. When only the images
differ the two sides run the same application revision, built from the same
tree, and the base being the same commit is then allowed rather than refused:
that is what makes the difference the database's. There is still one golden, so
one build wrote its data directory and the other opens it, and a build that
cannot open the other's data directory is reported as that finding rather than
as an environment that would not start. The report names which axis differed.

The base environment is torn down unless --keep says otherwise. The
environment for this build is left running whether or not this brought it up.

```
af load compare [flags]
```

```
af load compare
af load compare --baseline origin/main --duration 60s
af load compare --sql --concurrency 16
af load compare --sql --baseline-image postgres:17-alpine
af load compare --seed 7 --keep
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--baseline` | - | Revision to compare against, overriding load.comparison.base_ref. |
| `--baseline-image` | - | Database image the base side runs, overriding database.image. With --image this compares two database builds over one golden. |
| `--branch` | - | Branch to compare, defaulting to the checked out one. |
| `--concurrency` | `8` | Clients each side runs at once, overriding load.sql.clients. Needs --sql. |
| `--duration` | `0s` | How long to send for on each side, overriding the manifest. |
| `--image` | - | Database image this build runs, overriding database.image. The application is unchanged. |
| `--keep` | `false` | Leave the base environment up, for looking at a difference. |
| `--report` | - | Write the comparison here as well as to the terminal. |
| `--rounds` | `0` | Interleaved rounds per side, 16 when not set. 1 measures each side once, base first. |
| `--scale` | `0` | Fraction of production's arrival rate to send at each side, overriding the manifest. |
| `--seed` | `0` | Seed for the request sequence. The same seed is used on both sides. |
| `--sql` | `false` | Compare the SQL workload from load.sql instead of the HTTP mix. |
| `--think-time` | `0s` | How long a client waits between transactions, overriding load.sql.think_time. Needs --sql. |
| `--transactions` | `0` | Transactions each client runs, split across the rounds, overriding load.sql.transactions. Needs --sql. |
| `--warmup` | `0s` | Mix sent at each side and discarded before measuring, 5s when not set. 0s sends none. |

### `af load run`

Run the full load profile.

```
af load run [flags]
```

```
# A weighted mix, not one endpoint at a fixed rate.
af load run
af load run --duration 60s --scale 2
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to send at, defaulting to the checked out one. |
| `--duration` | `1m0s` | How long to send for. |
| `--scale` | `1` | Multiplier on production's rate. |
| `--seed` | `1` | Makes two runs send the same sequence. |

### `af load scenario`

Run the declared journeys against the environment.

A scenario is an ordered journey rather than a mix: open the billing page, ask
for the subscription, submit, and submit again three hundred milliseconds later
because the first one felt slow. Sessions walk it at once, and one scenario can
start after another so a burst arrives while something else is already running.

The requests are HTTP. Clicking a button is 'af test' and the browser agents;
this is what the load generator can send, at the concurrency load runs at.

Every step is checked against load.safe_routes before anything is sent, so a
scenario that names an undeclared route is blocked rather than run.

```
af load scenario [flags]
```

```
af load scenario
af load scenario --only checkout --concurrency 20
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to send at, defaulting to the checked out one. |
| `--concurrency` | `20` | Ceiling on requests in flight. |
| `--only` | - | Run just these scenarios, by name. |
| `--seed` | `1` | Makes two runs send the same schedule. |

### `af load smoke`

Send a short burst, to check the environment answers under any load at all.

```
af load smoke [flags]
```

```
af load smoke
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to send at, defaulting to the checked out one. |
| `--duration` | `10s` | How long to send for. |
| `--scale` | `0.1` | Multiplier on production's rate. |
| `--seed` | `1` | Makes two runs send the same sequence. |

### `af load sql`

Run a concurrent SQL workload against the branch's database.

Clients, each on its own connection, running whole transactions against the
database directly rather than through the application.

Everything else this engine sends goes over HTTP, so the number it reports is
the application's latency with the database somewhere inside it. That is the
right measurement for an application change and the wrong one for a database
change. Somebody changing an index, a lock, a storage parameter or a query
wants transactions per second and statement latency, and can only reach them
through whatever the application happens to do on a route they can reach.

The statements come from a document in the repository, or from
pg_stat_statements on the branch, which is the traffic that really ran weighted
by how often it ran. A derived mix cannot recover the values, because the
statistics normalise them away, so it asks the server for the parameter types
and generates values of those types. It refuses a write unless the manifest
allows one, and every run reports the rows its statements actually touched, so
a reader can tell a fast query from a query that found nothing.

The run reports how many of its own backends the server had inside a
transaction at one instant, read from pg_stat_activity while it was going. N
clients are not N concurrent sessions and that number is the evidence rather
than the claim.

The same connection asks pg_blocking_pids which of those backends were waiting
for a lock and which ones were in front of them, so a run reports the
contention it was under rather than only the deadlocks loud enough to end a
transaction. Sampled, so the counts are floors rather than totals, and a run
nobody watched reports nothing rather than zero.

```
af load sql [flags]
```

```
# Clients on their own connections, running transactions against the database.
af load sql
af load sql --concurrency 16 --duration 2m --think-time 20ms
af load sql --only 'read one order' --transactions 500
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to run against, defaulting to the checked out one. |
| `--concurrency` | `8` | How many clients run at once, each on its own connection. |
| `--duration` | `1m0s` | How long to run for. |
| `--only` | - | Run only these transactions, by name. Repeat the flag for several. |
| `--seed` | `1` | Makes two runs execute the same sequence. |
| `--think-time` | `0s` | How long a client waits between transactions. |
| `--transactions` | `0` | How many transactions each client runs, instead of a duration. |

### `af login`

Sign in to a control plane from this terminal.

Signs this machine in to a control plane using the device authorization grant.

af login prints a short code and opens a browser. Approve it there, and the
token arrives here over TLS and goes straight into the operating system's
credential store. The credential is never shown, never copied through a
clipboard, and never written to a shell history file.

By default the token can read environments and runs and write events, and
nothing else: it cannot manage members, change policy, or touch a provider key.

--scope asks for more. The scope is shown on the screen where the login is
approved, so nobody grants a capability without seeing the words:

  af login --scope providers.write

Nothing reads a key back. There is no scope for it, because storing a secret and
retrieving one are different capabilities and a terminal needs only the first.

Run af logout to remove it from this machine and revoke it everywhere.

```
af login [flags]
```

```
af login
af login --control-plane https://app.antifailure.dev --no-browser
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--control-plane` | - | The control plane to sign in to (default: AF_CONTROL_PLANE_URL, or the hosted instance). |
| `--no-browser` | `false` | Do not try to open a browser; print the address instead. |
| `--scope` | - | Ask for a capability beyond the default, e.g. providers.write. Repeatable. |

### `af logout`

Remove this machine's credential and revoke it.

Removes the stored token and tells the control plane to revoke it.

Both halves matter. Removing it locally stops this machine using it; revoking
it stops anybody who copied it. A logout that only deleted the local copy would
leave a working credential in whatever backup or screen recording captured it.

If the control plane cannot be reached, the local credential is still removed
and the command says the revocation did not happen, so nobody is left believing
a token is dead when it is not.

```
af logout [flags]
```

```
af logout
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--control-plane` | - | The control plane to sign out of (default: AF_CONTROL_PLANE_URL, or the hosted instance). |

### `af logs`

Show what the environment's services have written.

Output from every service, or from one if you name it.

Everything here goes through the redactor on the way out. A service's own log
is the second likeliest place for a secret to surface after a build log, and
this is the command people paste into issues.

```
af logs [service] [flags]
```

```
af logs
af logs web --tail 100
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to read, defaulting to the checked out one. |
| `--tail` | `200` | How many lines to show per service. |

### `af mask`

Plan, apply, and check the masking of this environment's data.

Masking is compiled from the live schema rather than from a list, because a
list of columns goes stale the moment somebody adds one and the failure mode is
silent: the new column holds real addresses and nothing says so.

A column no rule covers is reported rather than left alone. Left alone, for a
column called customer_notes, means the notes ship.

```
af mask
```

```
af mask plan
```

Subcommands:

- [`af mask apply`](#af-mask-apply) Rewrite this environment's data according to the plan.
- [`af mask crossstore`](#af-mask-crossstore) Check that one person masks to the same person in every store.
- [`af mask init`](#af-mask-init) Read the schema and write masking.yaml with a rule for every column.
- [`af mask plan`](#af-mask-plan) Show what masking would do, column by column.
- [`af mask preview`](#af-mask-preview) Show what a few rows would look like after masking.
- [`af mask verify`](#af-mask-verify) Read the data back and report anything that still looks real.

### `af mask apply`

Rewrite this environment's data according to the plan.

Applies the plan to the branch this environment is using.

This is irreversible: once a column is overwritten the original is gone. It is
safe here because the branch is a copy, and it is exactly how a golden is
produced, so trying it on a branch first is the way to iterate on rules.

```
af mask apply [flags]
```

```
# Rewrites this environment's data in place.
af mask apply
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to mask, defaulting to the checked out one. |

### `af mask crossstore`

Check that one person masks to the same person in every store.

Determinism inside one store has been enforced since the beginning, by the key
derivation. Across two stores it was a property of the construction that
nothing checked, and a property nothing checks is a property you have somebody's
word for.

The failure it exists to catch is silent. An empty ClickHouse beside a masked
Postgres is a twin that is visibly incomplete and somebody notices within a
minute of opening a chart. One identity masked into two different fake people
is a twin that is confidently wrong: every join across the two stores returns
nothing or returns the wrong person, every report built on it is plausible, and
nothing anywhere says so.

It reads schemas and no rows. The check masks its own probe values through both
stores' rules and compares the outputs, so what it needs from a store is the
catalog, which is why it is safe to point at production. Every store it reads
is named, every store it could not read is named with the reason, and a run
that reached one store reports that it proved nothing rather than reporting a
hundred percent of one.

Each datastore says where its schema is read from with source_url_env, which
names an environment variable and never the connection string. The primary
takes that from database.source_url_env and does not repeat it.

```
af mask crossstore [flags]
```

```
# Reads both stores' catalogs and no rows, which is what makes it safe
# to point at production.
af mask crossstore
af mask crossstore --branch main
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch context to use, defaulting to the checked out one. |

### `af mask init`

Read the schema and write masking.yaml with a rule for every column.

Reads the schema of the database source, or of this environment's branch when
one is up, decides every column the way the built in rules would, and writes
the result to masking.yaml as one explicit rule per column.

The file it writes leaves the plan with nothing to ask. A column a built in
rule recognises gets that rule restated with its reason. A column nothing
recognises gets a rule that empties it, with a reason saying it was
unrecognised and is emptied until somebody says otherwise. Numbers, times and
identifiers get no rule, because nothing is done to them.

It refuses to replace a file that is already there unless --force is passed,
because the rules somebody edited are the most valuable thing in it.

```
af mask init [flags]
```

```
# Reads the schema and writes masking.yaml with a rule for every column
# that needs one, so af mask plan has nothing left to ask.
af mask init
af mask init --force
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch whose environment to read, defaulting to the checked out one. |
| `--force` | `false` | Replace a masking file that is already there. |

### `af mask plan`

Show what masking would do, column by column.

```
af mask plan [flags]
```

```
af mask plan
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to plan against, defaulting to the checked out one. |

### `af mask preview`

Show what a few rows would look like after masking.

Reads a few rows, transforms them in memory, and writes nothing.

Somebody iterating on rules has to see the output before committing to it, and
the alternative, applying and then looking, is irreversible on a branch they may
want to keep.

```
af mask preview [flags]
```

```
af mask preview
af mask preview --table users --rows 5
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to read, defaulting to the checked out one. |
| `--rows` | `3` | How many rows to show. |
| `--table` | - | Preview one table, defaulting to the first being masked. |

### `af mask verify`

Read the data back and report anything that still looks real.

Reads a sample of every column it can read as text and runs the same detectors
that would find the data if it leaked. Strings, JSON, arrays and enums are read
through their text form; a bytea column is decoded as UTF-8 where it decodes.
A column of a type the scanner cannot read is listed as not readable rather
than passed over, and when no masking rule covers such a column and its name
says it holds a secret, the check fails.

The count of columns masking copied unchanged because no rule covered them is
printed beside the verdict, whichever way the verdict went.

Masking that is not checked is masking somebody believes in. A rule that missed
a column, a transform that failed on a null, a table added last week: each
produces data that looks masked and is not, and none of them announces itself.

```
af mask verify [flags]
```

```
af mask verify
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to check, defaulting to the checked out one. |

### `af mcp`

Serve the rehearsal tools to a model over the Model Context Protocol.

Serve this repository's rehearsal tools to an MCP client on standard input and
output.

The agent on the other end chooses what to rehearse. It does not choose how
safely the rehearsal runs: there is no argument on any tool that can disable
sanitization, widen the egress policy, lower a threshold or name a database.
Thresholds come from this project's manifest, and the verdict is decided by the
same evaluator af ci uses, so a tool call and a pull request check cannot
disagree about the same change.

The server serves exactly this checkout. A tool call may state which project it
believes it is talking to, and a call naming a different one is refused rather
than followed.

Standard output carries the protocol and nothing else. Progress, warnings and
errors go to standard error, where the client's log will show them.

Client setup differs by host. https://antifailure.dev/docs/reference/mcp has
the current command or configuration for each supported local client. This
release provides no hosted MCP URL. A browser client requires a separately
operated and authenticated Streamable HTTP bridge.

```
af mcp
```

```
# Started by an MCP client, not typed. It speaks the protocol on
# standard input and output, so running it in a terminal looks idle.
af mcp
# It serves exactly the checkout it starts in, so the client is
# configured to run it there. A client without a working directory
# setting passes -C with the absolute checkout in its configuration.
af mcp
```

### `af model`

The model key the agents use on this machine.

The agents can read a page and decide what a person would do next, which takes a
model. The key is yours: it is stored on this machine, the call goes straight to
the provider, and nothing hosted is involved.

With no key the deterministic planner runs instead. That is a supported mode,
not a broken one: workflows still run, still drive a real browser and still
produce a verdict. The model is what turns a workflow written as a sentence into
one the runner follows without being told every field.

  af model show     what is configured, and what a run will use
  af model test     prove the key works, with one cheap call
  af model set      store a key, without it touching the command line
  af model rm       remove a stored key

If you have a control plane, 'af provider' is the better place for a key: it
seals it, caps what may be spent on it per month, and checks that cap before the
key is ever decrypted. This command is the one that needs nothing but a terminal.

```
af model
```

```
af model show
```

Subcommands:

- [`af model rm`](#af-model-rm) Remove a stored key.
- [`af model set`](#af-model-set) Store a key, without it touching the command line.
- [`af model show`](#af-model-show) What is configured, and what a run will use.
- [`af model test`](#af-model-test) Prove the key works, with one cheap call.

### `af model rm`

Remove a stored key.

Removes the key from every place this command can write it, not from the first
one that answers. A key left in the encrypted store after the keyring entry was
removed is a key the next run silently uses, which is the exact failure somebody
is trying to prevent when they type this.

It cannot remove a key from a shell you exported it in or from a .env file, and
it says so when one is still there rather than reporting a removal that changed
nothing.

Removing a key that is not there is not an error. This is a command people run
in a hurry, and a retry after a timeout must not report failure for reaching the
state you asked for.

This does not reach the provider. If the key leaked, revoke it at Anthropic or
OpenAI as well: removing it here stops this machine using it and stops nobody
else.

```
af model rm <provider>
```

```
af model rm anthropic
```

### `af model set`

Store a key, without it touching the command line.

Stores a key in the system keyring where this platform has one, and in the
encrypted local store where it does not.

The key is never an argument. There is no --key flag, deliberately: a secret on
a command line is written to your shell's history file, is visible in ps to
every other user on the machine, and is captured by any recording of the
terminal. So there are three ways to give it, and none of them put it in the
argument vector:

  af model set anthropic                      asks, without echoing
  af model set anthropic --stdin < key.txt    reads one line
  af model set anthropic --from-env NAME      reads that environment variable

Where it lands is reported rather than assumed, because the two places are not
equivalent. macOS gates the keychain on the login keychain, Linux on the session
keyring daemon, and Windows on the user's credentials. The encrypted local store
is a file, and it is only as strong as the passphrase protecting it.

This does not reach the provider. Storing a key here does not create one and
removing it does not revoke one.

```
af model set <provider> [flags]
```

```
# The key is read from the environment or from stdin, so it never
# reaches the command line or the shell history.
af model set anthropic --from-env ANTHROPIC_API_KEY
af model set anthropic --stdin < key.txt
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--from-env` | - | Read the key from this environment variable instead of asking. |
| `--stdin` | `false` | Read the key from standard input, one line. |

### `af model show`

What is configured, and what a run will use.

Reports the provider, the model, the endpoint, where the key was found and when
it was last proven to work.

It does not show the key and there is no flag that would. The fingerprint
answers the question this is usually asked to answer, which is whether the key
here is the one you think it is, and it answers it without either person having
to read a secret out loud.

"Where it came from" is worth as much as the rest together. A key exported in
one shell and a key in the keyring look identical from a run's point of view
until they disagree, and then the only useful sentence is which one won.

```
af model show
```

```
af model show
af model show -o json
```

### `af model test`

Prove the key works, with one cheap call.

Sends one completion of a single token and reports what came back.

A real call rather than a check of the key's shape, because a well formed key
that was revoked this morning passes every shape check there is. It costs a
fraction of a cent, which is the point: this is meant to be run whenever you are
unsure, and a check people avoid because of the price is a check nobody runs.

What it can tell apart matters more than that it runs. A revoked key, an empty
balance, a model name that does not exist, a throttle, a provider outage and an
endpoint nothing answers on all fail, they all have different fixes, and being
told only that the call failed sends you to the wrong one first.

On success it writes down that this exact key worked, and 'af model show'
reports it. Rotating the key discards that, because a previous key's success
says nothing about the new one.

```
af model test [flags]
```

```
# One cheap call, so a broken key is found here and not mid run.
af model test
af model test --timeout 10s
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--timeout` | `30s` | How long to wait for the endpoint, which a local model may need more of. |

### `af net`

Inspect and explain the environment's network policy.

An environment reaches nothing on the network except the hosts in the manifest,
each in the mode named there. These commands say what that adds up to, without
needing an environment to be running.

```
af net
```

```
af net policy
```

Subcommands:

- [`af net explain`](#af-net-explain) Say what would happen to one request, and which rule decides it.
- [`af net log`](#af-net-log) Show what the environment tried to reach, and what happened.
- [`af net policy`](#af-net-policy) Show the effective policy, in the order that decides.

### `af net explain`

Say what would happen to one request, and which rule decides it.

Prints the decision, the rule that made it, and every other rule that also
matched, so a surprising answer is diagnosable rather than mysterious.

```
af net explain <method> <url>
```

```
af net explain GET https://api.stripe.com/v1/charges
af net explain POST https://api.resend.com/emails
```

### `af net log`

Show what the environment tried to reach, and what happened.

Every outbound request the environment made, allowed or refused, with the rule
that decided it.

The allowed ones are the point. A log of refusals answers "why was this
blocked"; a log of everything answers "did anything reach Stripe", which is the
question somebody asks after an incident.

```
af net log [flags]
```

```
af net log
af net log --blocked --limit 20
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--blocked` | `false` | Show only requests that were refused. |
| `--branch` | - | Branch to read, defaulting to the checked out one. |
| `--limit` | `200` | How many decisions to show, most recent last. |

### `af net policy`

Show the effective policy, in the order that decides.

Rules are printed most specific first, which is the order they are evaluated
in. An exact host beats a wildcard, a longer path beats a shorter one, and an
explicit method beats any, so where a rule sits in this list is where it sits
in the decision, no matter where it sits in the file.

```
af net policy
```

```
af net policy
```

### `af oracle`

Run this change beside the version it is replacing and diff what they did.

Brings a second environment up from a baseline revision, branches the same
golden for both so they start from identical rows, sends both the same requests
in the same order, and reports every difference in what came back and in what
ended up in the database.

Responses and database contents are compared. Events, outbound effects, traces
and query plans are not: two comparisons done completely are worth more than six
done shallowly, because the first one that cries wolf is the last one anybody
looks at.

Values that no two runs can agree on are normalised before they are compared:
two timestamps within an hour, two UUIDs, two numbers within a relative
tolerance. Everything the comparison declined to look at is printed, defaults
included, because an oracle that silently ignores a field is worse than one that
reports it.

The candidate environment is left running whether or not this command brought it
up. The baseline is torn down unless --keep says otherwise.

```
af oracle [flags]
```

```
# Runs this change beside the version it replaces and diffs both.
af oracle
af oracle --baseline origin/main --fail-on any
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--baseline` | - | Revision to compare against, overriding oracle.base_ref. |
| `--branch` | - | Branch to compare, defaulting to the checked out one. |
| `--fail-on` | - | Lowest severity that fails the command: none, minor, major, or critical. |
| `--keep` | `false` | Leave the baseline environment up, for looking at a difference. |
| `--report` | - | Write the report here as well as to the terminal. |

### `af provider`

Your own model provider keys and their monthly caps.

Stores your Anthropic and OpenAI keys on the control plane, sealed with a secret
that is not in its database, and caps what may be spent on each one per month.

Runs use your key. We never see it after you save it: what any screen or any
command here can read is the last four characters and a fingerprint.

These commands need a token that asked for the capability:

  af login --scope providers.write

A token from a plain af login cannot reach a key, which is deliberate. The scope
appears on the screen where the login is approved, so nobody grants this without
seeing the words.

```
af provider
```

```
af provider list
```

Subcommands:

- [`af provider budget`](#af-provider-budget) Cap what may be spent on a provider this month.
- [`af provider list`](#af-provider-list) What is stored, and what it may spend this month.
- [`af provider rm`](#af-provider-rm) Remove a stored key.
- [`af provider set`](#af-provider-set) Store or rotate a key, without it touching the command line.

### `af provider budget`

Cap what may be spent on a provider this month.

Sets the monthly cap in US dollars. The cap is checked BEFORE the key is
decrypted, so a run with no allowance never causes the key to exist in the
control plane's memory at all. That ordering is the difference between a cap and
a suggestion.

A provider with no cap cannot spend anything. A missing cap reads as zero rather
than as unlimited, because the alternative on somebody else's key is an
unbounded bill.

A cap of zero is allowed and means exactly that: spend nothing on this provider.

```
af provider budget <provider> <usd> [flags]
```

```
af provider budget anthropic 50
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--control-plane` | - | The control plane to use (default: AF_CONTROL_PLANE_URL, or the hosted instance). |

### `af provider list`

What is stored, and what it may spend this month.

Shows which providers have a key, the last four characters of each, and the
monthly cap against what has been spent.

It does not show a key, and there is no flag that would. The last four and the
fingerprint are enough to answer the question this is usually asked to answer:
whether the key here is the one you think it is.

```
af provider list [flags]
```

```
af provider list
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--control-plane` | - | The control plane to use (default: AF_CONTROL_PLANE_URL, or the hosted instance). |

### `af provider rm`

Remove a stored key.

Removes the stored key. Runs that need this provider are refused afterwards,
with a message saying why, rather than falling back to a key of ours.

This does not reach the provider. If the key leaked, revoke it there as well:
removing it here stops us using it and stops nobody else.

Removing a key that is not there is not an error. This is the command somebody
runs in a hurry, and a retry after a timeout must not report failure for
reaching the state they asked for.

```
af provider rm <provider> [flags]
```

```
af provider rm anthropic
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--control-plane` | - | The control plane to use (default: AF_CONTROL_PLANE_URL, or the hosted instance). |

### `af provider set`

Store or rotate a key, without it touching the command line.

Stores a key for anthropic or openai, replacing whatever was there.

The key is never an argument. There is no --key flag, deliberately: a secret on
a command line is in the shell's history file, is visible in ps to everybody
else on the machine, and is in any recording of the terminal. So there are three
ways to give it, and none of them put it in the argument vector: it is asked
for without echoing, read as one line from stdin, or read from an environment
variable this process already has.

Rotating stores the new key and revokes the old one together. If the key given
is the one already stored, that is reported rather than accepted quietly: it is
the mistake people make at the moment they believe they have replaced a leaked
key.

```
af provider set <provider> [flags]
```

```
# The key is read from the environment or stdin, never from a flag,
# so it does not land in shell history.
af provider set anthropic
af provider set anthropic --stdin
af provider set anthropic --from-env ANTHROPIC_API_KEY
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--control-plane` | - | The control plane to use (default: AF_CONTROL_PLANE_URL, or the hosted instance). |
| `--from-env` | - | Read the key from this environment variable instead of asking. |
| `--stdin` | `false` | Read the key from standard input, one line. |

### `af replay`

Reproduce an agent failure, then test a candidate in an independent branch.

```
af replay <scenario> [flags]
```

```
af replay billing --candidate HEAD
```

Subcommands:

- [`af replay inspect`](#af-replay-inspect) Read a replay attempt and its retained evidence.
- [`af replay recover`](#af-replay-recover) Reconcile an interrupted attempt's two environments.
- [`af replay retire`](#af-replay-retire) Delete a scenario's unreferenced content and retain its retirement reason.

| Flag | Default | What it does |
| --- | --- | --- |
| `--candidate` | `HEAD` | Candidate Git revision. |
| `--timeout` | `20m0s` | Shorten the 20-minute setup/replay cap; cleanup has its own budget. |

### `af replay inspect`

Read a replay attempt and its retained evidence.

```
af replay inspect <attempt>
```

```
af replay inspect rpl_example
```

### `af replay recover`

Reconcile an interrupted attempt's two environments.

```
af replay recover <attempt>
```

```
af replay recover rpl_example
```

### `af replay retire`

Delete a scenario's unreferenced content and retain its retirement reason.

```
af replay retire <scenario> [flags]
```

```
af replay retire billing --reason 'The billing workflow was removed'
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--reason` | - | Record why the regression case is retired. |

### `af runner`

Install and check the agent runner.

The runner drives a real browser, so it is a separate program in a separate
language. It is installed from a copy that ships with this engine rather than
downloaded, because the source a release was tested with is the source that
release should run.

```
af runner
```

```
af runner check
```

Subcommands:

- [`af runner check`](#af-runner-check) Say whether the runner can run.
- [`af runner install`](#af-runner-install) Put the runner where af test will find it.

### `af runner check`

Say whether the runner can run.

Reports each thing af test needs from the runner separately: the source, the
dependencies it declares, a node new enough to run it, and the browser.

It reports on the runner af test would actually use from here, which is the
nearest one that can run rather than the nearest one that exists. Any runner it
went past is named, with what is wrong with it, because a report about a
directory the reader did not mean is how this command came to say a runner was
ready while the run took a different copy and died on a module it could not
resolve.

It does not claim the runner executes. Knowing that means starting node and
launching a browser, which is what af test is. Anything this cannot determine
is reported as not checked rather than as ok, because a check that answers ok
about something it never examined is worse than one that admits the gap: this
command used to report "ok runner" whenever src/main.ts existed, which was true
of an install with no dependencies at all, and the real failure surfaced much
later inside af test as a node error about a module it could not resolve.

The verdict has three values and not two, for the same reason. Ready means
every question that decides whether af test can run was asked and answered ok,
and exits 0. Blocked means one of them was answered no, and exits 3. Undetermined
means one of them could not be answered at all, which is neither, and exits 9,
the code reference/errors.md publishes as "nothing was measured".
A runner whose package.json cannot be parsed used to land in the first of those
and report itself complete.

```
af runner check
```

```
af runner check
```

### `af runner install`

Put the runner where af test will find it.

```
af runner install [flags]
```

```
af runner install
af runner install --skip-browser
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--from` | - | Copy from this directory rather than the one beside the engine. |
| `--skip-browser` | `false` | Do not download the browser. |

### `af secret`

Store values in the encrypted local store.

The last place the engine looks for a declared variable, after this shell's
environment and after .env.

The file is encrypted with a key derived from a passphrase, and it lives under
.antifailure, which 'af init' adds to .gitignore. It is a convenience for a
workstation and it is not a secret manager for a team: a value here is as safe
as the passphrase and the disk it is on.

Set AF_SECRET_PASSPHRASE before using it. There is deliberately no default:
a store encrypted with a passphrase everybody knows is a store that only looks
encrypted.

```
af secret
```

```
af secret list
```

Subcommands:

- [`af secret list`](#af-secret-list) List the names in the store.
- [`af secret rm`](#af-secret-rm) Remove a value from the store.
- [`af secret set`](#af-secret-set) Store a value, read without echo.

### `af secret list`

List the names in the store.

Names only. There is no command that prints a stored value: a store that can
print its contents is one screenshot away from not being a store.

```
af secret list
```

```
af secret list
```

### `af secret rm`

Remove a value from the store.

```
af secret rm <name>
```

```
af secret rm STRIPE_SECRET_KEY
```

### `af secret set`

Store a value, read without echo.

Reads the value from the terminal without echoing it, or from stdin when there
is no terminal.

It is never taken as an argument. An argument is in the shell history, in the
process list, and in the CI log of whatever ran it.

```
af secret set <name> [flags]
```

```
# Prompts for the value, or reads it from stdin. Never a flag.
af secret set STRIPE_SECRET_KEY
af secret set STRIPE_SECRET_KEY --stdin
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--stdin` | `false` | Read the value from stdin rather than prompting. |

### `af start`

Say where you are on the first run, and what to run next.

The first run is a sequence, and a sequence can be interrupted. This reports
each step of it as observed on this machine right now, and names the one command
that moves you forward.

It runs nothing and writes nothing. Every answer comes from the machine rather
than from a record of what this command last did, so closing the laptop,
switching branches, or tearing an environment down by hand all move the answer
with you.

A step that cannot be answered without side effects is reported as not checked,
with the reason and the command that does answer them. That is the point rather
than a gap: a step reported as fine because nothing looked at it is how a green
run over nothing happens.

A step reported as a warning is missing and does not stop the next command. The
variable naming production is the one that earns it: when a verified golden for
this project already exists, af up branches that golden, and the variable is
needed by the next refresh rather than by you now.

Exit 0 means every step is either done or simply not reached yet, which is the
normal state of a first run in progress. Exit 3 means a step is broken and the
next command cannot work until it is fixed.

```
af start
```

```
# Where you are on the first run, and the one command that moves you on.
af start
af start -o json
```

### `af status`

Show what is running for this branch.

```
af status [flags]
```

```
af status
af status -o json
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to report on, defaulting to the checked out one. |

### `af support`

Collect a redacted diagnostic bundle.

```
af support
```

```
af support bundle
```

Subcommands:

- [`af support bundle`](#af-support-bundle) Write logs, decisions, the manifest, and doctor output, redacted.

### `af support bundle`

Write logs, decisions, the manifest, and doctor output, redacted.

Everything in the bundle goes through the redactor on the way in, and the
bundle lists exactly what it included so you can see what you are about to
send before you send it.

A bundle you have to trust is a bundle nobody sends, and a report nobody sends
is a bug nobody fixes.

```
af support bundle [flags]
```

```
# Redacted on the way in, with a list of what it included.
af support bundle
af support bundle --archive af-support.zip
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--archive` | - | Where to write the bundle. |
| `--branch` | - | Branch to collect, defaulting to the checked out one. |

### `af test`

Run the manifest's workflows against the environment.

Agents drive the application the way a person does, through the accessibility
tree, and return a verdict with a video, a trace, and steps to reproduce it.

The manifest's terminal workflows run in the same pass and are counted in the
same verdict. A terminal's rendered cells are its accessibility tree, so a
program that draws a full screen is driven on a real pseudo terminal and judged
on what it drew rather than on the bytes it wrote.

Five verdicts, not two. The one that matters is blocked: a browser that
crashed, a page that never loaded, or a persona with no password is not
evidence about the application, and charging it to the application is how
people learn to ignore the results. Only a real failure exits non zero.

```
af test [flags]
```

```
af test
af test --only checkout --headed
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--attempts` | `2` | How many times to try a workflow before deciding. |
| `--branch` | - | Branch to run against, defaulting to the checked out one. |
| `--headed` | `false` | Show the browser rather than running it hidden. |
| `--only` | - | Run just these workflows, by name, from either list. |
| `--runner` | - | Path to the runner's entry point. |

### `af token`

Engine tokens, which is what CI and a self-hosted engine present.

An engine token is what goes in AF_CONTROL_PLANE_TOKEN. It belongs to the
organization rather than to you, so it keeps working after you leave, and it
carries no identity: it can send events and read an environment back, and it
cannot reach a key, a member, or another token.

These commands need a token that asked for the capability:

  af login --scope tokens.manage

A token from a plain af login cannot mint one, which is deliberate. A credential
that can make more credentials is a credential worth stealing twice.

```
af token
```

```
af token list
```

Subcommands:

- [`af token create`](#af-token-create) Mint an engine token and show it once.
- [`af token list`](#af-token-list) What engine tokens exist, and when each was last used.
- [`af token rm`](#af-token-rm) Revoke an engine token.

### `af token create`

Mint an engine token and show it once.

Mints a token and prints it. Only its hash is stored, so this is the one and
only time it can be read: there is no command and no screen that will show it
again. If you lose it, mint another and revoke this one.

The name is a label you will read in a list months from now, so name it after
where it is going rather than after today.

```
af token create <name> [flags]
```

```
# Shown once, at creation. There is no command that prints it again.
af token create ci
af token create ci --control-plane https://app.antifailure.dev
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--control-plane` | - | The control plane to use (default: AF_CONTROL_PLANE_URL, or the hosted instance). |

### `af token list`

What engine tokens exist, and when each was last used.

Shows every engine token, revoked ones included. A revoked one is shown rather
than hidden, because the question this is usually asked is whether the token
that stopped working is the one you revoked.

It does not show a token and there is no flag that would. The prefix is what
tells two of them apart, and it is what af token rm accepts.

```
af token list [flags]
```

```
af token list
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--control-plane` | - | The control plane to use (default: AF_CONTROL_PLANE_URL, or the hosted instance). |

### `af token rm`

Revoke an engine token.

Revokes a token immediately. Anything presenting it stops being accepted on the
next request rather than at the end of a cache window.

Takes the prefix af token list shows, or the full id. Running it twice is not an
error: the second run says it was already revoked, because during an incident
the same command gets run twice and the second must not read as a new problem.

```
af token rm <id or prefix> [flags]
```

```
af token rm afe_1a2b3c4d
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--control-plane` | - | The control plane to use (default: AF_CONTROL_PLANE_URL, or the hosted instance). |

### `af traffic`

What production serves, and how much of it a load run actually sends.

A load run sends the routes safe_routes names. Without a traffic profile
nothing says how much of production that is, so four routes written by hand
report in the same words and with the same verdict as a mix read from a week of
production telemetry.

That is not a cosmetic gap. Measured on this repository on 2026-09-06: a
migration held an exclusive lock on nine relations for thirty seconds and the
load run over four hand written routes reported 0.0 percent failed, because
none of the four reads the locked table. A hand written route list cannot know
which routes touch which tables.

A profile is the endpoint mix, the arrival rate, the peak concurrency and the
per route p95, counted from telemetry a team already has. It carries no request
body, no header, no query string and no identifier. It is a count per route,
which is what makes it safe to commit beside the manifest, and committing it is
the point: the check running on a pull request cannot reach production.

Declare where it lives under load.traffic.profile, and how old it may be under
load.traffic.max_age. A profile past that age is refused rather than quoted.

```
af traffic
```

```
af traffic show
```

Subcommands:

- [`af traffic record`](#af-traffic-record) Count what production served from a trace export or an access log.
- [`af traffic show`](#af-traffic-show) Print what production serves and which of it this run sends.

### `af traffic record`

Count what production served from a trace export or an access log.

Reads the file load.source_config.path names, which --from overrides, and
writes the profile to the path load.traffic.profile names, which --out
overrides.

Two sources, both of them a file. An OpenTelemetry trace export in OTLP/JSON
answers every question the profile asks, because a span carries a start and an
end: the mix, the rate, the per route p95 a threshold compares against, and the
peak concurrency. A combined format access log answers the mix and the rate,
and says in the profile that it could answer neither of the others.

Nothing here opens a socket, and there is no agent to install. The file is one
a collector or a reverse proxy already wrote.

```
af traffic record [flags]
```

```
# Counts an OpenTelemetry export or an access log a collector already
# wrote. Nothing here opens a socket and there is no agent to install.
af traffic record
af traffic record --from telemetry/traces.json --out .antifailure/traffic.json
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch context to use, defaulting to the checked out one. |
| `--from` | - | Read this file instead of the one load.source_config.path names. |
| `--out` | - | Write the profile here instead of where the manifest says. |

### `af traffic show`

Print what production serves and which of it this run sends.

Reads the profile the manifest names and prints it, busiest route first, with a
mark against every route a load run would actually send.

The routes with no mark are the finding. They are what production serves and
this run never touches, so they are what a green run says nothing about, and
the safe_routes lines that would cover them are printed at the end for somebody
to read and paste. Nothing is written for you: this measures and states, and
the manifest confirms it.

A profile older than load.traffic.max_age is REFUSED rather than printed with a
warning beside it. A stale denominator is not a smaller number, it is an
unknown one.

```
af traffic show [flags]
```

```
af traffic show
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch context to use, defaulting to the checked out one. |

### `af up`

Create an environment for the current branch.

Build every service, branch the database from its masked golden, seal the
network, and bring the environment up.

The environment is created under a lock for this branch, so two invocations
cannot fight over it, and every resource is journaled before it is made, so an
interrupt at any point leaves something af down can clean up.

```
af up [flags]
```

```
af up
af up --rebuild --hud
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to create the environment for, defaulting to the checked out one. |
| `--hud` | `false` | Watch the run on a live dashboard, or a line per event where there is no terminal. |
| `--rebuild` | `false` | Build every image again, even when an identical one exists. |

### `af update`

Install the latest verified CLI release in place.

Downloads the latest stable community release for this platform, verifies the published SHA256 checksum, and replaces this binary and its bundled runner source. It leaves shell profiles and project files alone. Package-managed installations must be upgraded through their package manager; enterprise binaries must use their enterprise distribution. The check option reads the latest release without changing files. For a legacy installer with a separate binary directory, the prefix option names its original installation prefix.

```
af update [flags]
```

```
af update
# Check the latest release without replacing any file.
af update --check -o json
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--check` | `false` | Show the latest release without changing files. |
| `--prefix` | - | Installer prefix for a legacy custom binary directory. |

### `af version`

Print the version, commit, and edition.

```
af version [flags]
```

```
af version
af version --short
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--short` | `false` | Print only the version number. |

### `af volume`

What production holds, and what fraction of it this twin has.

A fidelity report can say a branch holds twelve tables over a hundred thousand
rows. Without a volume profile it has nothing to compare that against, so a
golden built from a staging database with two hundred rows in it reports as
reproducing a production holding four billion, in the same words and with the
same verdict as a full copy.

A profile is row counts, table and index sizes, partition counts and skew, and
the cardinality of every column anything joins on. It carries no data: every
figure comes from a catalog the planner already maintains, and no row is read.
That is what makes it safe to run against production itself and safe to commit
beside the manifest, which is where the check running on a pull request has to
read it from.

Declare where it lives under database.volume.profile, and how old it may be
under database.volume.max_age. A profile past that age is refused rather than
quoted, the same way a stale golden is refused rather than branched.

```
af volume
```

```
af volume show
```

Subcommands:

- [`af volume record`](#af-volume-record) Read production's shape over a read only connection and write the profile.
- [`af volume show`](#af-volume-show) Print the committed profile, or say why there is none to print.

### `af volume record`

Read production's shape over a read only connection and write the profile.

Reads the database named by database.source_url_env and writes the profile to
the path database.volume.profile names, which --out overrides.

Nothing here reads a row. It is pg_class, pg_stats and the partition catalogs,
which is why a read only role on a replica is enough and why the result is a
file somebody can read before committing it.

```
af volume record [flags]
```

```
# Reads pg_class and pg_stats over the connection database.source_url_env
# names. No row is read, so a read only role on a replica is enough.
af volume record
af volume record --out .antifailure/volume.json
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch context to use, defaulting to the checked out one. |
| `--out` | - | Write the profile here instead of where the manifest says. |

### `af volume show`

Print the committed profile, or say why there is none to print.

Reads the profile the manifest names and prints it, largest table first.

A profile older than database.volume.max_age is REFUSED rather than printed
with a warning beside it. A stale denominator is not a smaller number, it is an
unknown one, and the one thing a number in a report must never be is a figure
somebody quotes without knowing how old it is.

```
af volume show [flags]
```

```
af volume show
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch context to use, defaulting to the checked out one. |

### `af watch`

Watch the manifest's workflows run live in the terminal.

Runs the workflows and streams them as they happen, every agent on screen at
once, so you can see the swarm rather than read what it did afterwards.

Each pane names the personality driving that agent, the workflow it is running,
the account it signed in as, its state and its current step, and shows the
agent's most recent frame as a real picture in the terminal, about once a
second. Focus a pane with the number keys, the arrows or tab, press f to give
one agent the whole screen, and quit with q.

The picture needs a terminal that draws inline images, and the terminal is asked
rather than guessed at: iTerm2, kitty and anything that reports sixel graphics
all draw. A terminal that draws none of them gets the same panes with the
frame's own detail in place of the picture, and the footer says which terminal
you have. Set AF_IMAGES to iterm2, kitty, sixel or off when the question cannot
reach your terminal, which is what a multiplexer or a forwarded connection can
do to it. The frames never leave this machine for the control plane.

The verdict is the same one a plain run produces, printed when it finishes.

```
af watch [flags]
```

```
af watch
af watch --only checkout
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--attempts` | `0` | how many times to try a workflow. |
| `--branch` | - | the branch to watch, defaulting to the checkout's. |
| `--headed` | `false` | show the browser window as well. |
| `--no-images` | `false` | draw no pictures even on a terminal that would show them. |
| `--only` | - | watch just these workflows. |
| `--runner` | - | override where the runner lives. |

### `af webhook`

Send the inbound events a flow is waiting on.

Sends a provider's callback into the environment, signed the way that provider
signs it.

The signature is the point. An application that verifies signatures, which is
every application that should, will reject an unsigned event, and a simulator
that cannot get past the application's own verification simulates nothing.

```
af webhook
```

```
af webhook list
```

Subcommands:

- [`af webhook list`](#af-webhook-list) List the providers and events that can be sent.
- [`af webhook trigger`](#af-webhook-trigger) Send one signed event into the environment.

### `af webhook list`

List the providers and events that can be sent.

```
af webhook list
```

```
af webhook list
af webhook list stripe
```

### `af webhook trigger`

Send one signed event into the environment.

The path is taken from the manifest's webhook_path for that provider unless
--path says otherwise, and the signing secret from the same variable the
application reads, so both sides agree without anybody configuring twice.

```
af webhook trigger <provider> <event> [flags]
```

```
af webhook trigger stripe checkout.session.completed
af webhook trigger stripe invoice.paid --set id=in_123 --set amount_paid=4900
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--branch` | - | Branch to deliver to, defaulting to the checked out one. |
| `--path` | - | Path to deliver to, defaulting to the manifest's webhook_path. |
| `--secret` | - | Signing secret, defaulting to the provider's variable in this shell. |
| `--service` | - | Service to deliver to, defaulting to the first reachable one. |
| `--set` | - | Set a field on the event payload, as key=value. |

### `af whoami`

Who this machine is signed in as.

Asks the control plane who the stored token belongs to.

It asks rather than reading the stored copy, because the stored copy is what
this machine believed at login time and the control plane is what is true now.
A token whose membership has been removed still looks perfectly good on disk,
and reporting it would tell somebody they have access they do not have.

--offline reports the stored copy without a network call, and says so.

```
af whoami [flags]
```

```
af whoami
af whoami --offline
```

| Flag | Default | What it does |
| --- | --- | --- |
| `--control-plane` | - | The control plane to ask (default: AF_CONTROL_PLANE_URL, or the hosted instance). |
| `--offline` | `false` | Report the stored credential without asking the control plane. |
