Audit stream
Requires an enterprise license with the audit_stream feature.
Two streams, from two places, and they are configured separately because they run on different machines. The engine forwards the privileged things it does from wherever you run it. The control plane forwards its own audit log, the one with the hash chain in it, from wherever you run that. Neither replaces the other and neither replaces the log itself, which is written regardless: a sink that is unreachable loses forwarding and never loses the entry.
The control plane’s stream has two ways to choose a destination, and which one applies to you depends on who runs the control plane. An operator names one destination for the whole installation in the environment, which is the section below. An organization on a hosted control plane names its own, through an API, which is the section after it. An organization that has named one is delivered there and nowhere else, and the installation destination covers every organization that has not.
What the engine forwards
Section titled “What the engine forwards”Five actions:
| Action | When |
|---|---|
environment.refused |
organisation policy refused an environment, before anything was created |
environment.created |
an environment was brought up, with the outcome when it failed |
environment.torn_down |
an environment was removed, with what was left behind |
golden.published |
a masked copy of production was written to a shared store |
golden.pulled |
a published golden was restored onto this machine |
Egress decisions and build steps are not forwarded. They are high volume and are already reported through the event bus.
What one entry looks like
Section titled “What one entry looks like”One line of JSON, the same bytes at every destination, so a query written against your SIEM works against your archive:
{"occurred_at":"2026-09-07T11:22:33.456789Z","forwarded_at":"2026-09-07T11:22:33.481204Z","org":"acme","actor":"dana@acme.example","action":"golden.published","target_type":"golden","target_id":"gv_9f2c","origin":"engine","detail":{"repository":"acme/shop","store":"the bucket s3://acme-goldens/audit"}}occurred_at is when the action happened and forwarded_at is when a sink
succeeded in sending it, which a retry can put minutes later. An entry whose
producer did not say when it happened carries no occurred_at at all rather than
borrowing the sink’s clock.
org and actor come from AF_ORG and from AF_ACTOR, falling back to
GITHUB_ACTOR on a GitHub Actions runner. Neither is invented when it is
absent. The operating system user is never consulted: on a CI runner it is
runner for everybody, which reads as an attribution and is not one.
Turning the engine’s stream on
Section titled “Turning the engine’s stream on”AF_AUDIT_SINKS lists the destinations, in the order they are written:
export AF_AUDIT_SINKS=syslog,webhook,object_storeA sink named here that cannot be built stops the engine at startup with the reason. With the variable unset nothing is registered and nothing is printed. Nothing is ever detected automatically.
syslog over TLS
Section titled “syslog over TLS”export AF_AUDIT_SYSLOG_ADDRESS=collector.example.com:6514export AF_AUDIT_SYSLOG_CA_FILE=/etc/ssl/collector-ca.pem# Optional, for a collector that authenticates its senders:export AF_AUDIT_SYSLOG_CERT_FILE=/etc/ssl/engine.pemexport AF_AUDIT_SYSLOG_KEY_FILE=/etc/ssl/engine-key.pem# Optional, what the messages claim to come from. Defaults to the hostname.export AF_AUDIT_SYSLOG_HOSTNAME=runner-7RFC 5424 messages with RFC 5425 octet counted framing, at facility 13, “log audit”, so a receiver routing on facility files them as what they are. The action is the message id, which is what a receiver filters on. Port 6514 is assumed when the address carries none.
There is no plaintext option. An address written as syslog:// or tcp:// is
refused rather than downgraded.
HTTPS webhook
Section titled “HTTPS webhook”export AF_AUDIT_WEBHOOK_URL=https://siem.example/ingestexport AF_AUDIT_WEBHOOK_DEAD_LETTER_FILE=/var/lib/antifailure/audit-dead-letter.jsonl# Optional. Keys an HMAC-SHA256 over the exact bytes posted.export AF_AUDIT_WEBHOOK_SECRET=...# Optional, for a receiver that takes a bearer token.export AF_AUDIT_WEBHOOK_HEADER="Authorization: Bearer ..."With a secret set, every request carries Af-Audit-Signature: sha256=<hex> over
the body, in the same shape GitHub and Stripe use.
The dead letter file is required, and it is the reason the retry is allowed to
be short. Three attempts, pausing 200 ms and then 600 ms between them, and the
entry is appended to that file and flushed before the call returns, in the same
JSON the receiver would have been given. The measured total, round trips
included, is in the report just benchmark writes.
Object store
Section titled “Object store”export AF_AUDIT_OBJECT_STORE_URL=s3://acme-audit/antifailureOr a server that speaks the same API, as https://minio.example.com/bucket/prefix,
or an Azure Blob container URL carrying a shared access signature. The S3 form
signs its requests with AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY read
from the environment, by the names the AWS tools already use, so a machine set
up for the AWS CLI needs nothing else.
One object per entry, keyed by date:
antifailure/2026/09/07/112233.456789000-golden.published-9f2ca10b.jsonNot a batch and not an append: an object written once can be locked, and an object is never replaced. The date is a path so a lifecycle rule and a partitioned query both work without parsing a filename. Two entries in the same nanosecond are two objects, because the key carries eight random characters as well as the time.
What a sink cannot do
Section titled “What a sink cannot do”A sink observes. It cannot refuse an environment, cannot change an entry, and cannot see what another sink received. An error from one is recorded and the lifecycle continues.
A SIEM you cannot reach costs one progress line, carrying the sink’s own words:
audit sink: forwarding to syslog over TLS at collector.example.com:6514: dial tcp10.0.0.9:6514: i/o timeoutThe environment still comes up, and the teardown still finishes. What is lost is
the forwarding, and for the webhook not even that: an entry no receiver would
take is in the dead letter file before Write returns.
The control plane’s own audit log
Section titled “The control plane’s own audit log”The engine forwards five actions from a machine with no database.
The control plane forwards audit_entries, the organization log covering
actions including sign on, directory provisioning and administration. The
separate global operator log, admin_audit_entries, is forwarded only where its
writer also produces an organization entry. Each organization has its own hash
chain: each entry holds the hash
of the one before it, so altering an old entry breaks every entry after it.
What one batch looks like
Section titled “What one batch looks like”Batched rather than one entry per request, because the batch carries the proof.
The webhook posts this JSON field schema directly. Splunk and Event Hubs wrap
entries in their collector formats, described below. Organization identifiers
are UUID strings and occurredAt is an ISO timestamp:
interface AuditBatch { entries: Array<{ seq: number orgId: string actor: string action: string targetType: string targetId: string | null origin: string detail: Record<string, unknown> occurredAt: string entryHash: string }> manifest: { org: string count: number firstSeq: number lastSeq: number headHash: string digest: string signature: string }}headHash is the chain hash of the last entry in the batch, and digest is a
sha256 over the canonical batch body with signature an HMAC of that digest
under AF_AUDIT_STREAM_KEY. That is what lets a batch sitting in an archive be
checked without reaching back to the control plane that wrote it, which is the
situation an auditor is usually in.
For webhooks, x-antifailure-timestamp is the first entry’s event time, not
the delivery time. Catching up after an outage can deliver old events. Verify
the signature and deduplicate by organization and sequence; signature
verification alone does not reject replay.
One batch holds one organization.
Turning the control plane’s stream on
Section titled “Turning the control plane’s stream on”export AF_AUDIT_STREAM_SINK=webhookexport AF_AUDIT_STREAM_KEY="$(openssl rand -base64 32)"export AF_AUDIT_STREAM_WEBHOOK_URL=https://siem.example/ingestexport AF_AUDIT_STREAM_WEBHOOK_SECRET=...AF_AUDIT_STREAM_SINK takes splunk, event_hubs or webhook. Splunk reads
AF_AUDIT_STREAM_SPLUNK_URL and AF_AUDIT_STREAM_SPLUNK_TOKEN, with
AF_AUDIT_STREAM_SPLUNK_INDEX and AF_AUDIT_STREAM_SPLUNK_SOURCETYPE optional.
Event Hubs reads AF_AUDIT_STREAM_EVENT_HUBS_URL and
AF_AUDIT_STREAM_EVENT_HUBS_AUTHORIZATION, the second being a shared access
signature you generate, so no key reaches this process and managed identity
stays possible.
Event Hubs receives each entry as a JSON string in the event body and retains
the signed batch manifest in the antifailure_manifest application property.
Its batch API ignores properties supplied only through HTTP headers.
Splunk stores the same manifest in the indexed antifailure_manifest field,
alongside the audit entry’s event data.
AF_AUDIT_STREAM_KEY is required whenever a sink is named.
Remote collector URLs require HTTPS and cannot contain user information. Loopback HTTP is permitted for a local collector. Redirects are refused, each request carries a thirty second deadline, and response bodies are discarded without being buffered or included in error logs.
AF_AUDIT_STREAM_INTERVAL_MS is how often a pass runs, ten seconds by default.
AF_AUDIT_STREAM_BATCH is how many entries one pass reads, 500 by default, and
AF_AUDIT_STREAM_DELIVERY_BATCH is how many one request carries, defaulting to
the pass size. They are two numbers rather than one because how fast the
forwarder catches up and what your collector accepts in one request are
different questions.
A sink named with its variables missing stops the control plane at startup with the reason, for the same reason the engine’s does.
An object store sink exists in the code and cannot be turned on from the environment, because it needs a request signer this half of the product does not carry. Naming one is refused rather than accepted and then silently writing nowhere.
Choosing your own destination, per organization
Section titled “Choosing your own destination, per organization”On a hosted control plane the destination is an API instead. One destination per organization; to reach two collectors, use one collector that fans out after receiving.
curl -X PUT https://<your-control-plane>/enterprise/audit-stream \ -H "x-antifailure-csrf: $CSRF" -H 'content-type: application/json' \ --cookie "af_session=$SESSION" \ -d '{"kind":"webhook","url":"https://siem.example/ingest","credential":"..."}'GET returns the destination and what the stream has done for you. PUT
stores or replaces it. PATCH with {"enabled": false} stops delivery without
discarding the endpoint and the credential. DELETE removes it. kind takes
splunk, event_hubs or webhook, and Splunk additionally accepts
indexName and sourcetype so entries land where your existing searches
already look.
Only an owner or an admin may change it. Any member may read it, because the answer carries the endpoint, the last four characters of the credential and a fingerprint of it, and never the credential itself.
The credential is required on every save, including a change of endpoint. Changing where a credential is sent requires having it.
Your credential is stored sealed. It is encrypted with AES-256-GCM under a key held in the deployment’s key vault and never in the database, bound to your organization and to the kind of destination it was sealed for, so a copy of the row is useless anywhere else. A database backup on its own decrypts nothing. The only thing that ever holds the plaintext is the code putting it in a request header to your collector.
Delivery starts when you save, not at the beginning of your history. The organization’s current audit sequence is recorded with the destination, and entries above it are what get delivered, so configuring a collector does not replay months of entries into it as a surprise. The configuration change is itself an audit entry, written after that sequence is read, so the first thing your collector receives is the record of its own creation.
Switching a destination off stops delivery on the next pass and does not fall back to the installation destination: an organization that turned its stream off did not ask for its entries to go somewhere else instead. Switching it on starts from the moment of the switch, for the same reason a first save does.
Batch manifests are signed under a key derived from your own credential,
rather than under the operator’s AF_AUDIT_STREAM_KEY, which you do not hold. A
signature its reader cannot check is decoration. The key is the HMAC-SHA256 of
the label antifailure audit manifest v1 keyed by the credential you gave,
rendered as lowercase hexadecimal, so a receiver can derive it and verify every
batch without asking this control plane anything.
GET also reports what the stream has actually done: the sequence delivered so
far, when it last tried, when it last succeeded, how many passes have failed in
a row, and the collector’s own words about the last failure. A credential your
security team rotates or revokes shows up there as the status your collector
answered with.
What a destination may be, and what no URL check can see
Section titled “What a destination may be, and what no URL check can see”A destination you supply is an untrusted address from the control plane’s point of view, so it is held to a stricter rule than the installation destination in the section above.
- HTTPS always. There is no loopback exception, unlike the operator’s destination, which removes every plaintext service inside the deployment including the control plane’s own port.
- No credentials in the URL, because every proxy log on the way keeps them.
- No literal address that is not a public one: loopback, the private ranges, link local including the address cloud metadata services answer on, unique local, multicast, the unspecified address, and the IPv4 addresses that arrive wearing an IPv6 coat.
- No single label hostname and nothing under
.local, because those resolve inside a container network and nowhere else.
The rule is applied when you save and again when a batch is delivered, so a row written by any other path is refused too.
What it cannot see, stated rather than implied: a public hostname whose DNS resolves into a private network. No check on a URL can, and neither can a check made when the row is saved, because resolution can change between the save and the delivery. The control that would close it is egress policy on the control plane’s own network, which this deployment does not have today.
The deployment’s sealing secret can be rotated without your involvement. The control plane holds a set of sealing keys, each stored credential records which one sealed it, and the operator’s re-sealing run moves collector credentials and provider keys together, so a rotation done by the rotating secrets procedure changes nothing you can see. If your credential names a key the control plane has stopped holding, the stream holds its entries rather than dropping them, and the status names the missing key version instead of calling the credential altered. The operator fixes that by restoring the key. Saving the credential again also repairs it, because a fresh save is sealed under a key the control plane holds.
Delivery, and what happens when your collector is down
Section titled “Delivery, and what happens when your collector is down”Transient failures are retried with at least once delivery. Each organization’s
position advances after delivery, so a collector outage causes forwarding lag.
A crash after acceptance and before saving the position can redeliver a batch; use
orgId and seq to deduplicate. Positions are separate because transactions
from different organizations can commit in a different order from their
sequence numbers. The installation cursor is only an operational summary.
A batch your endpoint will never accept, meaning it answers 400, 401, 403, 404 or 413, is given up on rather than retried forever, because one batch nobody will ever take must not stop every entry behind it. The rest of the stream continues.
What is not forwarded, and it is stated rather than implied
Section titled “What is not forwarded, and it is stated rather than implied”An organization that is not entitled to audit_stream is skipped and the stream
moves on past it. It is not held for an entitlement that might arrive later, and
that is the same behaviour the engine has. Its delivery position advances so
these deliberately declined entries are not reconsidered on every pass.
The licence is asked per action, not at startup
Section titled “The licence is asked per action, not at startup”The engine checks audit_stream on every entry. The control plane checks its
process licence status and organization entitlement on each pass, so expiry or
an entitlement withdrawal takes effect on the next pass without a restart.
Organization entitlement grants also take effect on the next pass. Replacing
the control plane’s AF_LICENSE_KEY requires restarting the process, because
the key is parsed at startup. A configured sink on an installation without the
feature accepts every entry and writes none, so the engine says so once at
startup:
af: audit sink: configured, and audit_stream is not licensed on thisinstallation, so nothing is forwardedMeasuring the delay yourself
Section titled “Measuring the delay yourself”just benchmark writes a dated report of how long an action takes to reach each
destination, and how long an undeliverable entry takes to become durable on disk
while a receiver is down. With nothing configured it measures loopback. Point it
at your own collector and the number becomes the whole path:
AF_AUDIT_BENCHMARK_SYSLOG_ADDRESS=collector.example.com:6514 \ AF_AUDIT_BENCHMARK_SYSLOG_CA_FILE=/etc/ssl/collector-ca.pem \ AF_AUDIT_BENCHMARK_WEBHOOK_URL=https://siem.example/ingest \ just benchmarkRelated: licensing, policy, compliance.