Skip to content

Type to search pages.

View .md

Audit stream

Requires an enterprise license with the audit_stream feature.

Two streams, from two places, and they are configured separately because they run on different machines. The engine forwards the privileged things it does from wherever you run it. The control plane forwards its own audit log, the one with the hash chain in it, from wherever you run that. Neither replaces the other and neither replaces the log itself, which is written regardless: a sink that is unreachable loses forwarding and never loses the entry.

The control plane’s stream has two ways to choose a destination, and which one applies to you depends on who runs the control plane. An operator names one destination for the whole installation in the environment, which is the section below. An organization on a hosted control plane names its own, through an API, which is the section after it. An organization that has named one is delivered there and nowhere else, and the installation destination covers every organization that has not.

Five actions:

Action When
environment.refused organisation policy refused an environment, before anything was created
environment.created an environment was brought up, with the outcome when it failed
environment.torn_down an environment was removed, with what was left behind
golden.published a masked copy of production was written to a shared store
golden.pulled a published golden was restored onto this machine

Egress decisions and build steps are not forwarded. They are high volume and are already reported through the event bus.

One line of JSON, the same bytes at every destination, so a query written against your SIEM works against your archive:

{"occurred_at":"2026-09-07T11:22:33.456789Z","forwarded_at":"2026-09-07T11:22:33.481204Z","org":"acme","actor":"dana@acme.example","action":"golden.published","target_type":"golden","target_id":"gv_9f2c","origin":"engine","detail":{"repository":"acme/shop","store":"the bucket s3://acme-goldens/audit"}}

occurred_at is when the action happened and forwarded_at is when a sink succeeded in sending it, which a retry can put minutes later. An entry whose producer did not say when it happened carries no occurred_at at all rather than borrowing the sink’s clock.

org and actor come from AF_ORG and from AF_ACTOR, falling back to GITHUB_ACTOR on a GitHub Actions runner. Neither is invented when it is absent. The operating system user is never consulted: on a CI runner it is runner for everybody, which reads as an attribution and is not one.

AF_AUDIT_SINKS lists the destinations, in the order they are written:

Terminal window
export AF_AUDIT_SINKS=syslog,webhook,object_store

A sink named here that cannot be built stops the engine at startup with the reason. With the variable unset nothing is registered and nothing is printed. Nothing is ever detected automatically.

Terminal window
export AF_AUDIT_SYSLOG_ADDRESS=collector.example.com:6514
export AF_AUDIT_SYSLOG_CA_FILE=/etc/ssl/collector-ca.pem
# Optional, for a collector that authenticates its senders:
export AF_AUDIT_SYSLOG_CERT_FILE=/etc/ssl/engine.pem
export AF_AUDIT_SYSLOG_KEY_FILE=/etc/ssl/engine-key.pem
# Optional, what the messages claim to come from. Defaults to the hostname.
export AF_AUDIT_SYSLOG_HOSTNAME=runner-7

RFC 5424 messages with RFC 5425 octet counted framing, at facility 13, “log audit”, so a receiver routing on facility files them as what they are. The action is the message id, which is what a receiver filters on. Port 6514 is assumed when the address carries none.

There is no plaintext option. An address written as syslog:// or tcp:// is refused rather than downgraded.

Terminal window
export AF_AUDIT_WEBHOOK_URL=https://siem.example/ingest
export AF_AUDIT_WEBHOOK_DEAD_LETTER_FILE=/var/lib/antifailure/audit-dead-letter.jsonl
# Optional. Keys an HMAC-SHA256 over the exact bytes posted.
export AF_AUDIT_WEBHOOK_SECRET=...
# Optional, for a receiver that takes a bearer token.
export AF_AUDIT_WEBHOOK_HEADER="Authorization: Bearer ..."

With a secret set, every request carries Af-Audit-Signature: sha256=<hex> over the body, in the same shape GitHub and Stripe use.

The dead letter file is required, and it is the reason the retry is allowed to be short. Three attempts, pausing 200 ms and then 600 ms between them, and the entry is appended to that file and flushed before the call returns, in the same JSON the receiver would have been given. The measured total, round trips included, is in the report just benchmark writes.

Terminal window
export AF_AUDIT_OBJECT_STORE_URL=s3://acme-audit/antifailure

Or a server that speaks the same API, as https://minio.example.com/bucket/prefix, or an Azure Blob container URL carrying a shared access signature. The S3 form signs its requests with AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY read from the environment, by the names the AWS tools already use, so a machine set up for the AWS CLI needs nothing else.

One object per entry, keyed by date:

antifailure/2026/09/07/112233.456789000-golden.published-9f2ca10b.json

Not a batch and not an append: an object written once can be locked, and an object is never replaced. The date is a path so a lifecycle rule and a partitioned query both work without parsing a filename. Two entries in the same nanosecond are two objects, because the key carries eight random characters as well as the time.

A sink observes. It cannot refuse an environment, cannot change an entry, and cannot see what another sink received. An error from one is recorded and the lifecycle continues.

A SIEM you cannot reach costs one progress line, carrying the sink’s own words:

audit sink: forwarding to syslog over TLS at collector.example.com:6514: dial tcp
10.0.0.9:6514: i/o timeout

The environment still comes up, and the teardown still finishes. What is lost is the forwarding, and for the webhook not even that: an entry no receiver would take is in the dead letter file before Write returns.

The engine forwards five actions from a machine with no database. The control plane forwards audit_entries, the organization log covering actions including sign on, directory provisioning and administration. The separate global operator log, admin_audit_entries, is forwarded only where its writer also produces an organization entry. Each organization has its own hash chain: each entry holds the hash of the one before it, so altering an old entry breaks every entry after it.

Batched rather than one entry per request, because the batch carries the proof. The webhook posts this JSON field schema directly. Splunk and Event Hubs wrap entries in their collector formats, described below. Organization identifiers are UUID strings and occurredAt is an ISO timestamp:

interface AuditBatch {
entries: Array<{
seq: number
orgId: string
actor: string
action: string
targetType: string
targetId: string | null
origin: string
detail: Record<string, unknown>
occurredAt: string
entryHash: string
}>
manifest: {
org: string
count: number
firstSeq: number
lastSeq: number
headHash: string
digest: string
signature: string
}
}

headHash is the chain hash of the last entry in the batch, and digest is a sha256 over the canonical batch body with signature an HMAC of that digest under AF_AUDIT_STREAM_KEY. That is what lets a batch sitting in an archive be checked without reaching back to the control plane that wrote it, which is the situation an auditor is usually in.

For webhooks, x-antifailure-timestamp is the first entry’s event time, not the delivery time. Catching up after an outage can deliver old events. Verify the signature and deduplicate by organization and sequence; signature verification alone does not reject replay.

One batch holds one organization.

Terminal window
export AF_AUDIT_STREAM_SINK=webhook
export AF_AUDIT_STREAM_KEY="$(openssl rand -base64 32)"
export AF_AUDIT_STREAM_WEBHOOK_URL=https://siem.example/ingest
export AF_AUDIT_STREAM_WEBHOOK_SECRET=...

AF_AUDIT_STREAM_SINK takes splunk, event_hubs or webhook. Splunk reads AF_AUDIT_STREAM_SPLUNK_URL and AF_AUDIT_STREAM_SPLUNK_TOKEN, with AF_AUDIT_STREAM_SPLUNK_INDEX and AF_AUDIT_STREAM_SPLUNK_SOURCETYPE optional. Event Hubs reads AF_AUDIT_STREAM_EVENT_HUBS_URL and AF_AUDIT_STREAM_EVENT_HUBS_AUTHORIZATION, the second being a shared access signature you generate, so no key reaches this process and managed identity stays possible.

Event Hubs receives each entry as a JSON string in the event body and retains the signed batch manifest in the antifailure_manifest application property. Its batch API ignores properties supplied only through HTTP headers. Splunk stores the same manifest in the indexed antifailure_manifest field, alongside the audit entry’s event data.

AF_AUDIT_STREAM_KEY is required whenever a sink is named.

Remote collector URLs require HTTPS and cannot contain user information. Loopback HTTP is permitted for a local collector. Redirects are refused, each request carries a thirty second deadline, and response bodies are discarded without being buffered or included in error logs.

AF_AUDIT_STREAM_INTERVAL_MS is how often a pass runs, ten seconds by default. AF_AUDIT_STREAM_BATCH is how many entries one pass reads, 500 by default, and AF_AUDIT_STREAM_DELIVERY_BATCH is how many one request carries, defaulting to the pass size. They are two numbers rather than one because how fast the forwarder catches up and what your collector accepts in one request are different questions.

A sink named with its variables missing stops the control plane at startup with the reason, for the same reason the engine’s does.

An object store sink exists in the code and cannot be turned on from the environment, because it needs a request signer this half of the product does not carry. Naming one is refused rather than accepted and then silently writing nowhere.

Choosing your own destination, per organization

Section titled “Choosing your own destination, per organization”

On a hosted control plane the destination is an API instead. One destination per organization; to reach two collectors, use one collector that fans out after receiving.

Terminal window
curl -X PUT https://<your-control-plane>/enterprise/audit-stream \
-H "x-antifailure-csrf: $CSRF" -H 'content-type: application/json' \
--cookie "af_session=$SESSION" \
-d '{"kind":"webhook","url":"https://siem.example/ingest","credential":"..."}'

GET returns the destination and what the stream has done for you. PUT stores or replaces it. PATCH with {"enabled": false} stops delivery without discarding the endpoint and the credential. DELETE removes it. kind takes splunk, event_hubs or webhook, and Splunk additionally accepts indexName and sourcetype so entries land where your existing searches already look.

Only an owner or an admin may change it. Any member may read it, because the answer carries the endpoint, the last four characters of the credential and a fingerprint of it, and never the credential itself.

The credential is required on every save, including a change of endpoint. Changing where a credential is sent requires having it.

Your credential is stored sealed. It is encrypted with AES-256-GCM under a key held in the deployment’s key vault and never in the database, bound to your organization and to the kind of destination it was sealed for, so a copy of the row is useless anywhere else. A database backup on its own decrypts nothing. The only thing that ever holds the plaintext is the code putting it in a request header to your collector.

Delivery starts when you save, not at the beginning of your history. The organization’s current audit sequence is recorded with the destination, and entries above it are what get delivered, so configuring a collector does not replay months of entries into it as a surprise. The configuration change is itself an audit entry, written after that sequence is read, so the first thing your collector receives is the record of its own creation.

Switching a destination off stops delivery on the next pass and does not fall back to the installation destination: an organization that turned its stream off did not ask for its entries to go somewhere else instead. Switching it on starts from the moment of the switch, for the same reason a first save does.

Batch manifests are signed under a key derived from your own credential, rather than under the operator’s AF_AUDIT_STREAM_KEY, which you do not hold. A signature its reader cannot check is decoration. The key is the HMAC-SHA256 of the label antifailure audit manifest v1 keyed by the credential you gave, rendered as lowercase hexadecimal, so a receiver can derive it and verify every batch without asking this control plane anything.

GET also reports what the stream has actually done: the sequence delivered so far, when it last tried, when it last succeeded, how many passes have failed in a row, and the collector’s own words about the last failure. A credential your security team rotates or revokes shows up there as the status your collector answered with.

What a destination may be, and what no URL check can see

Section titled “What a destination may be, and what no URL check can see”

A destination you supply is an untrusted address from the control plane’s point of view, so it is held to a stricter rule than the installation destination in the section above.

  • HTTPS always. There is no loopback exception, unlike the operator’s destination, which removes every plaintext service inside the deployment including the control plane’s own port.
  • No credentials in the URL, because every proxy log on the way keeps them.
  • No literal address that is not a public one: loopback, the private ranges, link local including the address cloud metadata services answer on, unique local, multicast, the unspecified address, and the IPv4 addresses that arrive wearing an IPv6 coat.
  • No single label hostname and nothing under .local, because those resolve inside a container network and nowhere else.

The rule is applied when you save and again when a batch is delivered, so a row written by any other path is refused too.

What it cannot see, stated rather than implied: a public hostname whose DNS resolves into a private network. No check on a URL can, and neither can a check made when the row is saved, because resolution can change between the save and the delivery. The control that would close it is egress policy on the control plane’s own network, which this deployment does not have today.

The deployment’s sealing secret can be rotated without your involvement. The control plane holds a set of sealing keys, each stored credential records which one sealed it, and the operator’s re-sealing run moves collector credentials and provider keys together, so a rotation done by the rotating secrets procedure changes nothing you can see. If your credential names a key the control plane has stopped holding, the stream holds its entries rather than dropping them, and the status names the missing key version instead of calling the credential altered. The operator fixes that by restoring the key. Saving the credential again also repairs it, because a fresh save is sealed under a key the control plane holds.

Delivery, and what happens when your collector is down

Section titled “Delivery, and what happens when your collector is down”

Transient failures are retried with at least once delivery. Each organization’s position advances after delivery, so a collector outage causes forwarding lag. A crash after acceptance and before saving the position can redeliver a batch; use orgId and seq to deduplicate. Positions are separate because transactions from different organizations can commit in a different order from their sequence numbers. The installation cursor is only an operational summary.

A batch your endpoint will never accept, meaning it answers 400, 401, 403, 404 or 413, is given up on rather than retried forever, because one batch nobody will ever take must not stop every entry behind it. The rest of the stream continues.

What is not forwarded, and it is stated rather than implied

Section titled “What is not forwarded, and it is stated rather than implied”

An organization that is not entitled to audit_stream is skipped and the stream moves on past it. It is not held for an entitlement that might arrive later, and that is the same behaviour the engine has. Its delivery position advances so these deliberately declined entries are not reconsidered on every pass.

The licence is asked per action, not at startup

Section titled “The licence is asked per action, not at startup”

The engine checks audit_stream on every entry. The control plane checks its process licence status and organization entitlement on each pass, so expiry or an entitlement withdrawal takes effect on the next pass without a restart. Organization entitlement grants also take effect on the next pass. Replacing the control plane’s AF_LICENSE_KEY requires restarting the process, because the key is parsed at startup. A configured sink on an installation without the feature accepts every entry and writes none, so the engine says so once at startup:

af: audit sink: configured, and audit_stream is not licensed on this
installation, so nothing is forwarded

just benchmark writes a dated report of how long an action takes to reach each destination, and how long an undeliverable entry takes to become durable on disk while a receiver is down. With nothing configured it measures loopback. Point it at your own collector and the number becomes the whole path:

Terminal window
AF_AUDIT_BENCHMARK_SYSLOG_ADDRESS=collector.example.com:6514 \
AF_AUDIT_BENCHMARK_SYSLOG_CA_FILE=/etc/ssl/collector-ca.pem \
AF_AUDIT_BENCHMARK_WEBHOOK_URL=https://siem.example/ingest \
just benchmark

Related: licensing, policy, compliance.