Fault injection and crash recovery
A rehearsal tells you what a change does to a system that works. The chaos block tells you what the system does when it stops working, and then it proves the answer instead of reporting that everything came back.
chaos: enabled: true faults: - name: postgres-crash kind: process_kill target: database process: "postgres: checkpointer"Run it with af chaos against a running environment, or let af ci run it at
the end of a check.
What it proves
Section titled “What it proves”Around a fault aimed at the database, concurrent writers commit into a schema the engine owns, and the fault lands while they are committing. Afterwards the run establishes four things:
- No lost durable commit. Every transaction the client was told was committed is still there.
- No phantom commit. Nothing is there that no client ever tried to write.
- The write ahead log replayed. Recovery started at the position the control file named before the crash, and reached past the last flush a writer saw.
- The relations survived. A sequential scan and an index only scan count
the same rows, and
amcheckfinds an index entry for every live heap tuple.
The sequential scan also reads every page of the writers’ table, and on a
cluster with data checksums on, a page torn by the crash fails its checksum and
stops that read. af chaos prints the result on each crash fault’s pages
line. It covers the writers’ table and no other, and it says the pages were
not checked when checksums are off, when the control file could not be read
after the fault, or when the read did not finish. The amcheck line beside it
prints what the index verifier said, or that it did not run.
The first two need something the database cannot give you, because they are claims about what the database said rather than about what it holds. The engine keeps a ledger on the client side of the wire: an identifier goes in before the statement is sent, and moves to acknowledged only when the call returns without an error. A commit that returned success and is absent afterwards is a durability failure whatever caused it.
The faults
Section titled “The faults”| Kind | What happens | Undo |
|---|---|---|
process_kill |
SIGKILL to one process inside the container, matched by a substring of its command line. The container keeps running. |
None. The recovery is the system’s own, and that is the fault. |
container_kill |
SIGKILL to the container’s main process. The container stops. |
Starts it again. |
container_stop |
SIGTERM, then SIGKILL after a grace period. |
Starts it again. |
container_pause |
Freezes every process with the cgroup freezer. Nothing is killed and no connection closes. | Thaws it. |
network_partition |
Detaches the container from the environment’s network. | Attaches it again, with the aliases it had. |
read_only_data |
Removes write permission from the data directory. | Restores the mode it recorded. |
disk_fill |
Fills the filesystem holding the data directory to a stated headroom. Needs database.data_filesystem.size_bytes, below. |
Removes the file it wrote. |
process_kill and container_kill are the two kinds that stop Postgres
uncleanly, so they are the two the recovery proof expects a replay from. The
others are useful and they are honest about what they are: a container_stop
shuts the database down cleanly and replays nothing, and a run that declared it
as a crash reports that it could not establish a recovery rather than reporting
a clean one.
What it will not touch
Section titled “What it will not touch”A fault reaches the containers this environment created and nothing else. The target resolves from the labels the runtime stamped at create time, never from a name a fault supplied, and the ownership is read again from the daemon at the instant of the act. Three refusals have no override:
- a container carrying no
dev.antifailure.managedlabel is not ours - a container belonging to a different environment
- the egress sidecar and the emulators, whatever environment they belong to
The sidecar carries the egress policy. A fault that could stop it would switch off the control that decides what the environment may reach, and a chaos feature that can disable a safety control is a way out with a feature name. An emulator stands in for a third party the environment must not reach, so stopping one does not produce an outage: it produces a request that goes looking for the real host.
disk_fill carries a fourth refusal, and a declaration that lifts it.
A container’s writable layer is the daemon’s own disk, so filling a directory
on it fills the machine and every other container running on it. The fault
reads the mount at the data directory from the daemon and refuses unless it is
a volume this environment created with a size fixed when it was created. A
mount of its own is not enough on its own: a plain named volume is its own
mount and is still a slice of the daemon’s disk, so it would pass a device
check and take the machine down having satisfied the guard. Both refusals are
reported as chaos.fault.unsafe: the claim the fault was declared to establish
was not established, and nothing else in the run was touched by it.
Giving the data directory a filesystem of its own
Section titled “Giving the data directory a filesystem of its own”database: data_filesystem: size_bytes: 536870912chaos: enabled: true faults: - name: fill-the-data-volume kind: disk_fill target: database headroom_bytes: 8388608 max_fill_bytes: 536870912With that, the branch keeps its data directory on a filesystem of the declared
size and disk_fill lands: the fill writes one file until the stated headroom
is left, Postgres meets a real No space left on device on its next extend,
and the undo removes the file and the free space comes back. Without it the
data directory is on the writable layer and the fault is refused before it acts.
The filesystem is held in memory, and that is the containment argument rather than an implementation detail. A volume on the daemon’s disk cannot be filled without taking space from every other container on the machine; one in memory has a size fixed at creation and takes nothing from anything outside the environment. Three things follow, and they are the cost of the feature:
- The whole database lives in it, so the size has to hold the data directory with room left for the fault to fill. A copy that does not fit is refused by name, with both numbers, rather than truncated.
- A size of more than half the memory the Docker daemon reports is refused. A filesystem in memory larger than the machine moves the same problem from the disk to the memory, and a daemon killed for memory takes every other environment with it.
- The data directory does not survive the Docker daemon restarting.
af upbuilds it again from the golden.
The branch pays a copy of the data directory when it comes up, where an ordinary branch pays nothing because the daemon’s storage driver copies on write. So this is the layout for rehearsing a disk that fills, and not the one to measure how a disk performs.
The environment also runs one container that holds that filesystem mounted and
does nothing else. It is not decoration: the local volume driver unmounts a
memory backed volume when the last container using it stops, so without it a
container_kill or container_stop would delete the data directory rather
than crash the database, the undo would start a container that initialised an
empty one, and the durability proof would report every acknowledged commit
lost. Faults refuse to touch it for the same reason they refuse to touch the
egress sidecar.
Nothing that changed nothing counts as survived
Section titled “Nothing that changed nothing counts as survived”A fault that was applied and had no effect is refused, not reported. The reason is the whole point of the feature: every assertion after such a fault describes a system that never broke, and a recovery check that passes on one is a check that answers the same whether or not it ran.
So a process_kill whose pattern matches nothing is refused rather than
reported as a crash the database survived. A read_only_data fault probes a
write as the directory’s owner and refuses if the write still succeeds, which
is what happens on a directory owned by root, because root ignores the mode.
A container_pause that the daemon accepts and that leaves the container
running is refused.
The same discipline runs through the findings. A run that could not establish what it set out to is reported as unverified and never as a pass:
| Finding | Meaning |
|---|---|
chaos.durability.lost_commit |
A transaction the client was told was committed is gone. |
chaos.durability.phantom_commit |
A row is present that no client wrote. |
chaos.recovery.replay_short |
Recovery stopped before the last position the client saw flushed. |
chaos.recovery.timeline_moved |
The timeline changed, and crash recovery does not change it. |
chaos.integrity.relation_damaged |
The heap and its index disagree. |
chaos.invariant.broken_by_fault |
One of this project’s own invariants held before the fault and does not hold after the recovery. |
chaos.recovery.no_crash |
The fault was declared as a crash and nothing crashed. |
chaos.recovery.no_replay |
The database came back and the log records no replay. |
chaos.integrity.checksums_off |
Data page checksums are off, so a torn page would not be seen. |
chaos.integrity.amcheck_unavailable |
The index could not be verified. |
chaos.durability.inconsistent_ledger |
The engine’s own bookkeeping does not add up. |
chaos.invariant.already_violated |
One of this project’s own invariants did not hold before the fault either, so nothing after it is attributable to the fault. |
chaos.invariant.unevaluated |
One of this project’s own invariants could not be asked on one side or the other, which a database that did not come back is the loudest case of. |
chaos.fault.refused |
A fault tried to go in and failed, so it established nothing. |
chaos.fault.unsafe |
A fault was refused before it acted, because its effect would reach past this environment. It changed nothing the other faults measured. |
chaos.fault.not_undone |
A fault went in and its undo failed, so the environment is still broken and anything measured after it is suspect. |
The first six are failures and carry policy.chaos_failure, which defaults to
fail. The last ten are the ones the run could not look at, and they carry
policy.chaos_unverified, which defaults to warn. They are two keys because
a check that found a problem and a check that could not look are different
facts, and reporting the second as the first teaches a project to ignore both.
Your own rules, asked of the recovered database
Section titled “Your own rules, asked of the recovered database”Everything the durability proof asserts is about a schema of the engine’s own,
and that is deliberate: asserting that a table your application is writing did
not change, while it is writing it, is a claim about a moving target. That
reason stops applying the moment the writers stop and the database answers a
query again, and that is exactly when the invariants your manifest declares
are the right question. The ledger proves the engine’s commits survived. Only
your invariants can say whether your data still means what you say it means.
So around a fault with the durability proof on, every invariant the manifest declares is asked twice: once before anything is broken, and once against the recovered database. Both answers are printed, because one of them cannot be read on its own.
invariant no-negative-balance: before the fault held; after the recovery held invariant orders-have-a-customer: before the fault held; after the recovery violated, 2 rowsAn invariant that was already violated before the fault is reported and is
attributed to nothing: the rule is broken and this run is not what broke it, so
chaos.invariant.already_violated is unverified and never fails the run. A gate
that stopped a merge for a rule the change did not break would teach a project
to switch the whole arm off. Only a rule that held before the fault and does
not hold after the recovery is something the run can attribute to it, and that
one is chaos.invariant.broken_by_fault, which fails.
An invariant that could not be asked, on either side, is
chaos.invariant.unevaluated. A database that did not come back is the loudest
case of it, and it is the one where reporting nothing would be worst: an
absent arm reads as an arm with nothing to report.
A manifest that declares no invariants runs none of this and nothing about
it appears in any output.
Asking for a run that loses data
Section titled “Asking for a run that loses data”crash_recovery.synchronous_commit sets what the writers ask of the database.
With it off, Postgres acknowledges a commit before the write ahead log record
has left shared memory, so a crash that discards shared memory loses commits
the client was told were durable. That is the setting’s documented behavior and
the run reports the loss:
chaos: enabled: true crash_recovery: synchronous_commit: off faults: - name: prove-the-check-can-say-no kind: process_kill target: database process: "postgres: checkpointer"Leave it out unless you mean it. A manifest that sets it to off is asking for
a run that is expected to report lost commits, which is useful exactly once:
to see the check say no before you trust it saying yes.
Reading the numbers
Section titled “Reading the numbers”The unreachable line is measured by a probe that starts with the fault and
runs beside it. Every 100 milliseconds it opens a connection and runs
SELECT 1, and an attempt that gets no answer within a second counts as
unanswered. The outage runs from the first unanswered attempt to the first
answer after it, so it is known to the probe’s interval, which the line
prints: unreachable 110ms, probed every 100ms. The settle, the undo and
the stopping of the writers happen while the probe runs and are not part of the
number. When every attempt was answered the line says never rather than
printing a zero. A frozen database counts as unreachable: the kernel accepts
the connection and nothing answers it.
The first crash after af up can take noticeably longer to recover than later
ones. Before it replays anything, Postgres syncs every file in the data
directory to disk (recovery_init_sync_method, which defaults to fsync), and
on the first crash those files include every page written when the branch was
created. Measured on the demo ledger, that step took between 1.9 and 8.4
seconds on the first crash after bringing the environment up, and under 0.2
seconds on the crashes after it. The database’s log shows it between
database system was interrupted and redo starts at, and with
log_startup_progress_interval lowered it prints syncing data directory (fsync) as it goes. It is Postgres making the data directory durable before
trusting it, not the fault or the engine, and how long it takes depends on the
disk under the container.
Tuning
Section titled “Tuning”| Key | Default | What it is |
|---|---|---|
crash_recovery.writers |
8 | Connections committing at once. |
crash_recovery.commits_before_fault |
200 | Acknowledged commits before a fault lands. |
crash_recovery.recovery_timeout |
2m |
How long the database has to answer a query again. |
faults[].after |
5s |
A floor on how long the run waits before the fault: with the writers committing around a database fault, and as a plain wait before any other. |
faults[].hold |
3s |
How long the fault stays in place. |
Every fault reports how long it was in place, measured from the moment the
injection returned to the moment its undo began, beside the hold it declared:
It was in place for 5.001s (declared 5s), then undone. in the terminal and
the pull request comment, and in_place_ms, hold_declared_ms and in_place
in the MCP result. The duration_ms beside them is the whole step, including
the wait before the fault, and is not how long the fault lasted.
Around a database fault the fault is undone at its hold and the writers are stopped after it, so a freeze lasts as long as it declares. Commits the writers make after the undo are counted and checked like every other: each one the client was told was committed must still be there.
commits_before_fault counts commits rather than seconds on purpose. A second
on a loaded machine can be a second in which nothing committed, and a crash
with nothing to lose passes every durability assertion by having none to make.
Limits
Section titled “Limits”Faults run on the local runtime, against Docker containers. On Kubernetes the
run reports AF-CHS-007 rather than injecting anything.
Network latency and packet loss are not implemented. Shaping traffic needs
tc inside the target’s network namespace, which the database and application
images do not carry and which the environment cannot fetch, because everything
it reaches goes through a default deny egress policy. A declared fault that
silently did nothing would be worse than an absent one, so the kind does not
exist. network_partition is the network fault that does work.