Skip to content

Type to search pages.

View .md

Fault injection and crash recovery

A rehearsal tells you what a change does to a system that works. The chaos block tells you what the system does when it stops working, and then it proves the answer instead of reporting that everything came back.

chaos:
enabled: true
faults:
- name: postgres-crash
kind: process_kill
target: database
process: "postgres: checkpointer"

Run it with af chaos against a running environment, or let af ci run it at the end of a check.

Around a fault aimed at the database, concurrent writers commit into a schema the engine owns, and the fault lands while they are committing. Afterwards the run establishes four things:

  1. No lost durable commit. Every transaction the client was told was committed is still there.
  2. No phantom commit. Nothing is there that no client ever tried to write.
  3. The write ahead log replayed. Recovery started at the position the control file named before the crash, and reached past the last flush a writer saw.
  4. The relations survived. A sequential scan and an index only scan count the same rows, and amcheck finds an index entry for every live heap tuple.

The sequential scan also reads every page of the writers’ table, and on a cluster with data checksums on, a page torn by the crash fails its checksum and stops that read. af chaos prints the result on each crash fault’s pages line. It covers the writers’ table and no other, and it says the pages were not checked when checksums are off, when the control file could not be read after the fault, or when the read did not finish. The amcheck line beside it prints what the index verifier said, or that it did not run.

The first two need something the database cannot give you, because they are claims about what the database said rather than about what it holds. The engine keeps a ledger on the client side of the wire: an identifier goes in before the statement is sent, and moves to acknowledged only when the call returns without an error. A commit that returned success and is absent afterwards is a durability failure whatever caused it.

Kind What happens Undo
process_kill SIGKILL to one process inside the container, matched by a substring of its command line. The container keeps running. None. The recovery is the system’s own, and that is the fault.
container_kill SIGKILL to the container’s main process. The container stops. Starts it again.
container_stop SIGTERM, then SIGKILL after a grace period. Starts it again.
container_pause Freezes every process with the cgroup freezer. Nothing is killed and no connection closes. Thaws it.
network_partition Detaches the container from the environment’s network. Attaches it again, with the aliases it had.
read_only_data Removes write permission from the data directory. Restores the mode it recorded.
disk_fill Fills the filesystem holding the data directory to a stated headroom. Needs database.data_filesystem.size_bytes, below. Removes the file it wrote.

process_kill and container_kill are the two kinds that stop Postgres uncleanly, so they are the two the recovery proof expects a replay from. The others are useful and they are honest about what they are: a container_stop shuts the database down cleanly and replays nothing, and a run that declared it as a crash reports that it could not establish a recovery rather than reporting a clean one.

A fault reaches the containers this environment created and nothing else. The target resolves from the labels the runtime stamped at create time, never from a name a fault supplied, and the ownership is read again from the daemon at the instant of the act. Three refusals have no override:

  • a container carrying no dev.antifailure.managed label is not ours
  • a container belonging to a different environment
  • the egress sidecar and the emulators, whatever environment they belong to

The sidecar carries the egress policy. A fault that could stop it would switch off the control that decides what the environment may reach, and a chaos feature that can disable a safety control is a way out with a feature name. An emulator stands in for a third party the environment must not reach, so stopping one does not produce an outage: it produces a request that goes looking for the real host.

disk_fill carries a fourth refusal, and a declaration that lifts it.

A container’s writable layer is the daemon’s own disk, so filling a directory on it fills the machine and every other container running on it. The fault reads the mount at the data directory from the daemon and refuses unless it is a volume this environment created with a size fixed when it was created. A mount of its own is not enough on its own: a plain named volume is its own mount and is still a slice of the daemon’s disk, so it would pass a device check and take the machine down having satisfied the guard. Both refusals are reported as chaos.fault.unsafe: the claim the fault was declared to establish was not established, and nothing else in the run was touched by it.

Giving the data directory a filesystem of its own

Section titled “Giving the data directory a filesystem of its own”
database:
data_filesystem:
size_bytes: 536870912
chaos:
enabled: true
faults:
- name: fill-the-data-volume
kind: disk_fill
target: database
headroom_bytes: 8388608
max_fill_bytes: 536870912

With that, the branch keeps its data directory on a filesystem of the declared size and disk_fill lands: the fill writes one file until the stated headroom is left, Postgres meets a real No space left on device on its next extend, and the undo removes the file and the free space comes back. Without it the data directory is on the writable layer and the fault is refused before it acts.

The filesystem is held in memory, and that is the containment argument rather than an implementation detail. A volume on the daemon’s disk cannot be filled without taking space from every other container on the machine; one in memory has a size fixed at creation and takes nothing from anything outside the environment. Three things follow, and they are the cost of the feature:

  • The whole database lives in it, so the size has to hold the data directory with room left for the fault to fill. A copy that does not fit is refused by name, with both numbers, rather than truncated.
  • A size of more than half the memory the Docker daemon reports is refused. A filesystem in memory larger than the machine moves the same problem from the disk to the memory, and a daemon killed for memory takes every other environment with it.
  • The data directory does not survive the Docker daemon restarting. af up builds it again from the golden.

The branch pays a copy of the data directory when it comes up, where an ordinary branch pays nothing because the daemon’s storage driver copies on write. So this is the layout for rehearsing a disk that fills, and not the one to measure how a disk performs.

The environment also runs one container that holds that filesystem mounted and does nothing else. It is not decoration: the local volume driver unmounts a memory backed volume when the last container using it stops, so without it a container_kill or container_stop would delete the data directory rather than crash the database, the undo would start a container that initialised an empty one, and the durability proof would report every acknowledged commit lost. Faults refuse to touch it for the same reason they refuse to touch the egress sidecar.

Nothing that changed nothing counts as survived

Section titled “Nothing that changed nothing counts as survived”

A fault that was applied and had no effect is refused, not reported. The reason is the whole point of the feature: every assertion after such a fault describes a system that never broke, and a recovery check that passes on one is a check that answers the same whether or not it ran.

So a process_kill whose pattern matches nothing is refused rather than reported as a crash the database survived. A read_only_data fault probes a write as the directory’s owner and refuses if the write still succeeds, which is what happens on a directory owned by root, because root ignores the mode. A container_pause that the daemon accepts and that leaves the container running is refused.

The same discipline runs through the findings. A run that could not establish what it set out to is reported as unverified and never as a pass:

Finding Meaning
chaos.durability.lost_commit A transaction the client was told was committed is gone.
chaos.durability.phantom_commit A row is present that no client wrote.
chaos.recovery.replay_short Recovery stopped before the last position the client saw flushed.
chaos.recovery.timeline_moved The timeline changed, and crash recovery does not change it.
chaos.integrity.relation_damaged The heap and its index disagree.
chaos.invariant.broken_by_fault One of this project’s own invariants held before the fault and does not hold after the recovery.
chaos.recovery.no_crash The fault was declared as a crash and nothing crashed.
chaos.recovery.no_replay The database came back and the log records no replay.
chaos.integrity.checksums_off Data page checksums are off, so a torn page would not be seen.
chaos.integrity.amcheck_unavailable The index could not be verified.
chaos.durability.inconsistent_ledger The engine’s own bookkeeping does not add up.
chaos.invariant.already_violated One of this project’s own invariants did not hold before the fault either, so nothing after it is attributable to the fault.
chaos.invariant.unevaluated One of this project’s own invariants could not be asked on one side or the other, which a database that did not come back is the loudest case of.
chaos.fault.refused A fault tried to go in and failed, so it established nothing.
chaos.fault.unsafe A fault was refused before it acted, because its effect would reach past this environment. It changed nothing the other faults measured.
chaos.fault.not_undone A fault went in and its undo failed, so the environment is still broken and anything measured after it is suspect.

The first six are failures and carry policy.chaos_failure, which defaults to fail. The last ten are the ones the run could not look at, and they carry policy.chaos_unverified, which defaults to warn. They are two keys because a check that found a problem and a check that could not look are different facts, and reporting the second as the first teaches a project to ignore both.

Your own rules, asked of the recovered database

Section titled “Your own rules, asked of the recovered database”

Everything the durability proof asserts is about a schema of the engine’s own, and that is deliberate: asserting that a table your application is writing did not change, while it is writing it, is a claim about a moving target. That reason stops applying the moment the writers stop and the database answers a query again, and that is exactly when the invariants your manifest declares are the right question. The ledger proves the engine’s commits survived. Only your invariants can say whether your data still means what you say it means.

So around a fault with the durability proof on, every invariant the manifest declares is asked twice: once before anything is broken, and once against the recovered database. Both answers are printed, because one of them cannot be read on its own.

invariant no-negative-balance: before the fault held; after the recovery held
invariant orders-have-a-customer: before the fault held; after the recovery violated, 2 rows

An invariant that was already violated before the fault is reported and is attributed to nothing: the rule is broken and this run is not what broke it, so chaos.invariant.already_violated is unverified and never fails the run. A gate that stopped a merge for a rule the change did not break would teach a project to switch the whole arm off. Only a rule that held before the fault and does not hold after the recovery is something the run can attribute to it, and that one is chaos.invariant.broken_by_fault, which fails.

An invariant that could not be asked, on either side, is chaos.invariant.unevaluated. A database that did not come back is the loudest case of it, and it is the one where reporting nothing would be worst: an absent arm reads as an arm with nothing to report.

A manifest that declares no invariants runs none of this and nothing about it appears in any output.

crash_recovery.synchronous_commit sets what the writers ask of the database. With it off, Postgres acknowledges a commit before the write ahead log record has left shared memory, so a crash that discards shared memory loses commits the client was told were durable. That is the setting’s documented behavior and the run reports the loss:

chaos:
enabled: true
crash_recovery:
synchronous_commit: off
faults:
- name: prove-the-check-can-say-no
kind: process_kill
target: database
process: "postgres: checkpointer"

Leave it out unless you mean it. A manifest that sets it to off is asking for a run that is expected to report lost commits, which is useful exactly once: to see the check say no before you trust it saying yes.

The unreachable line is measured by a probe that starts with the fault and runs beside it. Every 100 milliseconds it opens a connection and runs SELECT 1, and an attempt that gets no answer within a second counts as unanswered. The outage runs from the first unanswered attempt to the first answer after it, so it is known to the probe’s interval, which the line prints: unreachable 110ms, probed every 100ms. The settle, the undo and the stopping of the writers happen while the probe runs and are not part of the number. When every attempt was answered the line says never rather than printing a zero. A frozen database counts as unreachable: the kernel accepts the connection and nothing answers it.

The first crash after af up can take noticeably longer to recover than later ones. Before it replays anything, Postgres syncs every file in the data directory to disk (recovery_init_sync_method, which defaults to fsync), and on the first crash those files include every page written when the branch was created. Measured on the demo ledger, that step took between 1.9 and 8.4 seconds on the first crash after bringing the environment up, and under 0.2 seconds on the crashes after it. The database’s log shows it between database system was interrupted and redo starts at, and with log_startup_progress_interval lowered it prints syncing data directory (fsync) as it goes. It is Postgres making the data directory durable before trusting it, not the fault or the engine, and how long it takes depends on the disk under the container.

Key Default What it is
crash_recovery.writers 8 Connections committing at once.
crash_recovery.commits_before_fault 200 Acknowledged commits before a fault lands.
crash_recovery.recovery_timeout 2m How long the database has to answer a query again.
faults[].after 5s A floor on how long the run waits before the fault: with the writers committing around a database fault, and as a plain wait before any other.
faults[].hold 3s How long the fault stays in place.

Every fault reports how long it was in place, measured from the moment the injection returned to the moment its undo began, beside the hold it declared: It was in place for 5.001s (declared 5s), then undone. in the terminal and the pull request comment, and in_place_ms, hold_declared_ms and in_place in the MCP result. The duration_ms beside them is the whole step, including the wait before the fault, and is not how long the fault lasted.

Around a database fault the fault is undone at its hold and the writers are stopped after it, so a freeze lasts as long as it declares. Commits the writers make after the undo are counted and checked like every other: each one the client was told was committed must still be there.

commits_before_fault counts commits rather than seconds on purpose. A second on a loaded machine can be a second in which nothing committed, and a crash with nothing to lose passes every durability assertion by having none to make.

Faults run on the local runtime, against Docker containers. On Kubernetes the run reports AF-CHS-007 rather than injecting anything.

Network latency and packet loss are not implemented. Shaping traffic needs tc inside the target’s network namespace, which the database and application images do not carry and which the environment cannot fetch, because everything it reaches goes through a default deny egress policy. A declared fault that silently did nothing would be worse than an absent one, so the kind does not exist. network_partition is the network fault that does work.