Operate · 03

Restore any moment.

Back up continuously to S3 or a compatible store, and restore the state as of any point in the last 30 days.

Set FLOWER_BACKUP_URL and the cluster backs itself up, continuously, to S3 or an S3-compatible store (Garage, MinIO, Ceph, Cloudflare R2), or to a directory. You can restore the state as of any moment within the retention window, 30 days by default, to the entry.

How backups work

  • The leader ships the log. About once a second it uploads the entries it applied since the last time, as one zstd-compressed object, each entry stamped with when it was applied. Followers ship nothing; a new leader picks up where the old one stopped.
  • Now and then it writes a base: the whole state, compressed, as a snapshot transfer encodes it. A base is written once a day, and sooner once the log shipped since the last one is as large as it (and at least 64 MiB). Writing one streams the state to the store in parts; nothing goes to local disk.
  • A restore starts from the newest base at or before the point you ask for and replays the log after it, up to that point, without running your code.
  • A generation is one unbroken history: a base and the log that follows it, with later bases. A new leader continues the generation if its log holds the generation's last shipped entry. Otherwise, for example after a restore, or after the store was unreachable for so long that the log it needed is gone, it starts a new generation with a base of its own.
  • What is lost if every node's disk is lost at once: at most the last second or so of writes (FLOWER_BACKUP_INTERVAL_MS), more while the store is unreachable.

Raft compacts its log aggressively. So that compaction never outruns shipping, each node keeps the entries its backup still needs, up to FLOWER_BACKUP_HOLD_MAX_BYTES (1 GiB): the leader until it has shipped them, a follower the newest ones, in case it becomes leader. If the store is unreachable long enough to fill that, the oldest held entries go, and once the store is back the leader starts a new generation.

Turn backups on

Set the same values on every node and restart them. For AWS S3:

Every node · AWS S3
export FLOWER_BACKUP_URL='s3://my-bucket/flower/production'
export AWS_REGION='eu-west-3'
export AWS_ACCESS_KEY_ID='AKIA…' AWS_SECRET_ACCESS_KEY='…'

For an S3-compatible store, give its endpoint. Path-style addressing is the default with an endpoint:

Every node · Garage, MinIO…
export FLOWER_BACKUP_URL='s3://flower-backup/production'
export FLOWER_BACKUP_S3_ENDPOINT='http://127.0.0.1:3900'
export FLOWER_BACKUP_S3_REGION='garage'
export FLOWER_BACKUP_S3_ACCESS_KEY_ID='GK…' FLOWER_BACKUP_S3_SECRET_ACCESS_KEY='…'

A mounted volume works too: FLOWER_BACKUP_URL=file:///mnt/backups/flower. A process hosting replicas of several groups (--replica) backs each up under its name: …/production/west/.

The key needs to list the bucket and to get, put and delete objects under the prefix, multipart uploads included. Give each cluster a prefix of its own. Add a lifecycle rule that aborts incomplete multipart uploads after a day: a base whose upload a crash interrupted leaves its parts behind.

Backup settings
SettingDefaultWhat it does
FLOWER_BACKUP_URLunset (no backups)s3://BUCKET/PREFIX or file:///PATH.
FLOWER_BACKUP_S3_ENDPOINTAWS for the regionAn S3-compatible store's http:// or https:// URL.
FLOWER_BACKUP_S3_REGIONAWS_REGION, else us-east-1The region requests are signed for.
FLOWER_BACKUP_S3_ADDRESSINGpath with an endpoint, else virtualpath (host/bucket/key) or virtual (bucket.host/key).
FLOWER_BACKUP_S3_ACCESS_KEY_ID, …_SECRET_ACCESS_KEY, …_SESSION_TOKENAWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKENStatic credentials; there is no instance-profile lookup.
FLOWER_BACKUP_INTERVAL_MS1,000 msHow often the leader ships the log.
FLOWER_BACKUP_RETENTION_MS30 daysHow far back restores reach.
FLOWER_BACKUP_BASE_INTERVAL_MS24 hA new base once the newest is this old.
FLOWER_BACKUP_BASE_AFTER_BYTES64 MiBA new base once this much log has shipped since the newest, or as much as that base holds if more.
FLOWER_BACKUP_SEGMENT_MAX_BYTES16 MiBLog bytes per uploaded object, at most.
FLOWER_BACKUP_PART_BYTES16 MiBMultipart upload part size of bases (at least 5 MiB).
FLOWER_BACKUP_HOLD_MAX_BYTES1 GiBLog each node keeps for its backup, at most.

Check on backups

GET /admin/backup (operator token) says what this node's backup is doing: its role (leader ships, follower holds), the generation, the newest entry shipped and when it was applied, how many applied entries are not shipped yet (lagEntries), the newest base, retention's last run, and errors. hold says how much log the node keeps for it. Failures are retried every interval and logged under flower::backup.

Operator terminal
curl -fsS -H "Authorization: Bearer $FLOWER_ADMIN_TOKEN" http://127.0.0.1:7101/admin/backup

# What can be restored, generation by generation:
./flower backup list

Restore to a point in time

A restore runs offline. It writes a new data directory holding the state as of the point you give, for a new single-node cluster, then exits. Start a node on it with the same ID and it elects itself; add other nodes with the membership API.

Restore · the same FLOWER_BACKUP_* settings
# The state as of 12:23 Pacific time (or --at 1790796180000, Unix ms):
./flower backup restore --data .flower/restored --id 1 \
  --advertise 127.0.0.1:7101 --at 2026-09-30T12:23:00-07:00

# Or through a log index (--index 123456), or everything the backups hold (neither).
./flower --id 1 --listen 127.0.0.1:7101 --data .flower/restored
  • A point in time restores every entry the old leader had applied by then. --from URL reads other backups than FLOWER_BACKUP_URL's, --replica NAME a hosted replica's, and --generation picks a generation other than the newest one covering the point.
  • The restored node backs up as a new generation, which says where it came from. The old generations stay, and restores still reach them, until retention lets them go.
  • The new cluster's Raft terms start far above the old history's, so its log never mixes with the old cluster's.
  • Stop the old cluster first, or keep clients away from it. Restoring rolls back retry history, transaction records and fencing counters with the data. If you use retry retention, give the restored service a new incarnation before it takes traffic, as RETENTION.md describes. Restoring one group of several doesn't reconcile transactions that span groups.
  • A restore needs a binary with the same snapshot and value formats as the one that wrote the backups, and state machine commands at least as new. Backup objects have a format of their own, which each generation records: a binary continues and restores only generations in its own, so one that changes it starts a new generation (reason format), and backup list shows the older ones as not restorable by it.

What the store holds

Under the prefix, each generation has a directory of its own. Every object ends with the SHA-256 of what precedes it, and a restore refuses a damaged one.

Object keys
generations/{created}-{random}/generation.json   why it began, after what, in which format
generations/{created}-{random}/tip.json          the newest shipped entry, every 10 s
generations/{created}-{random}/bases/{index}-{time}.base
generations/{created}-{random}/log/{first}-{last}-{time}.seg

Retention runs hourly and after each base. In each generation it keeps the newest base at or before the horizon and everything after it, so every point since the horizon stays restorable, and it deletes generations that ended before the horizon. Bases older than a day (FLOWER_BACKUP_BASE_INTERVAL_MS) thin out to one a day: restoring to an older point replays up to a day of log from the base before it. A busy cluster uploads about one log object a second; an idle one none.

What the store holds is mostly the log: every write of the retention window, compressed, plus about one base a day. To size it, watch how fast segmentBytesShipped in GET /admin/backup grows. When the store shares a disk with the cluster, as a local Garage or MinIO does, a longer FLOWER_BACKUP_INTERVAL_MS (say 30 s) costs nothing that store could save and packs the log better: an object compresses as a whole, and more entries in one compress smaller.