* fix(heal): persist pool/set outcomes and recovery scope * test(heal): verify selector receipts and process-crash recovery * fix(heal): scope replacement pool metadata to its owning set (cherry picked from commita72a75569c) * fix(heal): use domain imports in pool metadata regression test (cherry picked from commite8c5448e21) * test(heal): set pool metadata regression recursion limit --------- Co-authored-by: Hiroaki KAWAI <1468181+hkwi@users.noreply.github.com>
Documentation
Use the focused indexes rather than treating this directory as an unordered collection:
Operations
Operational runbooks live under operations/. Replication
operators should start with:
| Runbook | Use it for |
|---|---|
| Site replication operations | Health fields, pending operations, outage recovery, re-pair admission, IAM/SSE boundaries, and upgrades. |
| Replication target check | Validating an S3 destination and version fidelity before enabling replication. |
| Replication object size limits | Multipart routing, large-object limits, and retry characteristics. |
| Replication outbound transport | Integrity headers, generic target behavior, and transport knobs. |
For persisted administrator bucket tasks and bucket recreation, see Bucket heal recovery.
For disk replacement across VM restarts and schema 5/6 maintenance migration, see Replacement generation recovery.
For historical GET timeouts during PUT or Heal, see Object lock contention diagnostics.
Other runbooks remain grouped by filename in operations/;
architecture pages link to the relevant runbook where a cross-boundary
procedure is required.
For storage dashboards, see Storage metrics and observer selection: drive ownership, snapshot freshness, counter queries, and rolling upgrades.
For optional shard commitments, see Independent shard integrity rollout: activation, legacy repair results, multipart mode changes, and rollback limits.