Files
rustfs/docs
cxymds d0ce2f758b fix(ecstore): bound decommission target gate contention (#8061)
* fix(ecstore): bound decommission target gate contention

Target capacity gate contention during pool decommission escalated a
per-object transient into a durable bucket pause: each contended object
failed its bucket entry, which re-ran the whole bucket listing and amplified
attempts on the same objects.

- centralize the decommission capacity failure classification so gate
  contention, benign contention and fatal failures are decided once
- retry target-gate contention inline (bounded, jittered) before a mutation
  is admitted, covering put, part, complete, new-multipart and abort
- defer contended entries to the end of the round instead of failing the
  bucket entry, and require the deferred set to drain before a set completes
- treat a missing object or version, an overwrite and a superseded upload id
  as benign contention that is neither counted as a failure nor escalated
- use full-jitter exponential backoff, capped, for decommission retries
- expose the capacity pause reason, the waiting reason and a cumulative
  pause count in the admin pool status, plus gate-retry and per-object
  attempt metrics

Related: rustfs/backlog#2644

* fix(ecstore): correct deferred replay and metadata compatibility
2026-09-22 19:20:05 +08:00
..

Documentation

Use the focused indexes rather than treating this directory as an unordered collection:

Operations

Operational runbooks live under operations/. Replication operators should start with:

Runbook Use it for
Site replication operations Health fields, pending operations, outage recovery, re-pair admission, IAM/SSE boundaries, and upgrades.
Replication target check Validating an S3 destination and version fidelity before enabling replication.
Replication object size limits Multipart routing, large-object limits, and retry characteristics.
Replication outbound transport Integrity headers, generic target behavior, and transport knobs.

For the erasure-coded cluster lifecycle (planning, parity and EC:0, expansion, rebalance, decommission, heal, drive replacement, restart recovery, and the rc CLI mapping), start with Cluster and erasure-coding lifecycle operations.

For persisted administrator bucket tasks and bucket recreation, see Bucket heal recovery.

For disk replacement across VM restarts and schema 5/6 maintenance migration, see Replacement generation recovery.

For historical GET timeouts during PUT or Heal, see Object lock contention diagnostics.

Other runbooks remain grouped by filename in operations/; architecture pages link to the relevant runbook where a cross-boundary procedure is required.

For storage dashboards, see Storage metrics and observer selection: drive ownership, snapshot freshness, counter queries, and rolling upgrades.

For optional shard commitments, see Independent shard integrity rollout: activation, legacy repair results, multipart mode changes, and rollback limits.

For crates.io publication of workspace crates, see Workspace Cargo Publish: dependency ordering, dry-run, publish, and failure handling.