fix(ecstore): bound decommission target gate contention (#8061)

* fix(ecstore): bound decommission target gate contention

Target capacity gate contention during pool decommission escalated a
per-object transient into a durable bucket pause: each contended object
failed its bucket entry, which re-ran the whole bucket listing and amplified
attempts on the same objects.

- centralize the decommission capacity failure classification so gate
  contention, benign contention and fatal failures are decided once
- retry target-gate contention inline (bounded, jittered) before a mutation
  is admitted, covering put, part, complete, new-multipart and abort
- defer contended entries to the end of the round instead of failing the
  bucket entry, and require the deferred set to drain before a set completes
- treat a missing object or version, an overwrite and a superseded upload id
  as benign contention that is neither counted as a failure nor escalated
- use full-jitter exponential backoff, capped, for decommission retries
- expose the capacity pause reason, the waiting reason and a cumulative
  pause count in the admin pool status, plus gate-retry and per-object
  attempt metrics

Related: rustfs/backlog#2644

* fix(ecstore): correct deferred replay and metadata compatibility
This commit is contained in:
cxymds
2026-09-22 19:20:05 +08:00
committed by GitHub
parent 9c30cc8851
commit d0ce2f758b
4 changed files with 1026 additions and 88 deletions
@@ -92,7 +92,21 @@ When decommission metadata is present, `decommissionInfo` includes:
- progress counters: `objectsDecommissioned`, `objectsDecommissionedFailed`, `bytesDecommissioned`, and `bytesDecommissionedFailed`;
- current location: `bucket`, `prefix`, and `object`;
- queue/history lists: `queuedBuckets` and `decommissionedBuckets`;
- `waitingReason`: `queued` for queued entries and `waiting_for_worker` when metadata exists but no worker has started.
- `waitingReason`: `capacity` while the pool is paused on target capacity,
`queued` for queued entries, and `waiting_for_worker` when metadata exists but
no worker has started;
- `capacityBlockedReason`: the persisted detail for the active capacity pause,
absent once the pause clears.
A capacity pause is reported ahead of the worker states so operators can tell
"contending but progressing" from "genuinely out of target capacity". The existing `pool.bin` layout is unchanged;
cumulative pause history requires a separately versioned persistence contract.
The object-attempt metrics count entry passes and repeated version-copy attempts
within a listing and its deferred replay. Inline gate waits use the gate-retry
counter. A new listing after a durable pause starts a new per-object observation;
the maximum is the highest observation in the process, not a persisted lifetime
attempt count.
This makes queued pools and stalled metadata visible without requiring operators to inspect pool metadata files directly.