mirror of
https://github.com/rustfs/rustfs.git
synced 2026-07-26 08:18:18 +00:00
03af8e472b
The background tier free-version recovery walk pinned a hardcoded 60s total wall-clock timeout that overrides every operator knob, so any bucket whose healthy full walk exceeds 60s fails forever; the failed run's duration was also subtracted from the next 60s tick, restarting the walk immediately and pinning CPU and disk I/O. - Drop the total wall-clock budget on the recovery walk (walkdir_timeout: Duration::ZERO) and inherit the operator-tunable drive stall budget (RUSTFS_DRIVE_WALKDIR_STALL_TIMEOUT_SECS) for per-call progress, so hung disks still fail fast. - Back off failed runs from completion time: 60s doubling to a 600s cap, reset on success; never subtract the failed run's duration. - Add RUSTFS_TIER_FREE_VERSION_RECOVERY_ENABLED (default true) to opt out of the recovery worker on deployments with no remote tiers; invalid values warn and fail open. Fixes #5130 Co-authored-by: claude <claude@ehdtn.com>