mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-05 21:07:43 +00:00
5a372557e5
The whole disk traversal runs inside spawn_blocking with no timeout anywhere on the async side; the in-scan ProgressMonitor checks are cooperative and only run between walker entries. A stat/readdir blocked on a dying disk or hung NFS mount therefore never returns: the blocking thread leaks, the refresh singleflight stays running forever, every subsequent scheduled refresh is skipped as inflight, and every admin capacity query joins an unbounded wait until process restart (S02). - Wrap each disk scan in tokio::time::timeout with a hard wall-clock ceiling of 2x the cooperative budget (min 5s). On expiry the caller fails the disk scan (releasing the singleflight through the normal error path, where the degraded/partial machinery from backlog#1014 keeps the failed disk's last-known value) and a shared AtomicBool asks the blocking walker to exit at its next entry, bounding the thread leak to the single wedged syscall. - Bound refresh_or_join joiner waits at 5 minutes so admin queries degrade into a clear error instead of hanging if the leader wedges in a way the drop/panic guards don't cover. Ref: rustfs/backlog#1017 (S02 from audit rustfs/backlog#1010)