RUSTFS_CAPACITY_MAX_FILES_THRESHOLD=0 made every file count as
overflow with an empty exact prefix, so disks with fewer files than the
sample rate committed 0 bytes as the cluster total; STAT_TIMEOUT=0 with
dynamic timeout disabled made every scan fail at the first entry. Both
were accepted unvalidated, unlike the already-clamped sample_rate and
interval values.
Zero values for the threshold and the stat/min/max timeouts now fall
back to their defaults with a single structured warn at config load
(S11). The symlink-depth knob is left untouched here: it is removed
entirely by the backlog#1018 fix.
Ref: rustfs/backlog#1019 (S11 from audit rustfs/backlog#1010)
Progress/timeout checks only ran when file_count advanced, so trees of
directory-only or error-dense entries (dirs, traversal errors, symlinks
never increment file_count) bypassed the entire cooperative time budget
and held the blocking thread for unbounded time, while each error entry
logged its own warn — permission-dense trees flooded the log (S13).
The stall detector was structurally unreachable: it fired on 'no file
progress between checks', but checks themselves only ran when file
progress happened, so RUSTFS_CAPACITY_STALL_TIMEOUT never did anything
since its introduction (S09).
- Progress checks (timeout + early-sampling entry) are now driven by a
visited-entry counter that also advances on directories and errors, so
every tree shape reaches the budget checks.
- Remove the unreachable stall detector and its plumbing end to end
(ProgressMonitor fields, ScanLimits, config getter, env const, stall
metric). Genuine walker wedges are handled by the hard outer
wall-clock budget from backlog#1017; no decorative protection is left.
- Cap per-entry error warns at 10 per scan with an explicit suppression
notice; had_partial_errors still records the condition.
Ref: rustfs/backlog#1016 (S09+S13 from audit rustfs/backlog#1010)
Two dirty-mark lifecycle gaps (S19+S30):
- A commit cleared every dirty mark for the disks it scanned, even marks
recorded while the walker was already past the written prefix, so the
next dirty-subset refresh served stale cached bytes as exact until the
next full scan. Dirty marks now carry the instant they were last
recorded and a commit only clears marks that predate its scan start
(CapacityUpdate::scan_started_at, set by refresh_capacity_with_scope).
- The write side marks every disk of an EC set dirty, including remote
peers, while scans and clearing only cover local disks — remote or
removed disks stayed marked forever, keeping the dirty-disk gauge
permanently non-zero with bounded memory stuck behind it. The refresh
disk selection now prunes dirty entries outside the current local
topology before consulting them.
Ref: rustfs/backlog#1020 (S19+S30 from audit rustfs/backlog#1010)
Two defensive gaps left the refresh singleflight vulnerable to a
permanent wedge (running=true forever: scheduled refreshes silently
stop, admin capacity joiners hang until restart):
- spawn_refresh_if_needed's spawned task had no leader guard, unlike
refresh_or_join: a panic after the refresh future completes (commit,
metrics) killed the task before the reset block. The task now holds
the same RAII RefreshLeaderGuard, disarmed only after the state is
reset and the result published (S20).
- refresh_fn() was evaluated before AssertUnwindSafe wrapping in both
paths, so a panic while constructing the future escaped catch_unwind;
construction now happens inside the wrapped future.
- RefreshLeaderGuard::drop silently skipped the reset when try_lock was
contended and no tokio runtime was current; it now falls back to a
blocking reset, which is safe precisely because there is no executor
to stall in that context (S31).
Ref: rustfs/backlog#1021 (S20+S31 from audit rustfs/backlog#1010)
The whole disk traversal runs inside spawn_blocking with no timeout
anywhere on the async side; the in-scan ProgressMonitor checks are
cooperative and only run between walker entries. A stat/readdir blocked
on a dying disk or hung NFS mount therefore never returns: the blocking
thread leaks, the refresh singleflight stays running forever, every
subsequent scheduled refresh is skipped as inflight, and every admin
capacity query joins an unbounded wait until process restart (S02).
- Wrap each disk scan in tokio::time::timeout with a hard wall-clock
ceiling of 2x the cooperative budget (min 5s). On expiry the caller
fails the disk scan (releasing the singleflight through the normal
error path, where the degraded/partial machinery from backlog#1014
keeps the failed disk's last-known value) and a shared AtomicBool asks
the blocking walker to exit at its next entry, bounding the thread
leak to the single wedged syscall.
- Bound refresh_or_join joiner waits at 5 minutes so admin queries
degrade into a clear error instead of hanging if the leader wedges in
a way the drop/panic guards don't cover.
Ref: rustfs/backlog#1017 (S02 from audit rustfs/backlog#1010)
A full refresh with partial disk failures used to commit the surviving
subset's sum as a fresh exact cluster total (no field carried the
partial-failure fact), while the complete disk cache kept the failed
disk's old value — so the reported capacity oscillated between the
partial sum and the cache-merged total on alternating refreshes.
- Add a degraded flag to CapacityUpdate/CachedCapacity, set when the
scan behind the update had partial errors; expose it in refresh logs,
admin capacity logs and a new rustfs_capacity_degraded_readings_total
counter.
- On a degraded full refresh, surface only the disks whose own scan
fully succeeded and merge them over a complete disk cache, so failed
disks keep their last-known values and the published total no longer
dips and bounces back. The disk cache is never replaced from a
degraded refresh.
- Without a complete cache, keep the partial sum (unchanged #805
non-pollution behavior) but mark the reading degraded.
Ref: rustfs/backlog#1014 (S06 from audit rustfs/backlog#1010)
A dirty-subset refresh scans only the dirty disks, so its raw
`CapacityUpdate` carries just that subset's `total_used`/`file_count`.
`update_capacity` recomputed the correct cluster-wide total into a local
variable and wrote it to the cache, but never wrote it back to the
`CapacityUpdate` that `refresh_or_join`/`spawn_refresh_if_needed` return
and publish to joiners. The admin blocking path (`resolve_admin_used_capacity`
-> `refresh_or_join_admin_disks(allow_dirty_subset=true)`) consumes
`update.total_used`, so a single dirty disk in an N-disk cluster made
admin StorageInfo report only the scanned subset's bytes (a large
transient undercount for that request and same-cycle joiners).
`CachedDiskCapacity` also dropped each disk's `file_count`/`is_estimated`,
so subset refreshes could not recompute a correct cluster file count and
would launder an estimated per-disk value into an exact cluster total.
Fix:
- `update_capacity` now reconciles `total_used`, `file_count`, and
`is_estimated` from the full per-disk cache and returns the corrected
`CapacityUpdate`.
- `CachedDiskCapacity` stores `file_count`/`is_estimated` per disk;
`file_count` sums the cache and `is_estimated` is the OR across disks.
- `refresh_or_join` and `spawn_refresh_if_needed` rebind their result to
the reconciled update so the leader return value and the value
published to joiners both carry cluster totals.
Tests: extend the subset-refresh test to assert the returned update's
reconciled `total_used`/`file_count`/`is_estimated`, and add a
`refresh_or_join` dirty-subset test asserting the leader returns the
merged cluster total (not the subset sum) and matches the cache.
Refs: https://github.com/rustfs/backlog/issues/1011
* fix(object-capacity): stop cancelled/remote-disk refreshes from corrupting capacity
Two independent capacity-refresh bugs (backlog rustfs/backlog#805):
- refresh_or_join set the singleflight `running` flag then awaited the
refresh future with no drop guard. When the admin request that became
leader was cancelled (client disconnect) mid-await, `running` stayed true
forever: joiners blocked indefinitely and the 120s scheduled refresh could
never start again. catch_unwind covered panics but not cancellation. A
RefreshLeaderGuard now resets the state and publishes an error on drop.
- capacity_disk_refs mapped the cluster-wide storage_info disk list without
filtering non-local disks, so admin-triggered refreshes ran a local WalkDir
over remote disks' drive_path. On multi-node clusters this double-counted
local bytes (shared mount layout) or hit NotFound (per-node layouts), and
poisoned the per-disk cache the scheduled local-only refresh depends on,
making the cached total oscillate. Filter to local disks.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(rio): harden internode HTTP client build, cache, and PUT body integrity
Follow-ups from the internode HTTP review (backlog rustfs/backlog#805):
- handle_put_file accepted a truncated body as success: a HttpWriter dropped
mid-stream closes the chunked body cleanly, indistinguishable from EOF, and
the server never compared bytes copied against the declared size. Reject
size mismatches on the create path (append/unknown-size writes send size<=0
and are exempt).
- build_http_client used `.expect()` on ClientBuilder::build(), which runs
lazily on the first request and on every TLS generation bump (cert
rotation), so a build failure panicked a serving task. It now returns an
io::Error; get_http_client falls back to the previous TLS generation when a
rebuild fails instead of failing the request.
- CLIENT_CACHE was a tokio::Mutex taken on every stream open (data_shards
times per GET). Replaced with arc_swap::ArcSwapOption for lock-free reads on
the hot path; the generation-monotonic replacement guard is preserved via rcu.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Remove or consolidate 57 test cases that cannot catch regressions
(literal-constant asserts, construct-then-assert, derived-serde
round-trips, near-duplicate env/getter matrices) in common, config,
iam, madmin, and object-capacity, keeping all wire-format and
error-path guards. Add 13 tests for previously uncovered high-risk
behavior: filemeta version-sort determinism and merge resilience to
garbage headers, zip extraction path-traversal rejection and exact
limit boundaries, JWT tampered-signature rejection, and the bytes
variant of dual-key (rustfs/minio) metadata fallback and precedence.
Test-only change; no production code touched.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Treat empty ping bodies as liveness probes, add a startup cleanup barrier for early walk_dir calls, and delay immediate background cleanup/capacity timers to reduce transient restart-time VolumeNotFound noise.
Also downgrade expected missing-path producer results during startup from generic errors to warnings while preserving existing storage semantics.