* fix(core-storage): fix critical correctness defects from core-storage audit
Fixes verified defects found in a deep audit of the core storage path
(erasure coding, disk persistence, quorum, heal, replication resync):
- ecstore/disk: rewrite live xl.meta atomically (temp+rename) in
delete_versions_internal and write_metadata instead of in-place
truncate, which exposed torn metadata to concurrent readers and
crashes on the DeleteObjects hot path
- ecstore/erasure: allow heal to reconstruct from exactly data_shards
bitrot-verified sources; requiring data_shards+1 made objects
permanently unhealable after losing parity_shards disks
- ecstore/set_disk: direct-memory inline GET applied the erasure
distribution permutation twice (shuffled inputs re-indexed through
distribution), concatenating wrong shards into the response body in
degraded reads; collect from canonical disk-ordered inputs
- ecstore/set_disk: heal now preserves the committed inline layout
instead of recomputing it with a hardcoded unversioned threshold,
which split quorum identity of healed replicas and caused endless
re-heal churn
- ecstore/replication: resync results channel switched from
broadcast(1) to mpsc; a lagged broadcast receiver ended the stats
collector and every subsequent failure went uncounted, letting
failed resyncs be marked completed
- ecstore/replication: ignore an empty persisted resync checkpoint;
resuming with one skipped every object and marked the resync
completed without replicating anything
- ecstore/replication: fix inverted not-found error classification in
replicate_object/replicate_delete logging paths
- ecstore/erasure: guard decode paths against zero block_size or
data_shards from corrupt on-disk metadata (divide-by-zero panic)
- ecstore/disk: os::read_dir no longer consumes the entry limit on
entries it does not return (is_empty_dir misjudgment); create_file
opens with O_TRUNC to avoid stale trailing bytes
- filemeta: treat Some(nil) version id as a null version in
matches_not_strict; disk-loaded headers never store None, so the
mod_time quorum guard for unversioned overwrites never fired and an
interrupted overwrite could displace the committed version in merge
- filemeta: fix msgpack skip lengths for fixext (missed the ext type
byte) and ext16/32 (over-skipped) unknown fields
- filemeta: return FileCorrupt instead of usize underflow when
xl.meta is truncated inside the CRC trailer
- filemeta: surface delete-marker insertion failure in delete_version
instead of reporting success when the data dir is shared
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(replication): drop duplicate cfg(test) etag import from boundary module
The test module already imports content_matches_by_etag locally, so the
top-level cfg(test) import is unused under -D warnings and fails clippy.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Gate scanner-triggered deep heal behind a short new-object cooldown and retry transient bitrot shard size mismatches during local verification.
Add targeted tests for cooldown mode selection and size-mismatch classification, and log when deep scans are downgraded or size-mismatch retries occur.
* fix(ecstore): skip hidden metadata in walk limit
* fix(ecstore): match walk limit to visible versions
* fix(ecstore): avoid decoding all versions for limit
* fix(http): reduce header timeout log noise
Log peer_addr for HTTP connection errors so idle or probe traffic can be traced back to the source.
Downgrade HeaderTimeout connection logs from warn to info because these timeouts are often caused by incomplete probe traffic rather than a service-side fault.
* fix(ecstore): suppress spurious warn logs for missing trash cleanup sources
Extract `reliable_rename_inner` with a `warn_on_failure` flag so that
`rename_all_ignore_missing_source` can skip the warn log when the source
is absent. This completes the idempotent cleanup fix — the previous
implementation still emitted a `reliable_rename failed` warning before
the NotFound was caught and silenced.
Also restore the `recursive` branching in `move_to_trash` so that
non-recursive deletions use `fs::rename` directly (preserving original
semantics), while recursive deletions use the idempotent helper.
* fix(iam): prevent transient IAM walk timeout from crashing startup
IAM startup performs a blocking full metadata walk on `.rustfs.sys/config/iam/`.
When that distributed walk times out (e.g. disk pressure after cluster reboot),
the old code treated the failure as fatal and exited the process, causing a
systemd restart loop.
Changes:
- Add `startup_iam.rs`: attempt IAM init, enter degraded mode on failure,
spawn background retry task with exponential backoff (5s→10s→20s→30s cap)
- Log level escalates to ERROR after 12 retries (~5 min) to aid diagnosis
- `/health/ready` returns 503 until IAM recovers; IAM-dependent ops return
`IamSysNotInitialized` (existing fail-closed behavior preserved)
- Fix admin path boundary matching: `/minio/administrator` no longer falsely
matches as admin prefix
- Normalize Content-Length: 0 for admin GET requests with empty body
Fixes#3175
* fix(iam): move constant assertion into const block
Fixes clippy::assertions-on-constants warning on
IAM_RETRY_ESCALATION_THRESHOLD assertion.
* fix(iam): address PR review comments
- Replace OnceLock with AtomicU64 sentinel for test isolation;
add reset_test_failure_counter() for integration tests
- Use u32::try_from() instead of `as u32` narrowing cast in
compute_backoff_interval
- Rename misleading test; update to verify finalize retry behavior
- Restructure spawn_iam_recovery_task into init-retry and
finalize-retry phases so transient readiness failures are retried
instead of leaving the server permanently degraded
* fix(iam): gate test hooks behind debug_assertions
- reset_test_failure_counter() now stores sentinel (u64::MAX) to
correctly trigger env var re-read on next call
- RUSTFS_TEST_IAM_FAIL_INIT_ATTEMPTS only honored in debug builds
- RUSTFS_TEST_IAM_RETRY_INTERVAL_MS only honored in debug builds
* test(iam): cover deferred bootstrap recovery
Add a dedicated embedded deferred-IAM integration test in a separate test binary to avoid process-global startup collisions.
Strengthen startup IAM recovery coverage with focused unit tests and keep the existing embedded smoke test isolated while carrying the manual test license header update in the same change set.
* fix(startup): tighten deferred IAM recovery path
Adopt follow-up review feedback by silencing misleading app-context warnings after IAM recovery, reusing boundary-aware path prefix checks in the readiness gate, and tying deferred IAM recovery retries to server shutdown tokens.
Keep the deferred IAM embedded integration coverage and startup recovery unit coverage green after the follow-up hardening.
* refactor(startup): simplify IAM recovery task
Collapse the deferred IAM recovery implementation back to a concrete production flow instead of keeping boxed callback seams in the runtime path.
Keep only stable backoff unit coverage in startup_iam and rely on the embedded deferred bootstrap integration test for end-to-end recovery behavior.
* refactor(startup): trim IAM recovery test scaffolding
Keep the concrete deferred IAM recovery path intact while removing bulky test-only async loop scaffolding from startup_iam.
Retain the stable backoff unit checks and rely on the embedded deferred bootstrap integration test for end-to-end recovery coverage.
* fix: apply code review improvements from PR #3188 review
- Simplify RecoveryFuture type alias by removing unnecessary lifetime
- Fix finalize_iam_recovery to return Err if app context unavailable
- Update bootstrap_or_defer_iam_init doc comment to reflect Err case
- Use boundary-aware has_path_prefix for admin path matching in utils.rs
- Add test for adminx boundary rejection in utils.rs and layer.rs
- Improve embedded deferred IAM test with timeout wrapper
* style: merge has_path_prefix import into existing use block
* fix(iam): address final review follow-ups
- fix main startup readiness publication to pass ServiceStateManager correctly
- centralize IAM test env keys in rustfs_config and reuse them in runtime/tests
- keep deferred IAM bootstrap validation aligned with the final review fixes
* fix: isolate listing timeouts from drive health
Keep walk_dir scanner timeouts request-scoped instead of marking local drives faulty.
Add regression coverage for follow-up bucket info, set-level list_path, and system-prefix listings after prior walk timeouts.
* test(iam): gate deferred bootstrap test to debug
Align the deferred IAM embedded integration test with debug-only IAM fault injection hooks so release-profile runs do not assert deferred bootstrap behavior that cannot be triggered.
* test(ecstore): bound prior walk timeout regressions
- set walk_dir stall timeout explicitly in prior-timeout listing tests
- keep the system-prefix follow-up listing scoped to the same base dir
- assert the expected directory entry so the timeout regression test stays fast and stable
* fmt
* perf: reduce spawn_blocking contention in PUT path (~23% throughput gain)
Flame graph profiling identified tokio blocking pool mutex contention
as the #1 bottleneck (17.3% of CPU time). Each spawn_blocking call must
acquire parking_lot::raw_mutex to enqueue work. With 16 concurrent PUTs
× 4 disks × 3+ spawn_blocking per disk, this became a serialization point.
Optimizations applied:
- Merge make_dir_all + file write into single spawn_blocking
- Merge read_file + parse + write + rename into single spawn_blocking
for inline objects (small files)
- Optimize reliable_rename to try rename first, mkdir only on ENOENT
- Optimize remove/remove_std to try remove_file first, EISDIR fallback
- Add encode_inline_small fast path for small objects
- Parallelize bitrot writer creation with join_all
Benchmark (4KiB PUT, 4-disk EC, 16 concurrent, 8 rounds, randomized A/B):
Baseline: ~950 obj/s → Optimized: ~1173 obj/s (+23%)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix: address review comments
- Return old_data_dir for non-inline rename_data path (was incorrectly None)
- Restore delete_all cleanup of PUT temp data on failure paths
- Fix cargo fmt formatting
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* style: apply rustfmt from stable 1.96.0
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix: collapse nested if-let chains for clippy compliance
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix: address Copilot review comments
- fs.rs: handle macOS EPERM from remove_file on directories
- os.rs: restore NotFound=Ok(()) semantics on first rename attempt
- local.rs: use try-rename-then-mkdir pattern for inline rename_data
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* test: add unit tests for encode_inline_small fast path
* test: fix comment and line length in encode_inline_small tests
* fix: revert reliable_rename and write_all_internal to match original
Restore the original reliable_rename logic (check parent exists, then
rename in loop) and the original write_all_internal (make_dir_all outside
spawn_blocking). The optimization changes caused a CI-only test failure
in capacity_dirty_scope_test that could not be reproduced locally.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix: update comment and fix formatting for CI
- Fix macOS/BSD comment to accurately say macOS only
- Fix encode_inline_small test formatting to match rustfmt 1.96.0
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix: propagate inline rename errors
* fix: retry rename when parent missing
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Treat empty ping bodies as liveness probes, add a startup cleanup barrier for early walk_dir calls, and delay immediate background cleanup/capacity timers to reduce transient restart-time VolumeNotFound noise.
Also downgrade expected missing-path producer results during startup from generic errors to warnings while preserving existing storage semantics.
* perf(ecstore): use direct std writes for local disk
* fix(ecstore): avoid blocking disk writes in local disk path
* fix(ecstore): use tokio async write in local disk path
* chore(ecstore): remove unused AsyncWriteExt import
* fix(ecstore): preserve write semantics and avoid bytes copy
* fix(ecstore): optimize write_all_internal buffer paths
* fix(ecstore): sync write path in local disk
* fix(ecstore): align sync flag docs and avoid ref copy
* chore(rust): format local write path
---------
Co-authored-by: houseme <housemecn@gmail.com>
Move build_internode_data_transport_from_env() out of RemoteDisk::new()
into new_disk(), so RemoteDisk depends only on the
Arc<dyn InternodeDataTransport> trait object. TCP/HTTP remains the
default backend; all data paths unchanged.
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
* fix: keep scanner walk timeouts from offlining drives
Scanner walk operations can time out on large or slow directory listings without proving the backing drive is faulty. Keep the timeout error local to the scan while preserving failure marking for ordinary disk operations.
Constraint: Scanner walk_dir can include listing work that exceeds the drive timeout under slow storage.
Rejected: Disable timeout failure marking globally | real stuck disk operations must still affect drive health
Confidence: high
Scope-risk: narrow
Directive: Do not route scanner/listing timeout back into drive offline state without reproducing issue #2651
Tested: cargo test -p rustfs-ecstore timeout
Tested: cargo fmt --all --check
Tested: cargo clippy --workspace --all-features --all-targets -- -D warnings
Tested: make pre-commit with en_US.UTF-8 locale
Related: https://github.com/rustfs/rustfs/issues/2651
* fix: keep remote scanner walks from offlining drives
Remote walk_dir uses a streaming request that can hit total or stall timeouts during large scans without proving the remote drive is faulty. Route that scanner path through an explicit ignore action while preserving failure marking for ordinary remote operations.
Also treat Duration::ZERO as no operation timeout for remote health tracking, matching the existing local disk wrapper contract while still marking network-like errors on default paths.
Constraint: Scanner walk_dir streams can be slow because of directory size or consumer backpressure.
Rejected: Disable remote timeout failure marking globally | normal remote disk operations still need to evict failed connections and update drive health
Confidence: high
Scope-risk: narrow
Directive: Do not reintroduce remote scanner timeout failure marking without reproducing issue #2651 against remote disks.
Tested: cargo test -p rustfs-ecstore execute_with_timeout
Tested: cargo test -p rustfs-ecstore timeout
Tested: cargo fmt --all --check
Tested: cargo clippy --workspace --all-features --all-targets -- -D warnings
Tested: make pre-commit with en_US.UTF-8 locale
Related: https://github.com/rustfs/rustfs/issues/2651
* test: prove scanner walk backpressure keeps drives online
Add a local scanner walk regression test where the output writer never makes progress. The test confirms the walk timeout returns without marking the drive faulty, covering a stall source that is not caused by object count or network transfer speed.
Constraint: Scanner walk can block on downstream writer backpressure as well as disk or network IO.\nConfidence: high\nScope-risk: narrow\nTested: cargo test -p rustfs-ecstore walk_dir_writer_backpressure_timeout_does_not_mark_drive_failure\nTested: cargo test -p rustfs-ecstore timeout\nTested: cargo fmt --all --check\nTested: cargo clippy --workspace --all-features --all-targets -- -D warnings\nTested: make pre-commit
---------
Co-authored-by: houseme <housemecn@gmail.com>