* fix(iam): prevent transient IAM walk timeout from crashing startup
IAM startup performs a blocking full metadata walk on `.rustfs.sys/config/iam/`.
When that distributed walk times out (e.g. disk pressure after cluster reboot),
the old code treated the failure as fatal and exited the process, causing a
systemd restart loop.
Changes:
- Add `startup_iam.rs`: attempt IAM init, enter degraded mode on failure,
spawn background retry task with exponential backoff (5s→10s→20s→30s cap)
- Log level escalates to ERROR after 12 retries (~5 min) to aid diagnosis
- `/health/ready` returns 503 until IAM recovers; IAM-dependent ops return
`IamSysNotInitialized` (existing fail-closed behavior preserved)
- Fix admin path boundary matching: `/minio/administrator` no longer falsely
matches as admin prefix
- Normalize Content-Length: 0 for admin GET requests with empty body
Fixes#3175
* fix(iam): move constant assertion into const block
Fixes clippy::assertions-on-constants warning on
IAM_RETRY_ESCALATION_THRESHOLD assertion.
* fix(iam): address PR review comments
- Replace OnceLock with AtomicU64 sentinel for test isolation;
add reset_test_failure_counter() for integration tests
- Use u32::try_from() instead of `as u32` narrowing cast in
compute_backoff_interval
- Rename misleading test; update to verify finalize retry behavior
- Restructure spawn_iam_recovery_task into init-retry and
finalize-retry phases so transient readiness failures are retried
instead of leaving the server permanently degraded
* fix(iam): gate test hooks behind debug_assertions
- reset_test_failure_counter() now stores sentinel (u64::MAX) to
correctly trigger env var re-read on next call
- RUSTFS_TEST_IAM_FAIL_INIT_ATTEMPTS only honored in debug builds
- RUSTFS_TEST_IAM_RETRY_INTERVAL_MS only honored in debug builds
* test(iam): cover deferred bootstrap recovery
Add a dedicated embedded deferred-IAM integration test in a separate test binary to avoid process-global startup collisions.
Strengthen startup IAM recovery coverage with focused unit tests and keep the existing embedded smoke test isolated while carrying the manual test license header update in the same change set.
* fix(startup): tighten deferred IAM recovery path
Adopt follow-up review feedback by silencing misleading app-context warnings after IAM recovery, reusing boundary-aware path prefix checks in the readiness gate, and tying deferred IAM recovery retries to server shutdown tokens.
Keep the deferred IAM embedded integration coverage and startup recovery unit coverage green after the follow-up hardening.
* refactor(startup): simplify IAM recovery task
Collapse the deferred IAM recovery implementation back to a concrete production flow instead of keeping boxed callback seams in the runtime path.
Keep only stable backoff unit coverage in startup_iam and rely on the embedded deferred bootstrap integration test for end-to-end recovery behavior.
* refactor(startup): trim IAM recovery test scaffolding
Keep the concrete deferred IAM recovery path intact while removing bulky test-only async loop scaffolding from startup_iam.
Retain the stable backoff unit checks and rely on the embedded deferred bootstrap integration test for end-to-end recovery coverage.
* fix: apply code review improvements from PR #3188 review
- Simplify RecoveryFuture type alias by removing unnecessary lifetime
- Fix finalize_iam_recovery to return Err if app context unavailable
- Update bootstrap_or_defer_iam_init doc comment to reflect Err case
- Use boundary-aware has_path_prefix for admin path matching in utils.rs
- Add test for adminx boundary rejection in utils.rs and layer.rs
- Improve embedded deferred IAM test with timeout wrapper
* style: merge has_path_prefix import into existing use block
* fix(iam): address final review follow-ups
- fix main startup readiness publication to pass ServiceStateManager correctly
- centralize IAM test env keys in rustfs_config and reuse them in runtime/tests
- keep deferred IAM bootstrap validation aligned with the final review fixes
* fix: isolate listing timeouts from drive health
Keep walk_dir scanner timeouts request-scoped instead of marking local drives faulty.
Add regression coverage for follow-up bucket info, set-level list_path, and system-prefix listings after prior walk timeouts.
* test(iam): gate deferred bootstrap test to debug
Align the deferred IAM embedded integration test with debug-only IAM fault injection hooks so release-profile runs do not assert deferred bootstrap behavior that cannot be triggered.
* test(ecstore): bound prior walk timeout regressions
- set walk_dir stall timeout explicitly in prior-timeout listing tests
- keep the system-prefix follow-up listing scoped to the same base dir
- assert the expected directory entry so the timeout regression test stays fast and stable
* fmt
* fix(ecstore): retry transient walk dir stream errors
Retry RemoteDisk walk_dir once when the first failure looks transport-related so restart-time HttpReader stream errors do not immediately count as hard listing failures.
Improve HttpReader request and stream error messages with method and URL context, and add regression coverage for retry recovery and diagnostics.
* fix(ecstore): avoid retry after partial walk dir stream
Stop retrying walk_dir after copy_stream_with_buffer has already written partial bytes into the destination writer.
Keep the retry only for open_walk_dir failures and add a regression test that proves partial stream failures are returned without issuing a second stream.
* feat(admin): restore config admin compatibility
Co-authored-by: weisd <im@weisd.in>
* fix(admin): align config admin clean rebuild
Co-authored-by: weisd <im@weisd.in>
* fix(admin): align config history and peer signals
* fix(admin): harden config admin mutations
* fix(admin): tighten config review follow-ups
* perf(admin): reuse env snapshot in config render
* fix(ecstore): clean up config admin and listing error handling
Remove redundant is_all_volume_not_found check in list_merged, add
storage class encode/decode roundtrip tests, fresh boot integration
test, and config admin clean rebuild improvements.
Co-authored-by: hehutu <heihutu@gmail.com>
* fix(admin): sync global server config on mutation and reload
Change GLOBAL_SERVER_CONFIG from OnceLock to RwLock so config mutations
(set/del/restore/reload) are visible to readers without restart. Call
set_global_server_config after every store save and on snapshot reload.
Register storage_class as a dynamic config subsystem.
Co-authored-by: hehutu <heihutu@gmail.com>
* style: apply rustfmt to config and admin tests
Co-authored-by: hehutu <heihutu@gmail.com>
* fix(test): update signal_service test for storage_class dynamic subsystem
storage_class is now a valid dynamic config subsystem, so the
"requires object layer" test should expect "storage layer not initialized"
instead of "unsupported dynamic config subsystem".
Co-authored-by: hehutu <heihutu@gmail.com>
* fix(config): publish storage_class runtime config on dynamic reload
Change GLOBAL_STORAGE_CLASS from OnceLock to RwLock so runtime updates
are possible. apply_storage_class_runtime_config now actually publishes
the parsed config via set_global_storage_class instead of dropping it.
Addresses review feedback: storage_class was marked as dynamically
applied but the parsed result was discarded, so mc admin config set
returned config_applied=true while the runtime kept using stale parity
settings until restart.
Co-authored-by: hehutu <heihutu@gmail.com>
* fix(admin): harden config init, history ordering, and env redaction
- Change GLOBAL_SERVER_CONFIG from RwLock<Config> to RwLock<Option<Config>>
initialized with None, preserving "not initialized" detection via None
- Move save_server_config_history before save_server_config_to_store in
SetConfigKVHandler, DelConfigKVHandler, and SetConfigHandler so a
restore point exists before mutations are persisted
- Redact sensitive env override values with *redacted* instead of
silently omitting the line, improving admin visibility
- Add code comment explaining VolumeNotFound removal rationale in
list_merged for listing paths
Co-authored-by: hehutu <heihutu@gmail.com>
* fix(config): keep in-memory config in sync after set/restore/reload
GLOBAL_SERVER_CONFIG was a OnceLock set once at startup and never
updated. After mc admin config set writes to the store, any fallback
to get_global_server_config() returned stale init-time data. Similarly,
reload_runtime_config_snapshot read from the store but discarded the
result.
- Replace OnceLock with RwLock for GLOBAL_SERVER_CONFIG and
GLOBAL_STORAGE_CLASS so they can be updated at runtime
- Add set_global_server_config / set_global_storage_class setters
- Call set_global_server_config after every config save (set-kv,
del-kv, set-config, restore-history)
- Re-apply dynamic subsystems (storage_class, audit_webhook,
audit_mqtt) and signal peers in reload_runtime_config_snapshot
and full-config operations
- Fix render_selected_config scope boundary check: track per-scope
line count instead of checking global lines.is_empty()
- Include STORAGE_CLASS_SUB_SYS in is_dynamic_config_subsystem so
apply_storage_class_runtime_config is reachable
Co-authored-by: hehutu <heihutu@gmail.com>
* fix(storageclass): use CLASS_RRS key in lookup_config for RRS parity
lookup_config used kvs.get(RRS) where RRS="REDUCED_REDUNDANCY", but the
admin config path writes the key as CLASS_RRS="rrs". This caused RRS
values to never be read back, always falling back to default parity.
- Changed kvs.get(RRS) to kvs.get(CLASS_RRS) in lookup_config
- Added regression tests verifying RRS read/write consistency
Co-authored-by: hehutu <heihutu@gmail.com>
* fix(config): add peer-side logging and don't swallow apply errors
- Add tracing::warn! in reload_dynamic_config_runtime_state and
reload_runtime_config_snapshot when config read or subsystem apply
fails, so on-host diagnostics show which signal failed and why
- Change `let _ = apply_dynamic_config_for_subsystem(...)` to
`if let Err(err) = ... { warn!(...) }` in reload_runtime_config_snapshot
so per-subsystem failures are logged instead of silently swallowed
- Remove weak test global_server_config_returns_none_before_init that
had no meaningful assertion due to shared global state
Co-authored-by: hehutu <heihutu@gmail.com>
* style: apply rustfmt to config and storageclass tests
Co-authored-by: hehutu <heihutu@gmail.com>
---------
Co-authored-by: weisd <im@weisd.in>
Co-authored-by: hehutu <heihutu@gmail.com>
* fix(ecstore): send valid ping body in remote locker
Build ping requests with a flatbuffer payload so health checks remain compatible with the ping response parser after restart.
* fix(bench): use multi-host warp target during failover
Normalize comma-separated warp host lists in run_object_batch_bench and let four-node failover bench pass BENCH_WARP_HOSTS so rolling restart does not pin load to a single restarting node.
* feat(health): add compat health probes with busy/KMS checks
- Add /health/live liveness probe endpoint
- Add busy protection (429) for readiness probes, gated by RUSTFS_HEALTH_COMPAT_BUSY_CHECK_ENABLE
- Add KMS readiness check for /health/ready, gated by RUSTFS_HEALTH_COMPAT_KMS_READY_CHECK_ENABLE
- Add lock quorum status caching with TTL to reduce RPC pressure
- Consolidate health response building into build_health_response_parts
- Register /health/live in console router and readiness gate
- Remove MinIO references from newly added health code
* fix(health): decouple kms readiness from lock quorum
* fix(ecstore): harden issue3031 multipart validation path
- clear stale multipart part destinations before rename fan-out
- add repeated part overwrite regression coverage
- reduce remote disk startup false-fault escalation to suspect-first
- refine remote locker diagnostics and lower scanner leader-lock log noise
- add a dedicated 4-node issue3031 docker validation script
* refactor(admin): inline console version json macro
- drop the unused serde_json::json import in admin console
- call serde_json::json! inline in version_handler
- keep the console version response behavior unchanged
* fix(remote-disk): recover suspect health on probe success
- record probe success during remote disk health checks so suspect drives recover
- use async_with_vars for the remote disk health probe test
- make the missing-listener test assert the state transition more robustly
* refactor(config): centralize internode transport constants
* fix(bench): guard all ripgrep calls behind dry-run check
Move require_cmd rg and metrics collection inside the non-dry-run
path so that --dry-run works on hosts without rg installed.
* feat(tooling): cross-platform protoc setup for Linux and macOS
Make install-protoc.sh support Linux (x86_64, aarch64) alongside
macOS, and bump CI protoc from 29.3 to 33.1 to match the version
required by the gproto build script.
* fix(bench): record internode baseline error counts
* fix(skill): correct YAML frontmatter formatting for release-version-bump
* chore(ci): bump protoc version to 34.1
* fix(tooling): bump protoc 33.1 to 34.1 in install script, restore SKILL.md description
---------
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
Move build_internode_data_transport_from_env() out of RemoteDisk::new()
into new_disk(), so RemoteDisk depends only on the
Arc<dyn InternodeDataTransport> trait object. TCP/HTTP remains the
default backend; all data paths unchanged.
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
* feat(internode): p0 transport baseline and ci hardening
* fix(internode): avoid double wrapping transport errors
---------
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
Co-authored-by: houseme <housemecn@gmail.com>
* docs: add internode data transport RFC
* feat: add internode operation metrics
* fix feedback
* fix(ci): fallback protoc token to github.token
* feat(ecstore): add internode transport boundary and baseline runner
* feat(internode): harden data transport baseline
* Revert "feat(internode): harden data transport baseline"
This reverts commit 5b8d6b8aa4.
* fix(internode): address baseline review comments
* fix(ci): pin setup-protoc to stable release
* fix(ci): install protoc via apt on linux
* fix(ci): restore protoc install for macos and windows
---------
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
Co-authored-by: houseme <housemecn@gmail.com>
* fix: keep scanner walk timeouts from offlining drives
Scanner walk operations can time out on large or slow directory listings without proving the backing drive is faulty. Keep the timeout error local to the scan while preserving failure marking for ordinary disk operations.
Constraint: Scanner walk_dir can include listing work that exceeds the drive timeout under slow storage.
Rejected: Disable timeout failure marking globally | real stuck disk operations must still affect drive health
Confidence: high
Scope-risk: narrow
Directive: Do not route scanner/listing timeout back into drive offline state without reproducing issue #2651
Tested: cargo test -p rustfs-ecstore timeout
Tested: cargo fmt --all --check
Tested: cargo clippy --workspace --all-features --all-targets -- -D warnings
Tested: make pre-commit with en_US.UTF-8 locale
Related: https://github.com/rustfs/rustfs/issues/2651
* fix: keep remote scanner walks from offlining drives
Remote walk_dir uses a streaming request that can hit total or stall timeouts during large scans without proving the remote drive is faulty. Route that scanner path through an explicit ignore action while preserving failure marking for ordinary remote operations.
Also treat Duration::ZERO as no operation timeout for remote health tracking, matching the existing local disk wrapper contract while still marking network-like errors on default paths.
Constraint: Scanner walk_dir streams can be slow because of directory size or consumer backpressure.
Rejected: Disable remote timeout failure marking globally | normal remote disk operations still need to evict failed connections and update drive health
Confidence: high
Scope-risk: narrow
Directive: Do not reintroduce remote scanner timeout failure marking without reproducing issue #2651 against remote disks.
Tested: cargo test -p rustfs-ecstore execute_with_timeout
Tested: cargo test -p rustfs-ecstore timeout
Tested: cargo fmt --all --check
Tested: cargo clippy --workspace --all-features --all-targets -- -D warnings
Tested: make pre-commit with en_US.UTF-8 locale
Related: https://github.com/rustfs/rustfs/issues/2651
* test: prove scanner walk backpressure keeps drives online
Add a local scanner walk regression test where the output writer never makes progress. The test confirms the walk timeout returns without marking the drive faulty, covering a stall source that is not caused by object count or network transfer speed.
Constraint: Scanner walk can block on downstream writer backpressure as well as disk or network IO.\nConfidence: high\nScope-risk: narrow\nTested: cargo test -p rustfs-ecstore walk_dir_writer_backpressure_timeout_does_not_mark_drive_failure\nTested: cargo test -p rustfs-ecstore timeout\nTested: cargo fmt --all --check\nTested: cargo clippy --workspace --all-features --all-targets -- -D warnings\nTested: make pre-commit
---------
Co-authored-by: houseme <housemecn@gmail.com>