* fix(iam): load IAM bootstrap snapshot without namespace locks
IAM bootstrap (init_iam_sys -> load_all) read every config object with
default ObjectOptions (no_lock=false), so each read acquired a
distributed namespace read lock. Lock quorum is counted over cluster
nodes and unreachable peers are hard failures, so during a sequential
restart the very first read failed with
"Quorum not reached: required 2, achieved 0" and IAM could not come up
until enough peers' lock RPC surfaces converged - even when the storage
read quorum was already satisfiable (rustfs#4304).
Extend the startup contract from rustfs#4056 ("startup metadata I/O must
not require namespace locks") to the IAM bootstrap path:
- Introduce LoadMode {Locked, BootstrapNoLock} and plumb it through the
load_all chain (groups, users, policies, mapped policies and their
concurrent variants) down to the storage read options.
- load_all now performs all reads with no_lock=true; on-line
single-object loads via the Store trait keep locked semantics.
- Fail-closed behavior is unchanged: any loader error still aborts the
whole snapshot load.
Safety: config objects are atomic whole-object writes, so a lock-free
read only observes an old or a new value; staleness is bounded by the
existing periodic IAM reload. Listing (walk) never took namespace locks,
and maybe_schedule_lazy_rewrite stays a best-effort background task.
Verification:
- cargo test -p rustfs-iam --lib (153 passed, incl. new LoadMode tests)
- cargo clippy -p rustfs-iam --all-targets
- cargo check -p rustfs
- make pre-commit
Ref: rustfs#4304; tracking rustfs/backlog#884, rustfs/backlog#885
Co-Authored-By: heihutu <heihutu@gmail.com>
* test(iam): sequential-restart regression test for lock-free bootstrap
Add an integration test reproducing the rustfs#4304 failure mode against
a real 4-disk temp-dir ECStore:
- Seed IAM group data in single-node mode, then flip the runtime into
distributed-erasure mode. new_ns_lock now builds a distributed lock
over the set's (empty) lock-client list, so every namespace-locked
read fails exactly like a sequential restart with unreachable peers
(lock quorum unavailable, storage read quorum healthy).
- Assert the locked load_group path fails in that state, while the
bulk snapshot load_all (no_lock plumbing from the previous commit)
succeeds, and the data survives intact once single-node mode is
restored.
Reverse-verified: temporarily switching load_all back to the locked
mode makes the test fail, so it genuinely guards the contract.
Verification:
- cargo test -p rustfs-iam --test iam_bootstrap_no_lock_test
- cargo test -p rustfs-iam --lib
Ref: rustfs#4304; tracking rustfs/backlog#886
Co-Authored-By: heihutu <heihutu@gmail.com>
* test(iam): route test ECStore imports through ecstore_test_compat boundary
The new integration test imported rustfs_ecstore facade paths directly,
tripping three architecture migration rules. Move every ECStore import
behind crates/iam/tests/ecstore_test_compat/mod.rs (the sanctioned
test-compat pattern), and register that module as a reviewed test-only
global-facade boundary in check_architecture_migration_rules.sh: the
sequential-restart regression test needs api::global::update_erasure_type
to flip into distributed-erasure mode for lock-quorum fault injection.
Verification:
- ./scripts/check_architecture_migration_rules.sh
- cargo test -p rustfs-iam --test iam_bootstrap_no_lock_test
- make pre-commit
Co-Authored-By: heihutu <heihutu@gmail.com>
* feat(server): expose readiness blocking reason + rolling-restart runbook
Operators hitting the rustfs#4304 sequential cold start could not tell
from the outside why a node stayed unavailable. Three additions:
- The readiness gate's 503 now names the blocking dependency in both the
body ("Service not ready: waiting for storage_quorum") and a new
x-rustfs-readiness-pending header (storage_quorum | iam |
startup_finalization), derived from the current startup stage.
/health/ready already returned details + degradedReasons; this covers
the plain S3 requests that hit the gate.
- IAM bootstrap retry logs now carry an actionable `hint` field that
classifies the failure (storage read quorum vs lock quorum vs
uninitialized metadata) instead of only echoing the storage error.
- New docs/operations/rolling-restart.md runbook: correct rolling
restart procedure, sequential cold-start expectations (degraded ->
auto-recovery), readiness signal reference, and
RUSTFS_STARTUP_READINESS_MAX_WAIT_SECS guidance.
Verification:
- cargo test -p rustfs --lib -- hint_tests service_not_ready readiness_pending
- make pre-commit
Ref: rustfs#4304; tracking rustfs/backlog#887
Co-Authored-By: heihutu <heihutu@gmail.com>
* upgrade deps version and improve import
* feat(iam): notification-path cache refreshes read without namespace locks (#4368)
P3 step 1 of rustfs/backlog#884 (scoped down from full MinIO readConfig
alignment after review): cross-node notification handlers
(group/policy/policy-mapping/user) refresh the local IAM cache with
single-object reads that previously took distributed namespace read
locks. These refreshes are asynchronous, best-effort, and already
stale-tolerant (the periodic reload converges them), so a node-counted
lock quorum failure or lock RPC hiccup on a peer must not fail them —
the same rationale as the lock-free bootstrap load_all (rustfs#4304).
- Store trait: add load_user_no_lock / load_group_no_lock /
load_policy_doc_no_lock / load_mapped_policy_no_lock with defaults
forwarding to the locked variants, so existing implementations and
test mocks keep their behavior.
- ObjectStore overrides them via the existing LoadMode::BootstrapNoLock
plumbing. Deletions triggered by the handlers keep locked writes.
- manager.rs: the four *_notification_handler paths (8 call sites)
switch to the lock-free variants.
- Integration test: while the lock quorum is unavailable (DistErasure
with empty lockers), load_group_no_lock must succeed exactly where
the locked load_group fails.
Request-path loads (check_key, verify_temp_user_persistence) and admin
write-then-reload paths intentionally stay locked: load_user_identity
embeds expiry deletions, so those need the side-effect extraction
tracked in rustfs/backlog#884 before going lock-free.
Verification:
- cargo test -p rustfs-iam --lib (156 passed)
- cargo test -p rustfs-iam --test iam_bootstrap_no_lock_test
- make pre-commit
Ref: rustfs/backlog#884, rustfs#4304
Co-authored-by: heihutu <heihutu@gmail.com>
* fix(server): drop unused iam_bootstrap_failure_hint import in tests
The hint tests live in their own hint_tests module with a local import;
the stale re-import in mod tests failed clippy's -D warnings on the
Test and Lint CI variants.
Verification:
- cargo clippy -p rustfs --all-targets
Co-Authored-By: heihutu <heihutu@gmail.com>
* fix
---------
Co-authored-by: heihutu <heihutu@gmail.com>
5.0 KiB
Restarting a multi-node RustFS cluster
How to restart nodes of an erasure-coded multi-node cluster without losing availability, what to expect when several nodes are down at once (sequential cold start), and how to read the degraded-mode signals. Written for the failure pattern reported in rustfs/rustfs#4304.
Upgrading the binary or container image never changes the on-disk data format. Replacing the executable and restarting is safe; no migration step runs on startup.
TL;DR
- Rolling restart (no downtime): restart one node at a time, and wait
for the restarted node to report
200on/health/readybefore touching the next one. The remaining nodes keep serving traffic. - Sequential cold start (several nodes down): nodes started before the
cluster has quorum come up in degraded mode — the process stays alive,
answers
503with the blocking reason, and recovers automatically as soon as enough peers are online. Do not restart-loop them; just keep starting the remaining nodes.
Why a single node cannot serve alone
Erasure coding shards every object (including internal metadata such as IAM
users, groups, and policies under .rustfs.sys) across the drives of a set.
Reading an object back needs a read quorum of shards online. With the
drives of one set spread over several nodes, one node alone can never satisfy
the read quorum — this is a mathematical property of erasure coding, not a
bug. The cluster becomes readable once enough nodes are up (for internal
configuration objects, which are written with maximum parity, that is
typically about half the nodes of a set).
Distributed locking similarly needs a majority of nodes' lock RPC endpoints. The startup path no longer takes namespace locks while loading IAM (rustfs/rustfs#4363), so IAM recovery depends only on the storage read quorum.
Rolling restart procedure
For each node, in any order, one at a time:
-
Restart the node (upgrade the binary/image first if this is an upgrade).
-
Wait until the node reports ready:
curl -fsS http://<node>:9000/health/readyA ready node returns
200with"ready": truein the JSON body. -
Only then move on to the next node.
While one node is down, the rest of the cluster keeps quorum and serves all traffic. If you take a second node down before the first is back, some erasure sets may lose write or even read quorum and requests start failing — this is the situation to avoid.
Sequential cold start (multiple nodes down)
When the whole cluster (or several nodes) went down — power loss, host maintenance, crash-looping deployment — and nodes are brought back one at a time:
-
Early nodes come up degraded. The process does not exit. S3 requests receive
503 Service Unavailablewith aRetry-After: 5header, anx-rustfs-readiness-pendingheader, and a body naming the blocking dependency:storage_quorum— waiting for enough nodes/disks for the erasure read quorum;iam— storage is up, IAM cache is still loading;startup_finalization— last startup steps are being published.
-
Logs say what the node waits for. The IAM recovery loop retries with backoff and logs
event="iam_bootstrap_retry_failed"with an actionablehintfield (for example, "storage read quorum not met yet; waiting for enough cluster nodes/disks to come online"). After repeated failures the log level escalates from WARN to ERROR — this still does not kill the process. -
Recovery is automatic. As soon as enough peers are online for the storage read quorum, the pending nodes finish IAM bootstrap on the next retry and flip
/health/readyto200on their own. No manual restart is needed, and restarting them does not speed anything up. -
Check readiness detail while waiting.
/health/ready(and/minio/health/ready) return per-dependency detail during degradation:curl -s http://<node>:9000/health/ready | jqThe
detailsobject showsstorage/iam/lockreadiness, anddegradedReasonslists machine-readable causes such asstorage_quorum_unavailableorlock_quorum_unavailable.
Tuning
RUSTFS_STARTUP_READINESS_MAX_WAIT_SECS(default120): how long startup waits for full readiness before continuing in degraded mode with background recovery. Raising it delays the listener during genuinely slow starts; lowering it surfaces degraded mode sooner. Recovery retries continue regardless of this limit.
What is not normal
- A node process exiting with a fatal IAM/lock error during startup — that fatal path was removed after v1.0.0-beta.5 (rustfs/rustfs#4304); upgrade if you still see it.
- A node stuck degraded after the whole cluster is back: check network
reachability between nodes (peer RPC ports) and per-node clocks, then
inspect
degradedReasonsand thehintfield of the IAM retry logs. - A node shown offline in the console with no log output — tracked separately, see rustfs/backlog#888.