mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-04 20:37:43 +00:00
7001373316
* fix(iam): load IAM bootstrap snapshot without namespace locks
IAM bootstrap (init_iam_sys -> load_all) read every config object with
default ObjectOptions (no_lock=false), so each read acquired a
distributed namespace read lock. Lock quorum is counted over cluster
nodes and unreachable peers are hard failures, so during a sequential
restart the very first read failed with
"Quorum not reached: required 2, achieved 0" and IAM could not come up
until enough peers' lock RPC surfaces converged - even when the storage
read quorum was already satisfiable (rustfs#4304).
Extend the startup contract from rustfs#4056 ("startup metadata I/O must
not require namespace locks") to the IAM bootstrap path:
- Introduce LoadMode {Locked, BootstrapNoLock} and plumb it through the
load_all chain (groups, users, policies, mapped policies and their
concurrent variants) down to the storage read options.
- load_all now performs all reads with no_lock=true; on-line
single-object loads via the Store trait keep locked semantics.
- Fail-closed behavior is unchanged: any loader error still aborts the
whole snapshot load.
Safety: config objects are atomic whole-object writes, so a lock-free
read only observes an old or a new value; staleness is bounded by the
existing periodic IAM reload. Listing (walk) never took namespace locks,
and maybe_schedule_lazy_rewrite stays a best-effort background task.
Verification:
- cargo test -p rustfs-iam --lib (153 passed, incl. new LoadMode tests)
- cargo clippy -p rustfs-iam --all-targets
- cargo check -p rustfs
- make pre-commit
Ref: rustfs#4304; tracking rustfs/backlog#884, rustfs/backlog#885
Co-Authored-By: heihutu <heihutu@gmail.com>
* test(iam): sequential-restart regression test for lock-free bootstrap
Add an integration test reproducing the rustfs#4304 failure mode against
a real 4-disk temp-dir ECStore:
- Seed IAM group data in single-node mode, then flip the runtime into
distributed-erasure mode. new_ns_lock now builds a distributed lock
over the set's (empty) lock-client list, so every namespace-locked
read fails exactly like a sequential restart with unreachable peers
(lock quorum unavailable, storage read quorum healthy).
- Assert the locked load_group path fails in that state, while the
bulk snapshot load_all (no_lock plumbing from the previous commit)
succeeds, and the data survives intact once single-node mode is
restored.
Reverse-verified: temporarily switching load_all back to the locked
mode makes the test fail, so it genuinely guards the contract.
Verification:
- cargo test -p rustfs-iam --test iam_bootstrap_no_lock_test
- cargo test -p rustfs-iam --lib
Ref: rustfs#4304; tracking rustfs/backlog#886
Co-Authored-By: heihutu <heihutu@gmail.com>
* test(iam): route test ECStore imports through ecstore_test_compat boundary
The new integration test imported rustfs_ecstore facade paths directly,
tripping three architecture migration rules. Move every ECStore import
behind crates/iam/tests/ecstore_test_compat/mod.rs (the sanctioned
test-compat pattern), and register that module as a reviewed test-only
global-facade boundary in check_architecture_migration_rules.sh: the
sequential-restart regression test needs api::global::update_erasure_type
to flip into distributed-erasure mode for lock-quorum fault injection.
Verification:
- ./scripts/check_architecture_migration_rules.sh
- cargo test -p rustfs-iam --test iam_bootstrap_no_lock_test
- make pre-commit
Co-Authored-By: heihutu <heihutu@gmail.com>
* feat(server): expose readiness blocking reason + rolling-restart runbook
Operators hitting the rustfs#4304 sequential cold start could not tell
from the outside why a node stayed unavailable. Three additions:
- The readiness gate's 503 now names the blocking dependency in both the
body ("Service not ready: waiting for storage_quorum") and a new
x-rustfs-readiness-pending header (storage_quorum | iam |
startup_finalization), derived from the current startup stage.
/health/ready already returned details + degradedReasons; this covers
the plain S3 requests that hit the gate.
- IAM bootstrap retry logs now carry an actionable `hint` field that
classifies the failure (storage read quorum vs lock quorum vs
uninitialized metadata) instead of only echoing the storage error.
- New docs/operations/rolling-restart.md runbook: correct rolling
restart procedure, sequential cold-start expectations (degraded ->
auto-recovery), readiness signal reference, and
RUSTFS_STARTUP_READINESS_MAX_WAIT_SECS guidance.
Verification:
- cargo test -p rustfs --lib -- hint_tests service_not_ready readiness_pending
- make pre-commit
Ref: rustfs#4304; tracking rustfs/backlog#887
Co-Authored-By: heihutu <heihutu@gmail.com>
* upgrade deps version and improve import
* feat(iam): notification-path cache refreshes read without namespace locks (#4368)
P3 step 1 of rustfs/backlog#884 (scoped down from full MinIO readConfig
alignment after review): cross-node notification handlers
(group/policy/policy-mapping/user) refresh the local IAM cache with
single-object reads that previously took distributed namespace read
locks. These refreshes are asynchronous, best-effort, and already
stale-tolerant (the periodic reload converges them), so a node-counted
lock quorum failure or lock RPC hiccup on a peer must not fail them —
the same rationale as the lock-free bootstrap load_all (rustfs#4304).
- Store trait: add load_user_no_lock / load_group_no_lock /
load_policy_doc_no_lock / load_mapped_policy_no_lock with defaults
forwarding to the locked variants, so existing implementations and
test mocks keep their behavior.
- ObjectStore overrides them via the existing LoadMode::BootstrapNoLock
plumbing. Deletions triggered by the handlers keep locked writes.
- manager.rs: the four *_notification_handler paths (8 call sites)
switch to the lock-free variants.
- Integration test: while the lock quorum is unavailable (DistErasure
with empty lockers), load_group_no_lock must succeed exactly where
the locked load_group fails.
Request-path loads (check_key, verify_temp_user_persistence) and admin
write-then-reload paths intentionally stay locked: load_user_identity
embeds expiry deletions, so those need the side-effect extraction
tracked in rustfs/backlog#884 before going lock-free.
Verification:
- cargo test -p rustfs-iam --lib (156 passed)
- cargo test -p rustfs-iam --test iam_bootstrap_no_lock_test
- make pre-commit
Ref: rustfs/backlog#884, rustfs#4304
Co-authored-by: heihutu <heihutu@gmail.com>
* fix(server): drop unused iam_bootstrap_failure_hint import in tests
The hint tests live in their own hint_tests module with a local import;
the stale re-import in mod tests failed clippy's -D warnings on the
Test and Lint CI variants.
Verification:
- cargo clippy -p rustfs --all-targets
Co-Authored-By: heihutu <heihutu@gmail.com>
* fix
---------
Co-authored-by: heihutu <heihutu@gmail.com>
160 lines
6.9 KiB
Rust
160 lines
6.9 KiB
Rust
// Copyright 2024 RustFS Team
|
|
//
|
|
// Licensed under the Apache License, Version 2.0 (the "License");
|
|
// you may not use this file except in compliance with the License.
|
|
// You may obtain a copy of the License at
|
|
//
|
|
// http://www.apache.org/licenses/LICENSE-2.0
|
|
//
|
|
// Unless required by applicable law or agreed to in writing, software
|
|
// distributed under the License is distributed on an "AS IS" BASIS,
|
|
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
// See the License for the specific language governing permissions and
|
|
// limitations under the License.
|
|
|
|
//! Regression test for rustfs#4304: IAM bootstrap must not depend on the
|
|
//! distributed namespace-lock quorum.
|
|
//!
|
|
//! During a sequential cluster restart the peer lock RPC endpoints are
|
|
//! unreachable, so every namespace-locked read fails with
|
|
//! `QuorumNotReached` even though the storage read quorum is already
|
|
//! satisfiable. The bulk snapshot load (`Store::load_all`) therefore has to
|
|
//! read with `no_lock = true` (startup contract rustfs#4056).
|
|
//!
|
|
//! The test builds a real 4-disk ECStore over temp dirs, seeds IAM data in
|
|
//! single-node mode, then flips the runtime into distributed-erasure mode.
|
|
//! In that mode `SetDisks::new_ns_lock` builds a distributed lock over the
|
|
//! set's lock clients — which are empty for a locally-built store — so every
|
|
//! locked read fails exactly like the sequential-restart scenario, while
|
|
//! plain storage reads keep working. The old (locked) `load_group` path must
|
|
//! fail and the lock-free `load_all` path must succeed.
|
|
|
|
mod ecstore_test_compat;
|
|
|
|
use ecstore_test_compat::fixture::{
|
|
ECStore, Endpoint, EndpointServerPools, Endpoints, PoolEndpoints, SetupType, init_local_disks, update_erasure_type,
|
|
};
|
|
use rustfs_iam::cache::Cache;
|
|
use rustfs_iam::store::object::ObjectStore;
|
|
use rustfs_iam::store::{GroupInfo, Store};
|
|
use serial_test::serial;
|
|
use std::collections::HashMap;
|
|
use std::sync::Arc;
|
|
use tokio_util::sync::CancellationToken;
|
|
|
|
const TEST_GROUP: &str = "seq-restart-group";
|
|
const TEST_MEMBERS: [&str; 2] = ["alice", "bob"];
|
|
|
|
async fn build_local_ecstore(temp_dir: &std::path::Path) -> Arc<ECStore> {
|
|
let disk_paths: Vec<_> = (1..=4).map(|i| temp_dir.join(format!("disk{i}"))).collect();
|
|
for disk_path in &disk_paths {
|
|
tokio::fs::create_dir_all(disk_path).await.unwrap();
|
|
}
|
|
|
|
let mut endpoints = Vec::new();
|
|
for (i, disk_path) in disk_paths.iter().enumerate() {
|
|
let mut endpoint = Endpoint::try_from(disk_path.to_str().unwrap()).unwrap();
|
|
endpoint.set_pool_index(0);
|
|
endpoint.set_set_index(0);
|
|
endpoint.set_disk_index(i);
|
|
endpoints.push(endpoint);
|
|
}
|
|
|
|
let pool_endpoints = PoolEndpoints {
|
|
legacy: false,
|
|
set_count: 1,
|
|
drives_per_set: 4,
|
|
endpoints: Endpoints::from(endpoints),
|
|
cmd_line: "test".to_string(),
|
|
platform: format!("OS: {} | Arch: {}", std::env::consts::OS, std::env::consts::ARCH),
|
|
};
|
|
let endpoint_pools = EndpointServerPools::from(vec![pool_endpoints]);
|
|
|
|
init_local_disks(endpoint_pools.clone()).await.unwrap();
|
|
|
|
// Port 0 keeps this integration binary parallel-safe alongside other
|
|
// ECStore-backed tests.
|
|
let server_addr: std::net::SocketAddr = "127.0.0.1:0".parse().unwrap();
|
|
ECStore::new(server_addr, endpoint_pools, CancellationToken::new())
|
|
.await
|
|
.unwrap()
|
|
}
|
|
|
|
/// Restores single-node erasure mode even when an assertion panics, so a
|
|
/// failing run cannot poison later `#[serial]` tests in this process.
|
|
struct ErasureModeGuard;
|
|
|
|
impl Drop for ErasureModeGuard {
|
|
fn drop(&mut self) {
|
|
pollster::block_on(update_erasure_type(SetupType::Erasure));
|
|
}
|
|
}
|
|
|
|
#[tokio::test(flavor = "multi_thread")]
|
|
#[serial]
|
|
async fn load_all_bypasses_namespace_lock_quorum() {
|
|
// The lock acquire timeout is latched into a OnceLock on first use, so it
|
|
// must be shortened before the first locked operation of this process.
|
|
temp_env::async_with_vars([(rustfs_config::ENV_OBJECT_LOCK_ACQUIRE_TIMEOUT, Some("1"))], async {
|
|
let temp_dir = tempfile::TempDir::with_prefix("rustfs_iam_no_lock_test_").unwrap();
|
|
let ecstore = build_local_ecstore(temp_dir.path()).await;
|
|
let store = ObjectStore::new(ecstore);
|
|
|
|
// Seed IAM data while namespace locks still work (single-node mode).
|
|
store
|
|
.save_group_info(TEST_GROUP, GroupInfo::new(TEST_MEMBERS.iter().map(|m| m.to_string()).collect()))
|
|
.await
|
|
.expect("seeding group info in single-node mode must succeed");
|
|
|
|
let mut baseline = HashMap::new();
|
|
store
|
|
.load_group(TEST_GROUP, &mut baseline)
|
|
.await
|
|
.expect("locked load_group must succeed in single-node mode");
|
|
assert_eq!(baseline[TEST_GROUP].members, TEST_MEMBERS, "seeded group must round-trip");
|
|
|
|
// Flip into distributed-erasure mode: new_ns_lock now builds a
|
|
// distributed lock over the set's (empty) lock-client list, so every
|
|
// locked read fails — the sequential-restart failure mode of
|
|
// rustfs#4304 (lock quorum unavailable, storage quorum healthy).
|
|
update_erasure_type(SetupType::DistErasure).await;
|
|
let _mode_guard = ErasureModeGuard;
|
|
|
|
let mut locked_read = HashMap::new();
|
|
let locked_err = store
|
|
.load_group(TEST_GROUP, &mut locked_read)
|
|
.await
|
|
.expect_err("locked load_group must fail while the lock quorum is unavailable");
|
|
assert!(locked_read.is_empty(), "failed locked read must not populate results");
|
|
|
|
// The P0 fix: the bulk snapshot load reads with no_lock = true, so it
|
|
// must succeed in exactly the state where the locked path fails.
|
|
let cache = Cache::default();
|
|
store
|
|
.load_all(&cache)
|
|
.await
|
|
.unwrap_or_else(|err| panic!("load_all must bypass namespace locks (locked path failed with: {locked_err}): {err}"));
|
|
|
|
// P3 step 1: notification-path cache refreshes use the same lock-free
|
|
// reads, so a cross-node notification must also be able to refresh
|
|
// this group while the lock quorum is unavailable.
|
|
let mut notification_read = HashMap::new();
|
|
store
|
|
.load_group_no_lock(TEST_GROUP, &mut notification_read)
|
|
.await
|
|
.expect("lock-free notification-path load_group must succeed while the lock quorum is unavailable");
|
|
assert_eq!(notification_read[TEST_GROUP].members, TEST_MEMBERS);
|
|
|
|
// Back in single-node mode the locked path works again; verify the
|
|
// data survived the whole exercise intact (fail-closed integrity).
|
|
drop(_mode_guard);
|
|
let mut recovered = HashMap::new();
|
|
store
|
|
.load_group(TEST_GROUP, &mut recovered)
|
|
.await
|
|
.expect("locked load_group must succeed again after restoring single-node mode");
|
|
assert_eq!(recovered[TEST_GROUP].members, TEST_MEMBERS, "group data must be intact");
|
|
})
|
|
.await;
|
|
}
|