fix(iam): load IAM bootstrap snapshot without namespace locks (#4363)

* fix(iam): load IAM bootstrap snapshot without namespace locks

IAM bootstrap (init_iam_sys -> load_all) read every config object with
default ObjectOptions (no_lock=false), so each read acquired a
distributed namespace read lock. Lock quorum is counted over cluster
nodes and unreachable peers are hard failures, so during a sequential
restart the very first read failed with
"Quorum not reached: required 2, achieved 0" and IAM could not come up
until enough peers' lock RPC surfaces converged - even when the storage
read quorum was already satisfiable (rustfs#4304).

Extend the startup contract from rustfs#4056 ("startup metadata I/O must
not require namespace locks") to the IAM bootstrap path:

- Introduce LoadMode {Locked, BootstrapNoLock} and plumb it through the
  load_all chain (groups, users, policies, mapped policies and their
  concurrent variants) down to the storage read options.
- load_all now performs all reads with no_lock=true; on-line
  single-object loads via the Store trait keep locked semantics.
- Fail-closed behavior is unchanged: any loader error still aborts the
  whole snapshot load.

Safety: config objects are atomic whole-object writes, so a lock-free
read only observes an old or a new value; staleness is bounded by the
existing periodic IAM reload. Listing (walk) never took namespace locks,
and maybe_schedule_lazy_rewrite stays a best-effort background task.

Verification:
- cargo test -p rustfs-iam --lib (153 passed, incl. new LoadMode tests)
- cargo clippy -p rustfs-iam --all-targets
- cargo check -p rustfs
- make pre-commit

Ref: rustfs#4304; tracking rustfs/backlog#884, rustfs/backlog#885

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(iam): sequential-restart regression test for lock-free bootstrap

Add an integration test reproducing the rustfs#4304 failure mode against
a real 4-disk temp-dir ECStore:

- Seed IAM group data in single-node mode, then flip the runtime into
  distributed-erasure mode. new_ns_lock now builds a distributed lock
  over the set's (empty) lock-client list, so every namespace-locked
  read fails exactly like a sequential restart with unreachable peers
  (lock quorum unavailable, storage read quorum healthy).
- Assert the locked load_group path fails in that state, while the
  bulk snapshot load_all (no_lock plumbing from the previous commit)
  succeeds, and the data survives intact once single-node mode is
  restored.

Reverse-verified: temporarily switching load_all back to the locked
mode makes the test fail, so it genuinely guards the contract.

Verification:
- cargo test -p rustfs-iam --test iam_bootstrap_no_lock_test
- cargo test -p rustfs-iam --lib

Ref: rustfs#4304; tracking rustfs/backlog#886

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(iam): route test ECStore imports through ecstore_test_compat boundary

The new integration test imported rustfs_ecstore facade paths directly,
tripping three architecture migration rules. Move every ECStore import
behind crates/iam/tests/ecstore_test_compat/mod.rs (the sanctioned
test-compat pattern), and register that module as a reviewed test-only
global-facade boundary in check_architecture_migration_rules.sh: the
sequential-restart regression test needs api::global::update_erasure_type
to flip into distributed-erasure mode for lock-quorum fault injection.

Verification:
- ./scripts/check_architecture_migration_rules.sh
- cargo test -p rustfs-iam --test iam_bootstrap_no_lock_test
- make pre-commit

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(server): expose readiness blocking reason + rolling-restart runbook

Operators hitting the rustfs#4304 sequential cold start could not tell
from the outside why a node stayed unavailable. Three additions:

- The readiness gate's 503 now names the blocking dependency in both the
  body ("Service not ready: waiting for storage_quorum") and a new
  x-rustfs-readiness-pending header (storage_quorum | iam |
  startup_finalization), derived from the current startup stage.
  /health/ready already returned details + degradedReasons; this covers
  the plain S3 requests that hit the gate.
- IAM bootstrap retry logs now carry an actionable `hint` field that
  classifies the failure (storage read quorum vs lock quorum vs
  uninitialized metadata) instead of only echoing the storage error.
- New docs/operations/rolling-restart.md runbook: correct rolling
  restart procedure, sequential cold-start expectations (degraded ->
  auto-recovery), readiness signal reference, and
  RUSTFS_STARTUP_READINESS_MAX_WAIT_SECS guidance.

Verification:
- cargo test -p rustfs --lib -- hint_tests service_not_ready readiness_pending
- make pre-commit

Ref: rustfs#4304; tracking rustfs/backlog#887

Co-Authored-By: heihutu <heihutu@gmail.com>

* upgrade deps version and improve import

* feat(iam): notification-path cache refreshes read without namespace locks (#4368)

P3 step 1 of rustfs/backlog#884 (scoped down from full MinIO readConfig
alignment after review): cross-node notification handlers
(group/policy/policy-mapping/user) refresh the local IAM cache with
single-object reads that previously took distributed namespace read
locks. These refreshes are asynchronous, best-effort, and already
stale-tolerant (the periodic reload converges them), so a node-counted
lock quorum failure or lock RPC hiccup on a peer must not fail them —
the same rationale as the lock-free bootstrap load_all (rustfs#4304).

- Store trait: add load_user_no_lock / load_group_no_lock /
  load_policy_doc_no_lock / load_mapped_policy_no_lock with defaults
  forwarding to the locked variants, so existing implementations and
  test mocks keep their behavior.
- ObjectStore overrides them via the existing LoadMode::BootstrapNoLock
  plumbing. Deletions triggered by the handlers keep locked writes.
- manager.rs: the four *_notification_handler paths (8 call sites)
  switch to the lock-free variants.
- Integration test: while the lock quorum is unavailable (DistErasure
  with empty lockers), load_group_no_lock must succeed exactly where
  the locked load_group fails.

Request-path loads (check_key, verify_temp_user_persistence) and admin
write-then-reload paths intentionally stay locked: load_user_identity
embeds expiry deletions, so those need the side-effect extraction
tracked in rustfs/backlog#884 before going lock-free.

Verification:
- cargo test -p rustfs-iam --lib (156 passed)
- cargo test -p rustfs-iam --test iam_bootstrap_no_lock_test
- make pre-commit

Ref: rustfs/backlog#884, rustfs#4304

Co-authored-by: heihutu <heihutu@gmail.com>

* fix(server): drop unused iam_bootstrap_failure_hint import in tests

The hint tests live in their own hint_tests module with a local import;
the stale re-import in mod tests failed clippy's -D warnings on the
Test and Lint CI variants.

Verification:
- cargo clippy -p rustfs --all-targets

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix

---------

Co-authored-by: heihutu <heihutu@gmail.com>
This commit is contained in:
houseme
2026-07-07 20:07:09 +08:00
committed by GitHub
parent 2247823200
commit 7001373316
11 changed files with 735 additions and 155 deletions
+2 -1
View File
@@ -60,7 +60,8 @@ url = { workspace = true }
[dev-dependencies]
pollster.workspace = true
serial_test = { workspace = true }
temp-env = { workspace = true }
temp-env = { workspace = true, features = ["async_closure"] }
tempfile = { workspace = true }
[lib]
doctest = false
+12 -8
View File
@@ -1847,7 +1847,7 @@ where
pub async fn group_notification_handler(&self, group: &str) -> Result<()> {
let mut m = HashMap::new();
if let Err(err) = self.api.load_group(group, &mut m).await {
if let Err(err) = self.api.load_group_no_lock(group, &mut m).await {
if !is_err_no_such_group(&err) {
return Err(err);
}
@@ -1876,7 +1876,7 @@ where
pub async fn policy_notification_handler(&self, policy: &str) -> Result<()> {
let mut m = HashMap::new();
if let Err(err) = self.api.load_policy_doc(policy, &mut m).await {
if let Err(err) = self.api.load_policy_doc_no_lock(policy, &mut m).await {
if !is_err_no_such_policy(&err) {
return Err(err);
}
@@ -1927,7 +1927,7 @@ where
pub async fn policy_mapping_notification_handler(&self, name: &str, user_type: UserType, is_group: bool) -> Result<()> {
let mut m = HashMap::new();
if let Err(err) = self.api.load_mapped_policy(name, user_type, is_group, &mut m).await {
if let Err(err) = self.api.load_mapped_policy_no_lock(name, user_type, is_group, &mut m).await {
if !is_err_no_such_policy(&err) {
return Err(err);
}
@@ -1963,7 +1963,7 @@ where
}
pub async fn user_notification_handler(&self, name: &str, user_type: UserType) -> Result<()> {
let mut m = HashMap::new();
if let Err(err) = self.api.load_user(name, user_type, &mut m).await {
if let Err(err) = self.api.load_user_no_lock(name, user_type, &mut m).await {
if !is_err_no_such_user(&err) {
return Err(err);
}
@@ -2028,7 +2028,7 @@ where
let mut policies = HashMap::new();
if let Err(err) = self
.api
.load_mapped_policy(&parent_user, user_type, false, &mut policies)
.load_mapped_policy_no_lock(&parent_user, user_type, false, &mut policies)
.await
{
if !is_err_no_such_policy(&err) {
@@ -2040,7 +2040,11 @@ where
}
UserType::Reg => {
let mut policies = HashMap::new();
if let Err(err) = self.api.load_mapped_policy(name, user_type, false, &mut policies).await {
if let Err(err) = self
.api
.load_mapped_policy_no_lock(name, user_type, false, &mut policies)
.await
{
if !is_err_no_such_policy(&err) {
return Err(err);
}
@@ -2063,7 +2067,7 @@ where
let mut policies = HashMap::new();
if let Err(err) = self
.api
.load_mapped_policy(&parent_user, UserType::Reg, false, &mut policies)
.load_mapped_policy_no_lock(&parent_user, UserType::Reg, false, &mut policies)
.await
{
if !is_err_no_such_policy(&err) {
@@ -2076,7 +2080,7 @@ where
let mut policies = HashMap::new();
if let Err(err) = self
.api
.load_mapped_policy(&parent_user, UserType::Sts, false, &mut policies)
.load_mapped_policy_no_lock(&parent_user, UserType::Sts, false, &mut policies)
.await
{
if !is_err_no_such_policy(&err) {
+29
View File
@@ -71,6 +71,35 @@ pub trait Store: Clone + Send + Sync + 'static {
) -> Result<()>;
async fn load_all(&self, cache: &Cache) -> Result<()>;
// Lock-free variants used by the cross-node notification handlers.
//
// Notification-path cache refreshes are asynchronous, best-effort, and
// already tolerate stale values (the periodic reload converges them), so
// they must not depend on the node-counted namespace-lock quorum — the
// same rationale as the lock-free bootstrap `load_all` (rustfs#4304;
// MinIO's readConfig takes no distributed lock either). The defaults
// forward to the locked variants so existing `Store` implementations
// (including test mocks) keep their behavior; `ObjectStore` overrides
// them with lock-free reads.
async fn load_user_no_lock(&self, name: &str, user_type: UserType, m: &mut HashMap<String, UserIdentity>) -> Result<()> {
self.load_user(name, user_type, m).await
}
async fn load_group_no_lock(&self, name: &str, m: &mut HashMap<String, GroupInfo>) -> Result<()> {
self.load_group(name, m).await
}
async fn load_policy_doc_no_lock(&self, name: &str, m: &mut HashMap<String, PolicyDoc>) -> Result<()> {
self.load_policy_doc(name, m).await
}
async fn load_mapped_policy_no_lock(
&self,
name: &str,
user_type: UserType,
is_group: bool,
m: &mut HashMap<String, MappedPolicy>,
) -> Result<()> {
self.load_mapped_policy(name, user_type, is_group, m).await
}
}
#[derive(Debug, PartialEq, Eq, Clone, Copy)]
+272 -121
View File
@@ -95,6 +95,38 @@ fn get_mapped_policy_path(name: &str, user_type: UserType, is_group: bool) -> St
}
}
/// Lock semantics for IAM config reads.
///
/// Bootstrap/snapshot loads (`load_all`) must not require distributed
/// namespace locks: lock quorum is counted over cluster nodes, so during a
/// sequential restart the peers are unreachable and every locked read fails
/// with `QuorumNotReached` even though the storage read quorum is already
/// satisfiable (rustfs#4304). This mirrors the startup contract established
/// for pool metadata in rustfs#4056. Config objects are atomic whole-object
/// writes, so a lock-free read only ever observes an old or a new value;
/// staleness is bounded by the periodic IAM reload.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
enum LoadMode {
/// On-line load on behalf of a live request: takes namespace locks.
Locked,
/// Cache-refresh load (bootstrap/reload `load_all`, or a notification-path
/// single-object refresh): bypasses namespace locks.
BootstrapNoLock,
}
impl LoadMode {
fn no_lock(self) -> bool {
matches!(self, LoadMode::BootstrapNoLock)
}
fn read_opts(self) -> IamObjectOptions {
IamObjectOptions {
no_lock: self.no_lock(),
..Default::default()
}
}
}
#[derive(Debug)]
pub struct StringOrErr {
pub item: Option<String>,
@@ -359,9 +391,13 @@ impl ObjectStore {
Self::prepare_data_for_storage(data)
}
async fn load_iamconfig_bytes_with_metadata(&self, path: impl AsRef<str> + Send) -> Result<(Vec<u8>, IamObjectInfo)> {
async fn load_iamconfig_bytes_with_metadata(
&self,
path: impl AsRef<str> + Send,
mode: LoadMode,
) -> Result<(Vec<u8>, IamObjectInfo)> {
let path_ref = path.as_ref();
let (data, obj) = read_iam_config_with_metadata(self.object_api.clone(), path_ref, &IamObjectOptions::default()).await?;
let (data, obj) = read_iam_config_with_metadata(self.object_api.clone(), path_ref, &mode.read_opts()).await?;
let outcome = Self::decrypt_data_with_source(&data)?;
self.maybe_schedule_lazy_rewrite(path_ref, &outcome, &obj);
@@ -472,13 +508,13 @@ impl ObjectStore {
Ok(res)
}
async fn load_policy_doc_concurrent(&self, names: &[String]) -> Result<Vec<PolicyDoc>> {
async fn load_policy_doc_concurrent(&self, names: &[String], mode: LoadMode) -> Result<Vec<PolicyDoc>> {
let mut futures = Vec::with_capacity(names.len());
for name in names {
let policy_name = rustfs_utils::path::dir(name);
futures.push(async move {
match self.load_policy(&policy_name).await {
match self.load_policy_with(&policy_name, mode).await {
Ok(p) => Ok(p),
Err(err) => {
if !is_err_no_such_policy(&err) {
@@ -504,13 +540,13 @@ impl ObjectStore {
Ok(policies)
}
async fn load_user_concurrent(&self, names: &[String], user_type: UserType) -> Result<Vec<UserIdentity>> {
async fn load_user_concurrent(&self, names: &[String], user_type: UserType, mode: LoadMode) -> Result<Vec<UserIdentity>> {
let mut futures = Vec::with_capacity(names.len());
for name in names {
let user_name = rustfs_utils::path::dir(name);
futures.push(async move {
match self.load_user_identity(&user_name, user_type).await {
match self.load_user_identity_with(&user_name, user_type, mode).await {
Ok(res) => Ok(res),
Err(err) => {
if !is_err_no_such_user(&err) {
@@ -535,9 +571,155 @@ impl ObjectStore {
Ok(users)
}
async fn load_mapped_policy_internal(&self, name: &str, user_type: UserType, is_group: bool) -> Result<MappedPolicy> {
/// Parameterized core of [`Store::load_iam_config`]. `LoadMode::BootstrapNoLock`
/// keeps the bootstrap snapshot load free of namespace locks (rustfs#4304).
async fn load_iam_config_with<Item: DeserializeOwned>(&self, path: impl AsRef<str> + Send, mode: LoadMode) -> Result<Item> {
let path_ref = path.as_ref();
let (data, obj) = read_iam_config_with_metadata(self.object_api.clone(), path_ref, &mode.read_opts()).await?;
let outcome = match Self::decrypt_data_with_source(&data) {
Ok(v) => v,
Err(err) => {
warn!(path = %path_ref, error = %err, "IAM config decrypt failed; keeping file");
// keep the config file when decrypt failed - do not delete
return Err(Error::ConfigNotFound);
}
};
self.maybe_schedule_lazy_rewrite(path_ref, &outcome, &obj);
Ok(serde_json::from_slice(&outcome.plain)?)
}
/// Parameterized core of [`Store::load_user_identity`].
async fn load_user_identity_with(&self, name: &str, user_type: UserType, mode: LoadMode) -> Result<UserIdentity> {
let mut u: UserIdentity = self
.load_iam_config_with(get_user_identity_path(name, user_type), mode)
.await
.map_err(|err| {
if is_err_config_not_found(&err) {
warn!(name, user_type = ?user_type, "IAM user identity missing");
Error::NoSuchUser(name.to_owned())
} else {
warn!(name, user_type = ?user_type, error = ?err, "IAM user identity load failed");
err
}
})?;
if u.credentials.is_expired() {
let _ = self.delete_iam_config(get_user_identity_path(name, user_type)).await;
let _ = self.delete_iam_config(get_mapped_policy_path(name, user_type, false)).await;
warn!(name, user_type = ?user_type, "IAM user identity expired and was removed");
return Err(Error::NoSuchUser(name.to_owned()));
}
if u.credentials.access_key.is_empty() {
u.credentials.access_key = name.to_owned();
}
if !u.credentials.session_token.is_empty() {
let claims_result = if user_type == UserType::Svc && u.credentials.expiration.is_none() {
extract_jwt_claims_allow_missing_exp(&u)
} else {
extract_jwt_claims(&u)
};
match claims_result {
Ok(claims) => {
u.credentials.claims = Some(claims);
}
Err(err) => {
if u.credentials.is_temp() {
let _ = self.delete_iam_config(get_user_identity_path(name, user_type)).await;
let _ = self.delete_iam_config(get_mapped_policy_path(name, user_type, false)).await;
}
warn!(name, user_type = ?user_type, error = ?err, "IAM JWT claim extraction failed");
return Err(Error::NoSuchUser(name.to_owned()));
}
}
}
Ok(u)
}
/// Parameterized core of [`Store::load_user`].
async fn load_user_with(
&self,
name: &str,
user_type: UserType,
m: &mut HashMap<String, UserIdentity>,
mode: LoadMode,
) -> Result<()> {
self.load_user_identity_with(name, user_type, mode).await.map(|u| {
m.insert(name.to_owned(), u);
})
}
/// Parameterized core of [`Store::load_group`].
async fn load_group_with(&self, name: &str, m: &mut HashMap<String, GroupInfo>, mode: LoadMode) -> Result<()> {
let u: GroupInfo = self
.load_iam_config_with(get_group_info_path(name), mode)
.await
.map_err(|err| {
if is_err_config_not_found(&err) {
Error::NoSuchGroup(name.to_string())
} else {
err
}
})?;
m.insert(name.to_owned(), u);
Ok(())
}
/// Parameterized core of [`Store::load_policy`].
async fn load_policy_with(&self, name: &str, mode: LoadMode) -> Result<PolicyDoc> {
let (data, obj) = self
.load_iamconfig_bytes_with_metadata(get_policy_doc_path(name), mode)
.await
.map_err(|err| {
if is_err_config_not_found(&err) {
Error::NoSuchPolicy
} else {
err
}
})?;
let mut info = PolicyDoc::try_from(data)?;
if info.version == 0 {
info.create_date = obj.mod_time;
info.update_date = obj.mod_time;
}
Ok(info)
}
/// Parameterized core of [`Store::load_mapped_policy`].
async fn load_mapped_policy_with(
&self,
name: &str,
user_type: UserType,
is_group: bool,
m: &mut HashMap<String, MappedPolicy>,
mode: LoadMode,
) -> Result<()> {
let info = self.load_mapped_policy_internal(name, user_type, is_group, mode).await?;
m.insert(name.to_owned(), info);
Ok(())
}
async fn load_mapped_policy_internal(
&self,
name: &str,
user_type: UserType,
is_group: bool,
mode: LoadMode,
) -> Result<MappedPolicy> {
let info: MappedPolicy = self
.load_iam_config(get_mapped_policy_path(name, user_type, is_group))
.load_iam_config_with(get_mapped_policy_path(name, user_type, is_group), mode)
.await
.map_err(|err| {
if is_err_config_not_found(&err) {
@@ -555,13 +737,14 @@ impl ObjectStore {
names: &[String],
user_type: UserType,
is_group: bool,
mode: LoadMode,
) -> Result<Vec<MappedPolicy>> {
let mut futures = Vec::with_capacity(names.len());
for name in names {
let policy_name = name.trim_end_matches(".json");
futures.push(async move {
match self.load_mapped_policy_internal(policy_name, user_type, is_group).await {
match self.load_mapped_policy_internal(policy_name, user_type, is_group, mode).await {
Ok(p) => Ok(p),
Err(err) => {
if !is_err_no_such_policy(&err) {
@@ -615,21 +798,7 @@ impl Store for ObjectStore {
false
}
async fn load_iam_config<Item: DeserializeOwned>(&self, path: impl AsRef<str> + Send) -> Result<Item> {
let path_ref = path.as_ref();
let (data, obj) = read_iam_config_with_metadata(self.object_api.clone(), path_ref, &IamObjectOptions::default()).await?;
let outcome = match Self::decrypt_data_with_source(&data) {
Ok(v) => v,
Err(err) => {
warn!(path = %path_ref, error = %err, "IAM config decrypt failed; keeping file");
// keep the config file when decrypt failed - do not delete
return Err(Error::ConfigNotFound);
}
};
self.maybe_schedule_lazy_rewrite(path_ref, &outcome, &obj);
Ok(serde_json::from_slice(&outcome.plain)?)
self.load_iam_config_with(path, LoadMode::Locked).await
}
/// Saves IAM configuration with a retry mechanism on failure.
///
@@ -712,58 +881,10 @@ impl Store for ObjectStore {
Ok(())
}
async fn load_user_identity(&self, name: &str, user_type: UserType) -> Result<UserIdentity> {
let mut u: UserIdentity = self
.load_iam_config(get_user_identity_path(name, user_type))
.await
.map_err(|err| {
if is_err_config_not_found(&err) {
warn!(name, user_type = ?user_type, "IAM user identity missing");
Error::NoSuchUser(name.to_owned())
} else {
warn!(name, user_type = ?user_type, error = ?err, "IAM user identity load failed");
err
}
})?;
if u.credentials.is_expired() {
let _ = self.delete_iam_config(get_user_identity_path(name, user_type)).await;
let _ = self.delete_iam_config(get_mapped_policy_path(name, user_type, false)).await;
warn!(name, user_type = ?user_type, "IAM user identity expired and was removed");
return Err(Error::NoSuchUser(name.to_owned()));
}
if u.credentials.access_key.is_empty() {
u.credentials.access_key = name.to_owned();
}
if !u.credentials.session_token.is_empty() {
let claims_result = if user_type == UserType::Svc && u.credentials.expiration.is_none() {
extract_jwt_claims_allow_missing_exp(&u)
} else {
extract_jwt_claims(&u)
};
match claims_result {
Ok(claims) => {
u.credentials.claims = Some(claims);
}
Err(err) => {
if u.credentials.is_temp() {
let _ = self.delete_iam_config(get_user_identity_path(name, user_type)).await;
let _ = self.delete_iam_config(get_mapped_policy_path(name, user_type, false)).await;
}
warn!(name, user_type = ?user_type, error = ?err, "IAM JWT claim extraction failed");
return Err(Error::NoSuchUser(name.to_owned()));
}
}
}
Ok(u)
self.load_user_identity_with(name, user_type, LoadMode::Locked).await
}
async fn load_user(&self, name: &str, user_type: UserType, m: &mut HashMap<String, UserIdentity>) -> Result<()> {
self.load_user_identity(name, user_type).await.map(|u| {
m.insert(name.to_owned(), u);
})
self.load_user_with(name, user_type, m, LoadMode::Locked).await
}
async fn load_users(&self, user_type: UserType, m: &mut HashMap<String, UserIdentity>) -> Result<()> {
let base_prefix = match user_type {
@@ -823,16 +944,7 @@ impl Store for ObjectStore {
Ok(())
}
async fn load_group(&self, name: &str, m: &mut HashMap<String, GroupInfo>) -> Result<()> {
let u: GroupInfo = self.load_iam_config(get_group_info_path(name)).await.map_err(|err| {
if is_err_config_not_found(&err) {
Error::NoSuchGroup(name.to_string())
} else {
err
}
})?;
m.insert(name.to_owned(), u);
Ok(())
self.load_group_with(name, m, LoadMode::Locked).await
}
async fn load_groups(&self, m: &mut HashMap<String, GroupInfo>) -> Result<()> {
let ctx = CancellationToken::new();
@@ -875,25 +987,7 @@ impl Store for ObjectStore {
Ok(())
}
async fn load_policy(&self, name: &str) -> Result<PolicyDoc> {
let (data, obj) = self
.load_iamconfig_bytes_with_metadata(get_policy_doc_path(name))
.await
.map_err(|err| {
if is_err_config_not_found(&err) {
Error::NoSuchPolicy
} else {
err
}
})?;
let mut info = PolicyDoc::try_from(data)?;
if info.version == 0 {
info.create_date = obj.mod_time;
info.update_date = obj.mod_time;
}
Ok(info)
self.load_policy_with(name, LoadMode::Locked).await
}
async fn load_policy_doc(&self, name: &str, m: &mut HashMap<String, PolicyDoc>) -> Result<()> {
@@ -955,11 +1049,8 @@ impl Store for ObjectStore {
is_group: bool,
m: &mut HashMap<String, MappedPolicy>,
) -> Result<()> {
let info = self.load_mapped_policy_internal(name, user_type, is_group).await?;
m.insert(name.to_owned(), info);
Ok(())
self.load_mapped_policy_with(name, user_type, is_group, m, LoadMode::Locked)
.await
}
async fn load_mapped_policies(
&self,
@@ -1000,7 +1091,39 @@ impl Store for ObjectStore {
Ok(())
}
// Notification-path cache refreshes read lock-free (see the trait docs):
// they are best-effort and stale-tolerant, so they must not fail on the
// node-counted lock quorum. Deletions triggered by these refreshes keep
// their locked write semantics.
async fn load_user_no_lock(&self, name: &str, user_type: UserType, m: &mut HashMap<String, UserIdentity>) -> Result<()> {
self.load_user_with(name, user_type, m, LoadMode::BootstrapNoLock).await
}
async fn load_group_no_lock(&self, name: &str, m: &mut HashMap<String, GroupInfo>) -> Result<()> {
self.load_group_with(name, m, LoadMode::BootstrapNoLock).await
}
async fn load_policy_doc_no_lock(&self, name: &str, m: &mut HashMap<String, PolicyDoc>) -> Result<()> {
let info = self.load_policy_with(name, LoadMode::BootstrapNoLock).await?;
m.insert(name.to_owned(), info);
Ok(())
}
async fn load_mapped_policy_no_lock(
&self,
name: &str,
user_type: UserType,
is_group: bool,
m: &mut HashMap<String, MappedPolicy>,
) -> Result<()> {
self.load_mapped_policy_with(name, user_type, is_group, m, LoadMode::BootstrapNoLock)
.await
}
async fn load_all(&self, cache: &Cache) -> Result<()> {
// Bulk snapshot load: must stay free of namespace locks so IAM
// bootstrap only depends on the storage read quorum, never on the
// node-counted lock quorum (rustfs#4304; startup contract rustfs#4056).
const LOAD_ALL_MODE: LoadMode = LoadMode::BootstrapNoLock;
let cache_snapshot = cache.snapshot();
let listed_config_items = self.list_all_iamconfig_items().await?;
@@ -1011,7 +1134,7 @@ impl Store for ObjectStore {
loop {
if policies_list.len() < 32 {
let policy_docs = self.load_policy_doc_concurrent(&policies_list).await?;
let policy_docs = self.load_policy_doc_concurrent(&policies_list, LOAD_ALL_MODE).await?;
for (idx, p) in policy_docs.into_iter().enumerate() {
if p.policy.version.is_empty() {
@@ -1027,7 +1150,7 @@ impl Store for ObjectStore {
break;
}
let policy_docs = self.load_policy_doc_concurrent(&policies_list).await?;
let policy_docs = self.load_policy_doc_concurrent(&policies_list, LOAD_ALL_MODE).await?;
for (idx, p) in policy_docs.into_iter().enumerate() {
if p.policy.version.is_empty() {
@@ -1051,7 +1174,9 @@ impl Store for ObjectStore {
loop {
if item_name_list.len() < 32 {
let items = self.load_user_concurrent(&item_name_list, UserType::Reg).await?;
let items = self
.load_user_concurrent(&item_name_list, UserType::Reg, LOAD_ALL_MODE)
.await?;
for (idx, p) in items.into_iter().enumerate() {
if p.credentials.access_key.is_empty() {
@@ -1065,7 +1190,9 @@ impl Store for ObjectStore {
break;
}
let items = self.load_user_concurrent(&item_name_list, UserType::Reg).await?;
let items = self
.load_user_concurrent(&item_name_list, UserType::Reg, LOAD_ALL_MODE)
.await?;
for (idx, p) in items.into_iter().enumerate() {
if p.credentials.access_key.is_empty() {
@@ -1089,7 +1216,7 @@ impl Store for ObjectStore {
for item in item_name_list.iter() {
let name = rustfs_utils::path::dir(item);
debug!(group = %name, "IAM group loaded");
if let Err(err) = self.load_group(&name, &mut items_cache).await {
if let Err(err) = self.load_group_with(&name, &mut items_cache, LOAD_ALL_MODE).await {
return Err(Error::other(format!("load group failed: {err}")));
};
}
@@ -1107,7 +1234,7 @@ impl Store for ObjectStore {
loop {
if item_name_list.len() < 32 {
let items = self
.load_mapped_policy_concurrent(&item_name_list, UserType::Reg, false)
.load_mapped_policy_concurrent(&item_name_list, UserType::Reg, false, LOAD_ALL_MODE)
.await?;
for (idx, p) in items.into_iter().enumerate() {
@@ -1123,7 +1250,7 @@ impl Store for ObjectStore {
}
let items = self
.load_mapped_policy_concurrent(&item_name_list, UserType::Reg, false)
.load_mapped_policy_concurrent(&item_name_list, UserType::Reg, false, LOAD_ALL_MODE)
.await?;
for (idx, p) in items.into_iter().enumerate() {
@@ -1151,7 +1278,9 @@ impl Store for ObjectStore {
let name = item.trim_end_matches(".json");
debug!(group = %name, "IAM group policy loaded");
if let Err(err) = self.load_mapped_policy(name, UserType::Reg, true, &mut items_cache).await
if let Err(err) = self
.load_mapped_policy_with(name, UserType::Reg, true, &mut items_cache, LOAD_ALL_MODE)
.await
&& !is_err_no_such_policy(&err)
{
return Err(Error::other(format!("load group policy failed: {err}")));
@@ -1170,7 +1299,9 @@ impl Store for ObjectStore {
for item in item_name_list.iter() {
let name = rustfs_utils::path::dir(item);
debug!(user = %name, "IAM service user loaded");
if let Err(err) = self.load_user(&name, UserType::Svc, &mut items_cache).await
if let Err(err) = self
.load_user_with(&name, UserType::Svc, &mut items_cache, LOAD_ALL_MODE)
.await
&& !is_err_no_such_user(&err)
{
return Err(Error::other(format!("load svc user failed: {err}")));
@@ -1182,7 +1313,7 @@ impl Store for ObjectStore {
if !user_items_cache.contains_key(&parent) {
debug!(user = %parent, "IAM STS parent policy loaded");
if let Err(err) = self
.load_mapped_policy(&parent, UserType::Sts, false, &mut sts_policies_cache)
.load_mapped_policy_with(&parent, UserType::Sts, false, &mut sts_policies_cache, LOAD_ALL_MODE)
.await
&& !is_err_no_such_policy(&err)
{
@@ -1203,7 +1334,10 @@ impl Store for ObjectStore {
let name = rustfs_utils::path::dir(item);
debug!(user = %name, "IAM STS user loaded");
if let Err(err) = self.load_user(&name, UserType::Sts, &mut sts_items_cache).await {
if let Err(err) = self
.load_user_with(&name, UserType::Sts, &mut sts_items_cache, LOAD_ALL_MODE)
.await
{
debug!(user = %name, error = %err, "IAM STS user load failed");
};
}
@@ -1215,7 +1349,7 @@ impl Store for ObjectStore {
let name = item.trim_end_matches(".json");
debug!(user = %name, "IAM STS user policy loaded");
if let Err(err) = self
.load_mapped_policy(name, UserType::Sts, false, &mut sts_policies_cache)
.load_mapped_policy_with(name, UserType::Sts, false, &mut sts_policies_cache, LOAD_ALL_MODE)
.await
{
debug!(user = %name, error = %err, "IAM STS user policy load failed");
@@ -1250,12 +1384,29 @@ impl Store for ObjectStore {
#[cfg(test)]
mod tests {
use super::{DecryptSource, ObjectStore};
use super::{DecryptSource, LoadMode, ObjectStore};
use crate::keyring;
use rustfs_credentials::{Credentials, init_global_action_credentials};
use serial_test::serial;
use temp_env::with_vars;
/// rustfs#4304 / startup contract rustfs#4056: the bulk snapshot load
/// (`load_all`) must never depend on the node-counted namespace lock
/// quorum, so its reads have to carry `no_lock = true` down to the
/// storage layer.
#[test]
fn test_bootstrap_load_mode_bypasses_namespace_locks() {
assert!(LoadMode::BootstrapNoLock.no_lock());
assert!(LoadMode::BootstrapNoLock.read_opts().no_lock);
}
/// On-line single-object loads keep the locked read semantics.
#[test]
fn test_locked_load_mode_keeps_namespace_locks() {
assert!(!LoadMode::Locked.no_lock());
assert!(!LoadMode::Locked.read_opts().no_lock);
}
fn test_cred() -> Credentials {
if let Some(cred) = crate::root_credentials::credentials() {
return cred;
@@ -0,0 +1,35 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Test-only ECStore compatibility boundary for the IAM integration tests.
//!
//! All direct `rustfs_ecstore` facade imports used by tests in this crate
//! must go through this module (architecture migration rule:
//! `check_architecture_migration_rules.sh`). Keep the surface minimal —
//! only what the tests actually need to build a temp-disk ECStore fixture
//! and to flip the erasure setup type for lock-quorum fault injection.
#[allow(unused_imports)]
pub(crate) mod fixture {
pub(crate) use rustfs_ecstore::api::disk::endpoint::Endpoint;
pub(crate) use rustfs_ecstore::api::layout::{EndpointServerPools, Endpoints, PoolEndpoints, SetupType};
pub(crate) use rustfs_ecstore::api::storage::{ECStore, init_local_disks};
// `update_erasure_type` is a write-side global facade entry. Its use is
// restricted to reviewed storage_api boundaries; this test-only module is
// registered in check_architecture_migration_rules.sh precisely so the
// sequential-restart regression test can flip into distributed-erasure
// mode to make namespace-lock acquisition fail (rustfs#4304).
pub(crate) use rustfs_ecstore::api::global::update_erasure_type;
}
@@ -0,0 +1,159 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Regression test for rustfs#4304: IAM bootstrap must not depend on the
//! distributed namespace-lock quorum.
//!
//! During a sequential cluster restart the peer lock RPC endpoints are
//! unreachable, so every namespace-locked read fails with
//! `QuorumNotReached` even though the storage read quorum is already
//! satisfiable. The bulk snapshot load (`Store::load_all`) therefore has to
//! read with `no_lock = true` (startup contract rustfs#4056).
//!
//! The test builds a real 4-disk ECStore over temp dirs, seeds IAM data in
//! single-node mode, then flips the runtime into distributed-erasure mode.
//! In that mode `SetDisks::new_ns_lock` builds a distributed lock over the
//! set's lock clients — which are empty for a locally-built store — so every
//! locked read fails exactly like the sequential-restart scenario, while
//! plain storage reads keep working. The old (locked) `load_group` path must
//! fail and the lock-free `load_all` path must succeed.
mod ecstore_test_compat;
use ecstore_test_compat::fixture::{
ECStore, Endpoint, EndpointServerPools, Endpoints, PoolEndpoints, SetupType, init_local_disks, update_erasure_type,
};
use rustfs_iam::cache::Cache;
use rustfs_iam::store::object::ObjectStore;
use rustfs_iam::store::{GroupInfo, Store};
use serial_test::serial;
use std::collections::HashMap;
use std::sync::Arc;
use tokio_util::sync::CancellationToken;
const TEST_GROUP: &str = "seq-restart-group";
const TEST_MEMBERS: [&str; 2] = ["alice", "bob"];
async fn build_local_ecstore(temp_dir: &std::path::Path) -> Arc<ECStore> {
let disk_paths: Vec<_> = (1..=4).map(|i| temp_dir.join(format!("disk{i}"))).collect();
for disk_path in &disk_paths {
tokio::fs::create_dir_all(disk_path).await.unwrap();
}
let mut endpoints = Vec::new();
for (i, disk_path) in disk_paths.iter().enumerate() {
let mut endpoint = Endpoint::try_from(disk_path.to_str().unwrap()).unwrap();
endpoint.set_pool_index(0);
endpoint.set_set_index(0);
endpoint.set_disk_index(i);
endpoints.push(endpoint);
}
let pool_endpoints = PoolEndpoints {
legacy: false,
set_count: 1,
drives_per_set: 4,
endpoints: Endpoints::from(endpoints),
cmd_line: "test".to_string(),
platform: format!("OS: {} | Arch: {}", std::env::consts::OS, std::env::consts::ARCH),
};
let endpoint_pools = EndpointServerPools::from(vec![pool_endpoints]);
init_local_disks(endpoint_pools.clone()).await.unwrap();
// Port 0 keeps this integration binary parallel-safe alongside other
// ECStore-backed tests.
let server_addr: std::net::SocketAddr = "127.0.0.1:0".parse().unwrap();
ECStore::new(server_addr, endpoint_pools, CancellationToken::new())
.await
.unwrap()
}
/// Restores single-node erasure mode even when an assertion panics, so a
/// failing run cannot poison later `#[serial]` tests in this process.
struct ErasureModeGuard;
impl Drop for ErasureModeGuard {
fn drop(&mut self) {
pollster::block_on(update_erasure_type(SetupType::Erasure));
}
}
#[tokio::test(flavor = "multi_thread")]
#[serial]
async fn load_all_bypasses_namespace_lock_quorum() {
// The lock acquire timeout is latched into a OnceLock on first use, so it
// must be shortened before the first locked operation of this process.
temp_env::async_with_vars([(rustfs_config::ENV_OBJECT_LOCK_ACQUIRE_TIMEOUT, Some("1"))], async {
let temp_dir = tempfile::TempDir::with_prefix("rustfs_iam_no_lock_test_").unwrap();
let ecstore = build_local_ecstore(temp_dir.path()).await;
let store = ObjectStore::new(ecstore);
// Seed IAM data while namespace locks still work (single-node mode).
store
.save_group_info(TEST_GROUP, GroupInfo::new(TEST_MEMBERS.iter().map(|m| m.to_string()).collect()))
.await
.expect("seeding group info in single-node mode must succeed");
let mut baseline = HashMap::new();
store
.load_group(TEST_GROUP, &mut baseline)
.await
.expect("locked load_group must succeed in single-node mode");
assert_eq!(baseline[TEST_GROUP].members, TEST_MEMBERS, "seeded group must round-trip");
// Flip into distributed-erasure mode: new_ns_lock now builds a
// distributed lock over the set's (empty) lock-client list, so every
// locked read fails — the sequential-restart failure mode of
// rustfs#4304 (lock quorum unavailable, storage quorum healthy).
update_erasure_type(SetupType::DistErasure).await;
let _mode_guard = ErasureModeGuard;
let mut locked_read = HashMap::new();
let locked_err = store
.load_group(TEST_GROUP, &mut locked_read)
.await
.expect_err("locked load_group must fail while the lock quorum is unavailable");
assert!(locked_read.is_empty(), "failed locked read must not populate results");
// The P0 fix: the bulk snapshot load reads with no_lock = true, so it
// must succeed in exactly the state where the locked path fails.
let cache = Cache::default();
store
.load_all(&cache)
.await
.unwrap_or_else(|err| panic!("load_all must bypass namespace locks (locked path failed with: {locked_err}): {err}"));
// P3 step 1: notification-path cache refreshes use the same lock-free
// reads, so a cross-node notification must also be able to refresh
// this group while the lock quorum is unavailable.
let mut notification_read = HashMap::new();
store
.load_group_no_lock(TEST_GROUP, &mut notification_read)
.await
.expect("lock-free notification-path load_group must succeed while the lock quorum is unavailable");
assert_eq!(notification_read[TEST_GROUP].members, TEST_MEMBERS);
// Back in single-node mode the locked path works again; verify the
// data survived the whole exercise intact (fail-closed integrity).
drop(_mode_guard);
let mut recovered = HashMap::new();
store
.load_group(TEST_GROUP, &mut recovered)
.await
.expect("locked load_group must succeed again after restoring single-node mode");
assert_eq!(recovered[TEST_GROUP].members, TEST_MEMBERS, "group data must be intact");
})
.await;
}