mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-06 13:27:43 +00:00
98d3619613
* fix: address rc.1 release blockers
* fix: route release guards through architecture boundaries
* fix: close remaining rc.1 regression gaps
* refactor: group multipart listing options
* fix: resolve rc.1 CI regressions
* fix(ecstore): keep bucket-config writes off the caller's stack
A bucket-config write nests incarnation resolution (which can drive legacy
migration and a peer fan-out), a full metadata load, and `save` — itself an
object PUT that pulls in the whole erasure write path. Every request that
mutates bucket config is already several futures deep, so inlining all of
that into one state machine overflows the 2MiB worker stack in debug builds.
Two CI lanes aborted with SIGABRT on this:
ILM Integration (serial)
rustfs app::lifecycle_transition_api_test::
compensation_driven_complete_multipart_upload_still_transitions
Test and Lint (swift)
rustfs-protocols::swift_metadata_persistence::
swift_metadata_writes_are_durable
Neither test file is touched by this branch and both lanes are green on
main. Stack-pointer probing showed ~780KiB consumed between
`metadata_sys::update` and the config read alone, with single hops of
363KiB (`update` -> `acquire_config_write_guard_for_incarnation`), 125KiB
and 105KiB.
Box the deep sub-futures on both read-modify-write paths (`update` /
`update_checked` and `update_config_with` / `update_config_with_checked`)
so each guard's own state machine stays small. Behaviour is unchanged;
`update` -> guard drops to 253KiB and both tests pass on the default stack.
* fix(lifecycle): unbreak restore under the bucket generation fence
The ILM lane aborted on a stack overflow before reaching these, so they
were never reported; with that fixed, four restore tests fail. All four
are green on main and none of their test files are touched by this branch.
1. RestoreObject and ListMultipartUploads hard-required
`opts.expected_bucket_incarnation_id`, but `apply_bucket_generation_guard`
deliberately leaves it unset when no guard extension is present — only the
S3 access layer installs one. Every direct caller therefore got
`InternalError: ... bucket generation guard is missing`. Resolve the
current generation instead, the way the copy path already does. The fence
is unaffected: RestoreObject still re-reads the incarnation from disk and
compares before admitting the restore, and the multipart listing is
filtered by the value it resolves.
2. `restore_expiry_snapshot_matches` (new on this branch) rejected every
restored-copy expiry whose `restore_expires` had not already elapsed.
Whether the restored copy is due to expire is the ILM evaluator's
decision, made when it emitted DeleteRestoredAction; re-deriving it in
the set layer only adds a way for a legitimate action to be rejected.
The stale-event risk it appears to guard is already covered by the
surrounding snapshot match — a re-restore rewrites `restore_expires`,
so a replayed event fails the equality check. Drop the clause; the
fifteen identity clauses are unchanged.
Fixed:
rustfs app::lifecycle_transition_api_test::
restore_object_usecase_accepts_exactly_one_of_two_concurrent_restores
restore_object_usecase_completes_suspended_null_version_in_place
restore_object_usecase_reports_ongoing_conflict
rustfs-scanner::lifecycle_integration_test serial_tests::
test_restore_chain_local_read_expiry_keeps_remote_and_allows_re_restore
Verification: the CI ILM lane filter now runs 53/53 green locally.
* chore: address review follow-ups on this branch
Four items from the adversarial review that were still open.
- Restore the assertion `test_bucket_replication_replayed_delete_marker_
preserves_source_mtime_without_source_restart` is named for. The branch
had replaced the backlog#867 mtime check with `assert_replication_
converged`, which any successful replication satisfies, and deleted the
two helpers it needed — so the regression the test exists to catch would
now pass. This matters here specifically because the branch changes the
flag feeding `replication_delete_remove_options` and routes replay
through a new file and ordering.
- Drop `read_config_no_lock_preserve_empty`: zero production callers (the
one real consumer calls the `_with_metadata` variant directly). Its test
stanza now exercises that variant, so the coverage moves to live code
rather than being deleted.
- Revert the `bytesize` bump. It is a no-op: `Cargo.lock` already pinned
2.7.0 before this branch and is untouched, so the caret range already
resolved there. Nothing in the diff uses the crate.
- Split the AGENTS.md "Adversarial Validation" policy change out of this
branch. The edit is defensible on its own, but it relaxes the review gate
that this branch has to pass, so it should land as its own PR reviewed on
its own merits rather than bundled with the change that benefits from it.
The reverted hunks are unchanged and ready to re-apply.
Not changed, deliberately: the missing-sidecar path still fails closed.
`missing_bucket_incarnation_sidecar_for_new_metadata_fails_closed` pins
that on purpose, and serving a non-authoritative Object Lock state would
be the wrong trade. The residual concern stands and is recorded in review
— a crash between the two writes in `persist_new_and_set` leaves the
bucket unloadable until DeleteBucket+CreateBucket, and the repair branches
in `migrate_legacy_metadata` and `make_bucket` are unreachable dead code
for that case. Resolving it needs the read path and the (transaction-lock
holding) repair path to be separated, which is more than a follow-up edit.
* test(ci): serialize the new bucket-incarnation tests
The five tests this branch adds around the incarnation / lifecycle fence
drive `init_bucket_metadata_sys` and `bucket_metadata_sys_of` — process-global
OnceLock state that `serial_test`'s `#[serial]` cannot protect across
nextest's process boundary — and they delete+recreate buckets, the shape that
raced into InsufficientWriteQuorum in backlog#937.
Add them to the `ecstore-serial-flaky` group in both the default and ci
profiles (nextest evaluates a named profile's own overrides list, so the
ci mirror is required). Preventive serialization only, no retries.
Not a full fix for the review comment: `bucket_delete_waits_for_config_
mutation_fence` still proves liveness with a fixed 200ms sleep plus
`assert!(!delete.is_finished())`. Turning that into readiness polling needs
a production-side signal to wait on — asserting "still blocked" is inherently
a negative. Serializing the group removes the parallel-load pressure that
makes the window fragile; the sleep itself is left for a follow-up.
* test(ecstore): pin that a drained bucket is actually deletable
`DeleteBucket`'s emptiness check is `has_xlmeta_files`, a raw scan of the
bucket directory on local disks — not an S3-level listing. So "the client
drained the bucket" and "the bucket is deletable" are two different
contracts, and only the first one was covered.
That gap is what the `S3 Implemented Tests` lane is failing on: 219 cases,
all `BucketNotEmpty` on `nuke_prefixed_buckets`, with every test body
passing. The first one is `test_versioning_obj_suspend_versions`, reported
by pytest as PASSED followed by ERROR at teardown.
Add the missing assertion for the unversioned path: PUT, client DELETE,
then assert no `xl.meta` survives and `DeleteBucket` succeeds. It passes —
which is itself a result: the plain delete path leaves no residue, so the
s3-tests failure is not there.
The versioning-suspended path is the remaining suspect (the client DELETE
leaves a null delete marker, and draining means purging it by
`versionId=null`). It is not covered here: `BucketVersioningSys` resolves
through the ambient `get_bucket_metadata_sys()` OnceLock, which this unit
env cannot set, so the bucket never actually reports as suspended. That
repro belongs at the e2e layer where a real server owns the versioning
state.
* fix(ecstore): let an explicit null-version delete purge its delete marker
Root cause of the `S3 Implemented Tests` lane: 219 cases, all
`BucketNotEmpty` on `nuke_prefixed_buckets`, every test body passing.
On a versioning-suspended bucket a client DELETE leaves a null delete
marker — correct S3 semantics, and an `xl.meta` on disk. Draining the
bucket therefore means purging that marker as `?versionId=null`, which is
what `nuke_bucket` does before `DeleteBucket`. That purge was rejected:
explicit null-version purge of the null delete marker must succeed,
got [Some(MethodNotAllowed)]
so the marker survived, and `DeleteBucket`'s emptiness check — a raw
`has_xlmeta_files` scan of the bucket directory, not an S3 listing — kept
reporting the bucket as non-empty.
The two sides of the version comparison in the batch delete loop are in
different namespaces. `goi.version_id` is the client-facing identity, where
`from_file_info` synthesizes `Some(Uuid::nil())` for a null version on a
versioned *or versioning-suspended* bucket. `version_id` is the storage
identity, where `delete_file_info_version_id` maps an explicit
`?versionId=null` to `None`. Comparing them raw makes the purge look like a
version mismatch, so `explicit_delete_marker` is false and the
`MethodNotAllowed` from the lookup is recorded as a delete failure.
This only became reachable on this branch: previously `check_opts` did not
carry `dobj.version_id`, so `set_disk_delete_creates_delete_marker` was
true, `object_lock_check_required` was false, and the lookup that produces
`MethodNotAllowed` never ran. Adding the version id to `check_opts` lit up
a comparison that was already wrong.
Normalize both sides through `delete_file_info_version_id`.
The regression test injects a real Suspended bucket-config snapshot — the
delete path reads versioned/suspended from that snapshot, not from `opts`,
so without it `from_file_info` never synthesizes the null version id and
the branch is not reached. Mutation-checked: restoring the raw comparison
fails the test with the exact `MethodNotAllowed` above.
* fix(app): drop the now-needless struct update
Reverting `crates/replication` to main removed the extra `MrfReplicateEntry`
fields, so this literal specifies every field again and `..Default::default()`
trips `clippy::needless_update` under `-D warnings`.
Caught by CI, not locally: I had run `cargo check --workspace --all-targets`,
which does not see clippy-only lints. Ran `cargo clippy --workspace
--all-targets -- -D warnings` here — clean.
* test(e2e): assert the fresh-volume classification
four_node_empty_legacy_volumes_start_as_fresh only started the cluster and
listed buckets — no assertion, so any classification path that still permits
startup left it green without proving the pre-created empty `.minio.sys`
directories were treated as fresh volumes.
Pin what that classification actually leaves behind: no buckets adopted into
the namespace, `.rustfs.sys/format.json` written on every drive, and the empty
legacy directory left untouched rather than migrated into.
* fix(bucket): apply the requested Object Lock to existing buckets
Site replication replays make-with-versioning against the destination,
carrying the source's `lockEnabled`. When the destination bucket already
exists it takes `force_create`, and the whole option-application block was
gated on `confirmed_missing` — so the call returned success while the replica
stayed unlocked. Replicated versions could then be deleted without the
retention the source enforces.
Object Lock enable is one-way, so applying it to an existing bucket is safe:
move it out of the creation-only gate, keeping `created` and versioning-only
options creation-scoped as before.
An existing authoritative bucket takes the `cache_bucket_metadata_in` branch,
which only caches, so the enable would have been dropped on restart. Persist
instead when the enable actually changed something.
Mutation-checked: restoring the creation-only gate fails the new
`force_create_enables_object_lock_on_an_existing_bucket` with "Object Lock
must be enabled on the existing bucket".
cargo nextest run -p rustfs-ecstore --lib: 3633 passed.
* fix(ecstore): box the generation-checked config mutation paths too
The earlier stack fix boxed `update` and `delete`, but an authorized
bucket-config mutation carrying an incarnation takes `update_if_incarnation`
/ `delete_if_incarnation` instead — which were still inlining the whole
resolve/load/save chain into an already-deep request future. Same overflow,
sibling path.
* fix(restore): keep the nil-version normalization the strip removed
Reverting the replication subsystem to main took `set_disk/replication.rs`
with it, but one line in that file was this branch's own fix rather than
replication work:
- self.version_id.filter(|v| !v.is_nil()) == fi.version_id.filter(|v| !v.is_nil())
+ self.version_id == fi.version_id
For a versioning-suspended object the expected version is `Some(Uuid::nil())`
while the read-back `FileInfo` carries `None`, so the raw compare reports
every suspended restore as "restored object changed before restore metadata
finalization" and the copy-back never commits. Same nil-vs-None mismatch as
the null delete-marker purge fixed earlier on this branch.
Caught by `Test and Lint (rio-v2)`, not by my local runs: the test lives in
`transition_commit_failure_tests`, gated behind `feature = "test-util"`, so
the 3633-test suite I had been running never included it. Re-ran with
`--features rio-v2,test-util`: 3722 passed.
609 lines
27 KiB
Rust
609 lines
27 KiB
Rust
// Copyright 2024 RustFS Team
|
|
//
|
|
// Licensed under the Apache License, Version 2.0 (the "License");
|
|
// you may not use this file except in compliance with the License.
|
|
// You may obtain a copy of the License at
|
|
//
|
|
// http://www.apache.org/licenses/LICENSE-2.0
|
|
//
|
|
// Unless required by applicable law or agreed to in writing, software
|
|
// distributed under the License is distributed on an "AS IS" BASIS,
|
|
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
// See the License for the specific language governing permissions and
|
|
// limitations under the License.
|
|
|
|
//! System metadata compatibility: write both x-rustfs-internal-* and x-minio-internal-*
|
|
//! for MinIO interoperability. Read prefers RustFS, fallback to MinIO.
|
|
|
|
use std::collections::{BTreeMap, HashMap};
|
|
|
|
pub const RUSTFS_INTERNAL_PREFIX: &str = "x-rustfs-internal-";
|
|
pub const MINIO_INTERNAL_PREFIX: &str = "x-minio-internal-";
|
|
|
|
// Key suffixes (lowercase, no prefix)
|
|
pub const SUFFIX_INLINE_DATA: &str = "inline-data";
|
|
pub const SUFFIX_DATA_MOVED: &str = "data-moved";
|
|
/// Transient flag for data movement
|
|
pub const SUFFIX_DATA_MOV: &str = "data-mov";
|
|
/// Transient flag for healing
|
|
pub const SUFFIX_HEALING: &str = "healing";
|
|
pub const SUFFIX_COMPRESSION: &str = "compression";
|
|
pub const SUFFIX_COMPRESSION_SIZE: &str = "compression-size";
|
|
pub const SUFFIX_ACTUAL_SIZE: &str = "actual-size";
|
|
pub const SUFFIX_ACTUAL_OBJECT_SIZE: &str = "actual-object-size";
|
|
/// Used by replication; key stored with capital A
|
|
pub const SUFFIX_ACTUAL_OBJECT_SIZE_CAP: &str = "Actual-Object-Size";
|
|
pub const SUFFIX_CRC: &str = "crc";
|
|
pub const SUFFIX_TRANSITION_STATUS: &str = "transition-status";
|
|
pub const SUFFIX_TRANSITIONED_OBJECTNAME: &str = "transitioned-object";
|
|
pub const SUFFIX_TRANSITIONED_VERSION_ID: &str = "transitioned-versionID";
|
|
pub const SUFFIX_TRANSITIONED_VERSION_STATE: &str = "transitioned-version-state";
|
|
pub const SUFFIX_TRANSITION_TIER: &str = "transition-tier";
|
|
pub const SUFFIX_TRANSITION_TIER_DESTINATION_ID: &str = "transition-tier-destination-id";
|
|
pub const SUFFIX_TRANSITION_TRANSACTION_ID: &str = "transition-transaction-id";
|
|
pub const SUFFIX_RESTORE_OPERATION_ID: &str = "restore-operation-id";
|
|
pub const SUFFIX_BUCKET_INCARNATION_ID: &str = "bucket-incarnation-id";
|
|
pub const SUFFIX_FREE_VERSION: &str = "free-version";
|
|
pub const SUFFIX_PURGESTATUS: &str = "purgestatus";
|
|
pub const SUFFIX_REPLICA_STATUS: &str = "replica-status";
|
|
pub const SUFFIX_REPLICA_TIMESTAMP: &str = "replica-timestamp";
|
|
pub const SUFFIX_REPLICATION_STATUS: &str = "replication-status";
|
|
pub const SUFFIX_REPLICATION_TIMESTAMP: &str = "replication-timestamp";
|
|
pub const SUFFIX_TAGGING_TIMESTAMP: &str = "tagging-timestamp";
|
|
pub const SUFFIX_OBJECTLOCK_RETENTION_TIMESTAMP: &str = "objectlock-retention-timestamp";
|
|
pub const SUFFIX_OBJECTLOCK_LEGALHOLD_TIMESTAMP: &str = "objectlock-legalhold-timestamp";
|
|
pub const SUFFIX_REPLICATION_RESET: &str = "replication-reset";
|
|
/// Prefix for replication-reset-{arn} keys; use with internal_key_strip_suffix_prefix to extract arn.
|
|
pub const SUFFIX_REPLICATION_RESET_ARN_PREFIX: &str = "replication-reset-";
|
|
pub const SUFFIX_TIER_FV_ID: &str = "tier-free-versionID";
|
|
pub const SUFFIX_TIER_FV_MARKER: &str = "tier-free-marker";
|
|
pub const SUFFIX_TIER_SKIP_FV_ID: &str = "tier-skip-fvid";
|
|
|
|
/// Per-target delete-marker version ids are stored one key per target ARN.
|
|
pub const SUFFIX_REPLICATION_DELETE_MARKER_VERSION_ARN_PREFIX: &str = "replication-delete-marker-version-";
|
|
|
|
/// Case-insensitive (ASCII) check that `s` begins with `prefix`. Equivalent to
|
|
/// `s.to_lowercase().starts_with(prefix)` when `prefix` is ASCII (as both internal prefixes are),
|
|
/// but without allocating.
|
|
fn starts_with_ignore_ascii_case(s: &str, prefix: &str) -> bool {
|
|
s.len() >= prefix.len() && s.as_bytes()[..prefix.len()].eq_ignore_ascii_case(prefix.as_bytes())
|
|
}
|
|
|
|
/// Allocation-free equivalent of `key.to_lowercase() == format!("{prefix}{suffix}")`.
|
|
/// The common ASCII path matches the prefix case-insensitively and ASCII-lowercases the suffix
|
|
/// region byte-for-byte. Non-ASCII keys fall back to full Unicode lowercasing so the result is
|
|
/// identical to the original for every input (e.g. U+212A KELVIN SIGN lowercases to ASCII `k`).
|
|
fn internal_key_eq(key: &str, prefix: &str, suffix: &str) -> bool {
|
|
if !key.is_ascii() {
|
|
return key.to_lowercase() == format!("{prefix}{suffix}");
|
|
}
|
|
let key = key.as_bytes();
|
|
if key.len() != prefix.len() + suffix.len() {
|
|
return false;
|
|
}
|
|
let (key_prefix, key_suffix) = key.split_at(prefix.len());
|
|
key_prefix.eq_ignore_ascii_case(prefix.as_bytes())
|
|
&& key_suffix
|
|
.iter()
|
|
.zip(suffix.as_bytes())
|
|
.all(|(k, s)| k.to_ascii_lowercase() == *s)
|
|
}
|
|
|
|
/// Returns true if the key is an internal metadata key (x-rustfs-internal-* or x-minio-internal-*)
|
|
/// for xl.meta compatibility. Case-insensitive.
|
|
pub fn is_internal_key(key: &str) -> bool {
|
|
starts_with_ignore_ascii_case(key, RUSTFS_INTERNAL_PREFIX) || starts_with_ignore_ascii_case(key, MINIO_INTERNAL_PREFIX)
|
|
}
|
|
|
|
/// Returns true if the key matches the given suffix for either x-rustfs-internal-* or x-minio-internal-*.
|
|
pub fn has_internal_suffix(key: &str, suffix: &str) -> bool {
|
|
internal_key_eq(key, RUSTFS_INTERNAL_PREFIX, suffix) || internal_key_eq(key, MINIO_INTERNAL_PREFIX, suffix)
|
|
}
|
|
|
|
/// Strips x-rustfs-internal- or x-minio-internal- prefix from key. Returns the suffix part.
|
|
/// Case-insensitive. Returns None if key is not an internal key.
|
|
pub fn strip_internal_prefix(key: &str) -> Option<String> {
|
|
let rest = if starts_with_ignore_ascii_case(key, RUSTFS_INTERNAL_PREFIX) {
|
|
&key[RUSTFS_INTERNAL_PREFIX.len()..]
|
|
} else if starts_with_ignore_ascii_case(key, MINIO_INTERNAL_PREFIX) {
|
|
&key[MINIO_INTERNAL_PREFIX.len()..]
|
|
} else {
|
|
return None;
|
|
};
|
|
Some(rest.to_lowercase())
|
|
}
|
|
|
|
/// Returns true if key is internal and its suffix part starts with the given suffix_prefix.
|
|
/// E.g. internal_key_starts_with("x-rustfs-internal-replication-reset-arn1", "replication-reset") == true.
|
|
pub fn internal_key_starts_with(key: &str, suffix_prefix: &str) -> bool {
|
|
strip_internal_prefix(key).is_some_and(|s| s.starts_with(suffix_prefix))
|
|
}
|
|
|
|
/// For keys like x-rustfs-internal-replication-reset-{arn}, strips the internal prefix and suffix_prefix,
|
|
/// returning the remainder (e.g. "arn1"). Returns None if key does not match.
|
|
pub fn internal_key_strip_suffix_prefix(key: &str, suffix_prefix: &str) -> Option<String> {
|
|
let rest = strip_internal_prefix(key)?;
|
|
rest.strip_prefix(suffix_prefix).map(|s| s.to_string())
|
|
}
|
|
|
|
fn both_keys(suffix: &str) -> (String, String) {
|
|
(format!("{RUSTFS_INTERNAL_PREFIX}{suffix}"), format!("{MINIO_INTERNAL_PREFIX}{suffix}"))
|
|
}
|
|
|
|
/// Longest known suffix ("objectlock-retention-timestamp", 30 bytes) plus the 18-byte prefix fits
|
|
/// well under this; the cap leaves ample headroom so lookups never touch the heap in practice.
|
|
const INTERNAL_KEY_STACK_CAP: usize = 96;
|
|
|
|
/// Builds the internal key `{prefix}{suffix}` in a stack buffer and invokes `f` with it, avoiding
|
|
/// the per-lookup heap allocation that `format!` incurs. Falls back to an owned `String` only for
|
|
/// unusually long suffixes that do not fit the stack buffer.
|
|
fn with_internal_key<R>(prefix: &str, suffix: &str, f: impl FnOnce(&str) -> R) -> R {
|
|
let total = prefix.len() + suffix.len();
|
|
if total <= INTERNAL_KEY_STACK_CAP {
|
|
let mut buf = [0u8; INTERNAL_KEY_STACK_CAP];
|
|
buf[..prefix.len()].copy_from_slice(prefix.as_bytes());
|
|
buf[prefix.len()..total].copy_from_slice(suffix.as_bytes());
|
|
// `prefix` and `suffix` are both `&str`, so their concatenation is valid UTF-8.
|
|
match std::str::from_utf8(&buf[..total]) {
|
|
Ok(key) => f(key),
|
|
Err(_) => f(&format!("{prefix}{suffix}")),
|
|
}
|
|
} else {
|
|
f(&format!("{prefix}{suffix}"))
|
|
}
|
|
}
|
|
|
|
/// Builds the RustFS internal key for the given suffix. Use when a single key is needed (e.g. for
|
|
/// backward compat). Prefer insert_str/get_str when both keys should be written/read.
|
|
pub fn internal_key_rustfs(suffix: &str) -> String {
|
|
format!("{RUSTFS_INTERNAL_PREFIX}{suffix}")
|
|
}
|
|
|
|
// === String type (FileInfo.metadata, user_defined) ===
|
|
|
|
pub fn insert_str(map: &mut HashMap<String, String>, suffix: &str, value: String) {
|
|
let (k1, k2) = both_keys(suffix);
|
|
map.insert(k1, value.clone());
|
|
map.insert(k2, value);
|
|
}
|
|
|
|
pub fn get_str(map: &HashMap<String, String>, suffix: &str) -> Option<String> {
|
|
if let Some(v) = with_internal_key(RUSTFS_INTERNAL_PREFIX, suffix, |k1| map.get(k1).cloned()) {
|
|
return Some(v);
|
|
}
|
|
if let Some(v) = with_internal_key(MINIO_INTERNAL_PREFIX, suffix, |k2| map.get(k2).cloned()) {
|
|
return Some(v);
|
|
}
|
|
// Rare fallback: case-insensitive scan for non-canonical key casing.
|
|
let (k1, k2) = both_keys(suffix);
|
|
map.iter()
|
|
.find(|(key, _)| key.eq_ignore_ascii_case(&k1) || key.eq_ignore_ascii_case(&k2))
|
|
.map(|(_, value)| value.clone())
|
|
}
|
|
|
|
fn get_consistent_value<'a, V: AsRef<[u8]>>(map: &'a HashMap<String, V>, suffix: &str) -> Option<&'a V> {
|
|
let (rustfs_key, minio_key) = both_keys(suffix);
|
|
let mut value = None;
|
|
for (key, candidate) in map {
|
|
if !key.eq_ignore_ascii_case(&rustfs_key) && !key.eq_ignore_ascii_case(&minio_key) {
|
|
continue;
|
|
}
|
|
if candidate.as_ref().is_empty() || value.is_some_and(|current: &V| current.as_ref() != candidate.as_ref()) {
|
|
return None;
|
|
}
|
|
value = Some(candidate);
|
|
}
|
|
value
|
|
}
|
|
|
|
/// Returns a non-empty value when every compatibility key present for `suffix` agrees.
|
|
/// A single RustFS or MinIO key is accepted for backward compatibility; conflicting or empty
|
|
/// values return `None` so callers at destructive boundaries can fail closed.
|
|
pub fn get_consistent_str<'a>(map: &'a HashMap<String, String>, suffix: &str) -> Option<&'a str> {
|
|
get_consistent_value(map, suffix).map(String::as_str)
|
|
}
|
|
|
|
pub fn contains_key_str(map: &HashMap<String, String>, suffix: &str) -> bool {
|
|
if with_internal_key(RUSTFS_INTERNAL_PREFIX, suffix, |k1| map.contains_key(k1)) {
|
|
return true;
|
|
}
|
|
if with_internal_key(MINIO_INTERNAL_PREFIX, suffix, |k2| map.contains_key(k2)) {
|
|
return true;
|
|
}
|
|
let (k1, k2) = both_keys(suffix);
|
|
map.keys()
|
|
.any(|key| key.eq_ignore_ascii_case(&k1) || key.eq_ignore_ascii_case(&k2))
|
|
}
|
|
|
|
pub fn remove_str(map: &mut HashMap<String, String>, suffix: &str) {
|
|
with_internal_key(RUSTFS_INTERNAL_PREFIX, suffix, |k1| map.remove(k1));
|
|
with_internal_key(MINIO_INTERNAL_PREFIX, suffix, |k2| map.remove(k2));
|
|
let (k1, k2) = both_keys(suffix);
|
|
map.retain(|key, _| !key.eq_ignore_ascii_case(&k1) && !key.eq_ignore_ascii_case(&k2));
|
|
}
|
|
|
|
// === Vec<u8> type (meta_sys) ===
|
|
|
|
pub fn insert_bytes(map: &mut HashMap<String, Vec<u8>>, suffix: &str, value: Vec<u8>) {
|
|
let (k1, k2) = both_keys(suffix);
|
|
let v = value.clone();
|
|
map.insert(k1, value);
|
|
map.insert(k2, v);
|
|
}
|
|
|
|
pub fn get_bytes(map: &HashMap<String, Vec<u8>>, suffix: &str) -> Option<Vec<u8>> {
|
|
with_internal_key(RUSTFS_INTERNAL_PREFIX, suffix, |k1| map.get(k1).cloned())
|
|
.or_else(|| with_internal_key(MINIO_INTERNAL_PREFIX, suffix, |k2| map.get(k2).cloned()))
|
|
}
|
|
|
|
/// Byte-valued counterpart of [`get_consistent_str`].
|
|
pub fn get_consistent_bytes<'a>(map: &'a HashMap<String, Vec<u8>>, suffix: &str) -> Option<&'a [u8]> {
|
|
get_consistent_value(map, suffix).map(Vec::as_slice)
|
|
}
|
|
|
|
pub fn contains_key_bytes(map: &HashMap<String, Vec<u8>>, suffix: &str) -> bool {
|
|
with_internal_key(RUSTFS_INTERNAL_PREFIX, suffix, |k1| map.contains_key(k1))
|
|
|| with_internal_key(MINIO_INTERNAL_PREFIX, suffix, |k2| map.contains_key(k2))
|
|
}
|
|
|
|
pub fn remove_bytes(map: &mut HashMap<String, Vec<u8>>, suffix: &str) {
|
|
with_internal_key(RUSTFS_INTERNAL_PREFIX, suffix, |k1| map.remove(k1));
|
|
with_internal_key(MINIO_INTERNAL_PREFIX, suffix, |k2| map.remove(k2));
|
|
}
|
|
|
|
/// Strips an internal metadata prefix while preserving the suffix casing.
|
|
pub fn strip_internal_prefix_preserving_case(key: &str) -> Option<&str> {
|
|
if starts_with_ignore_ascii_case(key, RUSTFS_INTERNAL_PREFIX) {
|
|
key.get(RUSTFS_INTERNAL_PREFIX.len()..)
|
|
} else if starts_with_ignore_ascii_case(key, MINIO_INTERNAL_PREFIX) {
|
|
key.get(MINIO_INTERNAL_PREFIX.len()..)
|
|
} else {
|
|
None
|
|
}
|
|
}
|
|
|
|
/// Reads the bounded per-target delete-marker version map in one metadata scan.
|
|
/// The boolean is set when matching metadata is malformed or compatibility keys disagree.
|
|
pub fn target_delete_marker_versions(map: &HashMap<String, String>) -> (HashMap<String, String>, bool) {
|
|
const MAX_ENTRIES: usize = 1_000;
|
|
const MAX_ARN_LEN: usize = 1_024;
|
|
const MAX_VERSION_ID_LEN: usize = 1_024;
|
|
|
|
let mut versions = BTreeMap::<String, Option<String>>::new();
|
|
let mut corrupt = false;
|
|
for (key, value) in map {
|
|
let Some(suffix) = strip_internal_prefix_preserving_case(key) else {
|
|
continue;
|
|
};
|
|
let Some(prefix) = suffix.get(..SUFFIX_REPLICATION_DELETE_MARKER_VERSION_ARN_PREFIX.len()) else {
|
|
continue;
|
|
};
|
|
if !prefix.eq_ignore_ascii_case(SUFFIX_REPLICATION_DELETE_MARKER_VERSION_ARN_PREFIX) {
|
|
continue;
|
|
}
|
|
let arn = &suffix[SUFFIX_REPLICATION_DELETE_MARKER_VERSION_ARN_PREFIX.len()..];
|
|
if !arn.starts_with("arn:") || arn.len() > MAX_ARN_LEN || value.is_empty() || value.len() > MAX_VERSION_ID_LEN {
|
|
corrupt = true;
|
|
continue;
|
|
}
|
|
match versions.entry(arn.to_string()) {
|
|
std::collections::btree_map::Entry::Vacant(entry) => {
|
|
entry.insert(Some(value.clone()));
|
|
}
|
|
std::collections::btree_map::Entry::Occupied(mut entry) => {
|
|
if entry.get().as_deref() != Some(value.as_str()) {
|
|
entry.insert(None);
|
|
corrupt = true;
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
// Apply the cap after collecting, never during. Capping mid-iteration made
|
|
// the surviving subset depend on `HashMap` order, so two disks decoding the
|
|
// same metadata could keep different entries and hash differently — turning
|
|
// an over-cap object into a quorum failure instead of a reported corruption.
|
|
// `BTreeMap` order is total, so truncating here is identical everywhere.
|
|
if versions.len() > MAX_ENTRIES {
|
|
corrupt = true;
|
|
let retained = versions.keys().take(MAX_ENTRIES).cloned().collect::<Vec<_>>();
|
|
versions.retain(|arn, _| retained.binary_search(arn).is_ok());
|
|
}
|
|
|
|
(
|
|
versions
|
|
.into_iter()
|
|
.filter_map(|(arn, version_id)| version_id.map(|version_id| (arn, version_id)))
|
|
.collect(),
|
|
corrupt,
|
|
)
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
|
|
#[test]
|
|
fn test_is_internal_key() {
|
|
assert!(is_internal_key("x-rustfs-internal-healing"));
|
|
assert!(is_internal_key("x-rustfs-internal-purgestatus"));
|
|
assert!(is_internal_key("X-RustFS-Internal-purgestatus"));
|
|
assert!(is_internal_key("x-minio-internal-compression"));
|
|
assert!(is_internal_key("x-minio-internal-replication-status"));
|
|
assert!(is_internal_key("X-Minio-Internal-Compression"));
|
|
assert!(!is_internal_key("x-amz-meta-custom"));
|
|
assert!(!is_internal_key("content-type"));
|
|
assert!(!is_internal_key("x-rustfs-meta-custom"));
|
|
}
|
|
|
|
#[test]
|
|
fn test_has_internal_suffix() {
|
|
assert!(has_internal_suffix("x-rustfs-internal-purgestatus", SUFFIX_PURGESTATUS));
|
|
assert!(has_internal_suffix("X-Minio-Internal-purgestatus", SUFFIX_PURGESTATUS));
|
|
assert!(has_internal_suffix("x-minio-internal-compression", SUFFIX_COMPRESSION));
|
|
assert!(has_internal_suffix("x-rustfs-internal-healing", SUFFIX_HEALING));
|
|
assert!(has_internal_suffix("x-minio-internal-data-mov", SUFFIX_DATA_MOV));
|
|
assert!(!has_internal_suffix("x-rustfs-internal-purgestatus", SUFFIX_HEALING));
|
|
assert!(!has_internal_suffix("x-amz-meta-custom", SUFFIX_PURGESTATUS));
|
|
}
|
|
|
|
#[test]
|
|
fn test_str_lookup_accepts_minio_metadata_case() {
|
|
let mut metadata = HashMap::from([
|
|
("X-Minio-Internal-compression".to_string(), "klauspost/compress/s2".to_string()),
|
|
("X-Minio-Internal-actual-size".to_string(), "268435456".to_string()),
|
|
]);
|
|
|
|
assert!(contains_key_str(&metadata, SUFFIX_COMPRESSION));
|
|
assert_eq!(get_str(&metadata, SUFFIX_COMPRESSION).as_deref(), Some("klauspost/compress/s2"));
|
|
assert_eq!(get_str(&metadata, SUFFIX_ACTUAL_SIZE).as_deref(), Some("268435456"));
|
|
|
|
remove_str(&mut metadata, SUFFIX_COMPRESSION);
|
|
assert!(!contains_key_str(&metadata, SUFFIX_COMPRESSION));
|
|
assert!(!metadata.contains_key("X-Minio-Internal-compression"));
|
|
assert!(contains_key_str(&metadata, SUFFIX_ACTUAL_SIZE));
|
|
}
|
|
|
|
#[test]
|
|
fn test_str_prefers_rustfs_key_over_minio() {
|
|
let metadata = HashMap::from([
|
|
(internal_key_rustfs(SUFFIX_TRANSITION_TIER), "rustfs-tier".to_string()),
|
|
(format!("{MINIO_INTERNAL_PREFIX}{SUFFIX_TRANSITION_TIER}"), "minio-tier".to_string()),
|
|
]);
|
|
|
|
assert_eq!(get_str(&metadata, SUFFIX_TRANSITION_TIER).as_deref(), Some("rustfs-tier"));
|
|
}
|
|
|
|
#[test]
|
|
fn test_consistent_str_accepts_single_or_matching_values_and_rejects_conflicts() {
|
|
let rustfs_key = internal_key_rustfs(SUFFIX_TRANSITION_TIER_DESTINATION_ID);
|
|
let minio_key = format!("{MINIO_INTERNAL_PREFIX}{SUFFIX_TRANSITION_TIER_DESTINATION_ID}");
|
|
let mut metadata = HashMap::from([(rustfs_key, "identity-a".to_string())]);
|
|
assert_eq!(get_consistent_str(&metadata, SUFFIX_TRANSITION_TIER_DESTINATION_ID), Some("identity-a"));
|
|
|
|
metadata.insert(minio_key, "identity-a".to_string());
|
|
assert_eq!(get_consistent_str(&metadata, SUFFIX_TRANSITION_TIER_DESTINATION_ID), Some("identity-a"));
|
|
|
|
metadata.insert(
|
|
format!("{MINIO_INTERNAL_PREFIX}{SUFFIX_TRANSITION_TIER_DESTINATION_ID}"),
|
|
"identity-b".to_string(),
|
|
);
|
|
assert_eq!(get_consistent_str(&metadata, SUFFIX_TRANSITION_TIER_DESTINATION_ID), None);
|
|
}
|
|
|
|
#[test]
|
|
fn test_consistent_bytes_accepts_single_or_matching_values_and_rejects_conflicts() {
|
|
let rustfs_key = internal_key_rustfs(SUFFIX_TRANSITION_TIER_DESTINATION_ID);
|
|
let minio_key = format!("{MINIO_INTERNAL_PREFIX}{SUFFIX_TRANSITION_TIER_DESTINATION_ID}");
|
|
let mut metadata = HashMap::from([(rustfs_key, b"identity-a".to_vec())]);
|
|
assert_eq!(
|
|
get_consistent_bytes(&metadata, SUFFIX_TRANSITION_TIER_DESTINATION_ID),
|
|
Some(b"identity-a".as_slice())
|
|
);
|
|
|
|
metadata.insert(minio_key.clone(), b"identity-a".to_vec());
|
|
assert_eq!(
|
|
get_consistent_bytes(&metadata, SUFFIX_TRANSITION_TIER_DESTINATION_ID),
|
|
Some(b"identity-a".as_slice())
|
|
);
|
|
|
|
metadata.insert(minio_key, b"identity-b".to_vec());
|
|
assert_eq!(get_consistent_bytes(&metadata, SUFFIX_TRANSITION_TIER_DESTINATION_ID), None);
|
|
}
|
|
|
|
#[test]
|
|
fn test_bytes_lookup_falls_back_to_minio_key() {
|
|
let mut meta_sys =
|
|
HashMap::from([(format!("{MINIO_INTERNAL_PREFIX}{SUFFIX_TRANSITIONED_VERSION_ID}"), b"version-1".to_vec())]);
|
|
|
|
assert!(contains_key_bytes(&meta_sys, SUFFIX_TRANSITIONED_VERSION_ID));
|
|
assert_eq!(get_bytes(&meta_sys, SUFFIX_TRANSITIONED_VERSION_ID), Some(b"version-1".to_vec()));
|
|
|
|
remove_bytes(&mut meta_sys, SUFFIX_TRANSITIONED_VERSION_ID);
|
|
assert!(!contains_key_bytes(&meta_sys, SUFFIX_TRANSITIONED_VERSION_ID));
|
|
}
|
|
|
|
#[test]
|
|
fn test_bytes_prefers_rustfs_key_over_minio() {
|
|
let meta_sys = HashMap::from([
|
|
(internal_key_rustfs(SUFFIX_TRANSITIONED_VERSION_ID), b"rustfs-version".to_vec()),
|
|
(
|
|
format!("{MINIO_INTERNAL_PREFIX}{SUFFIX_TRANSITIONED_VERSION_ID}"),
|
|
b"minio-version".to_vec(),
|
|
),
|
|
]);
|
|
|
|
assert_eq!(get_bytes(&meta_sys, SUFFIX_TRANSITIONED_VERSION_ID), Some(b"rustfs-version".to_vec()));
|
|
}
|
|
|
|
// Reference implementations mirroring the original allocation-heavy logic, used to prove the
|
|
// optimized helpers are behavior-preserving across a battery of inputs.
|
|
fn is_internal_key_ref(key: &str) -> bool {
|
|
let lower = key.to_lowercase();
|
|
lower.starts_with(RUSTFS_INTERNAL_PREFIX) || lower.starts_with(MINIO_INTERNAL_PREFIX)
|
|
}
|
|
|
|
fn has_internal_suffix_ref(key: &str, suffix: &str) -> bool {
|
|
let lower = key.to_lowercase();
|
|
lower == format!("{RUSTFS_INTERNAL_PREFIX}{suffix}") || lower == format!("{MINIO_INTERNAL_PREFIX}{suffix}")
|
|
}
|
|
|
|
fn strip_internal_prefix_ref(key: &str) -> Option<String> {
|
|
let lower = key.to_lowercase();
|
|
lower
|
|
.strip_prefix(RUSTFS_INTERNAL_PREFIX)
|
|
.or_else(|| lower.strip_prefix(MINIO_INTERNAL_PREFIX))
|
|
.map(|s| s.to_string())
|
|
}
|
|
|
|
#[test]
|
|
fn test_classifiers_match_reference_impl() {
|
|
let keys = [
|
|
"x-rustfs-internal-inline-data",
|
|
"X-RustFS-Internal-Inline-Data",
|
|
"x-minio-internal-compression",
|
|
"X-MINIO-INTERNAL-actual-size",
|
|
"x-rustfs-internal-",
|
|
"x-rustfs-internal-replication-reset-arn:aws:s3:::bucket",
|
|
"x-rustfs-internal-Actual-Object-Size",
|
|
"x-amz-meta-custom",
|
|
"content-type",
|
|
"not-internal",
|
|
"",
|
|
"X",
|
|
"x-rustfs-interna", // one char short of the prefix
|
|
"x-rustfs-internal-tier-free-mar\u{212A}er", // U+212A KELVIN SIGN lowercases to ASCII 'k'
|
|
];
|
|
let suffixes = [
|
|
SUFFIX_INLINE_DATA,
|
|
SUFFIX_COMPRESSION,
|
|
SUFFIX_ACTUAL_SIZE,
|
|
SUFFIX_ACTUAL_OBJECT_SIZE_CAP,
|
|
SUFFIX_PURGESTATUS,
|
|
SUFFIX_TIER_FV_MARKER,
|
|
"inline-data",
|
|
"nonexistent",
|
|
"",
|
|
];
|
|
|
|
for key in keys {
|
|
assert_eq!(is_internal_key(key), is_internal_key_ref(key), "is_internal_key mismatch for {key:?}");
|
|
assert_eq!(
|
|
strip_internal_prefix(key),
|
|
strip_internal_prefix_ref(key),
|
|
"strip_internal_prefix mismatch for {key:?}"
|
|
);
|
|
for suffix in suffixes {
|
|
assert_eq!(
|
|
has_internal_suffix(key, suffix),
|
|
has_internal_suffix_ref(key, suffix),
|
|
"has_internal_suffix mismatch for key {key:?} suffix {suffix:?}"
|
|
);
|
|
}
|
|
}
|
|
}
|
|
|
|
#[test]
|
|
fn test_has_internal_suffix_kelvin_sign_equivalence() {
|
|
// U+212A KELVIN SIGN Unicode-lowercases to ASCII 'k'. The optimized ASCII fast path must
|
|
// fall back to full Unicode lowercasing for non-ASCII keys so it stays byte-for-byte
|
|
// equivalent to the original `key.to_lowercase()` implementation.
|
|
let key = "x-rustfs-internal-tier-free-mar\u{212A}er";
|
|
assert!(has_internal_suffix(key, SUFFIX_TIER_FV_MARKER));
|
|
assert_eq!(
|
|
has_internal_suffix(key, SUFFIX_TIER_FV_MARKER),
|
|
has_internal_suffix_ref(key, SUFFIX_TIER_FV_MARKER)
|
|
);
|
|
// A plain ASCII 'k' key still matches, and a non-matching suffix still fails.
|
|
assert!(has_internal_suffix("x-rustfs-internal-tier-free-marker", SUFFIX_TIER_FV_MARKER));
|
|
assert!(!has_internal_suffix(key, SUFFIX_COMPRESSION));
|
|
}
|
|
|
|
#[test]
|
|
fn test_get_str_case_insensitive_fallback() {
|
|
// Non-canonical mixed-case key must still be found via the case-insensitive scan.
|
|
let metadata = HashMap::from([("X-RustFS-Internal-Compression".to_string(), "s2".to_string())]);
|
|
assert_eq!(get_str(&metadata, SUFFIX_COMPRESSION).as_deref(), Some("s2"));
|
|
assert!(contains_key_str(&metadata, SUFFIX_COMPRESSION));
|
|
}
|
|
|
|
#[test]
|
|
fn test_get_bytes_no_case_insensitive_fallback() {
|
|
// get_bytes only checks the two canonical keys (matching the original behavior): a
|
|
// mixed-case key is not matched.
|
|
let meta_sys = HashMap::from([("X-RustFS-Internal-Crc".to_string(), b"z".to_vec())]);
|
|
assert_eq!(get_bytes(&meta_sys, SUFFIX_CRC), None);
|
|
|
|
let meta_sys = HashMap::from([(internal_key_rustfs(SUFFIX_CRC), b"z".to_vec())]);
|
|
assert_eq!(get_bytes(&meta_sys, SUFFIX_CRC), Some(b"z".to_vec()));
|
|
}
|
|
|
|
#[test]
|
|
fn test_long_suffix_falls_back_to_heap_key() {
|
|
// A suffix longer than the stack buffer must still round-trip via the heap fallback.
|
|
let long_suffix = "x".repeat(INTERNAL_KEY_STACK_CAP);
|
|
let mut meta_sys = HashMap::new();
|
|
insert_bytes(&mut meta_sys, &long_suffix, b"payload".to_vec());
|
|
|
|
assert!(contains_key_bytes(&meta_sys, &long_suffix));
|
|
assert_eq!(get_bytes(&meta_sys, &long_suffix), Some(b"payload".to_vec()));
|
|
remove_bytes(&mut meta_sys, &long_suffix);
|
|
assert!(!contains_key_bytes(&meta_sys, &long_suffix));
|
|
}
|
|
|
|
#[test]
|
|
fn target_delete_marker_versions_preserve_arn_case_and_report_conflicts() {
|
|
let arn = "arn:rustfs:replication::Target:Bucket";
|
|
let suffix = format!("{SUFFIX_REPLICATION_DELETE_MARKER_VERSION_ARN_PREFIX}{arn}");
|
|
let mut metadata = HashMap::new();
|
|
insert_str(&mut metadata, &suffix, "target-version".to_string());
|
|
|
|
let (versions, corrupt) = target_delete_marker_versions(&metadata);
|
|
assert_eq!(versions.get(arn).map(String::as_str), Some("target-version"));
|
|
assert!(!corrupt);
|
|
|
|
metadata.insert(format!("{MINIO_INTERNAL_PREFIX}{suffix}"), "other-version".to_string());
|
|
let (versions, corrupt) = target_delete_marker_versions(&metadata);
|
|
assert!(versions.is_empty());
|
|
assert!(corrupt);
|
|
}
|
|
|
|
#[test]
|
|
fn target_delete_marker_versions_bound_distinct_entries_during_scan() {
|
|
let metadata = (0..=1_000)
|
|
.map(|index| {
|
|
(
|
|
format!("{RUSTFS_INTERNAL_PREFIX}{SUFFIX_REPLICATION_DELETE_MARKER_VERSION_ARN_PREFIX}arn:target:{index:04}"),
|
|
format!("version-{index}"),
|
|
)
|
|
})
|
|
.collect();
|
|
|
|
let (versions, corrupt) = target_delete_marker_versions(&metadata);
|
|
|
|
assert_eq!(versions.len(), 1_000);
|
|
assert!(corrupt);
|
|
}
|
|
#[test]
|
|
fn target_delete_marker_versions_cap_is_deterministic_across_decodes() {
|
|
// Two decodes of the same oversized metadata must agree, or the two disks
|
|
// holding it hash differently and the object loses quorum instead of
|
|
// reporting corruption.
|
|
let mut metadata = HashMap::new();
|
|
for index in 0..1_050 {
|
|
metadata.insert(
|
|
format!("{RUSTFS_INTERNAL_PREFIX}{SUFFIX_REPLICATION_DELETE_MARKER_VERSION_ARN_PREFIX}arn:target:{index:05}"),
|
|
format!("version-{index}"),
|
|
);
|
|
}
|
|
|
|
let (first, first_corrupt) = target_delete_marker_versions(&metadata);
|
|
let (second, second_corrupt) = target_delete_marker_versions(&metadata);
|
|
|
|
assert!(first_corrupt, "exceeding the cap must be reported as corrupt");
|
|
assert_eq!(first_corrupt, second_corrupt);
|
|
assert_eq!(first, second, "the retained subset must not depend on map iteration order");
|
|
assert_eq!(first.len(), 1_000);
|
|
}
|
|
}
|