Two report-only defects on the heal result-reporting surface (they do not
affect the data path or heal decisions):
- default_heal_result (set_disk/ops/heal.rs): the offline-disk branch pushed
an Offline record but fell through into the unconditional push, emitting a
second (Corrupt) record for the same disk. This grew before/after.drives to
disk_count + offline_count and misaligned every entry after the first offline
slot. Add `continue` after the offline push, drive `disk_len` and the loop
from a single `self.disks` snapshot, and assert `errs.len() == disk_len`.
- Sets::heal_format (core/sets.rs): the before/after drive lists were
pre-filled with N default placeholders and then N real entries were pushed,
yielding a 2N list whose healed status updates (indexed 0..N) landed on the
blank placeholder half. Assign the lists directly from formats_to_drives_info
(mirroring the set-level heal_format) so the healed updates hit the real
entries.
Add regression tests covering offline/online record alignment and the
NoHealRequired and heal paths of the pool-level heal_format.
Co-authored-by: heihutu <heihutu@gmail.com>
* refactor(ecstore): add per-instance InstanceContext, migrate erasure setup type
Phase 5 of the global-singleton consolidation (backlog#939): begin moving
runtime identity state out of process globals so multiple ECStore instances
can coexist in one process. Isolation is carried by the object graph
(ECStore -> Sets -> SetDisks holding an Arc<InstanceContext>), not a
task-local, which does not propagate across the many internal tokio::spawn
boundaries in the data/background paths.
This first slice migrates the erasure setup type -- previously three
independent process-global bools -- into a single per-instance
RwLock<SetupType> that derives is_erasure / is_dist_erasure / is_erasure_sd,
removing a triple source of truth that could drift out of sync.
- New runtime::instance module: InstanceContext + process bootstrap context.
- The legacy free-function facade (is_erasure/update_erasure_type/...) keeps
its signatures and forwards to the current instance's context, falling back
to the bootstrap context before a store is published.
- ECStore gains a pub(crate) ctx field and setup_is_* accessors; its
constructors adopt the bootstrap context (never mint a fresh one) so startup
writes and post-construction reads share one cell -- single-instance
behavior is byte-for-byte unchanged.
Tests: erasure predicate derivation vs the legacy behavior, object-graph
carrier isolation across two ECStore instances, and bootstrap adoption.
Refs: backlog#939 (Phase 5, Slice 1), backlog#653 (item 8)
* refactor(ecstore): thread InstanceContext down the object graph (Phase 5 Slice 2) (#4415)
* refactor(ecstore): source the namespace lock manager per-instance (#4417)
refactor(ecstore): source the namespace lock manager per-instance (Phase 5 Slice 3)
Phase 5 Slice 3 (backlog#939): give each instance its own lock namespace by
sourcing SetDisks' lock manager from the instance context instead of the
process singleton. This removes the false cross-instance mutual exclusion (and
attendant ABBA risk) that a shared GlobalLockManager would cause once multiple
instances coexist.
- InstanceContext gains a `lock_manager: Arc<GlobalLockManager>`. `new()` mints
a fresh manager (independent per-instance); `bootstrap_ctx()` aliases the
process singleton via get_global_lock_manager(), so a single-instance
deployment keeps exactly one shared namespace.
- SetDisks::new sources `local_lock_manager` from `ctx.lock_manager()` (the ctx
it already adopts), not `runtime_sources::global_lock_manager()`. Single
instance: same Arc as before, so behavior is unchanged.
- Remove the now-unused `runtime_sources::global_lock_manager()` wrapper.
Tests: bootstrap lock manager aliases the process singleton; two fresh contexts
own distinct managers; a SetDisks' lock manager is the one from its context and
aliases the global singleton in a single-instance build.
Verification: cargo test -p rustfs-ecstore (10 Phase 5 + set_disk locking
regressions green), cargo clippy -p rustfs-ecstore --all-targets (clean),
make pre-commit (pass).
Refs: backlog#939 (Phase 5, Slice 3). Stacked on #4415 (Slice 2).
Second TODO-convergence round over the current tree (backlog#646). All
line numbers in the old inventory had gone stale after the set_disk /
diagnostics / cluster refactors, so this re-scans and reduces the marker
count from 144 to 99.
STALE removals (comment describes already-implemented behavior, or dead
commented-out blocks) across ecstore (set_disk ops/core, store,
cluster/rpc, bucket/metadata_sys, services), iam, filemeta, s3select and
rustfs auth/object_usecase. No behavior change.
Safe fills, each verified:
- filemeta: replication_info_equals now also compares
replication_state_internal (function currently has no callers; adds a
regression test).
- bitrot: drop the confirmed-unused `_want` parameter from bitrot_verify
and the now-unused `sum` on LocalDisk::bitrot_verify, removing a
Bytes::copy_from_slice allocation. Streaming verify uses the file's
embedded per-shard hash, never the passed sum.
- signer: rename v4_ignored_headers -> V4_IGNORED_HEADERS and drop the
non_upper_case_globals allow.
- admin/heal: test_decode was #[ignore]d and used serde_urlencoded on a
JSON body (would panic); rewire to serde_json::from_slice to match the
production decode path, add assertions, un-ignore.
Verified: cargo fmt; cargo check on touched crates; tests pass
(filemeta, signer, bitrot, heal::test_decode); arch guardrail scripts
pass.