mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-16 01:48:21 +00:00
refactor(ecstore): migrate the background-services cancel token into InstanceContext (Phase 5 Slice 13) (#4586)
* refactor(ecstore): migrate the background-services cancel token into InstanceContext (Phase 5 Slice 13) Phase 5 Slice 13 (backlog#939): move the background-services cancellation token out of the process static into the per-instance InstanceContext, so cancelling one instance's background workers (scanner/heal/tier/lifecycle) no longer touches another instance. - InstanceContext gains `background_cancel_token: OnceLock<CancellationToken>` with `init_background_cancel_token` (set-once) and `background_cancel_token()` returning an owned clone. - global.rs `init_/get_/create_/shutdown_` helpers keep their signatures and route through the current instance's context; the static is removed. The getter now returns an owned `Option<CancellationToken>` instead of a `Option<&'static _>`, which is what lets the token live in the context. - Callers adapt to the owned token: the metadata-refresh loop drops `.cloned()`; the lifecycle worker/loops take the owned token (a shared fallback is cloned when the token is somehow uninitialized). No Arc cycle is introduced — workers hold a token clone, not the instance context. Single-instance behavior is unchanged: startup creates one token in the bootstrap context the ECStore adopts, and shutdown cancels that same token. Tests: the token is set-once and cancelling one instance's token leaves a distinct instance uninitialized. Verification: cargo test -p rustfs-ecstore (23 instance-context tests green), cargo clippy -p rustfs-ecstore --all-targets (clean), make pre-commit (pass). Refs: backlog#939 (Phase 5, Slice 13) * test(ecstore): prove multi-instance isolation; document embedded guard retention (Phase 5 Slice 14) (#4588) Phase 5 (backlog#939) capstone. The prior 13 slices moved every piece of per-instance runtime state out of process globals into ECStore's InstanceContext. This slice proves the result and records the remaining work. - Add `two_instances_isolate_all_migrated_state`: an end-to-end acceptance test that constructs two independent InstanceContexts and verifies NONE of the migrated state is shared — erasure setup, lock manager, region, deployment id, the four service handles (tier/notifier/expiry/transition), the local disk registry, the bucket monitor, and the background cancel token. This is the object-graph isolation carrier working end to end. - Document why the embedded single-instance guard (EMBEDDED_SERVER_STARTED) is intentionally retained: storage startup still publishes into the process-level bootstrap context (write-once region/endpoints/deployment id) and the single GLOBAL_OBJECT_API handle, so a second startup would fail-fast on that shared state. Lifting the guard requires threading a per-instance context through startup — a follow-up beyond migrating the globals. The guard is NOT removed: rejecting the second start is safer than the panic it would otherwise become. Verification: cargo test -p rustfs-ecstore (acceptance test + all instance-context tests green), cargo clippy -p rustfs-ecstore --all-targets (clean), make pre-commit (pass). Refs: backlog#939 (Phase 5, Slice 14). Stacked on Slice 13 (#4586).
This commit is contained in:
@@ -41,6 +41,18 @@ const EVENT_SERVER_READY: &str = "server_ready";
|
||||
const EVENT_SERVER_SHUTDOWN_STATE: &str = "server_shutdown_state";
|
||||
const EVENT_EMBEDDED_SERVER_STATE: &str = "embedded_server_state";
|
||||
|
||||
/// Guards against a second embedded server starting in the same process.
|
||||
///
|
||||
/// Phase 5 (backlog#939) moved per-instance runtime state (erasure setup,
|
||||
/// region, deployment id, endpoints, service handles, disk registry, cancel
|
||||
/// token) into `ECStore`'s `InstanceContext`, so two `ECStore` object graphs
|
||||
/// can now stay isolated. This guard is intentionally **retained**: the startup
|
||||
/// path still publishes into the process-level *bootstrap* context (write-once
|
||||
/// region/endpoints/deployment id) and the single `GLOBAL_OBJECT_API` handle,
|
||||
/// so a second startup would fail-fast on that shared state. Lifting the guard
|
||||
/// requires threading a per-instance context through storage startup (so each
|
||||
/// server constructs its own context instead of sharing the bootstrap); until
|
||||
/// then, rejecting the second start is safer than the panic it would become.
|
||||
static EMBEDDED_SERVER_STARTED: AtomicBool = AtomicBool::new(false);
|
||||
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
|
||||
Reference in New Issue
Block a user