From 99f981a868191a0cd358e8081c3c02e5137dd9e4 Mon Sep 17 00:00:00 2001 From: Zhengchao An Date: Sun, 28 Jun 2026 21:55:10 +0800 Subject: [PATCH] docs: plan ecstore split boundaries (#4017) docs: plan ecstore module split boundaries --- .../architecture/ecstore-module-split-plan.md | 160 ++++++++++++++++++ .../global-state-crate-split-plan.md | 41 +++++ 2 files changed, 201 insertions(+) create mode 100644 docs/architecture/ecstore-module-split-plan.md diff --git a/docs/architecture/ecstore-module-split-plan.md b/docs/architecture/ecstore-module-split-plan.md new file mode 100644 index 000000000..490188cd0 --- /dev/null +++ b/docs/architecture/ecstore-module-split-plan.md @@ -0,0 +1,160 @@ +# ECStore Module Split Plan + +This plan records the remaining ECStore split work after the final audit +remediation pass. It is intentionally design-only: no crate movement should +start until each candidate boundary has explicit contracts, compatibility +coverage, dependency evidence, and rollback steps. + +## Current Shape + +| Area | Current owner | Size | Split status | +|---|---|---:|---| +| Bucket lifecycle | `crates/ecstore/src/bucket/lifecycle/` | 8,657 lines | Proposal only | +| Bucket replication | `crates/ecstore/src/bucket/replication/` | 8,730 lines | Proposal only | +| Set disks | `crates/ecstore/src/set_disk/` | state carrier plus operation modules | Keep in ECStore | +| Public ECStore facade | `crates/ecstore/src/api/mod.rs` | broad compatibility surface | Shrink only through guarded PRs | + +The file split inside `set_disk/` is already operation-oriented: read, write, +list, multipart, lock, heal, and replication code live in separate modules. +The remaining large surface is the shared `SetDisks` state and cross-cutting +contracts, not only file layout. + +## Non-Negotiable Rules + +- Do not split crates in the same PR that moves runtime state or changes + startup behavior. +- Do not change object placement, quorum, reader semantics, lifecycle queues, + replication queues, notification dispatch, audit events, or scanner repair + behavior during inventory and contract PRs. +- Do not expose new direct ECStore internals to outer crates; use the existing + storage-api and owner-local facade boundaries. +- Keep `rustfs_ecstore::api` compatibility visible until each consumer path has + compile coverage and an explicit replacement. + +## SetDisks Split Direction + +Do not replace `SetDisks` with several runtime structs in one change. The safe +path is: + +1. Keep `SetDisks` as the shared state carrier while operation modules continue + to own read/write/list/multipart/lock/heal/replication behavior. +2. Extract pure contracts first: shard source, disk error, bitrot IO, namespace + lock, metrics labels, and file metadata access. +3. Move one operation family only after its contracts are covered by focused + tests and the facade compatibility path is explicit. +4. Preserve the old `rustfs_ecstore::api::set_disk` surface until downstream + compatibility tests prove no caller depends on removed names. + +The first executable SetDisks follow-up should be an inventory or guardrail PR, +not a runtime split PR. + +## Lifecycle Candidate + +`bucket/lifecycle` is not ready for a standalone crate yet. + +Current coupling: + +- lifecycle workers and transition state read ECStore runtime sources for + object-store handles, expiry state, transition state, tier config, deployment + IDs, and local node names; +- stale multipart cleanup depends on `SetDisks` internals and bucket metadata; +- lifecycle expiry schedules bucket replication delete work directly; +- lifecycle evaluation shares S3 DTOs, object metadata, object lock, replication + config, scanner metrics, notification/audit side effects, and tier services. + +Required contracts before crate movement: + +- `LifecycleObjectStore`: object stat, delete, transition, restore, multipart + cleanup, and version-aware metadata operations needed by lifecycle workers. +- `LifecycleMetadataStore`: lifecycle, object-lock, replication, bucket + versioning, and stale multipart metadata lookups without importing ECStore + implementation modules. +- `LifecycleRuntime`: expiry state, transition state, tier config, deployment + ID, local node name, queue metrics, cancellation, and worker sizing. +- `LifecycleReplicationSink`: schedule lifecycle-originated replication deletes + without depending on the replication implementation module. +- `LifecycleAuditSink`: lifecycle audit and notification emission boundary. + +First safe PR: + +- add a lifecycle extraction inventory section or module-level README; +- list current ECStore/runtime dependencies and the target contract owner for + each dependency; +- add no code movement and no behavior changes. + +Focused verification for the first code-bearing lifecycle PR: + +- `cargo test -p rustfs-ecstore lifecycle --lib` +- `cargo check -p rustfs-ecstore --tests` +- `./scripts/check_architecture_migration_rules.sh` +- `git diff --check` + +## Replication Candidate + +`bucket/replication` is not ready for a standalone crate yet. + +Current coupling: + +- replication workers depend on `ReplicationStorage`, ECStore object APIs, + bucket target clients, bucket metadata, file metadata replication state, + scanner repair classification, runtime replication pool/stat handles, bucket + monitor, local node names, and notification events; +- resync and delete replication paths call metadata and bucket target systems + directly; +- lifecycle and heal paths schedule replication work through the current ECStore + module; +- global replication pool/stat initialization still lives with ECStore runtime + compatibility state. + +Required contracts before crate movement: + +- `ReplicationStorage`: keep the existing trait as the starting point, then + split object read/write/delete, walk, and metadata update responsibilities + only when call sites prove a narrower shape. +- `ReplicationMetadataStore`: replication config, bucket targets, target reset + headers, MRF/resync state, and status persistence. +- `ReplicationRuntime`: pool, stats, worker admission, bucket monitor, local + node identity, cancellation, and queue sizing. +- `ReplicationEventSink`: notification/audit events for skipped, failed, and + completed replication operations. +- `ReplicationLifecycleBridge`: lifecycle-originated delete and version-purge + scheduling without importing lifecycle internals. + +First safe PR: + +- add a replication extraction inventory section or module-level README; +- list current ECStore/runtime dependencies and the target contract owner for + each dependency; +- keep global pool/stat initialization and queue behavior unchanged. + +Focused verification for the first code-bearing replication PR: + +- `cargo test -p rustfs-ecstore replication --lib` +- `cargo check -p rustfs-ecstore --tests` +- `./scripts/check_architecture_migration_rules.sh` +- `git diff --check` + +## Facade Shrink Plan + +The broad `rustfs_ecstore::api` facade remains a compatibility boundary, not a +new architecture target. Shrinking it must be monotonic: + +1. Inventory every public facade group and consumer. +2. Add compile-time coverage before removing or narrowing a facade item. +3. Move outer consumers to storage-api or owner-local compatibility boundaries. +4. Remove one facade group per PR only after downstream compatibility tests pass. + +Do not delete facade groups only because the underlying module moved. Keep the +facade stable until the replacement path is visible and tested. + +## Ready-To-Split Checklist + +A candidate split is ready for code movement only when all items below are true: + +- dependency graph shows no cycle with ECStore, storage-api, runtime sources, or + owner-local compatibility modules; +- contract traits compile without importing ECStore implementation modules; +- old facade names have compatibility tests or explicit deprecation coverage; +- focused tests cover the changed owner path before any full gate is attempted; +- rollback preserves object IO, quorum, lifecycle/replication queues, scanner + repair, notification/audit events, and metadata compatibility. diff --git a/docs/architecture/global-state-crate-split-plan.md b/docs/architecture/global-state-crate-split-plan.md index fe058d2a7..507d199e2 100644 --- a/docs/architecture/global-state-crate-split-plan.md +++ b/docs/architecture/global-state-crate-split-plan.md @@ -82,6 +82,9 @@ to explicit AppContext ownership. `ecstore-erasure` and `storage-cluster` remain proposal-only until dependency cycles and hot-path risks are proven safe. The Phase 7 evaluation is complete for now: neither split is ready for code movement in this migration round. +The follow-up ECStore module split plan is recorded in +[`ecstore-module-split-plan.md`](ecstore-module-split-plan.md), including the +remaining `SetDisks`, lifecycle, replication, and facade-shrink boundaries. ### CRATE-001: `ecstore-erasure` @@ -142,6 +145,44 @@ Required evidence before proposing the split: - rollback plan that preserves quorum, remote disk IO, lock, peer health, and data movement behavior. +### CRATE-003: `bucket-lifecycle` + +Decision: do not split in code yet. Lifecycle remains coupled to ECStore object +operations, bucket metadata, `SetDisks` stale multipart cleanup, tier config, +runtime lifecycle state, scanner metrics, notification/audit side effects, and +replication delete scheduling. + +Required evidence before proposing the split: + +- contract sketch for lifecycle object operations, metadata access, runtime + state, replication delete scheduling, and audit/notification sinks; +- dependency graph showing the candidate crate can avoid importing ECStore + implementation modules; +- focused tests for lifecycle evaluation, expiry, transition, stale multipart + cleanup, tier journal recovery, and lifecycle-originated replication deletes; +- compatibility plan for `rustfs_ecstore::api::bucket::lifecycle` consumers; +- rollback plan that preserves lifecycle queues, scanner repair accounting, + tier transitions, object deletion behavior, and notification/audit events. + +### CRATE-004: `bucket-replication` + +Decision: do not split in code yet. Replication remains coupled to ECStore +object APIs, bucket target clients, metadata systems, file metadata replication +state, runtime replication pool/stat handles, bucket monitor state, scanner +repair classification, lifecycle-originated deletes, and notification events. + +Required evidence before proposing the split: + +- contract sketch for replication storage operations, metadata/target access, + runtime pool and stats, event sinks, and lifecycle/heal bridges; +- dependency graph showing the candidate crate can avoid importing ECStore + implementation modules; +- focused tests for object replication, delete replication, resync state, heal + repair queueing, target error handling, and queue admission; +- compatibility plan for `rustfs_ecstore::api::bucket::replication` consumers; +- rollback plan that preserves replication queues, MRF/resync state, target + client behavior, scanner repair, and event emission. + ## Preservation Rules - Do not reintroduce AppContext resolver fallback families in broad cleanup PRs.