mirror of
https://github.com/rustfs/rustfs.git
synced 2026-09-06 20:19:14 +00:00
docs(knowledge-base): prune stale content and add agent-facing index (#7035)
This commit is contained in:
+43
-52
@@ -1,66 +1,57 @@
|
||||
# Architecture Documentation
|
||||
|
||||
Durable architecture reference for RustFS: migration guardrails, runtime
|
||||
contracts, boundary rules, and support matrices.
|
||||
**Use this when:** you need the contract, invariant, or boundary rule that governs a change, and you want the one document that owns it.
|
||||
**Source of truth:** the code and the guards. `scripts/check_architecture_migration_rules.sh` enforces the CI-anchored documents below; `scripts/check_doc_paths.sh` fails the pre-commit gate when any doc under `docs/` cites a repository path that no longer exists.
|
||||
|
||||
Two rules keep this directory healthy:
|
||||
|
||||
1. **Durable reference only.** One-shot implementation plans, task trackers,
|
||||
and PR templates do not belong in the repository — keep them in the issue
|
||||
tracker or your local worktree. When their work closes, delete them rather
|
||||
than archiving them here.
|
||||
2. **No copies of other sources of truth.** Crate lists come from
|
||||
`Cargo.toml`, CI steps from `.github/workflows/ci.yml`, code structure from
|
||||
the code. `scripts/check_doc_paths.sh` fails the pre-commit gate when a
|
||||
doc here references a file path that no longer exists.
|
||||
1. **Durable reference only.** One-shot plans, task trackers, dated analyses, status snapshots, and PR-scoped notes do not belong in the repository; keep them in the issue tracker or a local worktree and delete them when the work closes.
|
||||
2. **No copies of other sources of truth.** Crate lists come from `Cargo.toml`, CI steps from `.github/workflows/`, code structure from the code. Cite a file path plus a symbol name, never a line number, and never paste counts or tables that a command can regenerate.
|
||||
|
||||
## Start here
|
||||
Every document starts with a `**Use this when:**` line so an agent can decide in one glance whether to read further. The index below repeats those lines.
|
||||
|
||||
- [overview.md](overview.md) — migration baseline, phase order, core principles
|
||||
## CI-anchored core
|
||||
|
||||
## CI-enforced core (required by `scripts/check_architecture_migration_rules.sh`)
|
||||
Required headings and strings in these files are asserted by `scripts/check_architecture_migration_rules.sh`; rename a heading only together with the guard.
|
||||
|
||||
- [crate-boundaries.md](crate-boundaries.md) — dependency direction, PR types, re-export contracts
|
||||
- [runtime-lifecycle.md](runtime-lifecycle.md) — startup/shutdown sequencing, readiness guarantees
|
||||
- [readiness-matrix.md](readiness-matrix.md) — request/dependency behavior, probe semantics
|
||||
- [storage-control-data-plane.md](storage-control-data-plane.md) — storage API contracts, control-plane boundaries
|
||||
- [global-state-crate-split-plan.md](global-state-crate-split-plan.md) — remaining global-state owners and split evaluation
|
||||
- [ecstore-module-split-plan.md](ecstore-module-split-plan.md) — ECStore decomposition rules and facade contracts
|
||||
| Document | Use this when |
|
||||
|---|---|
|
||||
| [crate-boundaries.md](crate-boundaries.md) | you add a crate dependency, move code across crates, touch a `storage_api.rs` boundary file, or need the change-type vocabulary the architecture guard enforces |
|
||||
| [runtime-lifecycle.md](runtime-lifecycle.md) | moving or reordering anything in `rustfs/src/startup_*.rs`, changing readiness publication, or touching shutdown ordering |
|
||||
| [readiness-matrix.md](readiness-matrix.md) | changing what a request surface does before storage or IAM is ready, changing probe semantics, or adding a runtime dependency that readiness must wait for |
|
||||
| [storage-control-data-plane.md](storage-control-data-plane.md) | adding a storage API surface, a cluster read model, or a background-service status/reconcile surface, and you need to know which layer owns it |
|
||||
| [global-state-crate-split-plan.md](global-state-crate-split-plan.md) | business logic needs runtime state (object store, endpoints, lock clients, lifecycle state, config) and you must pick the right boundary, or you are evaluating a crate split out of ECStore |
|
||||
| [global-state-inventory.md](global-state-inventory.md) | you meet a `GLOBAL_*` static or an `OnceLock` and need to know whether it is a runtime ownership handle, an owner-local static, or process-global by design |
|
||||
| [ecstore-module-split-plan.md](ecstore-module-split-plan.md) | you add lifecycle or replication logic and need to know which crate it belongs in, plan to move an operation family out of `SetDisks`, or the guard fails on one of the split rules |
|
||||
| [ecstore-api-facade-inventory.md](ecstore-api-facade-inventory.md) | you need something from `rustfs_ecstore` in another crate, you are narrowing a `rustfs_ecstore::api` facade group, or the guard reports a facade bypass |
|
||||
| [obs-ecstore-dependency-inventory.md](obs-ecstore-dependency-inventory.md) | adding, removing, or moving any `rustfs_ecstore` or `rustfs_storage_api` reference inside `crates/obs` |
|
||||
| [compat-cleanup-register.md](compat-cleanup-register.md) | you add, review, or remove a temporary compatibility path and need the `RUSTFS_COMPAT_TODO` marker format and its removal condition |
|
||||
| [overview.md](overview.md) | you need the historical framing of the architecture-migration program or the phase names that other contracts refer to |
|
||||
|
||||
## Contracts & invariants
|
||||
## Contracts and invariants
|
||||
|
||||
- [erasure-coding.md](erasure-coding.md) — normative erasure-coding algorithm and on-disk (`xl.meta`) compatibility contract; the frozen invariants for all user-data read/write, encode/decode, quorum, heal, and decode tolerance
|
||||
- [placement-repair-invariants.md](placement-repair-invariants.md)
|
||||
- [unified-object-generation.md](unified-object-generation.md) — single per-object generation authority (fencing epoch, transport/encoding/proto/mixed-version contracts)
|
||||
- [runtime-capability-contracts.md](runtime-capability-contracts.md)
|
||||
- [workload-admission-contracts.md](workload-admission-contracts.md)
|
||||
- [background-controller-contract.md](background-controller-contract.md)
|
||||
- [config-model-boundary-adr.md](config-model-boundary-adr.md)
|
||||
- [ecstore-layout-boundary.md](ecstore-layout-boundary.md)
|
||||
- [decommission-compatibility.md](decommission-compatibility.md)
|
||||
- [kms-bulk-rekey-contract.md](kms-bulk-rekey-contract.md) — object-side DEK re-wrap job: work unit, idempotency model, exclusion rules, and the never-destroy-old-key-versions constraint
|
||||
| Document | Use this when |
|
||||
|---|---|
|
||||
| [erasure-coding.md](erasure-coding.md) | changing anything under `crates/ecstore/src/erasure/`, `crates/filemeta/`, `crates/ecstore/src/set_disk/`, storage-class or layout code, or any decode, quorum, or heal boundary (normative spec) |
|
||||
| [placement-repair-invariants.md](placement-repair-invariants.md) | changing anything that resolves an object to a pool, set, or disk, or that admits scanner or heal work |
|
||||
| [heal-concurrency-model.md](heal-concurrency-model.md) | changing heal, PUT/multipart commit, delete, lifecycle expiry, or data-movement code that shares the `(bucket, object)` commit surface, or asking whether RustFS needs a persistent healing marker |
|
||||
| [unified-object-generation.md](unified-object-generation.md) | adding or changing anything that fences a commit, scopes a read lease, gates old-directory cleanup, binds prepared pool reads, or settles quota against the current object version |
|
||||
| [decommission-compatibility.md](decommission-compatibility.md) | changing pool decommission or rebalance behavior, its admin API shape, the persisted `PoolMeta` fields, or how tier free versions move between pools |
|
||||
| [ecstore-layout-boundary.md](ecstore-layout-boundary.md) | touching endpoint expansion, `FormatV3`, pool/set layout, or moving files between ECStore's internal directories |
|
||||
| [runtime-capability-contracts.md](runtime-capability-contracts.md) | changing the read-only observability or topology snapshot contracts in `rustfs-storage-api`, their providers, or the `storage_classes` payload of `GET /rustfs/admin/v4/runtime/capabilities` |
|
||||
| [workload-admission-contracts.md](workload-admission-contracts.md) | adding a workload class or snapshot provider, or consuming admission state from a background job |
|
||||
| [background-controller-contract.md](background-controller-contract.md) | adding a status snapshot or reconcile surface for a background service, or being tempted to fold several services into a generic controller |
|
||||
| [config-model-boundary-adr.md](config-model-boundary-adr.md) | touching the server-config model (`Config`, `KV`, `KVS`) or its persistence, or asking which crate owns which part of server configuration |
|
||||
| [admin-route-action-snapshot.md](admin-route-action-snapshot.md) | adding, moving, or re-authorizing an admin route and needing to know where the route → handler → `AdminAction` contract is enforced |
|
||||
| [kms-bulk-rekey-contract.md](kms-bulk-rekey-contract.md) | changing the bulk envelope re-wrap sweep, its admin endpoints, the re-wrap primitive, or which objects a rekey may touch |
|
||||
|
||||
## Support matrices (release-facing, keep current)
|
||||
## Support and compatibility matrices (release-facing, keep current)
|
||||
|
||||
- [s3-compatibility-matrix.md](s3-compatibility-matrix.md)
|
||||
- [s3-tables-support-matrix.md](s3-tables-support-matrix.md)
|
||||
- [minio-rustfs-router-compatibility.md](minio-rustfs-router-compatibility.md)
|
||||
- [minio-file-format-compat.md](minio-file-format-compat.md)
|
||||
| Document | Use this when |
|
||||
|---|---|
|
||||
| [s3-compatibility-matrix.md](s3-compatibility-matrix.md) | writing or checking a user-facing S3 compatibility claim, or moving a Ceph s3tests case between lists |
|
||||
| [s3-tables-support-matrix.md](s3-tables-support-matrix.md) | writing a release note or client-compatibility statement about S3 Tables / Iceberg REST Catalog (cutover procedure: [../operations/s3-tables-cutover-runbook.md](../operations/s3-tables-cutover-runbook.md)) |
|
||||
| [minio-rustfs-router-compatibility.md](minio-rustfs-router-compatibility.md) | a client or `mc` call that works against MinIO fails against RustFS and you need to know whether the endpoint is missing, stubbed, or deliberately different |
|
||||
| [minio-file-format-compat.md](minio-file-format-compat.md) | deciding whether a MinIO drive set, bucket-metadata blob, or SSE object can be read or imported by a given RustFS build, or before touching a listed version anchor |
|
||||
|
||||
## Inventories & baselines (snapshots that feed migration work)
|
||||
|
||||
- [global-state-inventory.md](global-state-inventory.md)
|
||||
- [ecstore-api-facade-inventory.md](ecstore-api-facade-inventory.md)
|
||||
- [ecstore-config-consumer-inventory.md](ecstore-config-consumer-inventory.md)
|
||||
- [obs-ecstore-dependency-inventory.md](obs-ecstore-dependency-inventory.md)
|
||||
- [background-services-inventory.md](background-services-inventory.md)
|
||||
- [scanner-heal-admission.md](scanner-heal-admission.md)
|
||||
- [admin-route-action-snapshot.md](admin-route-action-snapshot.md)
|
||||
- [compat-cleanup-register.md](compat-cleanup-register.md)
|
||||
|
||||
Historical plans and trackers (rebalance/decommission phases,
|
||||
migration-progress ledger, and the one-shot migration snapshots that fed it —
|
||||
startup timeline, scheduler baseline, profiling/NUMA capability inventory, KMS
|
||||
development defaults inventory) were retired in 2026-07 once the
|
||||
architecture-review ledger they served closed out (backlog#660/#665). Planning
|
||||
documents are no longer kept in the repository.
|
||||
Operations runbooks live in [../operations/](../README.md#operations) and testing references in [../testing/README.md](../testing/README.md).
|
||||
|
||||
@@ -1,145 +1,27 @@
|
||||
# Admin Route Action Snapshot
|
||||
|
||||
This snapshot records the current admin routing and authorization surface before
|
||||
directory moves or crate extraction. It is a migration guardrail: later pure
|
||||
move PRs must preserve the route, handler, authorization action, public
|
||||
exception, and compatibility alias semantics listed here unless the PR is
|
||||
explicitly scoped as a behavior change.
|
||||
**Use this when:** you add, move, or re-authorize an admin route and need to know where the route → handler → `AdminAction` contract is enforced.
|
||||
**Source of truth:** `rustfs/src/admin/route_policy.rs` (the `AdminRouteSpec` matrix, checked by `validate_admin_route_policy_specs`), `rustfs/src/admin/route_registration_test.rs` (registration coverage), `rustfs/src/admin/router.rs` (dispatch and credential checks), `rustfs/src/admin/handlers/*.rs` (handler-level authorization calls).
|
||||
|
||||
## Source Of Truth
|
||||
|
||||
- Router assembly: `rustfs/src/admin/mod.rs::make_admin_route`
|
||||
- Route registration coverage: `rustfs/src/admin/route_registration_test.rs`
|
||||
- Runtime dispatch: `rustfs/src/admin/router.rs`
|
||||
- Admin auth helpers: `rustfs/src/admin/auth.rs`
|
||||
- Handler route/action ownership: `rustfs/src/admin/handlers/*.rs`
|
||||
|
||||
The route registration test intentionally covers representative paths for every
|
||||
registered route family. This document uses route patterns from the registration
|
||||
functions and action names from the handler authorization calls.
|
||||
This page is a pointer, not a route table. The machine-checked matrix in `route_policy.rs` lists every admin route with its `AdminAction` and `RouteRiskLevel`; routes that are registered but answered by policy instead of a handler are declared there too through `DeferredRoutePolicyReason`. The `AdminRouteSpec` type lives in `crates/security-governance/src/admin_matrix.rs`.
|
||||
|
||||
## Prefix And Alias Contract
|
||||
|
||||
| Prefix | Current behavior | Migration rule |
|
||||
| Prefix | Behavior | Rule |
|
||||
|---|---|---|
|
||||
| `/rustfs/admin` | Canonical admin API prefix used by route registration | Keep as the single registered admin prefix |
|
||||
| `/minio/admin` | Compatibility alias accepted by `S3Router::is_match`; dispatch canonicalizes it to `/rustfs/admin` | Do not duplicate registrations; preserve canonicalization |
|
||||
| `/iceberg/v1` table catalog prefix | Registered through `table_catalog::register_table_catalog_route` and accepted by `is_admin_path` | Keep outside `/rustfs/admin` and document auth separately |
|
||||
| `/health` and `/health/ready` | Public health endpoints when `ENV_HEALTH_ENDPOINT_ENABLE` allows registration | Preserve unauthenticated health bypass |
|
||||
| `/profile/cpu` and `/profile/memory` | Registered by health handler but guarded by profile auth | Do not couple to health endpoint enablement |
|
||||
|
||||
The compatibility alias is not a second route table. `canonicalize_admin_path`
|
||||
maps `/minio/admin/...` to `/rustfs/admin/...` immediately before route lookup.
|
||||
|
||||
## Dispatch And Auth Shape
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
A["Incoming request"] --> B{"S3Router::is_match"}
|
||||
B -->|"Replication or misc extension"| X["Extension handler"]
|
||||
B -->|"Health path"| H["Public health"]
|
||||
B -->|"OIDC public path"| O["OIDC public handler"]
|
||||
B -->|"POST / STS form"| S["STS handler"]
|
||||
B -->|"Admin or console path"| C{"S3Router::check_access"}
|
||||
C -->|"public exception"| P["No SigV4 required"]
|
||||
C -->|"admin route"| D["Credential required"]
|
||||
D --> E["canonicalize /minio/admin to /rustfs/admin"]
|
||||
E --> F["matchit route lookup"]
|
||||
F --> G["AdminOperation handler"]
|
||||
G --> I["handler-level validate_admin_request"]
|
||||
```
|
||||
|
||||
Route-level credential presence and handler-level policy authorization are
|
||||
separate contracts. The router enforces credential presence for ordinary admin
|
||||
routes. Handler rows below record whether the current handler performs a
|
||||
precise `AdminAction` or `S3Action` check, or only repeats a credential
|
||||
presence check.
|
||||
| `/rustfs/admin` | Canonical admin prefix used by `make_admin_route` (`rustfs/src/admin/mod.rs`) | The only registered admin prefix |
|
||||
| `/minio/admin` | Compatibility alias accepted by `S3Router::is_match`; `canonicalize_admin_path` rewrites it to `/rustfs/admin` immediately before route lookup (`rustfs/src/admin/router.rs`) | Never register routes twice; preserve canonicalization |
|
||||
| `/iceberg/v1` | Table catalog prefix registered by `register_table_catalog_route` (`rustfs/src/admin/handlers/table_catalog/routes.rs`) and accepted by `is_admin_path` | Stays outside `/rustfs/admin`; table actions are authorized per handler |
|
||||
| `/health`, `/health/ready` | Public health endpoints, registered only when `ENV_HEALTH_ENDPOINT_ENABLE` allows | Preserve the unauthenticated bypass |
|
||||
| `/profile/cpu`, `/profile/memory` | Registered by the health handler but guarded by profile authorization | Never couple to health-endpoint enablement |
|
||||
|
||||
## Public Exceptions
|
||||
|
||||
| Method | Path pattern | Handler | Auth contract |
|
||||
|---|---|---|---|
|
||||
| `GET`, `HEAD` | `/health` | `HealthCheckHandler` | Public when health routes are registered |
|
||||
| `GET`, `HEAD` | `/health/ready` | `HealthCheckHandler` | Public when health routes are registered |
|
||||
| Registered as `GET`; auth bypass is path-based | `/rustfs/admin/v3/oidc/providers` and `/minio/admin/v3/oidc/providers` | `ListOidcProvidersHandler` | Public OIDC bootstrap path; `check_access` bypasses SigV4 for any method matching this path |
|
||||
| Registered as `GET`; auth bypass is path-prefix-based | `/rustfs/admin/v3/oidc/authorize/{provider_id}` and `/minio/admin/v3/oidc/authorize/{provider_id}` | `OidcAuthorizeHandler` | Public OIDC bootstrap path; `check_access` bypasses SigV4 for any method matching this path prefix |
|
||||
| Registered as `GET`; auth bypass is path-prefix-based | `/rustfs/admin/v3/oidc/callback/{provider_id}` and `/minio/admin/v3/oidc/callback/{provider_id}` | `OidcCallbackHandler` | Public OIDC bootstrap path; `check_access` bypasses SigV4 for any method matching this path prefix |
|
||||
| Registered as `GET`; auth bypass is path-based | `/rustfs/admin/v3/oidc/logout` and `/minio/admin/v3/oidc/logout` | `OidcLogoutHandler` | Public OIDC logout path; `check_access` bypasses SigV4 for any method matching this path |
|
||||
| `POST` | `/` with `application/x-www-form-urlencoded` | `AssumeRoleHandle` | Public only for unsigned STS web identity form requests; handler validates JWT/action |
|
||||
| Any matched method | `/favicon.ico` and `/rustfs/console...` | Console router | Public only when `console_enabled` is true; router bypasses SigV4 before handing off to the console router |
|
||||
Router-level credential checks (`S3Router::check_access`) are bypassed only for:
|
||||
|
||||
## Registered Route Families
|
||||
- health routes, when they are registered;
|
||||
- OIDC bootstrap paths matched by `is_oidc_path` (`providers`, `authorize/{provider_id}`, `callback/{provider_id}`, `logout`); the bypass is path-based, so it applies to any method on those paths;
|
||||
- unsigned STS web-identity form posts to `/` with `application/x-www-form-urlencoded`, which the STS handler validates itself;
|
||||
- console assets (`/favicon.ico`, `/rustfs/console...`), only while the console is enabled.
|
||||
|
||||
All rows with `/rustfs/admin` also accept the `/minio/admin` compatibility alias
|
||||
through router canonicalization unless the row explicitly says otherwise.
|
||||
|
||||
| Area | Methods and path patterns | Handler ownership | Authorization contract |
|
||||
|---|---|---|---|
|
||||
| STS and admin probe | `POST /`; `GET /rustfs/admin/v3/is-admin` | `sts.rs`, `is_admin.rs` | STS dispatch validates request action; is-admin checks `AllAdminActions` |
|
||||
| User lifecycle | `GET /v3/list-users`; `GET /v3/user-info`; `PUT /v3/add-user`; `PUT /v3/set-user-status`; `DELETE /v3/remove-user` | `user_lifecycle.rs`, `user.rs` | `ListUsersAdminAction`, `GetUserAdminAction`, `CreateUserAdminAction`, `EnableUserAdminAction`, `DeleteUserAdminAction` |
|
||||
| Group management | `GET /v3/groups`; `GET /v3/group`; `DELETE /v3/group/{group}`; `PUT /v3/set-group-status`; `PUT /v3/update-group-members` | `group.rs` | `ListGroupsAdminAction`, `GetGroupAdminAction`, `RemoveUserFromGroupAdminAction`, `EnableGroupAdminAction`, `AddUserToGroupAdminAction` |
|
||||
| Service accounts | `PUT /v3/add-service-account(s)`; `POST /v3/update-service-account`; `GET /v3/info-service-account`; `GET /v3/temporary-account-info`; `GET /v3/info-access-key`; `GET /v3/list-service-accounts`; `GET /v3/list-access-keys-bulk`; `DELETE /v3/delete-service-account(s)` | `service_account.rs` | create/update/list/temp-info/user-list/remove service account actions as checked in handler context |
|
||||
| IAM import/export | `GET /v3/export-iam`; `PUT /v3/import-iam` | `user_iam.rs`, `user.rs` | `ExportIAMAction`, `ImportIAMAction` |
|
||||
| IAM policies | `GET /v3/list-canned-policies`; `GET /v3/info-canned-policy`; `PUT /v3/add-canned-policy`; `DELETE /v3/remove-canned-policy`; `PUT /v3/set-user-or-group-policy`; `PUT /v3/set-policy`; `POST /v3/idp/builtin/policy/attach`; `POST /v3/idp/builtin/policy/detach`; `GET /v3/idp/builtin/policy-entities` | `policies.rs` | list/create/get/delete/attach policy actions; policy-entities combines list groups, users, and policies |
|
||||
| Account info | `GET /v3/accountinfo` | `account_info.rs` | S3 action checks for account-scoped bucket and object probes |
|
||||
| System info | `GET /v3/info`; `GET /v3/storageinfo`; `GET /v3/datausageinfo` | `system.rs` | `ServerInfoAdminAction`, `StorageInfoAdminAction`, `DataUsageInfoAdminAction` plus `ListBucketAction` for data usage |
|
||||
| Metrics stream | `GET /v3/metrics` | `metrics.rs` through `system.rs` | Router credential presence plus handler credential check; no handler-level `AdminAction` is currently enforced |
|
||||
| System service placeholders | `POST /v3/service`; `GET|POST /v3/inspect-data` | `system.rs` | Currently registered but handler returns `NotImplemented`; migration must preserve this unless behavior changes |
|
||||
| Pools | `GET /v3/pools/list`; `GET /v3/pools/status`; `POST /v3/pools/decommission`; `POST /v3/pools/cancel` | `pools.rs` | list/status accept server-info or decommission; decommission/cancel use `DecommissionAdminAction` |
|
||||
| Rebalance | `POST /v3/rebalance/start`; `GET /v3/rebalance/status`; `POST /v3/rebalance/stop` | `rebalance.rs` | `RebalanceAdminAction` |
|
||||
| Heal | `POST /v3/heal/`; `POST /v3/heal/{bucket}`; `POST /v3/heal/{bucket}/{prefix}`; `POST /v3/background-heal/status`; `GET /v4/heal/replacement-recovery` | `heal.rs` | `HealAdminAction` |
|
||||
| Tier | `GET /v3/tier`; `GET /v3/tier-stats`; `GET /v3/tier/{tier}`; `DELETE /v3/tier/{tiername}`; `PUT /v3/tier`; `POST /v3/tier/{tiername}`; `POST /v3/tier/clear` | `tier.rs` | `ListTierAction` for reads/status; `SetTierAction` for add/edit/remove/clear |
|
||||
| Quota legacy and bucket-scoped | `PUT /v3/set-bucket-quota`; `GET /v3/get-bucket-quota`; `PUT|GET|DELETE /v3/quota/{bucket}`; `GET /v3/quota-stats/{bucket}`; `POST /v3/quota-check/{bucket}` | `quota.rs` | `SetBucketQuotaAdminAction` for writes; `GetBucketQuotaAction` for bucket-scoped reads/stats/checks |
|
||||
| Bucket metadata | `GET /export-bucket-metadata`; `GET /v3/export-bucket-metadata`; `PUT /import-bucket-metadata`; `PUT /v3/import-bucket-metadata` | `bucket_meta.rs` | `ExportBucketMetadataAction`, `ImportBucketMetadataAction` |
|
||||
| Server config | `GET /v3/get-config-kv`; `PUT /v3/set-config-kv`; `DELETE /v3/del-config-kv`; `GET /v3/help-config-kv`; `GET /v3/list-config-history-kv`; `DELETE /v3/clear-config-history-kv`; `PUT /v3/restore-config-history-kv`; `GET|PUT /v3/config` | `config_admin.rs` | `ConfigUpdateAdminAction` helper path; read/write handlers preserve current per-handler checks |
|
||||
| Scanner | `GET /v3/scanner/status` | `scanner.rs` | `ServerInfoAdminAction` |
|
||||
| Notification targets | `GET /v3/target/list`; `GET /v3/target/arns`; `PUT /v3/target/{target_type}/{target_name}`; `DELETE /v3/target/{target_type}/{target_name}/reset` | `event.rs` through `user_policy_binding.rs` | `GetBucketTargetAction` for list/ARNs; `SetBucketTargetAction` for put/delete |
|
||||
| Audit targets | `GET /v3/audit/target/list`; `PUT /v3/audit/target/{target_type}/{target_name}`; `DELETE /v3/audit/target/{target_type}/{target_name}/reset` | `audit.rs` | `GetBucketTargetAction` for list; `SetBucketTargetAction` for put/delete |
|
||||
| Module switches | `GET|PUT /v3/module-switches` | `module_switch.rs` | `ServerInfoAdminAction` for get; `ConfigUpdateAdminAction` for update |
|
||||
| Plugin catalog | `GET /v4/plugins/catalog` | `plugins_catalog.rs` | `ServerInfoAdminAction` |
|
||||
| Plugin instances | `GET /v4/plugins/instances`; `GET|PUT|DELETE /v4/plugins/instances/{id}` | `plugins_instances.rs` | read uses `GetBucketTargetAction`; write/delete use `SetBucketTargetAction` |
|
||||
| Replication target list | `GET /v3/list-remote-targets` | `replication.rs` | Router credential presence plus handler credential check; no handler-level `AdminAction` is currently enforced |
|
||||
| Replication target metrics/mutation | `GET /v3/replicationmetrics`; `PUT /v3/set-remote-target`; `DELETE /v3/remove-remote-target` | `replication.rs` | `GetReplicationMetricsAction` for metrics; `SetBucketTargetAction` for target mutation |
|
||||
| Site replication | `PUT /v3/site-replication/add`; `PUT /v3/site-replication/remove`; `GET /v3/site-replication/info`; `GET /v3/site-replication/metainfo`; `GET /v3/site-replication/status`; `POST /v3/site-replication/devnull`; `POST /v3/site-replication/netperf`; `PUT /v3/site-replication/edit`; `PUT /v3/site-replication/peer/join`; `PUT /v3/site-replication/peer/bucket-ops`; `PUT /v3/site-replication/peer/iam-item`; `PUT /v3/site-replication/peer/bucket-meta`; `GET /v3/site-replication/peer/idp-settings`; `PUT /v3/site-replication/peer/edit`; `PUT /v3/site-replication/peer/remove`; `PUT /v3/site-replication/resync/op`; `PUT /v3/site-replication/state/edit` | `site_replication.rs` | add/remove/info/operation/resync actions selected per handler |
|
||||
| Admin profiling | `GET /rustfs/admin/debug/pprof/profile`; `GET /rustfs/admin/debug/pprof/status` | `profile_admin.rs`, `profile.rs` | `ProfilingAdminAction` |
|
||||
| TLS debug | `GET /rustfs/admin/debug/tls/status` | `tls_debug.rs`, `profile.rs` | `ProfilingAdminAction` via shared profile authorization |
|
||||
| KMS legacy management | `POST /v3/kms/create-key`; `POST /v3/kms/key/create`; `GET /v3/kms/describe-key`; `GET /v3/kms/key/status`; `GET /v3/kms/list-keys`; `POST /v3/kms/generate-data-key`; `GET|POST /v3/kms/status`; `GET /v3/kms/config`; `POST /v3/kms/clear-cache` | `kms_management.rs`, `kms_keys.rs` | dedicated `kms:*` actions throughout; `kms:ServiceControl` for the status paths, `kms:Configure` for config, `kms:ClearCache` for cache. No `ServerInfoAdminAction` fallback remains on any KMS route |
|
||||
| KMS dynamic control | `POST /v3/kms/configure`; `POST /v3/kms/start`; `POST /v3/kms/stop`; `GET /v3/kms/service-status`; `POST /v3/kms/reconfigure` | `kms_dynamic.rs` | `kms:Configure` for configure/reconfigure; `kms:ServiceControl` for start/stop/service-status |
|
||||
| KMS keys | `POST /v3/kms/keys`; `DELETE /v3/kms/keys/delete`; `POST /v3/kms/keys/cancel-deletion`; `GET /v3/kms/keys`; `GET /v3/kms/keys/{key_id}` | `kms_keys.rs` | dedicated `kms:*` actions per handler |
|
||||
| OIDC public | `GET /v3/oidc/providers`; `GET /v3/oidc/authorize/{provider_id}`; `GET /v3/oidc/callback/{provider_id}`; `GET /v3/oidc/logout` | `oidc.rs` | Public OIDC exception in `is_oidc_path` |
|
||||
| OIDC config | `GET /v3/oidc/config`; `PUT|DELETE /v3/oidc/config/{provider_id}`; `POST /v3/oidc/validate` | `oidc.rs` | `ServerInfoAdminAction` for read/validate; `ConfigUpdateAdminAction` for mutation |
|
||||
|
||||
## Table Catalog Routes
|
||||
|
||||
The table catalog API is registered by the admin router but is not under
|
||||
`/rustfs/admin`. It has its own prefix and Iceberg-style route shape.
|
||||
|
||||
| Method | Path pattern | Handler | Authorization action |
|
||||
|---|---|---|---|
|
||||
| `GET` | `/iceberg/v1/config` | `GET_CONFIG_HANDLER` | `GetTableCatalogAction` |
|
||||
| `GET` | `/iceberg/v1/{warehouse}/namespaces` | `LIST_NAMESPACES_HANDLER` | `GetTableNamespaceAction` |
|
||||
| `POST` | `/iceberg/v1/{warehouse}/namespaces` | `CREATE_NAMESPACE_HANDLER` | `SetTableNamespaceAction` |
|
||||
| `GET` | `/iceberg/v1/{warehouse}/namespaces/{namespace}` | `GET_NAMESPACE_HANDLER` | `GetTableNamespaceAction` |
|
||||
| `DELETE` | `/iceberg/v1/{warehouse}/namespaces/{namespace}` | `DROP_NAMESPACE_HANDLER` | `DeleteTableNamespaceAction` |
|
||||
| `GET` | `/iceberg/v1/{warehouse}/namespaces/{namespace}/tables` | `LIST_TABLES_HANDLER` | `GetTableAction` |
|
||||
| `POST` | `/iceberg/v1/{warehouse}/namespaces/{namespace}/tables` | `CREATE_TABLE_HANDLER` | `CreateTableAction` |
|
||||
| `POST` | `/iceberg/v1/{warehouse}/namespaces/{namespace}/register` | `REGISTER_TABLE_HANDLER` | `RegisterTableAction` |
|
||||
| `GET` | `/iceberg/v1/{warehouse}/namespaces/{namespace}/tables/{table}` | `LOAD_TABLE_HANDLER` | `GetTableAction` |
|
||||
| `POST` | `/iceberg/v1/{warehouse}/namespaces/{namespace}/tables/{table}` | `COMMIT_TABLE_HANDLER` | `CommitTableAction` |
|
||||
| `DELETE` | `/iceberg/v1/{warehouse}/namespaces/{namespace}/tables/{table}` | `DROP_TABLE_HANDLER` | `DeleteTableAction` |
|
||||
|
||||
## Migration Rules
|
||||
|
||||
1. Pure move PRs may move handler modules, but must not change registered
|
||||
methods, patterns, handler ownership, alias canonicalization, or public
|
||||
exception behavior.
|
||||
2. If an admin handler is wrapped to cut a dependency direction, the wrapper
|
||||
must preserve the same `AdminAction` or `S3Action` check and keep response
|
||||
compatibility unchanged.
|
||||
3. Do not duplicate `/minio/admin` registrations. The alias remains a router
|
||||
canonicalization concern.
|
||||
4. Do not move table catalog routes under `/rustfs/admin` during route cleanup.
|
||||
5. Registered-but-`NotImplemented` routes are behavior contracts too. Removing
|
||||
or implementing them requires a behavior-change PR type.
|
||||
6. Future route matrix automation should compare against this document and
|
||||
`route_registration_test.rs` before crate extraction begins.
|
||||
Every other admin route requires credentials at the router and a precise `AdminAction` or `S3Action` check in the handler (metrics routes, for example, authorize `GetMetricsAction`). The MinIO alias contract is specified in [minio-rustfs-router-compatibility.md](minio-rustfs-router-compatibility.md).
|
||||
|
||||
@@ -1,187 +1,55 @@
|
||||
# Background Controller Contract
|
||||
|
||||
This document defines `BGC-002` for
|
||||
[`rustfs/backlog#660`](https://github.com/rustfs/backlog/issues/660). It turns
|
||||
the background service inventory into a shared vocabulary for future read-only
|
||||
status work. It does not add a Rust trait, a scheduler, a service registry, or
|
||||
any worker start/stop behavior.
|
||||
**Use this when:** you add a status snapshot or reconcile surface for a background service (scanner, heal, lifecycle, replication, config reload, capacity, metrics, memory observability, allocator reclaim, auto-tuner), or you are tempted to fold several of them into a generic controller.
|
||||
**Source of truth:** the shipped reference surfaces — `MemoryObservabilityReconcilePlan` and `reconcile()` in `rustfs/src/memory_observability.rs`, `AllocatorReclaimControllerSnapshot` and `AllocatorReclaimReconcilePlan` in `rustfs/src/allocator_reclaim.rs`, `MetricsRuntimeReconcilePlan` in `crates/obs/src/metrics/scheduler.rs`. Startup and shutdown ordering is owned by [runtime-lifecycle.md](runtime-lifecycle.md); the plane-level overview is in [storage-control-data-plane.md](storage-control-data-plane.md).
|
||||
|
||||
## Scope
|
||||
There is no `BackgroundController` trait, scheduler, or service registry. Each service exposes its own typed snapshot and reconcile plan; this page fixes the vocabulary and the rules those surfaces follow.
|
||||
|
||||
- PR type: `docs-only`.
|
||||
- Baseline: `upstream/main` at
|
||||
`f9a5e6d7e67322ac6f626b6f437a5e722fbe22e2`.
|
||||
- Applies to future controller work for scanner, heal, lifecycle, replication,
|
||||
dynamic config reload, capacity, metrics, memory observability, allocator
|
||||
reclaim, and auto-tuning.
|
||||
- Out of scope: worker creation, worker shutdown, queue resizing, storage
|
||||
writes, readiness changes, peer signaling changes, scheduler replacement, and
|
||||
crate splitting.
|
||||
## Vocabulary
|
||||
|
||||
## Contract Vocabulary
|
||||
|
||||
| Term | Meaning | BGC-002 boundary |
|
||||
| Term | Meaning | Boundary |
|
||||
|---|---|---|
|
||||
| Desired | Static intent from env, persisted config, module switches, feature flags, bucket config, or admin configuration. | Read only. Do not normalize or mutate config while collecting desired state. |
|
||||
| Current | Observed local runtime state such as configured, disabled, running, degraded, stopping, or unknown. | Read only. Do not infer state by starting probes that create storage or network side effects. |
|
||||
| Status | Human-readable and machine-checkable snapshot of runtime counters, worker counts, queue pressure, last successful cycle, last error, cancellation source, and shutdown handle shape. | Side-effect-free. Missing status surfaces must be reported as `unknown`, not guessed. |
|
||||
| Reconcile | Future comparison between desired, current, and status that can produce a recommendation. | No action in `BGC-002`; future reconcile must not start or stop workers until a tested pilot PR allows it. |
|
||||
| Side effects | Writes, deletes, queue admission, target activation, external I/O, metrics emission, readiness publication, peer signal, or config reload fanout. | Must be declared before any controller migration touches that service. |
|
||||
| Desired | Static intent from env, persisted config, module switches, feature flags, bucket config, or admin configuration. | Read only; collecting desired state never normalizes or mutates config. |
|
||||
| Current | Observed local runtime state: configured, disabled, running, degraded, stopping, or unknown. | Read only; never inferred by probes that create storage or network side effects. |
|
||||
| Status | Machine-checkable snapshot of counters, worker counts, queue pressure, last cycle, last error, cancellation source, and shutdown-handle shape. | Side-effect-free; a missing surface is reported as `unknown`, never guessed. |
|
||||
| Reconcile | Comparison of desired, current, and status that yields a plan. | Shipped plans only report; the only worker mutation they may request is `none`. |
|
||||
| Side effects | Writes, deletes, queue admission, target activation, external I/O, metrics emission, readiness publication, peer signals, config reload fanout. | Declared per service before any controller touches it. |
|
||||
|
||||
## State Model
|
||||
|
||||
Future status snapshots should use the narrowest state that the current code can
|
||||
prove:
|
||||
Snapshots use the narrowest state the code can prove:
|
||||
|
||||
| State | Meaning | Notes |
|
||||
|---|---|---|
|
||||
| NotConfigured | No valid desired source exists for this service. | Use when config/module switches/features make the service absent. |
|
||||
| Disabled | Desired source exists and explicitly disables the service. | Do not use for missing config. |
|
||||
| Starting | Startup was requested and has not reached steady state. | Only expose when current code has a start boundary. |
|
||||
| Running | The service is active according to existing runtime state. | Do not use merely because config is enabled. |
|
||||
| Degraded | The service is active but current status exposes known error, partial, or stalled state. | Do not introduce new failure classification in docs-only work. |
|
||||
| Stopping | Shutdown was requested and the service has not fully exited. | Only expose where shutdown can be observed. |
|
||||
| Stopped | The service was started before and is now fully stopped. | Do not confuse with `Disabled` or `NotConfigured`. |
|
||||
| Unknown | Current code lacks a safe status surface. | Preferred over speculative status. |
|
||||
|
||||
## Lifecycle Boundary
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
D["Desired source"]
|
||||
C["Current runtime state"]
|
||||
S["Read-only status snapshot"]
|
||||
R["Future reconcile recommendation"]
|
||||
W["Workers and side effects"]
|
||||
|
||||
D --> S
|
||||
C --> S
|
||||
S --> R
|
||||
R -. "future tested pilot only" .-> W
|
||||
```
|
||||
|
||||
`BGC-002` stops at the read-only contract. The arrow from reconcile to workers is
|
||||
intentionally dotted because this PR does not allow any implementation to start,
|
||||
stop, resize, or reconfigure workers.
|
||||
|
||||
## Service Boundaries
|
||||
|
||||
| Service area | Desired source | Current/status inputs | Side effects to preserve |
|
||||
|---|---|---|---|
|
||||
| Data scanner | Scanner env and runtime scanner config. | Admin scanner status, scanner metrics, scanner cancellation token, checkpoint/yield/alert counters. | Data usage cache updates, lifecycle evaluation, replication heal admission, scanner heal admission, alerts, and scanner metrics. |
|
||||
| Heal/AHM | Heal enablement and scanner-driven heal admission. | Heal manager global channel, active task atomics, queue length atomics, AHM cancellation token. | Heal queue consumption, heal storage writes, and channel close semantics. |
|
||||
| Lifecycle expiry/transition | Bucket lifecycle config and scanner event source. | Lifecycle worker counts, active tasks, queue send timeouts, transition stats, expiry/transition queues. | Object deletes, transition queueing, stale multipart cleanup, and lifecycle metrics. |
|
||||
| Replication pool | Bucket/site replication config and resync admin requests. | Global replication stats, worker pool sizes, queue counters, persisted resync state, per-bucket cancel tokens. | Object replication, delete replication, queue resizing by channel close, persisted resync metadata, and admin-triggered cancel paths. |
|
||||
| Dynamic config reload | Persisted server config, admin config calls, and peer snapshot signals. | Last local reload result, per-subsystem reload errors, peer reload signal result. | Scanner/heal runtime config updates, audit reload, notification reload, peer signaling, and config snapshot fanout. |
|
||||
| Capacity manager | Local disk inventory and capacity feature state. | Capacity manager cache age, scheduled refresh state, last refresh result, runtime summary loop. | Global capacity cache refresh and runtime summary metrics/logging. |
|
||||
| Metrics runtime | Observability metrics feature state and collector configuration. | Collector intervals, last collection result, cancellation token state, collector grouping. | Metrics collection and emission only. |
|
||||
| Memory observability | Observability feature state and memory sampling config. | Sampler loop state, last sample time, last sample error, runtime cancellation token. | Memory metric emission. This is the preferred first BGC-003 status candidate. |
|
||||
| Allocator reclaim | Allocator reclaim env/config and backend support. | Enabled flag, idle streak, active request gauge, scanner/heal activity gauges, last reclaim result. | Backend-specific allocator reclaim and metrics. |
|
||||
| Auto-tuner | `RUSTFS_AUTOTUNER_ENABLED` and tuning inputs. | Last tuning attempt, last tuning error, 60-second loop state. | Runtime concurrency tuning. Treat as behavior-sensitive. |
|
||||
|
||||
The following areas stay outside the first controller migrations:
|
||||
|
||||
- deferred IAM recovery, because it can publish readiness;
|
||||
- optional protocol servers, because they already have protocol shutdown handles;
|
||||
- ECStore endpoint monitor and disk health monitor, because they are storage-
|
||||
adjacent and can affect disk state;
|
||||
- notification and audit runtime coupling, because live streams, replay, target
|
||||
activation, and reload behavior need dedicated preservation tests.
|
||||
| NotConfigured | No valid desired source exists. | Config, module switches, or features make the service absent. |
|
||||
| Disabled | A desired source exists and explicitly disables the service. | Not for missing config. |
|
||||
| Starting | Start requested, steady state not reached. | Only where a start boundary exists. |
|
||||
| Running | Active according to existing runtime state. | Not merely because config is enabled. |
|
||||
| Degraded | Active with known error, partial, or stalled status. | No new failure classification is invented for a snapshot. |
|
||||
| Stopping | Shutdown requested, not fully exited. | Only where shutdown is observable. |
|
||||
| Stopped | Started earlier, now fully stopped. | Distinct from `Disabled` and `NotConfigured`. |
|
||||
| Unknown | No safe status surface exists. | Preferred over speculation. |
|
||||
|
||||
## Read-Only Snapshot Requirements
|
||||
|
||||
Any future `BGC-003` status implementation must satisfy all of these:
|
||||
- Status collection never starts, stops, resizes, or wakes a worker.
|
||||
- Status collection never writes storage data, object metadata, target state, queue entries, persisted config, or resync metadata.
|
||||
- Status collection never publishes readiness or peer reload signals.
|
||||
- Missing fields are `unknown` or omitted with a documented reason.
|
||||
- Cancellation source and shutdown-handle shape are reported separately from desired enabled/disabled state.
|
||||
- Repeated `reconcile` calls over the same snapshot return the same plan.
|
||||
- Scanner, heal, lifecycle, and replication status must not hide their queue and admission coupling.
|
||||
|
||||
- status collection must not start, stop, resize, or wake a worker;
|
||||
- status collection must not write storage data, object metadata, target state,
|
||||
queue entries, persisted config, or resync metadata;
|
||||
- status collection must not publish readiness or peer reload signals;
|
||||
- missing fields must be represented as `unknown` or omitted with a documented
|
||||
reason;
|
||||
- cancellation source and shutdown handle shape must be reported separately from
|
||||
desired enabled/disabled state;
|
||||
- scanner, heal, lifecycle, and replication status must not hide their queue and
|
||||
admission coupling.
|
||||
## Coupling Notes
|
||||
|
||||
## BGC-003 Snapshot Pilot
|
||||
The services below share state or shutdown contracts and must not be folded into a generic controller without service-specific preservation tests:
|
||||
|
||||
The first read-only snapshot is memory observability status. It reports the
|
||||
service name, whether observability metrics currently enable the sampler, the
|
||||
configured sampler interval, runtime-token cancellation state, and the absence
|
||||
of a dedicated shutdown handle.
|
||||
|
||||
This snapshot intentionally does not define an admin route, scheduler, service
|
||||
registry, worker start/stop path, readiness signal, peer signal, storage write,
|
||||
or metrics emission change.
|
||||
|
||||
## BGC-004 Controller Pilot
|
||||
|
||||
The first controller pilot is also memory observability. It converts the
|
||||
existing desired inputs and status snapshot into a typed reconcile plan. The
|
||||
pilot reports desired state, current state, and worker mutation intent.
|
||||
|
||||
The only allowed worker mutation for this pilot is `none`. Repeated reconcile
|
||||
calls must return the same plan for the same snapshot and must not request a
|
||||
worker start, stop, resize, wakeup, storage write, readiness signal, peer
|
||||
signal, or metrics emission.
|
||||
|
||||
## BGC-005 Allocator Reclaim Status And Controller Surface
|
||||
|
||||
The second low-risk controller/status surface is allocator reclaim. It reports
|
||||
the service name, desired enablement, configured force flag, backend-specific
|
||||
effective force, idle interval settings, runtime-token cancellation state, and
|
||||
the absence of a dedicated shutdown handle.
|
||||
|
||||
The only allowed worker mutation for this surface is `none`. Reconcile output is
|
||||
read-only and must not start, stop, resize, wake, or otherwise drive the
|
||||
allocator reclaim loop. Existing backend-specific force handling, idle-streak
|
||||
logic, metrics emission, and runtime-token shutdown behavior remain owned by the
|
||||
current loop.
|
||||
|
||||
## BGC-006 Metrics Runtime Status And Controller Surface
|
||||
|
||||
The third low-risk controller/status surface is metrics runtime. It reports the
|
||||
service name, observability metrics enablement, collector task count, configured
|
||||
collector intervals, replication bandwidth zero-tombstone cycle count,
|
||||
runtime-token cancellation state, and the absence of a dedicated shutdown
|
||||
handle.
|
||||
|
||||
The only allowed worker mutation for this surface is `none`. Reconcile output is
|
||||
read-only and must not start, stop, resize, wake, or otherwise drive metrics
|
||||
collector tasks. Existing collector grouping, interval parsing, metrics
|
||||
emission, replication bandwidth tombstone handling, and runtime-token shutdown
|
||||
behavior remain owned by the current loops.
|
||||
|
||||
## Future Reconcile Rules
|
||||
|
||||
Future reconcile work is allowed only after a read-only status snapshot exists.
|
||||
The first reconcile pilot must:
|
||||
|
||||
- choose one low-risk service;
|
||||
- compare desired/current/status without side effects;
|
||||
- prove idempotence under repeated calls;
|
||||
- prove no duplicate workers are created;
|
||||
- preserve existing shutdown order and cancellation source;
|
||||
- include rollback guidance that removes the pilot without changing existing
|
||||
worker behavior.
|
||||
|
||||
Memory observability is the recommended first candidate because it already has a
|
||||
simple runtime cancellation loop and no storage writes. Scanner, heal,
|
||||
replication, lifecycle, disk health, deferred IAM recovery, and auto-tuning must
|
||||
wait for focused preservation tests.
|
||||
|
||||
## Verification Expectations
|
||||
|
||||
For this docs-only contract:
|
||||
|
||||
- architecture migration guard scripts must pass;
|
||||
- layer dependency and metrics reference guards must pass;
|
||||
- no Rust source, Cargo metadata, CI workflow, Makefile, or runtime config file
|
||||
may change.
|
||||
|
||||
For the next implementation PRs:
|
||||
|
||||
- add focused tests before changing behavior;
|
||||
- do not modify production logic only to make tests pass;
|
||||
- keep compatibility comments searchable with `RUSTFS_COMPAT_TODO(<task-id>)`
|
||||
whenever temporary old paths are retained for later deletion.
|
||||
- Scanner implies heal: the loop started by `init_data_scanner` (`rustfs/src/startup_lifecycle.rs`) enqueues heal work, so scanner status must separate scheduler state from work-source accounting.
|
||||
- Heal/AHM owns its own token: `create_ahm_services_cancel_token` and `init_heal_manager` run in `rustfs/src/startup_background.rs`; `shutdown_ahm_services` runs in `rustfs/src/startup_shutdown.rs`. Heal admission and channel-close semantics stay intact.
|
||||
- Replication has two shutdown contracts: the pool started by `init_background_replication` (`rustfs/src/startup_storage.rs`) stops workers by closing channels, while resync started by `init_resync` (`rustfs/src/startup_bucket_metadata.rs`) uses cancellation tokens, and admin-triggered resync uses per-bucket tokens.
|
||||
- Lifecycle expiry, transition, and stale-multipart cleanup are started by `ECStore::init` (`init_background_expiry`, `init_background_stale_multipart_upload_cleanup` in `crates/ecstore/src/store/init.rs`), which binds the runtime token through `bind_background_cancel_token`; the scanner is their event source, so they are not a separate periodic controller.
|
||||
- Notification and audit share a runtime pattern but not a lifecycle: `init_event_notifier` and `start_audit_system` (`rustfs/src/startup_audit.rs`), `shutdown_event_notifier` and `stop_audit_system` (`rustfs/src/startup_shutdown.rs`). Live event streams stay separate from target-delivery enablement.
|
||||
- Dynamic config reload is admin-triggered fanout (`apply_dynamic_config_for_subsystem`, `signal_dynamic_config_reload`, `signal_config_snapshot_reload` in `rustfs/src/admin/service/config.rs`), not a loop; per-subsystem validation and error boundaries are preserved.
|
||||
- Capacity refresh tasks are owned through `CapacityBackgroundTasks` returned by `init_capacity_management_managed` (`rustfs/src/capacity/capacity_integration.rs`, called from `rustfs/src/startup_entrypoint.rs`); scheduled interval defaults and singleflight refresh stay unchanged.
|
||||
- Storage-adjacent monitors (`monitor_and_connect_endpoints` in `crates/ecstore/src/core/sets.rs`, `enable_health_check` in `crates/ecstore/src/disk/disk_store.rs`) change disk state and stay outside controller work.
|
||||
- Deferred IAM recovery (`spawn_iam_recovery_task`, `rustfs/src/startup_iam.rs`) publishes readiness; optional protocol servers already own `ShutdownHandle`s; the auto-tuner (`init_auto_tuner` in `rustfs/src/init.rs`) changes runtime concurrency. All three stay outside generic controllers.
|
||||
|
||||
@@ -1,89 +0,0 @@
|
||||
# Background Services Inventory
|
||||
|
||||
This document records the current background service surface before
|
||||
BackgroundController work. It is a behavior-preservation inventory only; it does
|
||||
not define a new scheduler, controller framework, or shutdown contract.
|
||||
|
||||
## Scope
|
||||
|
||||
- Related migration task: `BGC-001`.
|
||||
- PR type: `docs-only`.
|
||||
- Baseline: `upstream/main` at
|
||||
`03eb10b07f5f968c531151ae667dfe218050493d`.
|
||||
- Out of scope: changing startup order, shutdown order, readiness, storage
|
||||
writes, heal admission, scanner scheduling, replication queues, config reload
|
||||
behavior, metrics intervals, or worker counts.
|
||||
|
||||
## Startup And Shutdown Owners
|
||||
|
||||
| Area | Startup owner | Shutdown owner | Current cancellation source |
|
||||
|---|---|---|---|
|
||||
| Main runtime token | `rustfs/src/main.rs::run` creates `ctx` after HTTP listeners start and before ECStore creation. | `rustfs/src/main.rs::handle_shutdown` calls `ctx.cancel()` before service-specific shutdown. | Shared `tokio_util::sync::CancellationToken`. |
|
||||
| Scanner | `rustfs/src/main.rs::run` calls `init_data_scanner(ctx.clone(), store.clone())` after successful startup log and global init time. | Main shutdown calls `ctx.cancel()`; if scanner was enabled it also calls `shutdown_background_services()`. | Scanner loop receives the main runtime token. |
|
||||
| Heal/AHM | Main creates `create_ahm_services_cancel_token()` before scanner/heal feature checks and calls `init_heal_manager(...)` when heal or scanner is enabled. | Main shutdown calls `shutdown_ahm_services()` when heal or scanner was enabled. | Global AHM token plus channel/worker-local state. |
|
||||
| Replication pool | Main calls `init_background_replication(store.clone())` after global config init, then `pool.init_resync(ctx.clone(), buckets.clone())` after bucket listing. | No direct main shutdown call for the replication pool; resync receives the main runtime token. | Resync routine uses the main runtime token; per-bucket resync uses registered cancel tokens. |
|
||||
| Lifecycle expiry/transition | `ECStore::init` calls `init_background_expiry(self.clone())` and `init_background_stale_multipart_upload_cleanup(self.clone())`. | Expiry workers read `get_background_services_cancel_token()` and fall back to a private token if none exists. Stale multipart cleanup exits when the weak ECStore reference cannot upgrade. | `ECStore::init` binds the main runtime token into the instance context with `bind_background_cancel_token(ctx)` before expiry starts, so the private-token fallback is a defensive path rather than the normal one. |
|
||||
| Notification runtime | Main calls `init_event_notifier()` after buffer profile init. | Main shutdown calls `shutdown_event_notifier().await`. | Notification runtime owns target/replay shutdown internally. |
|
||||
| Audit runtime | Main calls `start_audit_system().await`. | Main shutdown calls `stop_audit_system().await`. | Audit runtime owns target/replay shutdown internally. |
|
||||
| Metrics and memory loops | Main calls `init_metrics_runtime(ctx.clone())`, `init_memory_observability(ctx.clone())`, and `init_auto_tuner(ctx.clone())` when observability metrics are enabled. | Main shutdown only cancels the shared runtime token. | Shared runtime token. |
|
||||
| Allocator reclaim | Main calls `init_allocator_reclaim(ctx.clone())` unconditionally. | Main shutdown cancels the shared runtime token. | Shared runtime token. |
|
||||
| Capacity manager | Main calls `init_capacity_management().await` before HTTP listener startup and ECStore creation. | No direct main shutdown call. | Current scheduled capacity and metrics loops do not receive a shutdown token. |
|
||||
| Optional protocol servers | Main calls feature-gated FTP, FTPS, WebDAV, and SFTP init functions. | Main shutdown calls each stored `ShutdownHandle` and waits for all protocol shutdown futures. | Per-protocol broadcast shutdown handles. |
|
||||
| Deferred IAM recovery | `bootstrap_or_defer_iam_init(...)` may spawn a deferred recovery loop. | Main shutdown cancels the shared runtime token. | Shared runtime token. |
|
||||
|
||||
## Service Inventory
|
||||
|
||||
| Service | Trigger and workers | Side effects | Status and metrics | Migration notes |
|
||||
|---|---|---|---|---|
|
||||
| Capacity background refresh | `rustfs/src/capacity/capacity_integration.rs::init_capacity_management` delegates to `init_capacity_management_for_local_disks`, then `crates/object-capacity/src/capacity_manager.rs::start_background_task` spawns a scheduled refresh loop and a runtime summary loop. | Refreshes global capacity cache from local disks and logs runtime summaries. | Uses the object-capacity manager state and log summaries; no explicit shutdown status surface is exposed here. | Add read-only status before any controller migration. A future controller must not change scheduled interval defaults or singleflight refresh behavior. |
|
||||
| ECStore endpoint monitor | `crates/ecstore/src/core/sets.rs::new` spawns `monitor_and_connect_endpoints`. | Monitors endpoint connectivity and reconnect behavior for erasure sets. | Logs monitor start, cancellation, and exit. | This is storage-adjacent and must stay outside broad controller movement until storage shutdown semantics are explicitly covered. |
|
||||
| Local disk health monitor | `crates/ecstore/src/store/init.rs::init` enables disk health checks after store initialization; `crates/ecstore/src/disk/disk_store.rs::enable_health_check` spawns writable and recovery monitors. | Periodically probes disk writability, can create test objects named `health-check-*`, and updates disk runtime health state. | Disk info includes runtime health metrics and waiting counts. | Do not merge this with scanner/heal controller work; probes affect disk health semantics. |
|
||||
| Data scanner | `crates/scanner/src/scanner.rs::init_data_scanner` configures scanner defaults, applies runtime config, waits the initial scanner delay, then loops `run_data_scanner`. | Updates data usage cache, scans buckets/sets, evaluates lifecycle rules, queues replication heal, queues scanner heal, and emits scanner alerts. | Scanner runtime config/status is exposed through admin scanner status; scanner metrics record ILM, replication admission, heal admission, checkpoints, yields, and alerts. | Scanner implies heal because scanner can enqueue heal requests. Future controller status must separate scheduler state from scanner work-source accounting. |
|
||||
| Heal/AHM | `crates/heal/src/lib.rs::init_heal_manager` starts `HealManager`, initializes the shared heal channel, and spawns `HealChannelProcessor`. | Consumes heal requests from the global heal channel and drives heal work through the configured heal storage API. | Global active-task and queue-length atomics track current heal pressure. | Keep heal admission and channel semantics intact. Controller work should first expose queue/active status and shutdown state. |
|
||||
| Bucket replication pool | `crates/ecstore/src/bucket/replication/replication_pool.rs::init_background_replication` creates global replication stats and the global pool; pool resizing spawns regular, large-object, and failed-object workers. | Replicates object and delete operations, updates queue stats, and maintains replication worker pools. | Replication stats expose active worker counts and queue accounting. | Worker resize behavior currently closes channels to stop workers. Do not replace this with a generic controller until queue close semantics are captured by tests. |
|
||||
| Bucket replication resync | Main calls `get_global_replication_pool().init_resync(ctx.clone(), buckets.clone())`; the pool spawns `start_resync_routine`. Admin site-replication handlers can start or cancel per-bucket resync with dedicated tokens. | Loads persisted resync state, starts bucket resync, persists status, and can cancel per-target resync. | Admin site-replication status surfaces resync state. | Preserve the split between startup resync and admin-triggered resync operations. |
|
||||
| Lifecycle expiry and transition | `ECStore::init` calls `init_background_expiry(self.clone())`. Scanner evaluates lifecycle events and queues expiry/transition work through `apply_expiry_rule` and `apply_transition_rule`. | Deletes expired objects, queues transitions, updates lifecycle stats, and accounts scanner ILM actions only when work is queued. | Lifecycle state tracks worker counts, active tasks, queue send timeouts, compensation tasks, and transition stats. | This is not a separate periodic controller today; scanner is the main event source for object lifecycle evaluation. |
|
||||
| Stale multipart cleanup | `ECStore::init` calls `init_background_stale_multipart_upload_cleanup(self.clone())`. | Periodically deletes stale multipart upload data. | Logs cleanup passes when objects are deleted. | Current loop has no explicit cancellation token and exits when ECStore is dropped. Future controller work needs an explicit lifecycle decision before changing it. |
|
||||
| Notification runtime | `rustfs/src/server/event.rs::init_event_notifier` initializes live event stream support even when notification targets are disabled; when enabled, it loads server config and activates targets. Config reload uses `NotificationConfigManager::reload_config`. | Installs ECStore event dispatch hook, activates notification targets, manages replay/runtime target state, and supports live event streams. | Notification module state is refreshed from persisted module switches; target health is available through runtime target status. | Keep live event stream support separate from target delivery enablement. Reload must remain admin-triggered and peer-signaled. |
|
||||
| Audit runtime | `rustfs/src/server/audit.rs::start_audit_system` starts audit only when module switches and configured targets allow it. `AuditSystem::reload_config` replaces runtime targets. | Dispatches audit events to configured targets and manages replay workers. | Audit observability records config reloads and target delivery metrics. | Do not couple audit lifecycle to notification lifecycle even though the runtime patterns are similar. |
|
||||
| Dynamic config reload | Admin config handlers call `apply_dynamic_config_for_subsystem`, then `signal_dynamic_config_reload` or `signal_config_snapshot_reload` through the global notification system. | Applies scanner/heal runtime config, audit reloads, notification reloads, and peer reload signals. | Logs local and peer reload failures. Audit reload increments audit config reload metrics. | This is admin-triggered fanout, not a background scheduler. Controller work should preserve per-subsystem validation and error boundaries. |
|
||||
| Metrics runtime | `crates/obs/src/metrics/scheduler.rs::init_metrics_runtime` spawns multiple interval loops for cluster, bucket, node, resource, audit, notification, and replication bandwidth metrics. | Periodically collects and reports metrics. | Reports through the metrics runtime, logs cancellation warnings, and exposes a typed read-only status snapshot plus a no-op reconcile plan for enablement, collector task count, intervals, replication bandwidth tombstone cycles, cancellation source, and shutdown handle shape. | Keep intervals and collector grouping stable. The current controller surface does not mutate workers. |
|
||||
| Memory observability | `rustfs/src/memory_observability.rs::init_memory_observability` spawns a token-cancelled sampler. | Periodically records memory snapshots. | Emits memory observability metrics and exposes a read-only status snapshot plus a no-op reconcile plan for metrics enablement, interval, cancellation source, and shutdown handle shape. | This is the first low-risk pilot for controller status because it already has a simple token loop and the pilot does not mutate workers. |
|
||||
| Allocator reclaim | `rustfs/src/allocator_reclaim.rs::init_allocator_reclaim` spawns a token-cancelled reclaim loop when enabled. | Observes reclaimable work and may run allocator reclaim after idle intervals. | Emits reclaim enabled/backend counters, active-request gauges, scanner/heal activity gauges, and reclaim result counters. Exposes a typed read-only status snapshot plus a no-op reconcile plan for enablement, backend, effective force, intervals, cancellation source, and shutdown handle shape. | A controller must preserve idle-streak logic and backend-specific force behavior. The current controller surface does not mutate workers. |
|
||||
| Auto-tuner | `rustfs/src/init.rs::init_auto_tuner` optionally spawns a 60-second loop when `RUSTFS_AUTOTUNER_ENABLED` is true. | Tunes concurrency manager settings from performance metrics. | Logs iteration success/failure. | Treat as behavior-sensitive; a future controller needs explicit rollback because it can change runtime concurrency. |
|
||||
| Update check | `rustfs/src/init.rs::init_update_check` spawns one async task with a 30-second timeout when update checks are enabled. | Performs version check network I/O and logs available updates. | Logs result only. | This is a one-shot task, not a controller candidate for the first BGC PRs. |
|
||||
| Deferred IAM recovery | `rustfs/src/startup_iam.rs::spawn_iam_recovery_task` retries IAM init with backoff and finalizes readiness when successful. | Can initialize IAM later, initialize AppContext if needed, mark `IamReady`, and publish `FullReady`. | Readiness state reflects deferred recovery progress. | Keep this lifecycle-critical path separate from generic background controllers. |
|
||||
| Optional protocol servers | `rustfs/src/init.rs` starts FTP, FTPS, WebDAV, and SFTP with per-protocol `ShutdownHandle`s when features and config enable them. | Serve protocol traffic in background tasks. | Shutdown logs per protocol. | Protocol servers already have explicit handles; do not fold them into BGC until the service registry owns shutdown ordering. |
|
||||
|
||||
## Current Gaps To Preserve Before Controller Work
|
||||
|
||||
- ECStore background-service cancellation has a public global token API, but this
|
||||
inventory found no current startup call to create that token. Lifecycle expiry
|
||||
workers therefore use their fallback token when no global token exists.
|
||||
- Capacity manager loops do not receive the main runtime cancellation token.
|
||||
- Replication worker pools stop some workers by closing channels, while resync
|
||||
uses cancellation tokens. These are different shutdown contracts.
|
||||
- Scanner, lifecycle, replication, and heal are coupled by work queues and
|
||||
metrics. Moving one without status snapshots for the others risks hiding work
|
||||
admission failures.
|
||||
- Dynamic config reload is admin-triggered and peer-signaled, not a periodic
|
||||
background loop.
|
||||
|
||||
## BGC-002 Contract Inputs
|
||||
|
||||
These inputs are formalized in
|
||||
[`background-controller-contract.md`](background-controller-contract.md).
|
||||
|
||||
Future controller contract work should start with a read-only shape:
|
||||
|
||||
- `desired`: enabled/disabled plus static config source.
|
||||
- `current`: started, stopped, running, degraded, or disabled.
|
||||
- `status`: worker counts, queue lengths, last cycle/reload time, and last error.
|
||||
- `shutdown`: cancellation source and whether the service has an explicit stop
|
||||
handle.
|
||||
- `side_effects`: storage writes, target activation, external I/O, metrics, and
|
||||
readiness changes.
|
||||
|
||||
The first pilot should use a service with an existing simple cancellation loop
|
||||
and no storage writes, such as memory observability. Scanner, heal, replication,
|
||||
lifecycle, and disk health must wait for focused preservation tests.
|
||||
@@ -1,8 +1,7 @@
|
||||
# Compatibility Cleanup Register
|
||||
|
||||
Use this file to track temporary compatibility code introduced by architecture
|
||||
migration PRs. Entries are required only for compatibility paths that are planned
|
||||
for later deletion.
|
||||
**Use this when:** you add, review, or remove a temporary compatibility path (fallback, wrapper, re-export, legacy codec) and need the required marker and its removal condition.
|
||||
**Source of truth:** the `RUSTFS_COMPAT_TODO(<id>)` source markers, matched in both directions against `## Open Items` by `scripts/check_architecture_migration_rules.sh`. Entries exist only for compatibility paths planned for later deletion.
|
||||
|
||||
## Required Source Marker
|
||||
|
||||
@@ -14,7 +13,7 @@ for later deletion.
|
||||
|
||||
- `tokio-tar-extension-limits` bounded archive parser hardening: Snowball extraction depends on per-entry and cumulative GNU long-name, GNU long-link, and PAX extension limits; physical-entry, GNU sparse-map, and sparse-continuation limits; cancellation-safe sparse parsing; and fused entry streams after parser errors. The released tokio-tar API does not provide this complete boundary. Keep the reviewed fork pin until astral-sh/tokio-tar#118 is merged and one published tokio-tar release contains every listed capability with the Snowball regression fixtures passing against that release.
|
||||
- `backlog-2102` rc.2/rc.3 empty scanner usage floor recovery: old DeleteBucket cleanup could synthesize an empty incomplete v2 usage primary/backup before leadership added an epoch, while newer scanners require a durable authoritative baseline identity. New scanners recognize only that exact serialized empty-fence shape, preserve its epoch through a CAS-protected recovery marker, and rebuild namespace coverage without treating zero usage as authoritative. Remove this recovery path and marker after rc.2 and rc.3 are no longer supported direct-upgrade sources.
|
||||
- `s3gate-metadata-xml` persisted bucket XML migration: mixed-version site-replication peers, retained `.metadata.bin` objects, and backup archives can all carry XML written by the s3s codec, so the gateway migration must keep the legacy codec available until every stored form has crossed a verified rewrite boundary. Remove the legacy s3s parser and serializer only after the minimum supported direct-upgrade release reads and writes every persisted XML configuration family through the gateway codec, the four-way D1-D5 gate has remained clean for one full support window, every supported mixed-version site-replication topology has completed its writer upgrade, and migration tooling has verified or rewritten every retained bucket metadata object and restorable backup archive.
|
||||
- `s3gate-metadata-xml` persisted bucket XML migration: mixed-version site-replication peers, retained `.metadata.bin` objects, and backup archives can all carry XML written by the s3s codec, so the gateway migration must keep the legacy codec available until every stored form has crossed a verified rewrite boundary. Remove the legacy s3s parser and serializer only after the minimum supported direct-upgrade release reads and writes every persisted XML configuration family through the gateway codec, every supported mixed-version site-replication topology has completed its writer upgrade, and migration tooling has verified or rewritten every retained bucket metadata object and restorable backup archive.
|
||||
- `rustfs-6339` legacy bucket policy ID casing: earlier RustFS releases persisted the top-level policy identifier as "ID", while current writes use the S3-compatible "Id" spelling. Readers accept both spellings so retained bucket metadata remains usable after upgrade. Remove the legacy alias after migration tooling has rewritten every retained bucket policy using "ID".
|
||||
- `table-publication-fence-v1` table publication fencing: nodes that predate table and table-bucket publication fences can mutate live files while a new node is publishing a catalog pointer. New nodes retain exact object guards until the operator confirms that every serving node uses the new fences. Fleet confirmation also requires non-overlapping active warehouse prefixes and lifecycle workers that exclude table buckets. Remove the exact live-file fallback and the fleet-confirmation gate after the minimum supported RustFS release acquires table fences for registered-table mutations and table-bucket fences for unresolved-prefix mutations.
|
||||
- `table-catalog-strong-snapshot-v1` durable strong catalog snapshot compatibility: version 1 writes continue during mixed-version rollout until operators confirm that every serving node reads version 2, and version 1 table/view identifier collisions remain available only for cleanup. Remove version 1 writes and collision cleanup after the minimum supported RustFS release reads version 2 and every retained durable strong snapshot is collision-free and has been upgraded to version 2.
|
||||
@@ -29,7 +28,7 @@ for later deletion.
|
||||
- `rustfs-5416-zero-retry-delay` startup retry-delay validation: releases before bounded topology convergence accept RUSTFS_STARTUP_TOPOLOGY_RETRY_MAX_DELAY values of 0 or 0ms. New servers replace those values with the safe nonzero default so a direct upgrade neither fails startup nor enters a busy loop. Reject zero after the minimum supported direct-upgrade release validates or rewrites this setting before rollout.
|
||||
- `scanner-usage-v2` persisted scanner usage migration: pre-v2 scanners write `.usage.json`, so upgraded clusters read that primary/backup pair only while `.usage.v2.json` is absent and continue removing deleted buckets from legacy copies that still exist. The additive usage_snapshot_complete field in `.usage.v2.json` must remain optional while mixed-version clusters are supported; a missing field means the snapshot is not authoritative. The legacy read also feeds the degraded quota-admission baseline (issue #5716): while no authoritative usage exists, quota checks admit against the pre-discard sizes of the last loaded snapshot, including a legacy one. Remove the legacy object fallback and cleanup only after every supported direct-upgrade source writes `.usage.v2.json`; the baseline then feeds from incomplete v2 snapshots alone.
|
||||
- `ns-scanner-rpc-v3` namespace scanner capability and activity handshake: old peers and legacy internode transports lack the authenticated startup-epoch handshake. The oldest peers send an empty activity request and receive a field-empty protocol-0 response. Protocol v4 binds the challenge and response topology but cannot authenticate distributed dirty-usage state. Protocol v5 binds the request version, acknowledgement target and generation, and the response dirty-usage state, but predates set-scoped scanner cache locks. Protocol v6 additionally fences scanner cache lock-domain changes. Current protocol v7 binds the storage-owned movement generation and publication-blocked state, so distributed scanner cycles publish usage only after every peer reports a complete v7 activity proof; v6 responses remain readable but are treated as unverified for publication. Servers retain protocol-0, protocol-v4, and protocol-v6 codecs alongside the current v7 codec for rolling upgrades, while protocol-v5 peers are treated as previous-version peers that cannot safely participate in the new cache lock domain. Scanner selection treats HTTP 404/405/426 and the legacy MethodNotAllowed default as an explicit lack of remote scanner v3 support and assigns those disks to coordinator-driven workers; transient capability failures remain incomplete and do not activate the fallback. Remove the coordinator fallback after the minimum supported RustFS peer version implements namespace scanner protocol v3, remove protocol-0 activity requests and responses after every supported peer implements authenticated scanner activity protocol v4, remove the protocol-v4 activity codec after every supported peer implements protocol v5, and remove protocol-v5 previous-version rejection after every supported peer implements protocol v6; future protocol revisions must keep the same dual-version server/codec window before changing the advertised version.
|
||||
- `#4648` walk-dir stream completion capability: old clients can append fallback output to an already-used metacache writer after a terminal body error, so servers emit terminal walk errors only to clients that sign the `walk_dir_stream_completion=error-v1` query capability and its request-body digest. Remove the legacy clean-EOF path after the minimum supported RustFS peer version always advertises this capability.
|
||||
- `rustfs-4648` walk-dir stream completion capability: old clients can append fallback output to an already-used metacache writer after a terminal body error, so servers emit terminal walk errors only to clients that sign the `walk_dir_stream_completion=error-v1` query capability and its request-body digest. Remove the legacy clean-EOF path after the minimum supported RustFS peer version always advertises this capability.
|
||||
- `heal-rpc-auth-v2` internode gRPC authentication: servers temporarily accept legacy prefix signatures so old peers remain available during rolling upgrades. Remove the legacy fallback after the minimum supported RustFS peer version sends v2 authentication on every internode gRPC request.
|
||||
- `put-file-auth-epoch-strict` internode put_file epoch compatibility: rc.2 peers can cache a remote put_file capability before that remote node restarts, then continue sending v1 authenticated uploads with the old server epoch; those peers cannot recover from the 409 conflict used by newer clients to trigger a re-probe. Servers temporarily accept signed, non-nil stale put_file epochs while legacy put_file auth remains non-strict so mixed-version rolling upgrades can finish multipart/object writes. Remove the stale-epoch fallback after the minimum supported RustFS peer version re-probes put_file capability after server-epoch conflicts and legacy put_file auth is no longer accepted.
|
||||
- `disk-mutation-body-digest` internode mutating disk RPCs: servers temporarily accept mutating disk RPCs (RenameData, DeleteVersion, DeleteVersions, WriteMetadata, UpdateMetadata, WriteAll, Delete, DeletePaths, RenameFile, RenamePart, DeleteVolume, MakeVolume, MakeVolumes) that carry no signature-bound canonical body digest, so peers from releases that predate body-digest signing remain available during rolling upgrades. Accepted digestless mutations increment the internode body-digest fallback counter; that counter must read zero fleet-wide across a release window before RUSTFS_INTERNODE_RPC_BODY_DIGEST_STRICT is enabled. Because body-bound requests now consume replay-cache nonces on the receiver, deploy the raised RUSTFS_INTERNODE_RPC_REPLAY_CACHE_CAPACITY default fleet-wide before enabling strict mode, and watch the internode replay-cache overflow counter for undersized capacity during the rollout. Remove the digestless fallback after the minimum supported RustFS peer version body-binds every mutating disk RPC.
|
||||
|
||||
@@ -1,185 +1,49 @@
|
||||
# Config Model Boundary ADR
|
||||
|
||||
Related issue: [`rustfs/backlog#660`](https://github.com/rustfs/backlog/issues/660)
|
||||
|
||||
Task: `CFG-002`
|
||||
**Use this when:** you touch the server-config model (`Config`, `KV`, `KVS`) or its persistence, or you need to know which crate owns which part of server configuration.
|
||||
**Source of truth:** `crates/config/src/server_config.rs` (model, default registration, process-global snapshot) and `crates/ecstore/src/config/` (`ConfigSys`, persistence, migration, storage-class runtime state).
|
||||
|
||||
## Decision
|
||||
|
||||
Use the existing `crates/config` package (`rustfs-config`) as the target owner
|
||||
for the pure server-config model. Do not create a new config-model crate for
|
||||
the first extraction.
|
||||
`rustfs-config` (`crates/config`) owns the pure server-config model and the process-global server-config snapshot. ECStore keeps config persistence, migration, default-registration wiring, startup initialization, and storage-class runtime state. There is no separate config-model crate, and `rustfs_ecstore::config` does not re-export the model or the snapshot accessors.
|
||||
|
||||
The next model extraction PR should introduce the model under:
|
||||
|
||||
```text
|
||||
crates/config/src/server_config.rs
|
||||
```
|
||||
|
||||
The exported path should be:
|
||||
|
||||
```rust
|
||||
rustfs_config::server_config::{Config, KV, KVS}
|
||||
```
|
||||
|
||||
The extraction kept the existing path available through a temporary
|
||||
compatibility re-export:
|
||||
|
||||
```rust
|
||||
rustfs_ecstore::config::{Config, KV, KVS}
|
||||
```
|
||||
|
||||
That re-export included `RUSTFS_COMPAT_TODO(CFG-004)` and a matching entry in
|
||||
[`compat-cleanup-register.md`](compat-cleanup-register.md) until the model
|
||||
consumers were migrated. The CFG-004 cleanup removed this old model path after
|
||||
code scans showed consumers import the model directly from `rustfs-config`.
|
||||
|
||||
Follow-up `CFG-008` moved the process-global server-config snapshot accessors
|
||||
to `rustfs_config::server_config` after the model path stabilized. Its temporary
|
||||
`rustfs_ecstore::config::{get_global_server_config, set_global_server_config}`
|
||||
compatibility re-export was removed after in-repo runtime consumers migrated to
|
||||
the `rustfs-config` owner.
|
||||
Import path: `rustfs_config::server_config::{Config, KV, KVS}`. The model sits behind the `server-config-model` feature of `rustfs-config` (`crates/config/Cargo.toml`), which enables `serde` and `serde_json`.
|
||||
|
||||
## Why `rustfs-config`
|
||||
|
||||
`rustfs-config` is already the lowest RustFS crate for configuration constants
|
||||
and subsystem identifiers used by ECStore, notify, audit, targets, scanner, IAM,
|
||||
and admin code. The current `ecstore::config::{Config, KV, KVS}` model already
|
||||
uses `rustfs-config` constants, so moving the pure model upward to
|
||||
`rustfs-config` cuts the wrong dependency direction without adding another crate.
|
||||
- It is already the lowest RustFS crate for configuration constants and subsystem identifiers used by ECStore, notify, audit, targets, scanner, IAM, and admin code, and the model needs only those constants.
|
||||
- Moving the model upward removes the wrong-direction dependency (outer crates importing ECStore for a plain data type) without adding another crate or a second config namespace.
|
||||
|
||||
Creating a new crate now would add a second config namespace before consumers
|
||||
are migrated. That would increase re-export and compatibility surface while not
|
||||
removing any storage or runtime dependency by itself.
|
||||
## Ownership
|
||||
|
||||
## Allowed Dependencies
|
||||
| Item | Owner | Notes |
|
||||
|---|---|---|
|
||||
| `KV`, `KVS`, `Config` and their methods (`get_value`, `set_defaults`, `marshal`, `unmarshal`, `merge`) | `crates/config/src/server_config.rs` | Pure data model with serde roundtrip |
|
||||
| `DEFAULT_KVS`, `register_default_kvs` | `crates/config/src/server_config.rs` | Registration surface; ECStore still calls it from `init()` in `crates/ecstore/src/config/mod.rs` |
|
||||
| `GLOBAL_SERVER_CONFIG`, `get_global_server_config`, `set_global_server_config` | `crates/config/src/server_config.rs` | Process-global snapshot accessors |
|
||||
| `ConfigSys`, `init()`, `try_migrate_server_config` | `crates/ecstore/src/config/mod.rs` | Startup order and caller unchanged |
|
||||
| `read_config_without_migrate`, `save_server_config`, other config-object helpers | `crates/ecstore/src/config/com.rs` | Persistence over the object store |
|
||||
| `GLOBAL_STORAGE_CLASS` and storage-class parsing | `crates/ecstore/src/config/mod.rs`, `crates/ecstore/src/config/storageclass.rs` | Storage behavior stays in ECStore |
|
||||
|
||||
The server-config model module may use only:
|
||||
## Allowed Dependencies Of The Model Module
|
||||
|
||||
- `std::collections::HashMap`
|
||||
- `std::sync::{LazyLock, OnceLock, RwLock}` for the default `KVS` registration
|
||||
surface and process-global server-config snapshot
|
||||
- `serde` for `KV` and `KVS` serialization compatibility
|
||||
- `serde_json` for `Config::marshal` and `Config::unmarshal`
|
||||
- existing `rustfs-config` constants and subsystem modules
|
||||
- `std::collections::HashMap` and `std::sync::{LazyLock, OnceLock, RwLock}` for `DEFAULT_KVS` and `GLOBAL_SERVER_CONFIG`;
|
||||
- `serde` for `KV`/`KVS` and `serde_json` for `Config::marshal` / `Config::unmarshal`, gated by `server-config-model`;
|
||||
- existing `rustfs-config` constants and subsystem modules.
|
||||
|
||||
If `serde` and `serde_json` are added to `rustfs-config`, they should be attached
|
||||
only to a model feature such as `server-config-model` unless the implementation
|
||||
PR proves that making them non-optional is simpler and harmless for downstream
|
||||
builds.
|
||||
## Forbidden Dependencies Of The Model Module
|
||||
|
||||
## Forbidden Dependencies
|
||||
- `rustfs-ecstore`, `rustfs`, storage-api traits, or object persistence helpers;
|
||||
- notify, audit, targets, IAM, scanner, KMS, or admin handler crates;
|
||||
- async runtimes, HTTP/router crates, object-store crates, or runtime lifecycle state;
|
||||
- `ConfigSys`, `read_config_without_migrate`, `save_server_config`, or any `com.rs` helper.
|
||||
|
||||
The model module must not depend on:
|
||||
## Shape Preservation
|
||||
|
||||
- `rustfs-ecstore`
|
||||
- `rustfs`
|
||||
- `StorageAPI` or object persistence helpers
|
||||
- notify, audit, targets, IAM, scanner, KMS, or admin handler crates
|
||||
- async runtimes, HTTP/router crates, object-store crates, or runtime lifecycle
|
||||
state
|
||||
- unrelated runtime global state outside the process-global server-config
|
||||
snapshot
|
||||
- `ConfigSys`, `read_config_without_migrate`, `save_server_config`, or any
|
||||
`com.rs` persistence helper
|
||||
Persisted server-config JSON must keep decoding unchanged:
|
||||
|
||||
## Boundary Split
|
||||
|
||||
Move in the first extraction:
|
||||
|
||||
- `KV`
|
||||
- `KVS`
|
||||
- `Config`
|
||||
- `DEFAULT_KVS`
|
||||
- `register_default_kvs`
|
||||
- `Config::new`
|
||||
- `Config::get_value`
|
||||
- `Config::set_defaults`
|
||||
- `Config::marshal`
|
||||
- `Config::unmarshal`
|
||||
- `Config::merge`
|
||||
|
||||
Keep in `ecstore`:
|
||||
|
||||
- `ConfigSys`
|
||||
- `init_global_config_sys`
|
||||
- `try_migrate_server_config`
|
||||
- `read_config_without_migrate`
|
||||
- `save_server_config`
|
||||
- generic `com.rs` config-object helpers
|
||||
- storage-class runtime global state
|
||||
|
||||
Keep default registration wiring in `ecstore::config::init` until a later PR
|
||||
extracts a dedicated default-registration contract. The values may be registered
|
||||
through the moved `rustfs_config::server_config::register_default_kvs`, but the
|
||||
startup order and caller remain unchanged.
|
||||
|
||||
Move in `CFG-008`:
|
||||
|
||||
- `GLOBAL_SERVER_CONFIG`
|
||||
- `get_global_server_config`
|
||||
- `set_global_server_config`
|
||||
|
||||
The temporary ECStore compatibility re-export for these accessors was removed
|
||||
after code scans showed in-repo consumers use `rustfs_config::server_config`
|
||||
directly.
|
||||
|
||||
## Required Shape Preservation
|
||||
|
||||
The extraction PR must preserve:
|
||||
|
||||
- `KV { key, value, hidden_if_empty }`
|
||||
- `#[serde(default, alias = "hiddenIfEmpty")]` on `KV::hidden_if_empty`
|
||||
- `KVS(pub Vec<KV>)`
|
||||
- `Config(pub HashMap<String, HashMap<String, KVS>>)`
|
||||
- `KVS::new`, `get`, `lookup`, `is_empty`, `keys`, `insert`, and `extend`
|
||||
- `Config::new`, `get_value`, `set_defaults`, `marshal`, `unmarshal`, and
|
||||
`merge`
|
||||
- `Config::new()` default application after `ecstore::config::init()`
|
||||
- existing persisted server-config JSON shape
|
||||
- existing target, notify, audit, scanner, OIDC, and admin interpretation of
|
||||
`Config` and `KVS`
|
||||
|
||||
## Next PR Requirements
|
||||
|
||||
`CFG-003` should be a pure model extraction or narrow `api-extraction` PR. It
|
||||
must not migrate consumers, change persistence helpers, or alter runtime
|
||||
behavior.
|
||||
|
||||
`CFG-004` kept the old `rustfs_ecstore::config::*` path as a temporary
|
||||
compatibility shim, registered its removal condition, and removed the shim after
|
||||
all in-repo consumers migrated.
|
||||
|
||||
`CFG-005` should migrate external consumers one group at a time after the model
|
||||
and compatibility path are stable.
|
||||
|
||||
`CFG-008` moves only the global server-config snapshot accessors to
|
||||
`rustfs-config` and migrates in-repo direct consumers. It must not move
|
||||
`ConfigSys`, storage-class global state, persistence helpers, default
|
||||
registration wiring, startup order, or storage behavior.
|
||||
|
||||
## Verification Gate
|
||||
|
||||
Before pushing an extraction PR, run:
|
||||
|
||||
- serde roundtrip tests for old and new paths
|
||||
- tests for `hiddenIfEmpty` alias compatibility
|
||||
- tests for `KVS` insertion, lookup, extension, and keys behavior
|
||||
- tests for `Config::new`, `set_defaults`, `marshal`, `unmarshal`, and `merge`
|
||||
- a cleanup scan proving in-repo consumers no longer use the old
|
||||
`rustfs_ecstore::config::{Config, KV, KVS}` model path before removing the
|
||||
compatibility shim
|
||||
- `cargo tree -p rustfs-config --edges normal`
|
||||
- `cargo tree -p rustfs-ecstore --edges normal`
|
||||
- `./scripts/check_layer_dependencies.sh`
|
||||
- `./scripts/check_architecture_migration_rules.sh`
|
||||
- `cargo fmt --all --check`
|
||||
- `make pre-commit`
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No consumer migration in `CFG-002`.
|
||||
- No code movement in `CFG-002`.
|
||||
- No new crate in `CFG-002`.
|
||||
- No `com.rs` or `StorageAPI` movement in the first model extraction.
|
||||
- No global server-config state migration until the model path is stable.
|
||||
- `KV { key, value, hidden_if_empty }` with `#[serde(default, alias = "hiddenIfEmpty")]` on `hidden_if_empty`;
|
||||
- `KVS(pub Vec<KV>)` and `Config(pub HashMap<String, HashMap<String, KVS>>)`;
|
||||
- `KVS::{get, lookup, is_empty, keys, insert, extend}` and `Config::{get_value, set_defaults, marshal, unmarshal, merge}` keep their semantics;
|
||||
- `Config::new()` applies the defaults registered by `ecstore::config::init()`;
|
||||
- target, notify, audit, scanner, OIDC, and admin code keep interpreting `Config` and `KVS` the same way.
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# Crate Boundaries And Migration Guardrails
|
||||
|
||||
These rules apply to architecture-migration PRs linked to
|
||||
[`rustfs/backlog#660`](https://github.com/rustfs/backlog/issues/660).
|
||||
**Use this when:** you add a crate dependency, move code across crates, touch a `storage_api.rs` boundary file, or need the change-type vocabulary the architecture guard enforces.
|
||||
**Source of truth:** `scripts/check_architecture_migration_rules.sh` (the enumerated rules; this file is its boundary document) and `scripts/check_layer_dependencies.sh` (layer and edge checks). Extend those guards instead of adding a parallel system.
|
||||
|
||||
## PR Types
|
||||
|
||||
@@ -22,275 +22,51 @@ Do not mix directory movement, security tightening, and behavior changes in one
|
||||
|
||||
## Dependency Direction
|
||||
|
||||
Contract crates must stay below implementation crates. Initial forbidden edges:
|
||||
Contract crates stay below implementation crates. Forbidden edges:
|
||||
|
||||
- `storage-api -> ecstore`
|
||||
- `security-governance -> rustfs`
|
||||
- `extension-schema -> rustfs`
|
||||
- `extension-schema -> ecstore`
|
||||
|
||||
`rustfs-storage-api` may only expose storage-facing replication status/state
|
||||
contracts through `crates/storage-api/src/replication.rs` while the underlying
|
||||
wire types still live in `rustfs-filemeta`. This keeps the temporary dependency
|
||||
centralized until those wire contracts can move without introducing a
|
||||
`rustfs-replication` / `rustfs-storage-api` cycle.
|
||||
|
||||
Leaf crates carry exactly one adjudicated allowed edge:
|
||||
`io-metrics -> rustfs-s3-ops` (transitively `rustfs-s3-types`). Both are pure
|
||||
contract crates — types and enums only, no I/O, no global state, no non-contract
|
||||
internal dependencies — so `io-metrics` reuses the `S3Operation` vocabulary
|
||||
instead of copying it. `madmin` is no longer counted a leaf: since #6166 it is
|
||||
the SigV4-signed admin SDK client and deliberately depends on `rustfs-signer`;
|
||||
the guard pins its internal dependency surface to exactly that edge so it cannot
|
||||
quietly grow storage-side dependencies. The leaf-crate allowlist in
|
||||
`scripts/check_architecture_migration_rules.sh` fails any other `rustfs-*`
|
||||
dependency in `config`, `credentials`, `crypto`, `io-metrics`, or `madmin`, in
|
||||
either TOML spelling (`rustfs-x = ...` or `rustfs-x.workspace = true`).
|
||||
Adjudicated in
|
||||
[`rustfs/backlog#1834`](https://github.com/rustfs/backlog/issues/1834); a further
|
||||
leaf exception must meet the pure-contract criterion — types and enums only, no
|
||||
I/O, no globals, no non-contract internal dependencies — and land its guard
|
||||
allowlist entry alongside the dependency.
|
||||
|
||||
Dependency direction also applies to compile-time source reads:
|
||||
`include_str!`/`include!` of a `.rs` file must not resolve outside the
|
||||
including crate's own directory (`scripts/check_layer_dependencies.sh`
|
||||
enforces this). A source-text tripwire belongs in the crate that owns the
|
||||
asserted file; shared expectations move into a contract surface such as
|
||||
`rustfs_protos::compat_manifest` and are asserted by each owning crate.
|
||||
|
||||
Existing migration checks live in:
|
||||
|
||||
- `scripts/check_layer_dependencies.sh`
|
||||
- `scripts/check_architecture_migration_rules.sh`
|
||||
|
||||
Extend these guardrails instead of adding a parallel system.
|
||||
|
||||
## Required Architecture Documents
|
||||
|
||||
The migration guard must keep these baseline documents present and anchored to
|
||||
their required sections:
|
||||
|
||||
- `docs/architecture/overview.md`: Baseline, Core Principle, Phase Order.
|
||||
- `docs/architecture/runtime-lifecycle.md`: Startup And Readiness, Shutdown
|
||||
Lifecycle Boundary, AppContext Foundation.
|
||||
- `docs/architecture/storage-control-data-plane.md`: Storage API Contracts,
|
||||
Cluster Control Plane, Background Controllers.
|
||||
- `docs/architecture/crate-boundaries.md`: PR Types, Dependency Direction,
|
||||
Required Architecture Documents.
|
||||
- `docs/architecture/readiness-matrix.md`: Request Behavior Matrix, Runtime
|
||||
Dependency Matrix, Probe Semantics.
|
||||
- `docs/architecture/global-state-crate-split-plan.md`: Remaining Global
|
||||
Owners, Runtime Source Boundaries, Fallback Removal Plan, Crate Split
|
||||
Evaluation.
|
||||
|
||||
## Pre-Push Expert Review
|
||||
|
||||
Before pushing any PR branch, record three expert reviews in the task notes:
|
||||
|
||||
| Expert | Required focus |
|
||||
| Edge | Why |
|
||||
|---|---|
|
||||
| Quality/architecture | Structure, naming, dependency direction, PR type, scope, and over-abstraction risk |
|
||||
| Migration preservation | Startup order, readiness, quorum, reader semantics, AppContext/global fallback, notify/audit lifecycle, IAM/KMS boundaries, and compatibility |
|
||||
| Testing/verification | Focused tests, regression tests, commands run, missing coverage, and whether tests are forcing business-logic drift |
|
||||
| `storage-api -> ecstore` | Storage contracts must not depend on the storage implementation |
|
||||
| `security-governance -> rustfs` | Governance contracts stay below the binary crate |
|
||||
| `extension-schema -> rustfs` | The extension schema is consumed by the binary, never the reverse |
|
||||
| `extension-schema -> ecstore` | The extension schema must not reach storage internals |
|
||||
|
||||
Push is allowed only when all three experts return `pass` or
|
||||
`pass-with-nonblocking-follow-up`. Any `blocker` prevents push until the issue is
|
||||
fixed and the relevant review is repeated.
|
||||
- `rustfs-storage-api` exposes storage-facing replication status/state contracts only through `crates/storage-api/src/replication.rs`, so its temporary dependency on `rustfs-filemeta` wire types stays centralized and no `rustfs-replication` / `rustfs-storage-api` cycle appears.
|
||||
- Leaf crates (`config`, `credentials`, `crypto`, `io-metrics`, `madmin`) may not depend on other `rustfs-*` crates, in either TOML spelling, except the adjudicated edges pinned in the guard's leaf allowlist: `io-metrics -> rustfs-s3-ops` (pure contract crates sharing the `S3Operation` vocabulary) and `madmin -> rustfs-signer` (the SigV4-signed admin SDK client). A new leaf exception must be a pure contract dependency (types and enums only, no I/O, no globals, no non-contract internal dependencies) and land together with its allowlist entry.
|
||||
- Compile-time source reads follow the same direction: `include_str!` / `include!` of a `.rs` file must not resolve outside the including crate (`scripts/check_layer_dependencies.sh`). Shared source-text expectations belong in a contract surface such as `rustfs_protos::compat_manifest` (`crates/protos/src/compat_manifest.rs`) and are asserted by each owning crate.
|
||||
|
||||
## Temporary Compatibility Code
|
||||
## ECStore Access Boundary
|
||||
|
||||
Temporary compatibility code that must be removed later must include a searchable
|
||||
source comment and a cleanup-register entry.
|
||||
Outer crates reach ECStore only through `rustfs_ecstore::api`, and only from one local boundary file per owner (`storage_api.rs`). Boundary files and facade groups are inventoried in [ecstore-api-facade-inventory.md](ecstore-api-facade-inventory.md).
|
||||
|
||||
Use this source-comment format:
|
||||
|
||||
```rust
|
||||
// RUSTFS_COMPAT_TODO(API-005): keep old ecstore::store_api path during storage-api migration. Remove after all consumers use rustfs-storage-api.
|
||||
```
|
||||
|
||||
Rules:
|
||||
|
||||
- Add the marker only to temporary compatibility paths, not permanent APIs.
|
||||
- Include the task ID in the marker.
|
||||
- State why the compatibility path exists and when it can be removed.
|
||||
- Use this for temporary re-exports, wrappers, fallbacks, legacy action mappings,
|
||||
and old endpoint compatibility layers.
|
||||
- Delete compatibility layers in their own cleanup PR.
|
||||
|
||||
## Config Model First
|
||||
|
||||
`ecstore::config::{Config, KV, KVS}` should move before extension config adapters
|
||||
or config-schema work. First inventory consumers, then decide whether existing
|
||||
`crates/config` is enough or whether a smaller model crate is required.
|
||||
|
||||
The current decision is recorded in
|
||||
[`config-model-boundary-adr.md`](config-model-boundary-adr.md): use the existing
|
||||
`rustfs-config` package for the pure server-config model and global
|
||||
server-config snapshot accessors, while ECStore keeps config persistence,
|
||||
storage-class global state, default wiring, and startup initialization.
|
||||
|
||||
The old `rustfs_ecstore::config::{Config, KV, KVS, register_default_kvs,
|
||||
get_global_server_config, set_global_server_config}` compatibility path must
|
||||
not be restored after the Phase 1a cleanup. Consumers use
|
||||
`rustfs_config::server_config` for the moved model and accessors; ECStore public
|
||||
facades must not re-export those symbols.
|
||||
- Inside a boundary file, raw `rustfs_ecstore::api::...` paths are centralized behind local `ecstore_*` module aliases; code outside the boundary sees local type aliases, constants, traits, or wrapper functions, never the raw facade path.
|
||||
- Non-trait ECStore surfaces (metadata, object-lock, lifecycle journal, monitor, notification types) stay behind local aliases; boundary function signatures do not expose raw ECStore facade types once narrowed. Object and error aliases anchor on storage-api associated object types and a local `StorageError`.
|
||||
- Outer consumers use `rustfs-storage-api` operation traits (`ObjectIO`, `ObjectOperations`, `ListOperations`, `MultipartOperations`, `HealOperations`, `NamespaceLocking`) and generic list responses (`ListObjectsV2Info`, `ListObjectVersionsInfo`, `ObjectInfoOrErr`) directly; ECStore keeps concrete aliases only for internal implementation and compatibility.
|
||||
- Bucket lifecycle, replication, versioning, object-lock, restore-request, disk, RPC peer client, and warm-backend trait methods are reached through owner-local compatibility traits or wrapper functions, not by importing ECStore traits outside the boundary.
|
||||
- The old `StorageAPI` aggregate facade must not reappear in production `crates/ecstore/src` or `rustfs/src` code.
|
||||
- Facade-covered ECStore root modules (layout, `endpoints`, `disks_layout`, bitrot, erasure, object DTO/reader, event, list, batch processor, `global`) stay crate-private; public access goes through the matching `rustfs_ecstore::api::*` group.
|
||||
- Cluster control-plane read models stay owned by the crate-private `cluster` module and are published through `rustfs_ecstore::api::cluster`; pool-state, local-node storage, and peer-health projections are read-only.
|
||||
- RustFS startup internals are crate-private: only `startup_entrypoint` is a public startup module of the `rustfs` library (`rustfs/src/lib.rs`), and items inside the other `startup_*` modules use crate visibility.
|
||||
- The observability dependency baseline is [obs-ecstore-dependency-inventory.md](obs-ecstore-dependency-inventory.md); observability extraction updates it together with the guard.
|
||||
|
||||
## Loss-Prevention Coverage
|
||||
|
||||
Architecture migration checks must keep public contract re-exports and ECStore
|
||||
compatibility coverage from silently drifting during cleanup PRs.
|
||||
The guard pins specific public re-export lines (its `require_source_line` entries) so contract surfaces cannot silently disappear during cleanup. The canonical lists are the guard script and the owning files, not this page:
|
||||
|
||||
Required `rustfs-storage-api` public re-exports:
|
||||
- `crates/storage-api/src/lib.rs`: admin, bucket, capability, error, multipart, observability, object, and topology contract re-exports;
|
||||
- `crates/concurrency/src/lib.rs`: workload admission contract re-exports;
|
||||
- `rustfs/src/lib.rs`: `pub mod startup_entrypoint;`.
|
||||
|
||||
- `pub use admin::{DiskSetSelector, StorageAdminApi};`
|
||||
- `pub use bucket::{BucketInfo, BucketOperations, BucketOptions, DeleteBucketOptions, MakeBucketOptions, SRBucketDeleteOp};`
|
||||
- `pub use capability::{CapabilitySnapshotError, CapabilityState, CapabilityStatus};`
|
||||
- `pub use error::{StorageErrorCode, StorageResult};`
|
||||
- `pub use multipart::{CompletePart, ListMultipartsInfo, ListPartsInfo, MultipartInfo, MultipartUploadResult, PartInfo};`
|
||||
- `pub use observability::{MemorySamplingState, ObservabilitySnapshot, ObservabilitySnapshotProvider, PlatformSupport, UserspaceProfilingCapability};`
|
||||
- `pub use object::{HTTPPreconditions, HTTPRangeError, HTTPRangeSpec, ObjectLockRetentionOptions};`
|
||||
- `pub use object::{ExpirationOptions, TransitionedObject};`
|
||||
- `pub use object::{HealOperations, MultipartOperations, NamespaceLocking, ObjectIO, ObjectOperations};`
|
||||
- `pub use object::{ListObjectVersionsInfo, ListObjectsInfo, ListObjectsV2Info, ListOperations, ObjectInfoOrErr};`
|
||||
- `pub use object::{ObjectPreconditionError, ObjectPreconditionPart, ObjectPreconditionState};`
|
||||
- `pub use object::{VersionMarker, WalkOptions, WalkVersionsSortOrder};`
|
||||
- `pub use topology::{DiskCapabilities, TopologyCapabilities, TopologyDisk, TopologyLabels, TopologyPool, TopologySet, TopologySnapshot, TopologySnapshotProvider};`
|
||||
ECStore keeps compile-time coverage for `StorageAdminApi`, `HealOperations`, and the separate `NamespaceLocking` operation group (`crates/ecstore/tests/ecstore_contract_compat_test.rs`), and its internal consumers use the `rustfs-storage-api` lifecycle DTOs `ExpirationOptions` and `TransitionedObject` directly.
|
||||
|
||||
Required `rustfs-concurrency` public workload admission contract re-exports:
|
||||
## Temporary Compatibility Code
|
||||
|
||||
- `pub use workload::{AdmissionState, WorkloadAdmissionRegistrySnapshot, WorkloadAdmissionSnapshot, WorkloadAdmissionSnapshotProvider, WorkloadClass};`
|
||||
Every temporary compatibility path carries a `RUSTFS_COMPAT_TODO(<id>)` source marker with a removal condition and a matching entry in [compat-cleanup-register.md](compat-cleanup-register.md); the guard enforces the match in both directions. Compatibility layers are deleted in their own cleanup change, never bundled with new migration logic.
|
||||
|
||||
ECStore must keep compile-time coverage for `StorageAdminApi`, `HealOperations`,
|
||||
and the separate `NamespaceLocking` operation group.
|
||||
## Config Model
|
||||
|
||||
The old `StorageAPI` aggregate facade must not reappear in production
|
||||
`crates/ecstore/src` or `rustfs/src` code after the storage operation groups
|
||||
have been made explicit.
|
||||
The server-config model (`Config`, `KV`, `KVS`) and the global server-config snapshot accessors are owned by `rustfs_config::server_config`; ECStore keeps persistence, storage-class state, and startup wiring, and its public facades must not re-export those symbols. See [config-model-boundary-adr.md](config-model-boundary-adr.md).
|
||||
|
||||
Outer RustFS/IAM consumers must use `rustfs-storage-api` generic list response
|
||||
contracts directly for `ListObjectsV2Info`, `ListObjectVersionsInfo`, and
|
||||
`ObjectInfoOrErr`; ECStore keeps the concrete aliases only for internal
|
||||
implementation and compatibility.
|
||||
## Required Architecture Documents
|
||||
|
||||
Outer RustFS/scanner consumers must use `rustfs-storage-api` operation traits
|
||||
directly for `ObjectIO`, `ObjectOperations`, `ListOperations`,
|
||||
`MultipartOperations`, `HealOperations`, and `NamespaceLocking`; ECStore keeps
|
||||
the concrete compatibility traits only for internal implementation and
|
||||
downstream compatibility.
|
||||
Outer consumers must not import ECStore directly outside compatibility
|
||||
boundaries except for temporary trait imports needed for method resolution or
|
||||
local test trait implementations. Non-trait ECStore surfaces must stay behind
|
||||
local aliases, constants, or wrapper functions.
|
||||
|
||||
Outer compatibility boundary modules must use `rustfs_ecstore::api` for ECStore
|
||||
public facade surfaces such as layout, storage owner, admin, metrics,
|
||||
notification, capacity, bucket/config helpers, disk/error contracts, global
|
||||
state accessors, RPC constants/clients, reader helpers, tier helpers, and
|
||||
rebalance status contracts. Any non-ECStore `storage_compat.rs` import from
|
||||
`rustfs_ecstore` must route through the `rustfs_ecstore::api` facade.
|
||||
The legacy ECStore root `endpoints` and `disks_layout` compatibility modules
|
||||
must remain crate-private; public layout access goes through
|
||||
`rustfs_ecstore::api::layout`.
|
||||
Facade-covered ECStore root modules must remain crate-private after this
|
||||
boundary is established; outer crates should use `rustfs_ecstore::api::*`
|
||||
instead of legacy root module paths. This includes storage/layout surfaces as
|
||||
well as remaining bitrot, erasure coding, object DTO/reader, event, list, and
|
||||
batch processor root modules once their facade groups exist.
|
||||
ECStore root `global` re-exports must also stay removed once consumers use
|
||||
`rustfs_ecstore::api::global` or crate-internal `crate::global` paths.
|
||||
RustFS root `storage_compat.rs` must expose bucket metadata and quota contracts
|
||||
as explicit aliases only. Broad `metadata`, `metadata_sys`, and `quota` module
|
||||
passthroughs are reserved to narrower app/admin/storage compatibility
|
||||
boundaries that still need module-local owner cleanup.
|
||||
Root runtime storage config initialization and disk endpoint contracts must also
|
||||
stay explicit aliases. The root compatibility boundary must not restore `com`,
|
||||
bare `init`, or grouped `endpoint::Endpoint` passthroughs.
|
||||
RustFS root `storage_compat.rs` must not re-export ECStore API symbols directly;
|
||||
remaining root runtime compatibility symbols must be local type aliases,
|
||||
constants, traits, or wrapper functions so ownership stays visible at the
|
||||
boundary.
|
||||
RustFS admin `storage_compat.rs` must expose config IO and default
|
||||
initialization through explicit aliases. The admin compatibility boundary must
|
||||
not restore broad `com` or bare `init` passthroughs.
|
||||
RustFS admin and app `storage_compat.rs` bucket-facing compatibility contracts
|
||||
must stay explicitly whitelisted. They must not restore broad bucket module,
|
||||
client object API, client transition API, or storage-class module passthroughs
|
||||
once a local compatibility boundary has narrowed them to specific aliases.
|
||||
RustFS storage `storage_compat.rs` must expose bucket metadata, object-lock,
|
||||
policy, replication, tagging, versioning, object API, and test-only
|
||||
storage-class config contracts through explicit aliases. The storage
|
||||
compatibility boundary must not restore broad `metadata`, `metadata_sys`,
|
||||
`object_lock`, `policy_sys`, `replication`, `tagging`, `utils`, `versioning`,
|
||||
`versioning_sys`, `object_api_utils`, or `com` passthroughs.
|
||||
RustFS storage owner `storage_compat.rs` must not re-export ECStore API symbols
|
||||
directly except temporary trait imports needed for method resolution. Remaining
|
||||
storage-owner compatibility symbols must be local constants, type aliases, or
|
||||
wrapper functions so storage-owned global state and helper access stays visible
|
||||
at the boundary.
|
||||
RustFS app, admin, and storage outer `storage_compat.rs` object and error
|
||||
facade aliases must stay anchored on storage-api associated object types and
|
||||
local `StorageError` aliases. They must not reintroduce raw
|
||||
`rustfs_ecstore::api::object::{ObjectInfo,ObjectOptions}` or
|
||||
`rustfs_ecstore::api::error::{Error,Result}` references.
|
||||
Outer compatibility function signatures must also use local aliases for ECStore
|
||||
metadata, object-lock, lifecycle journal, monitor, and notification facade
|
||||
types. The boundary may define the local alias, but call signatures must not
|
||||
expose the raw ECStore facade path once narrowed.
|
||||
The RustFS storage owner compatibility boundary must keep raw ECStore facade
|
||||
paths centralized behind local `ecstore_*` module aliases rather than scattering
|
||||
`rustfs_ecstore::api::...` references through its aliases and wrappers.
|
||||
The RustFS app/admin storage compatibility boundaries must likewise route raw
|
||||
ECStore facade access through their local `ecstore_*` module aliases instead of
|
||||
scattering `rustfs_ecstore::api::...` paths through compatibility wrappers.
|
||||
Peripheral consumer storage compatibility boundaries must follow the same
|
||||
pattern. IAM, heal, scanner, notify, observability, Swift, S3 Select, test, and
|
||||
fuzz storage compatibility modules keep raw ECStore facade access centralized
|
||||
behind local `ecstore_*` module aliases.
|
||||
RustFS root runtime and e2e storage compatibility boundaries must follow the
|
||||
same pattern, keeping raw ECStore facade access centralized behind local
|
||||
`ecstore_*` module aliases.
|
||||
Outer bucket lifecycle, replication, versioning, object-lock, and
|
||||
restore-request trait method access must stay behind local compatibility traits
|
||||
or wrapper functions. Non-compat sources must not import those ECStore bucket
|
||||
API traits directly after the wrapper boundary is established. Disk, RPC peer
|
||||
client, and warm-backend method-resolution access must follow the same pattern:
|
||||
non-compat sources use owner-local compatibility traits or test aliases instead
|
||||
of importing ECStore traits directly.
|
||||
Scanner, notify, observability, and e2e `storage_compat.rs` boundaries must
|
||||
also stay narrow. Scanner must not restore grouped bucket compatibility exports
|
||||
for target, lifecycle, metadata, replication, or versioning modules. Notify
|
||||
must not restore broad `config`/`global` module imports. Observability must
|
||||
consume data usage through a local DTO projection instead of re-exporting the
|
||||
ECStore data-usage loader. The e2e harness must not restore grouped RPC
|
||||
passthroughs.
|
||||
Test and fuzz `storage_compat.rs` harnesses must also stay narrow. Heal and
|
||||
scanner test harnesses must expose ECStore contracts through direct aliases or
|
||||
local wrappers, and fuzz harnesses must wrap bucket utility entrypoints instead
|
||||
of restoring grouped ECStore passthrough exports.
|
||||
External ECStore API facade imports must stay inside local `storage_api`
|
||||
boundary files after the external runtime, test, and fuzz consumers have been
|
||||
narrowed. IAM, heal, scanner, notify, observability, Swift, S3 Select, e2e, and
|
||||
fuzz code must not reintroduce direct `rustfs_ecstore::api::...` references
|
||||
outside those boundary files.
|
||||
The observability ECStore dependency baseline is tracked in
|
||||
[`obs-ecstore-dependency-inventory.md`](obs-ecstore-dependency-inventory.md);
|
||||
future observability extraction PRs must update that inventory with the guard.
|
||||
|
||||
ECStore ClusterControlPlane read models must stay owned by the crate-private
|
||||
`cluster` module. Public access goes through `rustfs_ecstore::api::cluster` so
|
||||
outer crates cannot depend on ECStore root control-plane internals.
|
||||
Pool-state, local-node storage, and peer-health status projections are part of
|
||||
the same facade boundary and must remain read-only until a later controller
|
||||
slice explicitly wires dynamic health or membership behavior.
|
||||
|
||||
RustFS startup internals must stay crate-private after the startup owner split.
|
||||
Only `startup_entrypoint` remains a public startup module for the binary
|
||||
entrypoint; IAM bootstrap, optional runtime, and profiling startup shims must
|
||||
not be re-exported as public library modules. Items inside crate-private
|
||||
startup modules must also use crate visibility rather than bare public
|
||||
visibility.
|
||||
|
||||
ECStore internal consumers must use `rustfs-storage-api` lifecycle helper DTOs
|
||||
directly for `ExpirationOptions` and `TransitionedObject`; ECStore keeps the
|
||||
old lifecycle paths only as downstream compatibility re-exports.
|
||||
The guard requires the documents and section headings listed in its `require_source_contains` entries (`scripts/check_architecture_migration_rules.sh`); the directory index is [README.md](README.md).
|
||||
|
||||
@@ -1,47 +1,31 @@
|
||||
# Decommission Compatibility Scope
|
||||
|
||||
This note records the current RustFS decommission contract for admin/API
|
||||
compatibility reviews.
|
||||
**Use this when:** you change pool decommission or rebalance behavior, its admin API shape, the persisted `PoolMeta` decommission fields, or how tier free versions move between pools.
|
||||
**Source of truth:** `crates/ecstore/src/core/pools.rs` (queue, recovery, cleanup predicates), `crates/ecstore/src/services/rebalance/worker.rs` (rebalance predicates), `rustfs/src/admin/handlers/pools.rs` plus the `pools/*` rows of `rustfs/src/admin/route_policy.rs` (admin surface), `crates/ecstore/src/data_movement/` and `crates/ecstore/src/set_disk/` (free-version movement).
|
||||
|
||||
## Current Contract
|
||||
|
||||
RustFS supports queued multi-pool decommission start requests on multi-pool
|
||||
deployments.
|
||||
|
||||
The admin handler accepts the request shape used by the MinIO-compatible admin
|
||||
API, including comma-separated pool targets. An empty target list is rejected.
|
||||
Single-pool deployments reject decommission because there is no destination pool.
|
||||
On multi-pool deployments, one or more valid target pools are accepted as a
|
||||
single queued operation.
|
||||
RustFS supports queued multi-pool decommission start requests on multi-pool deployments. The admin handler accepts the MinIO-compatible request shape, including comma-separated pool targets. An empty target list is rejected; single-pool deployments reject decommission because there is no destination pool; on multi-pool deployments one or more valid target pools are accepted as a single queued operation.
|
||||
|
||||
### Request Semantics
|
||||
|
||||
`POST /v3/pools/decommission` with comma-separated pool targets is treated as a
|
||||
queue submission:
|
||||
`POST /v3/pools/decommission` with comma-separated pool targets is a queue submission:
|
||||
|
||||
- validate all requested pool identifiers before mutating metadata;
|
||||
- reject duplicate target pools in the same request;
|
||||
- reject active or queued target pools;
|
||||
- reject completed decommission targets because completion means the pool can be
|
||||
removed from the deployment configuration;
|
||||
- reject completed decommission targets, because completion means the pool can be removed from the deployment configuration;
|
||||
- allow failed or canceled targets to be retried;
|
||||
- persist queued metadata before starting workers;
|
||||
- start only the local-leader prefix of the queue on the receiving node.
|
||||
|
||||
The local-leader-prefix rule keeps the active worker on the leader for the pool
|
||||
being moved while still allowing a request to contain later targets whose leaders
|
||||
are different nodes. Later queued targets are recovered or promoted by the
|
||||
leader that owns that target.
|
||||
The local-leader-prefix rule keeps the active worker on the leader for the pool being moved while still allowing a request to contain later targets whose leaders are different nodes. Later queued targets are recovered or promoted by the leader that owns that target.
|
||||
|
||||
Admin start, cancel, and clear requests may arrive on any cluster node. When the
|
||||
target pool first endpoint is remote, RustFS forwards the operation over the
|
||||
authenticated internode RPC channel to that first endpoint. The receiving node
|
||||
still enforces the local-leader rule before mutating decommission state.
|
||||
Start, cancel (`POST /v3/pools/cancel`), and clear (`POST /v3/pools/clear`) requests may arrive on any cluster node. When the target pool's first endpoint is remote, RustFS forwards the operation over the authenticated internode RPC channel to that endpoint; the receiving node still enforces the local-leader rule before mutating decommission state.
|
||||
|
||||
### Persisted Metadata Shape
|
||||
|
||||
The queue is persisted in pool metadata and decoded with the rest of
|
||||
`PoolMeta`. Each pool entry can distinguish:
|
||||
The queue is persisted in pool metadata and decoded with the rest of `PoolMeta`. Each pool entry can distinguish:
|
||||
|
||||
- `active`: at most one pool currently moving data;
|
||||
- `queued`: validated pools waiting for the active entry to finish;
|
||||
@@ -49,310 +33,112 @@ The queue is persisted in pool metadata and decoded with the rest of
|
||||
- `failed`: pools whose worker reached terminal failure;
|
||||
- `canceled`: pools canceled before or during execution.
|
||||
|
||||
Legacy metadata without queue fields decodes as a non-queued decommission entry,
|
||||
preserving restart behavior for already deployed clusters.
|
||||
Legacy metadata without queue fields decodes as a non-queued decommission entry, preserving restart behavior for already deployed clusters.
|
||||
|
||||
### Serial Scheduling And Recovery
|
||||
|
||||
Only one queued entry may own a decommission worker at a time. Startup recovery:
|
||||
|
||||
- loads pool metadata before rebalance recovery;
|
||||
- resumes the first local non-terminal active/queued entry;
|
||||
- skips a durably completed prefix and promotes the next queued entry only after
|
||||
successful completion;
|
||||
- treats failed or canceled terminal entries as an automatic-promotion barrier,
|
||||
leaving later queued pools visible but stopped until an operator retries,
|
||||
clears, or otherwise resolves the terminal entry;
|
||||
- keeps queued pools out of active worker scheduling until promotion, while still
|
||||
making their future state visible in admin status.
|
||||
- computes the resumable entries with `resumable_decommission_queue_indices` (`crates/ecstore/src/core/pools.rs`): every pool that has decommission state and is not terminal (`complete`, `failed`, or `canceled`). Terminal predecessors are skipped, not treated as barriers, so a queued pool behind a failed or canceled attempt is still resumable (`test_resumable_decommission_queue_indices_skip_terminal_predecessors`);
|
||||
- starts workers only for the local-leader prefix of those entries; later queued pools stay out of worker scheduling until promotion while their state remains visible in admin status.
|
||||
|
||||
Promotion is persisted before worker execution. If cancellation is already
|
||||
requested immediately after promotion, RustFS persists a canceled terminal state
|
||||
instead of leaving the promoted pool active without a worker.
|
||||
Promotion is persisted before worker execution. If cancellation is already requested immediately after promotion, RustFS persists a canceled terminal state instead of leaving the promoted pool active without a worker.
|
||||
|
||||
### Cancel Semantics
|
||||
|
||||
Cancel separates active and queued behavior:
|
||||
|
||||
- canceling the active entry requests worker cancellation and persists terminal
|
||||
metadata;
|
||||
- canceling the active entry requests worker cancellation and persists terminal metadata;
|
||||
- canceling a queued entry marks that entry canceled before it becomes active;
|
||||
- failed or canceled terminal entries can be cleared explicitly when the operator
|
||||
chooses to abandon the decommission attempt;
|
||||
- peer reload failures during cancel must be surfaced in status and logs.
|
||||
- failed or canceled terminal entries can be cleared explicitly (`POST /v3/pools/clear`) when the operator abandons the decommission attempt;
|
||||
- peer reload failures during cancel are surfaced in status and logs.
|
||||
|
||||
Cancel requests can be accepted on non-leader nodes as remote cancel intent; the
|
||||
leader observes the pending cancel and applies it to the active worker.
|
||||
Cancel requests can be accepted on non-leader nodes as remote cancel intent; the leader observes the pending cancel and applies it to the active worker.
|
||||
|
||||
### Status Response Shape
|
||||
|
||||
`GET /v3/pools/list` and `GET /v3/pools/status?pool=...` expose per-pool
|
||||
machine-readable decommission state. The `status` field can report `active`,
|
||||
`running`, `queued`, `complete`, `failed`, or `canceled`.
|
||||
`GET /v3/pools/list` and `GET /v3/pools/status?pool=...` expose per-pool machine-readable decommission state. The `status` field can report `active`, `running`, `queued`, `complete`, `failed`, or `canceled`.
|
||||
|
||||
When decommission metadata is present, `decommissionInfo` includes:
|
||||
|
||||
- queue and terminal flags: `queued`, `complete`, `failed`, `canceled`;
|
||||
- progress counters: `objectsDecommissioned`,
|
||||
`objectsDecommissionedFailed`, `bytesDecommissioned`, and
|
||||
`bytesDecommissionedFailed`;
|
||||
- progress counters: `objectsDecommissioned`, `objectsDecommissionedFailed`, `bytesDecommissioned`, and `bytesDecommissionedFailed`;
|
||||
- current location: `bucket`, `prefix`, and `object`;
|
||||
- queue/history lists: `queuedBuckets` and `decommissionedBuckets`;
|
||||
- `waitingReason`, currently `queued` for queued entries and
|
||||
`waiting_for_worker` when metadata exists but no worker has started.
|
||||
- `waitingReason`: `queued` for queued entries and `waiting_for_worker` when metadata exists but no worker has started.
|
||||
|
||||
This makes queued pools and stalled metadata visible without requiring operators
|
||||
to inspect pool metadata files directly.
|
||||
This makes queued pools and stalled metadata visible without requiring operators to inspect pool metadata files directly.
|
||||
|
||||
## MinIO Divergence Decisions
|
||||
|
||||
This section records the current product decisions for behavior that is close to
|
||||
MinIO but not always byte-for-byte identical.
|
||||
Behavior that is close to MinIO but not byte-for-byte identical. Changing either decision requires an operator compatibility note and updated characterization tests.
|
||||
|
||||
### Empty Delete Markers
|
||||
|
||||
MinIO decommission documentation states that empty delete markers, meaning delete
|
||||
markers with no successor object versions, are not transitioned to another pool.
|
||||
MinIO decommission documentation states that empty delete markers (delete markers with no successor object versions) are not transitioned to another pool. RustFS follows that behavior for decommission when the bucket has no replication configuration: a lone remaining delete marker is cleanup-only metadata and is skipped. When replication is configured, RustFS keeps the delete marker eligible for movement so delete-marker replication and purge state are not lost.
|
||||
|
||||
RustFS follows that behavior for decommission when the bucket has no replication
|
||||
configuration: a lone remaining delete marker is treated as cleanup-only metadata
|
||||
and is skipped. When replication is configured, RustFS intentionally keeps the
|
||||
delete marker eligible for movement so delete-marker replication and purge state
|
||||
are not lost.
|
||||
|
||||
RustFS rebalance uses the same predicate as decommission: skip only a lone delete
|
||||
marker without replication. This is intentional even though MinIO's public
|
||||
documentation calls out the decommission case more explicitly than the rebalance
|
||||
case.
|
||||
|
||||
Regression guards:
|
||||
|
||||
- `should_skip_decommission_delete_marker_characterizes_empty_marker_without_replication`
|
||||
- `should_skip_decommission_delete_marker_characterizes_replication_configured`
|
||||
- `test_should_skip_rebalance_delete_marker_characterizes_empty_marker_without_replication`
|
||||
- `test_should_skip_rebalance_delete_marker_characterizes_replication_configured`
|
||||
Rebalance uses the same predicate as decommission (`should_skip_decommission_delete_marker` in `crates/ecstore/src/core/pools.rs`, `should_skip_rebalance_delete_marker` in `crates/ecstore/src/services/rebalance/worker.rs`), even though MinIO's public documentation calls out the decommission case more explicitly than the rebalance case.
|
||||
|
||||
### Lifecycle-Expired Versions During Cleanup
|
||||
|
||||
MinIO decommission ignores versions that are already expired by lifecycle rules.
|
||||
RustFS follows that decommission behavior by allowing safely expired versions to
|
||||
count toward source cleanup completion.
|
||||
|
||||
RustFS rebalance is intentionally stricter. Expired versions do not prove that a
|
||||
target pool received an equivalent version, so rebalance cleanup requires actual
|
||||
rebalance completion for the source entry instead of treating lifecycle-expired
|
||||
versions as moved.
|
||||
|
||||
Regression guards:
|
||||
|
||||
- `test_should_cleanup_decommission_source_entry_accepts_migrated_and_safely_expired_versions`
|
||||
- `test_should_cleanup_decommission_source_entry_accepts_versions_only_safely_expired_by_lifecycle`
|
||||
- `test_should_cleanup_rebalance_source_entry_rejects_versions_only_expired_by_lifecycle`
|
||||
|
||||
No migration step is required for these decisions because this note documents the
|
||||
current RustFS behavior. Changing either decision later requires an operator
|
||||
compatibility note and updated characterization tests.
|
||||
MinIO decommission ignores versions already expired by lifecycle rules. RustFS applies the same rule to decommission and rebalance: a source entry is cleanup-complete when moved versions plus safely expired versions equal the total version count (`should_cleanup_decommission_source_entry` in `crates/ecstore/src/core/pools.rs`, `should_cleanup_rebalance_source_entry` in `crates/ecstore/src/services/rebalance/worker.rs`). Versions retained by object lock or pending replication are not counted as safely expired by the callers, so an entry with such versions is retained. Both predicates accept an entry whose versions are all lifecycle-expired (`test_should_cleanup_decommission_source_entry_accepts_versions_only_safely_expired_by_lifecycle`, `test_should_cleanup_rebalance_source_entry_accepts_versions_only_expired_by_lifecycle`).
|
||||
|
||||
## Tier Free Versions During Decommission
|
||||
|
||||
A tier free version is an internal xl.meta record (`rustfs_filemeta::FREE_VERSION`,
|
||||
flagged `XL_FLAG_FREE_VERSION`) shaped like a delete marker. It is created by
|
||||
`MetaObject::init_free_version` when a version whose remote transition completed is
|
||||
deleted locally: the visible version is removed and the record keeps the remote-tier
|
||||
identity (tier, object name, version id, state, destination id) needed for an
|
||||
idempotent remote delete. Free versions are not user-visible versions; `num_versions`
|
||||
and all listing/GET paths exclude them.
|
||||
A tier free version is an internal xl.meta record (`rustfs_filemeta::FREE_VERSION`, flagged `XL_FLAG_FREE_VERSION`) shaped like a delete marker. It is created by `MetaObject::init_free_version` when a version whose remote transition completed is deleted locally: the visible version is removed and the record keeps the remote-tier identity (tier, object name, version id, state, destination id) needed for an idempotent remote delete. Free versions are not user-visible versions; `num_versions` and all listing/GET paths exclude them.
|
||||
|
||||
### Lifecycle And Consumers
|
||||
|
||||
Creation: a local delete that removes a version whose transition status is
|
||||
`complete` normally appends the record via `MetaObject::delete_version` →
|
||||
`init_free_version`. User-facing single and batch deletes always retain that
|
||||
historical owner when they actually remove a transitioned source; they do not
|
||||
create a tier journal, probe a fleet capability, or issue a peer mutation RPC.
|
||||
`TransitionVersionState::Unknown` and incomplete destination identities remain
|
||||
on the same conservative free-version path. Delete-marker creation on an Enabled
|
||||
bucket remains unchanged and does not schedule remote deletion.
|
||||
Creation: a local delete that removes a version whose transition status is `complete` normally appends the record via `MetaObject::delete_version` → `init_free_version` (skipped only when `skip_tier_free_version` is set, as on data-movement copies). User-facing single and batch deletes always retain that historical owner when they actually remove a transitioned source; they do not create a tier journal, probe a fleet capability, or issue a peer mutation RPC. `TransitionVersionState::Unknown` and incomplete destination identities stay on the same conservative free-version path. Delete-marker creation on an Enabled bucket is unchanged and does not schedule remote deletion.
|
||||
|
||||
Recursive prefix/delete-all cannot preserve per-object markers across its
|
||||
physical directory purge, so it requires a v6 recoverable journal for every
|
||||
transitioned visible source plus a durable dispatch manifest for the complete
|
||||
operation. It fails closed before mutation on legacy metadata or any existing
|
||||
hidden tier free-version under the prefix. Its internal streaming walk
|
||||
discovers logical keys, then exact-loads every key from its authoritative set in
|
||||
every pool, including free versions; the S3 listing merge is never treated as a
|
||||
complete physical-owner inventory. Tier-operation leases remain held from that
|
||||
preflight through journal prepare and physical deletion. Once physical deletion
|
||||
starts, any error is mutation-ambiguous: authorized/dispatched journals remain
|
||||
for recovery to commit owners only after all physical sets prove both the source
|
||||
and exact free-version identity absent; uncertain owners are retained.
|
||||
If a retry discovers a later transitioned source after the manifest reached
|
||||
`DispatchAuthorized`, it replays only the manifest's immutable predecessor set,
|
||||
completes that operation, and leaves the newcomer for a successor dispatch.
|
||||
Operators may retry after the legacy free-version worker has durably completed
|
||||
remote and local cleanup. Journal-less internal deletes and older nodes retain
|
||||
their established marker behavior.
|
||||
Recursive prefix/delete-all cannot preserve per-object markers across its physical directory purge, so it requires a v6 recoverable journal for every transitioned visible source plus a durable dispatch manifest for the whole operation. It fails closed before mutation on legacy metadata or on any existing hidden tier free-version under the prefix. Its internal streaming walk discovers logical keys, then exact-loads every key from its authoritative set in every pool, including free versions; the S3 listing merge is never treated as a complete physical-owner inventory. Tier-operation leases stay held from that preflight through journal prepare and physical deletion. Once physical deletion starts, any error is mutation-ambiguous: authorized/dispatched journals remain for recovery to commit owners only after all physical sets prove both the source and the exact free-version identity absent; uncertain owners are retained. If a retry discovers a later transitioned source after the manifest reached `DispatchAuthorized`, it replays only the manifest's immutable predecessor set, completes that operation, and leaves the newcomer for a successor dispatch. Operators may retry after the legacy free-version worker has durably completed remote and local cleanup. Journal-less internal deletes and older nodes keep their established marker behavior.
|
||||
|
||||
Consumption while the record exists: the background recovery loop started by
|
||||
`init_background_expiry` (spawned by `spawn_tier_free_version_recovery_once`,
|
||||
enabled by default) scans disks for pending records and re-enqueues them; the
|
||||
usage scanner does the same; the lifecycle worker then deletes the remote tier
|
||||
object idempotently and only afterwards removes the local record. Heal walks
|
||||
include free-version records in metadata healing. Transition planning,
|
||||
replication, restore, GET, listings, and usage aggregation never depend on
|
||||
them.
|
||||
Consumption while the record exists: the background recovery loop started by `init_background_expiry` (spawned by `spawn_tier_free_version_recovery_once`, enabled by default) scans disks for pending records and re-enqueues them; the usage scanner does the same; the lifecycle worker then deletes the remote tier object idempotently and only afterwards removes the local record. Heal walks include free-version records in metadata healing. Transition planning, replication, restore, GET, listings, and usage aggregation never depend on them.
|
||||
|
||||
### Decommission Handling
|
||||
|
||||
The exact decommission inventory loader (`load_file_info_versions_exact` via
|
||||
`get_all_file_info_versions`) keeps free-version records inline in `versions`.
|
||||
The migration loop handles them before lifecycle expiry and delete-marker
|
||||
shortcuts. It selects a target pool using the free-version-aware lookup, then
|
||||
writes the original free record to every target disk with the normal metadata
|
||||
write quorum. The free-version marker, local version id, transition identity,
|
||||
transition state, and destination id are preserved at the FileInfo/metadata
|
||||
boundary.
|
||||
The exact decommission inventory loader (`load_file_info_versions_exact` via `get_all_file_info_versions`) keeps free-version records inline in `versions`. The migration loop handles them before lifecycle expiry and delete-marker shortcuts. It selects a target pool using the free-version-aware lookup, then writes the original free record to every target disk with the normal metadata write quorum. The free-version marker, local version id, transition identity, transition state, and destination id are preserved at the FileInfo/metadata boundary.
|
||||
|
||||
The source record is physically removed only after the target write quorum has
|
||||
committed and the source cleanup preflight still matches the exact inventory.
|
||||
If the lifecycle worker has already completed the remote delete and removed the
|
||||
source record before decommission acquires the source lock, decommission records
|
||||
that identity as already consumed and treats the missing source record as safe.
|
||||
If target capacity, metadata validation, lock fencing, or quorum fails, the
|
||||
source record remains and the entry records `state = "free_version_retained"`
|
||||
with reason `tier_free_version_migration_failed`; the worker retries the
|
||||
operation on a later pass. A target record with the same version id is accepted
|
||||
only when its free-version identity matches; a conflicting ordinary version or
|
||||
different free record is an overwrite error. This makes retries idempotent and
|
||||
prevents a free record from replacing a user-visible version.
|
||||
The source record is physically removed only after the target write quorum has committed and the source cleanup preflight still matches the exact inventory. If the lifecycle worker has already completed the remote delete and removed the source record before decommission acquires the source lock, decommission records that identity as already consumed and treats the missing source record as safe. If target capacity, metadata validation, lock fencing, or quorum fails, the source record remains and the entry records `state = "free_version_retained"` with reason `tier_free_version_migration_failed`; the worker retries the operation on a later pass. A target record with the same version id is accepted only when its free-version identity matches; a conflicting ordinary version or different free record is an overwrite error. This makes retries idempotent and prevents a free record from replacing a user-visible version.
|
||||
|
||||
`TransitionVersionState::Unknown` records are migrated unchanged rather than discarded; the lifecycle worker retains them if remote identity validation cannot make a delete request. Only an authorized recursive prefix/delete-all v6 transaction may use a per-source journal as the sole retry source; ordinary single/batch deletes never take that path, and a journal discovered alongside an older or fallback free-version never authorizes dropping the xl.meta record.
|
||||
|
||||
### Remote-Tuple Publication Fence
|
||||
|
||||
Cross-pool capability v3 includes a commit-late publication contract for every
|
||||
path that can copy an existing transition tuple to a new physical owner. This
|
||||
capability version is independent of the tier-mutation RPC protocol version.
|
||||
A mixed fleet whose minimum cross-pool capability is below v3 cannot authorize
|
||||
journal-v6 remote deletion.
|
||||
Cross-pool capability v3 adds a commit-late publication contract for every path that can copy an existing transition tuple to a new physical owner. This capability version is independent of the tier-mutation RPC protocol version; a mixed fleet whose minimum cross-pool capability is below v3 cannot authorize journal-v6 remote deletion.
|
||||
|
||||
Data movement captures a non-cloneable, process-local source capability before
|
||||
copying, but it does not hold a namespace write lock or tier-operation lease
|
||||
while reading a large body or uploading multipart parts. `NewMultipartUpload`
|
||||
and `UploadPart` are staging only. Immediately before single-PUT rename,
|
||||
Multipart Complete, or a pure-remote/free-version metadata quorum write, the
|
||||
final consumer acquires the exact tier generation (when a remote tuple exists),
|
||||
then fixed/source/target write domains in stable order. The fixed domain is used
|
||||
only for a real remote-tuple decommission publisher; an ordinary local object
|
||||
keeps the lighter source/target commit scope.
|
||||
Data movement captures a non-cloneable, process-local source capability before copying, but it does not hold a namespace write lock or tier-operation lease while reading a large body or uploading multipart parts (`NewMultipartUpload` and `UploadPart` are staging only). Immediately before single-PUT rename, Multipart Complete, or a pure-remote/free-version metadata quorum write, the final consumer acquires the exact tier generation (when a remote tuple exists), then the fixed/source/target write domains in stable order. The fixed domain is used only for a real remote-tuple decommission publisher; an ordinary local object keeps the lighter source/target commit scope.
|
||||
|
||||
While that owned scope is held, the publisher re-reads the exact source pool and
|
||||
compares version, data directory, modification time, ETag, checksums, transition
|
||||
tuple, transition-version state, and destination identity. A missing or changed
|
||||
source, changed/revoked tier generation, bucket incarnation change, or lost lock
|
||||
fails before target rename. The scope remains owned through rename quorum and
|
||||
the existing rename-tail guard handoff. Consequently, recovery-first ordering
|
||||
cannot delete the remote object and then have a stale restored-transitioned
|
||||
rebalance recreate its tuple; publisher-first ordering makes recovery wait and
|
||||
rescan the newly committed owner.
|
||||
While that owned scope is held, the publisher re-reads the exact source pool and compares version, data directory, modification time, ETag, checksums, transition tuple, transition-version state, and destination identity. A missing or changed source, a changed or revoked tier generation, a bucket incarnation change, or a lost lock fails before target rename. The scope stays owned through rename quorum and the rename-tail guard handoff, so recovery-first ordering cannot delete the remote object and then let a stale restored-transitioned rebalance recreate its tuple, and publisher-first ordering makes recovery wait and rescan the newly committed owner.
|
||||
|
||||
Full cross-key S3 Copy is not an ownership-sharing operation: it materializes
|
||||
local data and strips transition, destination, transaction, and free-version
|
||||
keys. Same-key metadata/version-only updates preserve the existing protected
|
||||
state. Admin heal keeps the legacy `nolock` request field for wire compatibility
|
||||
but ignores it as lock authority; final heal writes enter the normal locked
|
||||
path. Restore similarly ignores ambient `ObjectOptions.no_lock`, acquires its
|
||||
own commit-late PUT/Complete lock, validates the restore operation id, and keeps
|
||||
an exact tier generation lease through the local commit.
|
||||
Full cross-key S3 Copy is not an ownership-sharing operation: it materializes local data and strips transition, destination, transaction, and free-version keys. Same-key metadata/version-only updates preserve the protected state. Admin heal keeps the legacy `nolock` request field for wire compatibility but ignores it as lock authority; final heal writes enter the normal locked path. Restore likewise ignores ambient `ObjectOptions.no_lock`, acquires its own commit-late PUT/Complete lock, validates the restore operation id, and keeps an exact tier generation lease through the local commit.
|
||||
|
||||
### Reference-Audit Result
|
||||
### Tier Mutation Protocol And Journal v6 Rollout
|
||||
|
||||
After migration, user-facing GET/list/transition/replication/restore paths still
|
||||
exclude the record. Recovery, usage scanning, lifecycle tier cleanup, and heal
|
||||
continue to see a legacy/fallback record when they request free versions, so an
|
||||
unresolved remote delete remains actionable on the target pool. Only an
|
||||
authorized recursive prefix/delete-all v6 transaction may instead use a
|
||||
per-source journal as the sole retry source; ordinary single/batch deletes never
|
||||
take that path. A journal discovered
|
||||
alongside an older or fallback free-version does not authorize dropping the
|
||||
record. In particular, `Unknown` transition state records are migrated unchanged
|
||||
rather than discarded: the lifecycle worker retains them if remote identity
|
||||
validation cannot make a delete request.
|
||||
Tier edit/remove/clear reference proof uses the internal walk with `include_free_versions = true`, in addition to persisted journal and transition-transaction checks. Protocol v3 peer Prepare blocks new reference creators and drains existing tier-operation leases before this proof; protocol v4 preserves that state machine and adds a signed failure classification. Abort carries the canonical Prepare intent, so a peer can create an identity-bound `Aborted` tombstone even when Abort overtakes Prepare; a delayed matching Prepare then converges on `Aborted` instead of reinstalling the block, and a conflicting intent with the same mutation id fails closed. The tombstone stays durable until intent expiry plus the configured clock-skew allowance, including across reload and coordinator-record cleanup. After expiry, a missing-record replay of the original signed Prepare is rejected and cannot recreate a peer-only runtime fence. Abort checks an existing same-identity terminal record before consulting mutable current-config proof, and recovery reconstructs the original Prepared revision for Abort fanout.
|
||||
|
||||
Tier edit/remove/clear reference proof uses the internal walk with
|
||||
`include_free_versions = true`, in addition to persisted journal and transition
|
||||
transaction checks. Protocol v3 peer Prepare blocks new reference creators and
|
||||
drains existing tier-operation leases before this proof; protocol v4 preserves
|
||||
that state machine and adds a signed failure classification. Abort carries the
|
||||
canonical Prepare intent, so a peer can create an identity-bound `Aborted`
|
||||
tombstone even when Abort overtakes Prepare. A delayed matching Prepare then
|
||||
converges on `Aborted` instead of reinstalling the block; a conflicting intent
|
||||
with the same mutation id fails closed. The tombstone remains durable until the
|
||||
intent expiry plus the configured clock-skew allowance, including across reload
|
||||
and coordinator-record cleanup. After
|
||||
expiry, a missing-record replay of the original signed Prepare is rejected and
|
||||
cannot recreate a peer-only runtime fence. Abort checks an existing same-identity
|
||||
terminal record before consulting mutable current-config proof, and recovery
|
||||
reconstructs the original Prepared revision for Abort fanout.
|
||||
A new server accepts both v3 and v4 requests and selects the matching canonical response proof. During a mixed rollout an older v3 server rejects a v4 request with an authenticated, byte-exact unsupported-version status before dispatch; the v4 coordinator treats only that exact rejection as definitely-not-installed, fails the admin mutation, and does not send the peer an incompatible Abort. There is deliberately no automatic v3 retry: `Unimplemented`, near-text, timeouts, missing or unknown failure classes, and other ambiguous outcomes still receive Abort and retain the coordinator retry record if Abort cannot be proven. Operators must pause and drain tier edit/remove/clear operations before starting a rolling upgrade, leave them disabled while any v3-only peer remains, and resume only after every topology member advertises the v4-capable release. Ordinary object I/O and free-version cleanup stay available; `xl.meta` is unchanged by a rejected mutation.
|
||||
|
||||
A new server accepts both v3 and v4 requests and selects the matching canonical
|
||||
response proof. During a mixed rollout, an older v3 server rejects a v4 request
|
||||
with an authenticated, byte-exact unsupported-version status before dispatch;
|
||||
the v4 coordinator treats only that exact rejection as definitely not installed,
|
||||
fails the admin mutation, and does not send the peer an incompatible Abort.
|
||||
There is deliberately no automatic v3 retry. `Unimplemented`, near-text,
|
||||
timeouts, missing/unknown failure classes, and other ambiguous outcomes still
|
||||
receive Abort and retain the coordinator retry record if Abort cannot be proven.
|
||||
Operators must pause and drain tier edit/remove/clear operations before starting
|
||||
the rolling upgrade, leave them disabled while any v3-only peer remains, and
|
||||
resume only after every topology member advertises the v4-capable release.
|
||||
Ordinary object I/O and free-version cleanup remain available; xl.meta is
|
||||
unchanged by the rejected mutation.
|
||||
Sole-owner transactions use journal v6: v5-and-older readers reject and retain those records, so an old recovery worker cannot bypass the all-pool proof. Older nodes may keep creating fallback free-versions until the rollout is homogeneous. Do not downgrade every v6-aware recovery worker while any v6 record remains; drain the journal first or keep at least one v6-aware worker until cleanup converges.
|
||||
|
||||
Sole-owner transactions use journal v6: v5-and-older readers reject and retain
|
||||
those records, so an old recovery worker cannot bypass the all-pool proof. Older
|
||||
nodes may continue to create fallback free-versions until the rollout is
|
||||
homogeneous. A deployment must not downgrade every v6-aware recovery worker
|
||||
while any v6 record remains; drain the journal first or keep at least one v6-aware
|
||||
worker until cleanup converges.
|
||||
### Disposition Events
|
||||
|
||||
Each migrated record emits `state = "free_version_migrated"` with reason
|
||||
`tier_free_version_migrated`. A record consumed before migration emits
|
||||
`state = "free_version_consumed"` with reason
|
||||
`tier_free_version_already_consumed`. Each failed record emits the retained state
|
||||
and failure reason above. The entry also emits a disposition summary with
|
||||
migrated, consumed, retained, and total counts. The final decommission sweep uses
|
||||
the exact loader, counts free records still present, and emits one retained
|
||||
record/reason for each unresolved free version before failing the sweep. This
|
||||
makes successful migration, completed cleanup, and retained cleanup obligations
|
||||
visible instead of silently omitting free records.
|
||||
Free versions remain internal, so no S3-visible version or admin response field is added. The structured `decommission_entry` events are the operational status surface:
|
||||
|
||||
No new S3-visible version or admin response field is needed: free versions remain
|
||||
internal and are never counted as user-visible versions. The structured
|
||||
`decommission_entry` events are the operational status surface for the
|
||||
free-version disposition; the existing decommission item/failed counters still
|
||||
report the enclosing object migration result.
|
||||
| Outcome | `state` | `reason` |
|
||||
|---|---|---|
|
||||
| Record migrated to the target pool | `free_version_migrated` | `tier_free_version_migrated` |
|
||||
| Record consumed by the lifecycle worker before migration | `free_version_consumed` | `tier_free_version_already_consumed` |
|
||||
| Migration failed, source retained for retry | `free_version_retained` | `tier_free_version_migration_failed` |
|
||||
|
||||
Regression guard:
|
||||
The entry also emits a disposition summary with migrated, consumed, retained, and total counts. The final decommission sweep uses the exact loader, counts free records still present, and emits one retained record/reason per unresolved free version before failing the sweep. The existing decommission item/failed counters still report the enclosing object migration result.
|
||||
|
||||
- `decommission_tier_free_version_preserves_remote_identity`
|
||||
- `decommission_tier_free_version_resume_requires_write_quorum`
|
||||
- `decommission_tier_free_version_commit_rejects_lost_fence`
|
||||
- `test_decommission_cleanup_preflight_accepts_migrated_free_version_consumed_from_source`
|
||||
- `decommission_entry_skips_cleanup_only_marker_when_free_version_is_present`
|
||||
- `decommission_entry_rejects_subquorum_free_version_conflict_and_retains_source`
|
||||
## Regression Guards
|
||||
|
||||
## Regression Guard
|
||||
Test names drift; locate the current guards instead of copying them:
|
||||
|
||||
The queued multi-pool contract is guarded by:
|
||||
|
||||
- `test_contextualized_decommission_start_request_allows_multiple_target_pools`
|
||||
- `test_decommission_start_local_leader_allows_remote_queued_pool`
|
||||
- `test_local_decommission_queue_prefix_stops_at_remote_leader`
|
||||
- `test_decommission_peer_target_returns_none_for_local_first_endpoint`
|
||||
- `test_pool_meta_queued_decommission_is_not_suspended_until_promoted`
|
||||
- `test_pool_meta_promoted_queued_decommission_can_be_canceled`
|
||||
- `test_first_resumable_decommission_queue_indices_stops_at_failed_or_canceled_state`
|
||||
- `test_first_resumable_decommission_queue_indices_allows_after_completed_prefix`
|
||||
- `admin_pool_list_item_exposes_queued_decommission_state`
|
||||
|
||||
These tests live in `crates/ecstore/src/core/pools.rs` and
|
||||
`rustfs/src/app/admin_usecase.rs`.
|
||||
```bash
|
||||
rg -n 'fn [a-z_]*decommission[a-z_]*\(' crates/ecstore/src/core/pools.rs crates/ecstore/src/set_disk/mod.rs crates/ecstore/src/data_movement/mod.rs crates/ecstore/src/store/init.rs rustfs/src/admin/handlers/pools.rs rustfs/src/app/admin_usecase.rs
|
||||
rg -n 'fn test_should_[a-z_]*rebalance[a-z_]*\(' crates/ecstore/src/services/rebalance/rebalance_unit_tests.rs
|
||||
```
|
||||
|
||||
@@ -1,76 +1,65 @@
|
||||
# ECStore API Facade Inventory
|
||||
|
||||
This inventory records the current `rustfs_ecstore::api` compatibility surface
|
||||
before any ECStore split PR removes or narrows re-exports. It is a planning and
|
||||
guardrail document only. It must not be used as approval to move lifecycle,
|
||||
replication, or `SetDisks` runtime behavior.
|
||||
**Use this when:** you need something from `rustfs_ecstore` in another crate, you are narrowing a `rustfs_ecstore::api` facade group, or the architecture guard reports a facade bypass.
|
||||
**Source of truth:** `crates/ecstore/src/api/mod.rs` (facade groups), the boundary files listed below, and the facade rules in `scripts/check_architecture_migration_rules.sh`.
|
||||
|
||||
The broad `rustfs_ecstore::api` facade is a compatibility boundary, not an architecture target. It shrinks monotonically and only through guarded changes; it is never approval to move lifecycle, replication, or `SetDisks` runtime behavior.
|
||||
|
||||
## Facade Group Inventory
|
||||
|
||||
| Facade group | Current role | Shrink posture |
|
||||
| Facade group | Role | Shrink posture |
|
||||
|---|---|---|
|
||||
| `storage`, `layout`, `error`, `runtime`, `cluster`, `rpc` | Compatibility spine for storage, topology, runtime handles, cluster control, and internode calls. | Keep until replacement contracts compile in downstream boundary files. |
|
||||
| `bucket` | Domain facade consumed through owner-local `storage_api` boundaries. The public API keeps compatibility paths but exposes explicit submodules and symbol lists instead of whole bucket owner modules. | Keep explicit lists aligned with owner-boundary consumers; do not restore whole-module passthroughs. |
|
||||
| `client`, `config`, `disk`, `tier` | Compatibility paths consumed through owner-local `storage_api` boundaries. The public API keeps existing path names but exposes explicit nested submodules and symbol lists instead of whole owner modules. | Keep explicit lists aligned with owner-boundary consumers; do not restore whole-module passthroughs. |
|
||||
| `data_usage`, `capacity`, `notification`, `metrics`, `rebalance` | Domain and service facades still consumed through owner-local `storage_api` boundaries. | Narrow one group at a time after explicit aliases or wrappers exist. |
|
||||
| `set_disk`, `object`, `rio`, `bitrot`, `erasure`, `compression`, `cache`, `store_list` | Low-level object IO, reader, erasure, cache, and migration helper compatibility. | Keep stable while `SetDisks` remains the shared state carrier. |
|
||||
| `admin`, `event`, `global` | Admin, event hook, and legacy global compatibility. | Keep `global` limited to bootstrap writes and lifecycle controls; read-only runtime access must use runtime-source contracts. |
|
||||
| `bucket` | Domain facade consumed through owner-local `storage_api` boundaries; explicit submodules and symbol lists, never whole bucket owner modules. | Keep lists aligned with boundary consumers; never restore whole-module passthroughs. |
|
||||
| `config`, `disk`, `tier` | Compatibility paths with explicit nested submodules and symbol lists. | Same as `bucket`. |
|
||||
| `data_usage`, `capacity`, `notification`, `metrics`, `rebalance` | Domain and service facades consumed through owner-local boundaries. | Narrow one group at a time after explicit aliases or wrappers exist. |
|
||||
| `set_disk`, `object`, `object_api_utils`, `rio`, `bitrot`, `erasure`, `compression`, `cache`, `store_list` | Low-level object IO, reader, erasure, cache, and migration helper compatibility. | Keep stable while `SetDisks` remains the shared state carrier. |
|
||||
| `admin`, `event`, `global` | Admin, event hook, and bootstrap-global compatibility. | `global` is limited to bootstrap writes and lifecycle controls; read-only runtime access goes through `runtime`. |
|
||||
|
||||
The S3 client is no longer a facade group: it lives in `crates/s3-client` (`rustfs_s3_client`). Regenerate the group list with:
|
||||
|
||||
```bash
|
||||
rg -n '^pub mod ' crates/ecstore/src/api/mod.rs
|
||||
```
|
||||
|
||||
## External Consumer Boundaries
|
||||
|
||||
External `rustfs_ecstore::api` imports must stay in these local boundary files:
|
||||
External `rustfs_ecstore::api` imports stay in these local boundary files:
|
||||
|
||||
| Boundary file | Current facade families |
|
||||
| Boundary file | Facade families consumed |
|
||||
|---|---|
|
||||
| `rustfs/src/storage/storage_api.rs` | Broad RustFS storage owner bridge for admin, explicit bucket facade submodules, capacity, client, compression, cluster, config, data usage, disk, error, event, global bootstrap controls, runtime-source getters, layout, metrics, notification, rebalance, rio, rpc, set disk, storage, and tier. Replication pool/stat handles are projected into RustFS-local wrapper types here. |
|
||||
| `crates/scanner/src/storage_api.rs` | Scanner bridge for bucket lifecycle, replication, metadata, capacity, config, data usage, disk, error, runtime, set disk, storage, and tier. Replication queue config, admission, and heal object DTOs are projected into scanner-local types here. |
|
||||
| `crates/obs/src/metrics/storage_api.rs` | Metrics bridge for bucket bandwidth, lifecycle, replication, quota, capacity, data usage, error, runtime, and storage. |
|
||||
| `crates/iam/src/storage_api.rs` | IAM bridge for config, error, notification, runtime, and storage. |
|
||||
| `crates/heal/src/heal/storage_api.rs` | Heal bridge for data usage, disk, error, runtime, and storage. |
|
||||
| `crates/notify/src/storage_api.rs` | Notification bridge for config, runtime, and storage. |
|
||||
| `crates/protocols/src/swift/storage_api.rs` | Swift bridge for bucket metadata, bucket metadata system, error, runtime, and storage. |
|
||||
| `crates/s3select-api/src/storage_api.rs` | S3 Select bridge for error, runtime, set disk, and storage. |
|
||||
| `crates/e2e_test/src/storage_api.rs` | E2E harness bridge for bucket targets, disk walking, and RPC helpers. |
|
||||
| `crates/heal/tests/*/storage_api.rs`, `crates/scanner/tests/storage_api/mod.rs`, `fuzz/fuzz_targets/*_storage_api.rs` | Test and fuzz bridges for the same compatibility seams under test. |
|
||||
| `rustfs/src/storage/storage_api.rs` | Broad storage-owner bridge: admin, bucket submodules, capacity, compression, cluster, config, data usage, disk, error, event, global bootstrap controls, runtime getters, layout, metrics, notification, rebalance, rio, rpc, set disk, storage, tier. Replication pool/stat handles are projected into RustFS-local wrapper types here. |
|
||||
| `rustfs/src/storage_api.rs`, `rustfs/src/admin/storage_api.rs`, `rustfs/src/app/storage_api.rs` | Root, admin, and app owner boundaries: explicit aliases only, no `metadata`, `metadata_sys`, `quota`, `com`, or bare `init` module passthroughs; object and error aliases anchor on storage-api associated types and a local `StorageError`. |
|
||||
| `crates/scanner/src/storage_api.rs` | Bucket lifecycle, replication, metadata, capacity, config, data usage, disk, error, runtime, set disk, storage, tier. Replication queue config, admission, and heal object DTOs are projected into scanner-local types. |
|
||||
| `crates/obs/src/metrics/storage_api.rs` | Bucket bandwidth, lifecycle, replication, quota, capacity, data usage, error, runtime, storage; data usage is consumed as a local DTO projection. |
|
||||
| `crates/iam/src/storage_api.rs` | Config, error, notification, runtime, storage. |
|
||||
| `crates/heal/src/heal/storage_api.rs` | Data usage, disk, error, runtime, storage. |
|
||||
| `crates/notify/src/storage_api.rs` | Config, runtime, storage; no broad `config` or `global` module imports. |
|
||||
| `crates/protocols/src/swift/storage_api.rs` | Bucket metadata, bucket metadata system, error, runtime, storage. |
|
||||
| `crates/s3select-api/src/storage_api.rs` | Error, runtime, set disk, storage. |
|
||||
| `crates/e2e_test/src/storage_api.rs` | E2E harness bridge for bucket targets, disk walking, and RPC helpers; no grouped RPC passthroughs. |
|
||||
| `crates/ecstore/tests/storage_api.rs`, `crates/heal/tests/storage_api.rs`, `crates/scanner/tests/storage_api/mod.rs`, `fuzz/fuzz_targets/*_storage_api.rs` | Test and fuzz bridges: direct aliases or local wrappers; fuzz harnesses wrap bucket utility entrypoints instead of grouped passthroughs. |
|
||||
| `crates/test-utils/src/ecstore_test_compat.rs`, `crates/iam/tests/ecstore_test_compat/mod.rs`, `crates/protocols/tests/ecstore_test_compat/mod.rs` | Test-only compatibility harnesses that import the facade directly for fixture setup. |
|
||||
|
||||
New production imports outside these boundary files are migration drift. Add a
|
||||
local boundary or storage-api contract first, then route consumers through it.
|
||||
`crates/replication/src/storage_api.rs` shares the file name but is not an ECStore boundary: it owns the delete work DTOs of `rustfs-replication`, which imports neither `rustfs_ecstore` nor `rustfs-storage-api`.
|
||||
|
||||
Regenerate the boundary list with:
|
||||
|
||||
```bash
|
||||
rg -l 'rustfs_ecstore::api' crates rustfs/src fuzz -g '*.rs' -g '!crates/ecstore/src/**'
|
||||
```
|
||||
|
||||
New production imports outside these files are migration drift. Do not add direct `rustfs_ecstore::api` imports outside the boundary files; add a local boundary or a storage-api contract first, then route consumers through it.
|
||||
|
||||
## Split Dependency Inventory
|
||||
|
||||
| Candidate | ECStore dependencies that block a crate split | Required owner contracts before movement |
|
||||
|---|---|---|
|
||||
| Lifecycle | Object API, `ECStore`, `SetDisks`, runtime sources/globals, bucket metadata/versioning/object lock/replication, disk, config, notification, audit, and tier services. | `LifecycleObjectStore`, `LifecycleMetadataStore`, `LifecycleRuntime`, `LifecycleReplicationSink`, and `LifecycleAuditSink`. |
|
||||
| Replication | Bucket target and metadata systems, bucket target client config, disk, object API, runtime sources, notification, and SetDisks lock timing. | `ReplicationStorage`, `ReplicationMetadataStore`, `ReplicationRuntime`, `ReplicationEventSink`, and `ReplicationLifecycleBridge`. |
|
||||
| SetDisks | Shared disks, endpoints, format state, namespace locks, cache, and implementations for object IO, namespace locking, bucket, object, list, multipart, and heal operations. | Pure shard source, disk error, bitrot IO, namespace lock, metrics label, and file metadata contracts before any operation family moves. |
|
||||
|
||||
`crates/ecstore/tests/ecstore_contract_compat_test.rs` keeps compile-time
|
||||
coverage for `ECStore` and `SetDisks` storage-api trait compatibility before
|
||||
any facade shrink or operation-family movement.
|
||||
Lifecycle, replication, and `SetDisks` split blockers, extracted contracts, and guard rule names are tracked in [ecstore-module-split-plan.md](ecstore-module-split-plan.md) and the module inventories `crates/ecstore/src/bucket/lifecycle/README.md` and `crates/ecstore/src/bucket/replication/README.md`. `crates/ecstore/tests/ecstore_contract_compat_test.rs` keeps compile-time coverage for `ECStore` and `SetDisks` storage-api trait compatibility before any facade shrink or operation-family movement.
|
||||
|
||||
## Shrink Rules
|
||||
|
||||
1. Do not remove a facade item until its downstream boundary has compile-time
|
||||
coverage or a documented replacement.
|
||||
2. Do not add direct `rustfs_ecstore::api` imports outside the boundary files
|
||||
listed above.
|
||||
3. Do not split lifecycle or replication into crates while they depend on
|
||||
ECStore runtime state, queues, notification, audit, scanner, or SetDisks
|
||||
internals.
|
||||
4. Do not replace `SetDisks` with multiple runtime structs in one PR. Move one
|
||||
operation family only after contracts and focused tests exist.
|
||||
5. Remove or narrow one facade group per PR so rollback preserves object IO,
|
||||
quorum, lifecycle/replication queues, scanner repair, notification/audit
|
||||
events, and metadata compatibility.
|
||||
6. Keep `api::bucket`, `api::client`, `api::config`, `api::disk`, and
|
||||
`api::tier` on explicit submodules and symbol lists; do not restore
|
||||
`pub use crate::<owner>::{...}` whole-module passthroughs for those groups.
|
||||
|
||||
## First PR Checklist
|
||||
|
||||
- inventory the facade group and all external boundary consumers;
|
||||
- add explicit aliases or wrappers before deleting any broad passthrough;
|
||||
- run `./scripts/check_architecture_migration_rules.sh`;
|
||||
- run focused compile or tests for the touched owner boundary;
|
||||
- keep runtime behavior unchanged unless the PR is explicitly a code-bearing
|
||||
follow-up with its own rollback plan.
|
||||
1. Do not remove a facade item until its downstream boundary has compile-time coverage or a documented replacement.
|
||||
2. Do not add direct `rustfs_ecstore::api` imports outside the boundary files listed above.
|
||||
3. Do not split lifecycle or replication into crates while they depend on ECStore runtime state, queues, notification, audit, scanner, or `SetDisks` internals.
|
||||
4. Do not replace `SetDisks` with multiple runtime structs in one change; move one operation family only after contracts and focused tests exist.
|
||||
5. Remove or narrow one facade group per change so rollback preserves object IO, quorum, lifecycle/replication queues, scanner repair, notification/audit events, and metadata compatibility.
|
||||
6. Keep `api::bucket`, `api::config`, `api::disk`, and `api::tier` on explicit submodules and symbol lists; do not restore `pub use crate::<owner>::{...}` whole-module passthroughs for those groups.
|
||||
|
||||
@@ -1,199 +0,0 @@
|
||||
# ECStore Config Consumer Inventory
|
||||
|
||||
This inventory is the Phase 0 baseline for moving
|
||||
`rustfs_ecstore::config::{Config, KV, KVS}` safely. It records the current
|
||||
definitions, persistence helpers, global accessors, and direct consumers before
|
||||
any contract extraction, global-state migration, or crate split.
|
||||
|
||||
Related issue: [`rustfs/backlog#660`](https://github.com/rustfs/backlog/issues/660)
|
||||
|
||||
## Scope
|
||||
|
||||
In scope:
|
||||
|
||||
- `rustfs_ecstore::config::KV`
|
||||
- `rustfs_ecstore::config::KVS`
|
||||
- `rustfs_ecstore::config::Config`
|
||||
- `rustfs_ecstore::config::DEFAULT_KVS`
|
||||
- `rustfs_ecstore::config::{get_global_server_config, set_global_server_config}`
|
||||
- `rustfs_ecstore::config::com::{read_config_without_migrate, save_server_config}`
|
||||
- Consumers that persist, clone, inspect, mutate, or pass these types across
|
||||
runtime boundaries.
|
||||
- Selected adjacent users of `rustfs_ecstore::config::com::{read_config,
|
||||
save_config, delete_config}` and related helper variants are listed separately
|
||||
when they appear outside the core `Config`, `KV`, and `KVS` consumer map.
|
||||
This is not a complete `com.rs` move inventory; any future `com.rs` move must
|
||||
first inventory ECStore-internal persistence helper users too.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Unrelated `Config` types from `rustfs::config`, SDKs, TLS, SSH, KMS, OIDC
|
||||
client libraries, or local module-specific config structs.
|
||||
- Storage-class-only imports are not treated as `Config`, `KV`, or `KVS`
|
||||
consumers unless they also use the server-config model.
|
||||
- Pure route/action snapshot work already covered by
|
||||
[`admin-route-action-snapshot.md`](admin-route-action-snapshot.md).
|
||||
|
||||
## Current Shape
|
||||
|
||||
Arrows show current source dependency or call direction: the left node imports
|
||||
or calls the right node.
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
EC["crates/ecstore/src/config"]
|
||||
Store["crates/ecstore/src/store/mod.rs"]
|
||||
AppCtx["rustfs/src/app/context.rs"]
|
||||
Server["rustfs/src/server/{event,audit}.rs"]
|
||||
Admin["rustfs/src/admin"]
|
||||
Notify["crates/notify"]
|
||||
Audit["crates/audit"]
|
||||
Targets["crates/targets"]
|
||||
IAM["crates/iam/src/oidc.rs"]
|
||||
Scanner["crates/scanner/src/{runtime_config,scanner}.rs"]
|
||||
|
||||
Store --> EC
|
||||
AppCtx --> EC
|
||||
Server --> AppCtx
|
||||
Admin --> EC
|
||||
Notify --> EC
|
||||
Audit --> EC
|
||||
Targets --> EC
|
||||
Notify --> Targets
|
||||
Audit --> Targets
|
||||
IAM --> EC
|
||||
Scanner --> EC
|
||||
```
|
||||
|
||||
The config model is currently both a persisted server-config representation and
|
||||
the runtime carrier for notify, audit, target-plugin, scanner, and OIDC
|
||||
settings. Any move must preserve that dual role until consumers are migrated
|
||||
behind narrower contracts.
|
||||
|
||||
## Core Model And Global State
|
||||
|
||||
| Item | Current owner | Current role | Migration note |
|
||||
|---|---|---|---|
|
||||
| `KV` | `crates/ecstore/src/config/mod.rs` | Key/value entry with `hidden_if_empty` metadata and serde compatibility. | Preserve field names, aliases, defaults, and redaction behavior before any model move. |
|
||||
| `KVS(Vec<KV>)` | `crates/ecstore/src/config/mod.rs` | Ordered key/value set used by server config, target factories, admin rendering, tests, and examples. | Preserve tuple shape and methods: `new`, `get`, `lookup`, `is_empty`, `keys`, `insert`, `extend`. |
|
||||
| `Config(HashMap<String, HashMap<String, KVS>>)` | `crates/ecstore/src/config/mod.rs` | Server config map by subsystem and target. | Direct `.0` access is widespread; add wrappers only after preserving the current public shape. |
|
||||
| `DEFAULT_KVS` | `crates/ecstore/src/config/mod.rs` | Registry for defaults across storage class, scanner, notify, audit, and OIDC. | Move defaults only after an explicit registration contract exists. |
|
||||
| `GLOBAL_SERVER_CONFIG` | `crates/ecstore/src/config/mod.rs` | Process-wide mutable server config snapshot. | Migrate readers behind `AppContext` or a server-config provider before changing storage. |
|
||||
| `ConfigSys::init` | `crates/ecstore/src/config/mod.rs` | Reads persisted config, looks up derived config, and stores the global snapshot. | Startup order must remain unchanged until the lifecycle contract owns this dependency. |
|
||||
| `read_config_without_migrate` | `crates/ecstore/src/config/com.rs` | Loads persisted server config through ECStore-owned object I/O and storage-admin contracts. | Persistence stays in `ecstore` until pure model and persistence are separated. |
|
||||
| `save_server_config` | `crates/ecstore/src/config/com.rs` | Persists the canonical server config object. | Preserve external object shape and config-history behavior. |
|
||||
| `get_global_server_config` / `set_global_server_config` | `crates/ecstore/src/config/mod.rs` | Clone/read and replace the global server-config snapshot. | Do not remove until all runtime readers have an injected provider path. |
|
||||
|
||||
## Consumer Map
|
||||
|
||||
### ECStore Ownership, Persistence, And Defaults
|
||||
|
||||
| Files | Current usage |
|
||||
|---|---|
|
||||
| `crates/ecstore/src/config/mod.rs` | Defines `KV`, `KVS`, `Config`, defaults, global snapshot, initialization, and tests. |
|
||||
| `crates/ecstore/src/config/com.rs` | Encodes, decodes, reads, writes, creates, and normalizes server config objects through ECStore-local persistence helpers. |
|
||||
| `crates/ecstore/src/config/{notify,audit,oidc,scanner,storageclass}.rs` | Register default `KVS` values and subsystem-specific parsing helpers. |
|
||||
| `crates/ecstore/src/store/mod.rs` | Exposes store-level server-config accessors that delegate to the global config snapshot. |
|
||||
|
||||
### App Context And Server Startup Consumers
|
||||
|
||||
| Files | Current usage |
|
||||
|---|---|
|
||||
| `rustfs/src/app/context.rs` | Defines `ServerConfigInterface`, keeps an `AppContext` server-config handle, and still falls back to `get_global_server_config`. |
|
||||
| `rustfs/src/server/event.rs` | Resolves server config through app context/global fallback before starting the notification runtime. |
|
||||
| `rustfs/src/server/audit.rs` | Resolves server config through app context/global fallback before starting the audit runtime. |
|
||||
|
||||
### Admin Control-Plane Readers And Writers
|
||||
|
||||
| Files | Current usage |
|
||||
|---|---|
|
||||
| `rustfs/src/admin/handlers/config_admin.rs` | Reads active/persisted server config, validates against `DEFAULT_KVS`, mutates `KVS`, saves config history, saves server config, and updates the global snapshot. |
|
||||
| `rustfs/src/admin/handlers/oidc.rs` | Reads and writes OIDC provider `KVS`, saves server config, and compares persisted config against the global snapshot for restart signaling. |
|
||||
| `rustfs/src/admin/handlers/audit_runtime_config.rs` | Reads persisted config, applies audit runtime target changes, saves server config, and reloads audit runtime state. |
|
||||
| `rustfs/src/admin/handlers/notify_runtime_access.rs` | Reads notification runtime config snapshots and passes `KVS` target changes into the notification system. |
|
||||
| `rustfs/src/admin/handlers/{event,audit}.rs` | Lists and validates notification/audit targets from `Config`; tests build `KV` and `KVS` fixtures. |
|
||||
| `rustfs/src/admin/handlers/plugins_instances.rs` | Maps target plugin `KVS` to response payloads and applies runtime target edits. |
|
||||
| `rustfs/src/admin/handlers/target_descriptor.rs` | Converts descriptor payloads into `KVS` for target plugin instances. |
|
||||
| `rustfs/src/admin/handlers/site_replication.rs` | Reads global server config for LDAP settings and parses LDAP `KVS` fixtures. |
|
||||
| `rustfs/src/admin/service/config.rs` | Reads persisted server config, validates storage-class `KVS`, derives target state, and updates global config/storage-class state. |
|
||||
| `rustfs/src/admin/router.rs` | Reads persisted/global server config for admin route behavior; route tests construct `Config`, `KV`, and `KVS`. |
|
||||
|
||||
### Adjacent ECStore Config-Object Helper Users
|
||||
|
||||
| Files | Current usage |
|
||||
|---|---|
|
||||
| `rustfs/src/admin/handlers/kms_dynamic.rs` | Uses generic `read_config` and `save_config` for dynamic KMS config objects. |
|
||||
| `rustfs/src/site_replication/state.rs` | Uses generic `read_config`, `save_config`, and `delete_config` (via the root storage facade) for site-replication state objects. |
|
||||
| `rustfs/src/admin/service/site_replication.rs` | Uses generic `read_config` and `save_config` for site-replication state normalization. |
|
||||
| `rustfs/src/server/module_switch.rs` | Uses generic `read_config` and `save_config` for module-switch config objects. |
|
||||
| `crates/iam/src/store/object.rs` | Uses generic `read_config_no_lock`, `read_config_with_metadata`, `save_config`, `save_config_with_opts`, and `delete_config` helper variants for IAM object-store persistence paths. |
|
||||
| `crates/scanner/src/{scanner,data_usage_define}.rs` | Uses generic `read_config` and `save_config` for scanner metadata and cache persistence paths. |
|
||||
|
||||
### Runtime Target, Notify, And Audit Crates
|
||||
|
||||
| Files | Current usage |
|
||||
|---|---|
|
||||
| `crates/notify/src/{global,integration,services,registry}.rs` | Carries `Config` into notification runtime startup/reload and target creation. |
|
||||
| `crates/notify/src/config_manager.rs` | Mutates `Config`, reads persisted server config with `read_config_without_migrate`, persists changes with `save_server_config`, and applies per-target `KVS` updates. |
|
||||
| `crates/notify/src/factory.rs` | Builds notification target arguments from `KVS`. |
|
||||
| `crates/notify/examples/{full_demo,full_demo_one}.rs` | Constructs `Config`, `KV`, and `KVS` directly for examples. |
|
||||
| `crates/audit/src/{global,system,registry}.rs` | Carries `Config` into audit runtime startup/reload and target creation. |
|
||||
| `crates/audit/src/factory.rs` | Builds audit target arguments from `KVS`. |
|
||||
| `crates/audit/tests/*.rs` | Constructs `Config` and `KVS` directly for runtime and parsing tests. |
|
||||
| `crates/audit/README.md` | Documents current direct `Config` usage. |
|
||||
| `crates/targets/src/plugin.rs` | Creates plugin targets from `Config` and merged `KVS`. |
|
||||
| `crates/targets/src/catalog/builtin.rs` | Declares builtin target descriptors and default `KVS` fields. |
|
||||
| `crates/targets/src/config/{common,target_args,loader,instance}.rs` | Collects, normalizes, redacts, and materializes target configs from `Config` and `KVS`, including environment overrides. |
|
||||
|
||||
### Identity, Scanner, Tests, And Fixtures
|
||||
|
||||
| Files | Current usage |
|
||||
|---|---|
|
||||
| `crates/iam/src/oidc.rs` | Reads global server config and parses OIDC provider `KVS`. |
|
||||
| `crates/scanner/src/{runtime_config,scanner}.rs` | Reads the global server-config snapshot and resolves scanner runtime config from `Config` and `KVS`. |
|
||||
| `rustfs/src/admin` handler/router tests, `crates/audit/tests/*.rs`, and selected in-crate tests in `crates/{targets,scanner}/src` | Build direct tuple-struct fixtures; use them as candidate regression guards during a pure model move. |
|
||||
|
||||
## Dependency Risk Classification
|
||||
|
||||
| Risk | Why it matters | Guardrail |
|
||||
|---|---|---|
|
||||
| `Config` is both persistence model and runtime input | A move can accidentally change persisted JSON/object shape or runtime target behavior. | Separate pure model contract from persistence helpers before moving `com.rs`. |
|
||||
| Direct `.0` map access is common | Replacing the tuple struct too early would create broad churn and likely behavior drift. | Preserve tuple shape in the first move, then add typed readers in later PRs. |
|
||||
| `KVS` is the effective target config carrier | Notify, audit, and target factories consume `KVS` after file/env merge. | Keep `KVS` API stable until target descriptor and runtime crates are behind a shared contract. |
|
||||
| `DEFAULT_KVS` registration is global | Defaults are initialized centrally and used by admin validation/rendering. | Add a registration contract before changing initialization order. |
|
||||
| Global snapshot readers still exist | Server, admin, IAM, scanner, and site-replication paths can still read global config. | Migrate readers through `AppContext`/provider paths in small steps after the model contract is stable. |
|
||||
| Persistence helpers depend on ECStore storage contracts | Moving them with the pure model would pull storage implementation dependencies upward. | Keep read/write helpers in `ecstore` until a storage-facing persistence contract is explicit. |
|
||||
|
||||
## Recommended Migration Order
|
||||
|
||||
1. Keep this inventory current while Phase 0 guardrails land.
|
||||
2. Add a focused contract surface for `KV`, `KVS`, and `Config` without changing
|
||||
serialization, tuple-struct shape, or method names.
|
||||
3. Add compile-time or scripted checks for temporary compatibility markers and
|
||||
config-model re-export coverage.
|
||||
4. Move only the pure model and defaults registration surface after targeted
|
||||
regression checks cover unchanged persisted object shape, target `KVS` merge
|
||||
behavior, and representative admin config rendering paths.
|
||||
5. Migrate global `Config` readers behind `ServerConfigInterface` or a narrower
|
||||
provider in small PRs.
|
||||
6. Move persistence helpers only after object-I/O and storage-admin dependencies
|
||||
can stay below the model contract.
|
||||
7. Evaluate crate split only after consumers no longer need old paths except
|
||||
explicit `RUSTFS_COMPAT_TODO(<task-id>)` compatibility shims.
|
||||
|
||||
## Do-Not-Change Contract
|
||||
|
||||
The first migration steps must preserve:
|
||||
|
||||
- `KV { key, value, hidden_if_empty }` serde behavior and redaction semantics.
|
||||
- `KVS(Vec<KV>)` tuple shape and public methods.
|
||||
- `Config(HashMap<String, HashMap<String, KVS>>)` tuple shape and public methods.
|
||||
- `Config::set_defaults`, `Config::unmarshal`, `Config::marshal`, and
|
||||
`Config::merge` behavior.
|
||||
- `read_config_without_migrate` fallback/creation behavior for missing server
|
||||
config objects.
|
||||
- `save_server_config` external object shape and config-history compatibility.
|
||||
- Existing notify, audit, scanner, OIDC, and target-plugin enable/disable
|
||||
interpretation from `Config`/`KVS` inputs; business-rule changes stay out of
|
||||
migration PRs.
|
||||
- AppContext/global fallback behavior until all readers are explicitly migrated.
|
||||
@@ -1,61 +1,39 @@
|
||||
# ECStore Layout Boundary
|
||||
|
||||
This document records the `E-001` and `E-SET-001` foundation slice for the
|
||||
architecture migration.
|
||||
**Use this when:** you touch endpoint expansion, `FormatV3`, pool/set layout, or move files between ECStore's internal directories.
|
||||
**Source of truth:** `crates/ecstore/src/layout/` (static layout), `crates/ecstore/src/core/sets.rs` (`Sets`) and `crates/ecstore/src/set_disk/mod.rs` (`SetDisks`) for runtime orchestration, `crates/ecstore/src/api/mod.rs` (`pub mod layout`) for the public surface.
|
||||
|
||||
## Directory Skeleton
|
||||
## Directory Ownership
|
||||
|
||||
The ECStore migration uses these internal ownership buckets before any pure
|
||||
file moves:
|
||||
Ownership buckets under `crates/ecstore/src` (a subset; list the rest with `ls crates/ecstore/src`):
|
||||
|
||||
- `api`: facade and compatibility re-export ownership.
|
||||
- `core`: store facade, object, bucket, list, multipart, and heal paths.
|
||||
- `layout`: static endpoint, disk, pool, and set layout descriptions.
|
||||
- `disk`: local disk, format, health, and disk error ownership.
|
||||
- `erasure`: erasure coding and bitrot ownership.
|
||||
- `metadata`: bucket metadata, config object-store, and data-usage ownership.
|
||||
- `cluster`: remote disk, peer, lock, membership, and health-control-plane
|
||||
ownership.
|
||||
- `services`: lifecycle, replication, tier, notification, rebalance, and
|
||||
metrics service ownership.
|
||||
| Directory | Owns |
|
||||
|---|---|
|
||||
| `api` | Facade and compatibility re-exports |
|
||||
| `core` | Store facade, pools, sets, and object/bucket/list/multipart/heal orchestration |
|
||||
| `layout` | Static endpoint, disk, pool, and set layout (`disks_layout`, `endpoint`, `endpoints`, `format`, `pool_space`, `set_heal`, `set_layout`) |
|
||||
| `disk` | Local disk, format compatibility, health, disk errors |
|
||||
| `erasure` | Erasure coding and bitrot |
|
||||
| `metadata` | Bucket metadata, config object store, data usage |
|
||||
| `cluster` | Remote disk, peer, lock, membership, health control plane |
|
||||
| `services` | Lifecycle, replication, tier, notification, rebalance, metrics services |
|
||||
| `set_disk`, `store`, `data_movement`, `data_usage`, `object_api`, `runtime` | Set-level operations, store init, pool data movement, usage accounting, object API helpers, runtime state owners |
|
||||
|
||||
## Static Set Layout
|
||||
|
||||
Static layout is derived from persisted `FormatV3` data and input endpoint
|
||||
expansion. It may describe:
|
||||
Static layout is derived from persisted `FormatV3` data (`crates/ecstore/src/layout/format.rs`) and endpoint expansion (`crates/ecstore/src/layout/disks_layout.rs`). It may describe the deployment id, set count and drives per set, disk UUID positions inside `format.erasure.sets`, the distribution algorithm, and endpoint grouping produced before runtime disk initialization. It must not own disk handles, lock clients, reconnect loops, repair state, or shutdown signaling.
|
||||
|
||||
- deployment id;
|
||||
- set count and drives per set;
|
||||
- disk UUID positions inside `format.erasure.sets`;
|
||||
- distribution algorithm;
|
||||
- endpoint grouping produced before runtime disk initialization.
|
||||
## Visibility
|
||||
|
||||
Static layout must not own disk handles, lock clients, reconnect loops, repair
|
||||
state, or shutdown signaling.
|
||||
|
||||
## Format And Disk Layout Ownership
|
||||
|
||||
`layout::format` owns persisted format structures and disk UUID position lookup.
|
||||
`layout::disks_layout` owns command-line volume expansion into pool/set layout.
|
||||
|
||||
Compatibility paths remain available through `disk::format` and `disks_layout`
|
||||
until downstream callers are moved or compatibility coverage allows removal.
|
||||
`layout::*` modules are `pub(crate)`; public access goes through `rustfs_ecstore::api::layout` (`DisksLayout`, `EndpointServerPools`, `Endpoints`, `PoolEndpoints`, `SetupType`). `disk::format` re-exports `layout::format` for crate-internal callers. Outer crates must not reach the root `endpoints` or `disks_layout` modules.
|
||||
|
||||
## Runtime Set Orchestration
|
||||
|
||||
Runtime orchestration remains owned by `Sets` and `SetDisks` until a later pure
|
||||
move. It may describe:
|
||||
|
||||
- flat disk index to `(set_index, disk_index)` mapping;
|
||||
- per-set local disk replacement after distributed setup detection;
|
||||
- per-set lock-client host deduplication;
|
||||
- endpoint reconnect monitoring and runtime shutdown signaling;
|
||||
- read/write/heal/list orchestration over initialized disks.
|
||||
`Sets` and `SetDisks` own the flat disk index to `(set_index, disk_index)` mapping, per-set local disk replacement after distributed setup detection, per-set lock-client host deduplication, endpoint reconnect monitoring and runtime shutdown signaling, and read/write/heal/list orchestration over initialized disks.
|
||||
|
||||
## Preservation Rules
|
||||
|
||||
- Object-to-set hashing and distribution algorithm selection must not change.
|
||||
- Format `sets` ordering and disk UUID position lookup must not change.
|
||||
- Local disk replacement and lock-client mapping stay runtime-only.
|
||||
- Later file moves must keep old public paths or add explicit compatibility
|
||||
coverage before deleting them.
|
||||
- File moves keep old public paths or add explicit compatibility coverage before deleting them.
|
||||
|
||||
@@ -1,353 +1,85 @@
|
||||
# ECStore Module Split Plan
|
||||
|
||||
This plan records the remaining ECStore split work after the final audit
|
||||
remediation pass. Runtime movement must still wait until each candidate
|
||||
boundary has explicit contracts, compatibility coverage, dependency evidence,
|
||||
and rollback steps.
|
||||
**Use this when:** you add lifecycle or replication logic and need to know which crate it belongs in, you plan to move an operation family out of `SetDisks`, or the guard fails on one of the split rules named below.
|
||||
**Source of truth:** `scripts/check_architecture_migration_rules.sh` (the rules), `crates/ecstore/src/bucket/lifecycle/README.md` and `crates/ecstore/src/bucket/replication/README.md` (module-level contract inventories, completion criteria, milestones), and [ecstore-api-facade-inventory.md](ecstore-api-facade-inventory.md) (facade groups and boundary files).
|
||||
|
||||
## Current Shape
|
||||
|
||||
| Area | Current owner | Size | Split status |
|
||||
|---|---|---:|---|
|
||||
| Bucket lifecycle | `crates/lifecycle/` + `crates/ecstore/src/bucket/lifecycle/` | core contracts + ECStore runtime | Core contract extracted |
|
||||
| Bucket replication | `crates/ecstore/src/bucket/replication/` | 15,619 lines | Contracts extracted; runtime move pending |
|
||||
| Set disks | `crates/ecstore/src/set_disk/` | state carrier plus operation modules | Keep in ECStore |
|
||||
| Public ECStore facade | `crates/ecstore/src/api/mod.rs` | broad compatibility surface | Shrink only through guarded PRs |
|
||||
| Embedded S3 client | `crates/s3-client/` (`rustfs-s3-client`) | ~8.4K lines | Extracted (rustfs/backlog#1842) |
|
||||
| Area | Owner | Split status |
|
||||
|---|---|---|
|
||||
| Bucket lifecycle | `crates/lifecycle/` (`rustfs-lifecycle`, pure contracts) + `crates/ecstore/src/bucket/lifecycle/` (runtime) | Core contracts extracted; runtime stays in ECStore |
|
||||
| Bucket replication | `crates/replication/` (`rustfs-replication`, contracts and wire formats) + `crates/ecstore/src/bucket/replication/` (worker runtime) | Contracts extracted; runtime move pending |
|
||||
| Set disks | `crates/ecstore/src/set_disk/` | Shared state carrier plus operation modules; stays in ECStore |
|
||||
| Public facade | `crates/ecstore/src/api/mod.rs` | Shrinks only through guarded changes |
|
||||
| S3 client | `crates/s3-client/` (`rustfs-s3-client`) | Extracted |
|
||||
|
||||
Measured 2026-08-12: the whole crate is 265 files / ~288K lines (roughly half
|
||||
is inline `#[cfg(test)]` code). The largest single files are `disk/local.rs`
|
||||
(21,063 lines), `bucket/lifecycle/bucket_lifecycle_ops.rs` (11,961 lines), and
|
||||
`set_disk/mod.rs` (11,151 lines). Reproduce with:
|
||||
Measure size instead of trusting numbers in a document:
|
||||
|
||||
```bash
|
||||
find crates/ecstore/src -name '*.rs' | xargs wc -l | sort -rn | head
|
||||
find crates/ecstore/src/bucket/replication -name '*.rs' | xargs wc -l | tail -1
|
||||
```
|
||||
|
||||
No split step has landed since the contract-extraction PRs of 2026-07-04,
|
||||
while the `bucket/replication` runtime grew from 8,730 to 15,619 lines (+79%)
|
||||
through feature work (e.g. SSE-C ciphertext passthrough replication #5898,
|
||||
delete-marker purge retry/replay #5864). To keep the gap from widening: in
|
||||
domains that already have a contract crate, new replication runtime logic that
|
||||
does not need ECStore runtime state must land in `rustfs-replication`, not in
|
||||
`crates/ecstore/src/bucket/replication/`.
|
||||
Rule for new code: in a domain that already has a contract crate, new logic that does not need ECStore runtime state lands in that crate (`rustfs-lifecycle`, `rustfs-replication`), not under `crates/ecstore/src/bucket/`.
|
||||
|
||||
The file split inside `set_disk/` is already operation-oriented: read, write,
|
||||
list, multipart, lock, heal, and replication code live in separate modules.
|
||||
The remaining large surface is the shared `SetDisks` state and cross-cutting
|
||||
contracts, not only file layout.
|
||||
|
||||
## Completed: S3 Client Extraction (rustfs/backlog#1842)
|
||||
|
||||
`crates/ecstore/src/client/` was a ~8.4K-line hand-written S3 HTTP client the engine uses to *consume* remote S3-compatible endpoints (ILM tier warm backends, transition targets). It was a legitimate engine capability misfiled inside the engine: it pulled `s3s`/`hyper` wire types into ecstore against ARCHITECTURE.md invariant 4, which distinguishes serving the S3 wire protocol (forbidden in ecstore) from consuming it (allowed, but in a dedicated crate).
|
||||
|
||||
The extraction landed as: pure move of the 21 client modules to `crates/s3-client` (`rustfs-s3-client`) with a temporary re-export shim, then direct `rustfs_s3_client::` imports and shim deletion. The two server-side modules historically misfiled under `client/` stayed in ecstore and moved to their real homes: `object_api_utils.rs` under `object_api/`, `object_handlers_common.rs` under `bucket/lifecycle/` (behind the `replication_sink` boundary). The remaining serving-side `s3s` references in ecstore are ratcheted shrink-only by the `S3S_ECSTORE_FILES_BASELINE` counter in `scripts/check_s3s_footprint.sh`; per-module conversions to storage-level types (first: `bucket/object_lock/`) lower the baseline in the same change.
|
||||
The S3 client extraction is complete: the former `client/` directory moved to `crates/s3-client`, its two server-side modules moved to `crates/ecstore/src/object_api/object_api_utils.rs` and `crates/ecstore/src/bucket/lifecycle/object_handlers_common.rs`, and the remaining serving-side `s3s` references in ECStore are ratcheted shrink-only by `S3S_ECSTORE_FILES_BASELINE` in `scripts/check_s3s_footprint.sh`.
|
||||
|
||||
## Non-Negotiable Rules
|
||||
|
||||
- Do not split crates in the same PR that moves runtime state or changes
|
||||
startup behavior.
|
||||
- Do not change object placement, quorum, reader semantics, lifecycle queues,
|
||||
replication queues, notification dispatch, audit events, or scanner repair
|
||||
behavior during inventory and contract PRs.
|
||||
- Do not expose new direct ECStore internals to outer crates; use the existing
|
||||
storage-api and owner-local facade boundaries.
|
||||
- Keep `rustfs_ecstore::api` compatibility visible until each consumer path has
|
||||
compile coverage and an explicit replacement.
|
||||
- Do not split crates in the same change that moves runtime state or changes startup behavior.
|
||||
- Do not change object placement, quorum, reader semantics, lifecycle queues, replication queues, notification dispatch, audit events, or scanner repair behavior during inventory and contract work.
|
||||
- Do not expose new direct ECStore internals to outer crates; use storage-api and owner-local facade boundaries.
|
||||
- Keep `rustfs_ecstore::api` compatibility visible until each consumer path has compile coverage and an explicit replacement.
|
||||
|
||||
## SetDisks Split Direction
|
||||
## Guarded Split Rules
|
||||
|
||||
Do not replace `SetDisks` with several runtime structs in one change. The safe
|
||||
path is:
|
||||
Each rule is enforced by `scripts/check_architecture_migration_rules.sh`; the name is the vocabulary used in reviews and guard failures.
|
||||
|
||||
1. Keep `SetDisks` as the shared state carrier while operation modules continue
|
||||
to own read/write/list/multipart/lock/heal/replication behavior.
|
||||
2. Extract pure contracts first: shard source, disk error, bitrot IO, namespace
|
||||
lock, metrics labels, and file metadata access.
|
||||
3. Move one operation family only after its contracts are covered by focused
|
||||
tests and the facade compatibility path is explicit.
|
||||
4. Preserve the old `rustfs_ecstore::api::set_disk` surface until downstream
|
||||
compatibility tests prove no caller depends on removed names.
|
||||
| Rule | What the guard checks |
|
||||
|---|---|
|
||||
| `LifecycleCrateCoreIndependence` | `crates/lifecycle` (rule validation, filtering, event evaluation, transition/expiration options, tag decoding, object-lock metadata checks, expiry-time rounding) imports no ECStore internals, `rustfs-filemeta`, or `rustfs-utils`; ECStore owns the `ObjectInfo` adapter in `crates/ecstore/src/bucket/lifecycle/core.rs`. |
|
||||
| `ReplicationCrateFileMetaIndependence` | Replication status, decision, MRF, resync, and target-reset wire contracts live in `crates/replication/src/filemeta.rs`; `rustfs-replication` neither imports nor depends on `rustfs-filemeta`. |
|
||||
| `ReplicationCrateStorageApiIndependence` | Delete work DTOs live in `crates/replication/src/storage_api.rs`; ECStore converts storage-api delete DTOs at its replication storage boundary; `rustfs-replication` does not depend on `rustfs-storage-api`. |
|
||||
| `ReplicationCrateUtilsIndependence` | HTTP metadata keys, S3 header labels, ETag trimming, and prefix matching used by replication wire contracts live in `crates/replication/src/http.rs`; `rustfs-replication` does not depend on `rustfs-utils`. |
|
||||
| `EcstoreReplicationBoundaryImports` | ECStore-side `rustfs_replication` imports are confined to the `*_boundary.rs` modules under `crates/ecstore/src/bucket/replication/`; grouped queue, stats, resync, and object-decision symbols each have one owning boundary file. |
|
||||
| `RuntimeReplicationFacadeConsumers` | Scanner, admin, storage-owner, and app code consume replication status/DTO/helper contracts through the `rustfs_ecstore` facade; the `rustfs` and `rustfs-scanner` crates do not depend on `rustfs-replication` directly. |
|
||||
| `StorageApiReplicationContracts` | Owner-facing storage-api delete DTO replication state/status helpers stay in `crates/storage-api/src/replication.rs`; replication worker DTOs stay in `rustfs-replication`. |
|
||||
|
||||
The first executable SetDisks follow-up should be an inventory or guardrail PR,
|
||||
not a runtime split PR.
|
||||
## Lifecycle
|
||||
|
||||
## Lifecycle Candidate
|
||||
`rustfs-lifecycle` owns the pure rule, event, evaluator, tag-filter, object-lock metadata check, and expiry-time contracts. ECStore keeps the object-store runtime, queues, tiering, audit/notification, metadata access (`crates/ecstore/src/bucket/lifecycle/metadata_boundary.rs`), and replication-delete scheduling adapters.
|
||||
|
||||
`rustfs-lifecycle` now owns the pure lifecycle rule, event, evaluator, tag
|
||||
filtering, object-lock metadata check, and expiry-time contracts. ECStore keeps
|
||||
the object-store runtime, queues, tiering, audit/notification, metadata, and
|
||||
replication scheduling adapters.
|
||||
Coupling that still blocks a runtime move: lifecycle workers read ECStore runtime sources (object store, expiry and transition state, tier config, deployment id, local node name); stale multipart cleanup depends on `SetDisks` internals and bucket metadata; expiry schedules replication deletes through the replication lifecycle bridge; the lifecycle runtime coordinates scanner metrics and notification/audit side effects. The contract list and the next step live in `crates/ecstore/src/bucket/lifecycle/README.md`.
|
||||
|
||||
Current coupling:
|
||||
## Replication
|
||||
|
||||
- lifecycle workers and transition state read ECStore runtime sources for
|
||||
object-store handles, expiry state, transition state, tier config, deployment
|
||||
IDs, and local node names;
|
||||
- stale multipart cleanup depends on `SetDisks` internals and bucket metadata
|
||||
through the lifecycle metadata boundary;
|
||||
- lifecycle expiry schedules bucket replication delete work through the
|
||||
replication lifecycle bridge contract;
|
||||
- lifecycle evaluation uses S3 DTOs and replication status contracts from the
|
||||
independent `rustfs-lifecycle`/`rustfs-replication` crates, while ECStore maps
|
||||
`ObjectInfo` into lifecycle object options at the compatibility boundary;
|
||||
- lifecycle runtime still coordinates scanner metrics, notification/audit side
|
||||
effects, metadata access, replication delete scheduling, and tier services.
|
||||
`rustfs-replication` owns resync status contracts, the persisted resync status wire format, filemeta-derived wire contracts, delete work DTOs, and HTTP helper contracts. ECStore keeps the worker runtime, error mapping, MRF persistence, and global pool/stat initialization.
|
||||
|
||||
Current extracted contracts:
|
||||
Boundary layout inside `crates/ecstore/src/bucket/replication/`: `*_boundary.rs` modules concentrate imports from `rustfs-replication`, storage-api, filemeta, config, target, error, lock, msgp, versioning, tagging, bandwidth, queue, stats, resync, and object-decision surfaces; `replication_*_bridge.rs` modules (lifecycle, scanner, object, migration, target-config) expose replication scheduling to other owners without leaking DTO construction; `replication_config_store.rs` exposes config persistence and storage-class labels. Modules inside the directory use relative self-imports, and the facade in `mod.rs` uses explicit symbol lists, never wildcard re-exports.
|
||||
|
||||
- `LifecycleCrateCoreIndependence`: lifecycle rule validation, filtering,
|
||||
event evaluation, transition/expiration options, tag decoding, object-lock
|
||||
metadata checks, and ILM expiry-time rounding live in `rustfs-lifecycle`.
|
||||
`rustfs-lifecycle` must not import ECStore internals, file metadata, or
|
||||
`rustfs-utils`; ECStore owns the `ObjectInfo` adapter in
|
||||
`crates/ecstore/src/bucket/lifecycle/core.rs`.
|
||||
Consumers outside ECStore: RustFS runtime code receives pool/stat handles through storage-owner wrapper types in `rustfs/src/storage/storage_api.rs`; scanner code receives scanner-local config/admission/heal DTOs from `crates/scanner/src/storage_api.rs`; observability reads replication metrics through obs-local snapshot DTOs in `crates/obs/src/metrics/storage_api.rs`; app object and multipart writes call object-replication bridge helpers instead of constructing replication work DTOs.
|
||||
|
||||
Required contracts before crate movement:
|
||||
Completion criteria, the milestone order, and the per-dependency contract inventory live in `crates/ecstore/src/bucket/replication/README.md` (sections "Completion Criteria" and "Milestones"). Remaining work starts from moving resyncer pure decision logic.
|
||||
|
||||
- `LifecycleObjectStore`: object stat, delete, transition, restore, multipart
|
||||
cleanup, and version-aware metadata operations needed by lifecycle workers.
|
||||
- `LifecycleMetadataStore`: lifecycle, object-lock, replication, bucket
|
||||
versioning, and stale multipart metadata lookups without importing ECStore
|
||||
implementation modules. Current lifecycle config reads are concentrated in
|
||||
`crates/ecstore/src/bucket/lifecycle/metadata_boundary.rs`.
|
||||
- `LifecycleRuntime`: expiry state, transition state, tier config, deployment
|
||||
ID, local node name, queue metrics, cancellation, and worker sizing.
|
||||
- `LifecycleReplicationSink`: schedule lifecycle-originated replication deletes
|
||||
without depending on the replication implementation module.
|
||||
- `LifecycleAuditSink`: lifecycle audit and notification emission boundary.
|
||||
## SetDisks
|
||||
|
||||
Next safe PR:
|
||||
Do not replace `SetDisks` with several runtime structs in one change:
|
||||
|
||||
- move one runtime-facing dependency behind a trait or adapter owned by
|
||||
`rustfs-lifecycle` without changing queue, transition, or delete behavior;
|
||||
- keep ECStore compatibility shims until scanner and RustFS app consumers stop
|
||||
depending on `rustfs_ecstore::api::bucket::lifecycle` paths;
|
||||
- add focused tests for the moved contract and keep architecture guard coverage.
|
||||
1. Keep `SetDisks` as the shared state carrier while operation modules own read/write/list/multipart/lock/heal/replication behavior.
|
||||
2. Extract pure contracts first: shard source, disk error, bitrot IO, namespace lock, metrics labels, and file metadata access.
|
||||
3. Move one operation family only after its contracts are covered by focused tests and the facade compatibility path is explicit.
|
||||
4. Preserve the `rustfs_ecstore::api::set_disk` surface until downstream compatibility tests prove no caller depends on removed names.
|
||||
|
||||
The module-level inventory lives in
|
||||
`crates/ecstore/src/bucket/lifecycle/README.md`.
|
||||
## Facade Shrink
|
||||
|
||||
Focused verification for the first code-bearing lifecycle PR:
|
||||
|
||||
- `cargo test -p rustfs-ecstore lifecycle --lib`
|
||||
- `cargo check -p rustfs-ecstore --tests`
|
||||
- `./scripts/check_architecture_migration_rules.sh`
|
||||
- `git diff --check`
|
||||
|
||||
## Replication Candidate
|
||||
|
||||
`rustfs-replication` now owns the resync status contracts and persisted resync
|
||||
status wire format. The remaining `bucket/replication` worker runtime is not
|
||||
ready for a full standalone crate yet.
|
||||
|
||||
The completion criteria and milestone sequence for this candidate (when the
|
||||
split counts as done, the target end state, and the order of the remaining
|
||||
moves) live in the module inventory:
|
||||
`crates/ecstore/src/bucket/replication/README.md`, sections "Completion
|
||||
Criteria" and "Milestones". The originally proposed first code-bearing step
|
||||
(event sink / runtime contracts) has landed; remaining work starts from moving
|
||||
resyncer pure decision logic.
|
||||
|
||||
Current coupling:
|
||||
|
||||
- replication workers depend on `ReplicationStorage`, ECStore object APIs and
|
||||
owner storage-api contracts through the replication storage boundary, bucket
|
||||
target clients, bucket metadata, file metadata replication state through the
|
||||
filemeta boundary, config-derived storage class labels through the config store, scanner repair
|
||||
classification, runtime replication pool/stat handles, bucket monitor and
|
||||
bandwidth reader access through local boundaries, local node names, and
|
||||
notification events;
|
||||
- resync and delete replication paths call metadata paths through the metadata
|
||||
boundary, while bucket target system access, target config types, and target
|
||||
operation types are concentrated behind the replication target boundary;
|
||||
- lifecycle delete paths schedule replication work through
|
||||
`ReplicationLifecycleBridge`, while scanner heal paths schedule replication
|
||||
work through `ReplicationScannerBridge`, and app/SetDisks object write/delete
|
||||
paths use `ReplicationObjectBridge`;
|
||||
- bucket metadata migration and bucket target removal checks use local
|
||||
replication bridges instead of importing resyncer codec or config helper
|
||||
internals;
|
||||
- resync options, bucket/target resync status DTOs, status display labels, and
|
||||
the persisted resync status wire format live in `crates/replication`, with
|
||||
ECStore retaining only error mapping and MRF persistence locally;
|
||||
- `ReplicationCrateFileMetaIndependence`: replication status, decision, MRF,
|
||||
resync, and target-reset wire contracts are owned inside `rustfs-replication`
|
||||
instead of importing `rustfs-filemeta`;
|
||||
- `ReplicationCrateStorageApiIndependence`: delete work DTOs are owned inside
|
||||
`rustfs-replication`; ECStore converts storage-api delete DTOs at the
|
||||
replication storage boundary instead of `rustfs-replication` importing
|
||||
`rustfs-storage-api`;
|
||||
- `ReplicationCrateUtilsIndependence`: HTTP metadata keys, S3 header labels,
|
||||
ETag trimming, and case-insensitive prefix matching used by replication wire
|
||||
contracts are owned inside `rustfs-replication` instead of importing
|
||||
`rustfs-utils`;
|
||||
- direct ECStore replication imports from `rustfs-replication` are limited to
|
||||
`*_boundary.rs` modules;
|
||||
- storage-api delete replication status/state helpers use the local
|
||||
`crates/storage-api/src/replication.rs` contract boundary; ECStore converts
|
||||
those owner DTOs at the replication storage boundary before queueing work;
|
||||
- admin replication extension target filtering and resync request construction
|
||||
stay behind the admin storage boundary instead of exposing replication work
|
||||
DTO construction to handlers;
|
||||
- scanner, admin, storage-owner, and app storage replication status/DTO/helper
|
||||
consumers import those contracts through the ECStore replication facade;
|
||||
- app object and multipart writes call object-replication boundary helpers
|
||||
instead of constructing replication work DTOs or choosing object replication
|
||||
operation types at the use-case layer;
|
||||
- RustFS runtime consumers receive replication pool/stat handles through
|
||||
storage-owner wrapper types instead of carrying ECStore replication handles
|
||||
through app, admin, startup, or workload-admission layers;
|
||||
- global replication pool/stat initialization still lives with ECStore runtime
|
||||
compatibility state;
|
||||
- modules inside `bucket/replication` use local relative paths rather than the
|
||||
ECStore owner path for replication self-imports;
|
||||
- replication runtime source access uses storage/bandwidth boundary aliases for
|
||||
ECStore object store and bucket monitor implementation types;
|
||||
- the ECStore replication facade in `mod.rs` uses explicit compatibility
|
||||
exports instead of wildcard re-exports from implementation modules.
|
||||
|
||||
Required contracts before crate movement:
|
||||
|
||||
- `ReplicationObjectIO`: object read/write primitives for config, MRF, resync
|
||||
status, and multipart replication paths. ECStore object API reader/writer
|
||||
types and storage-api object IO contracts are concentrated in
|
||||
`crates/ecstore/src/bucket/replication/replication_storage_boundary.rs`.
|
||||
- `ReplicationStorage`: keep the existing trait as the starting point, then
|
||||
split object read/write/delete, walk, and metadata update responsibilities
|
||||
only when call sites prove a narrower shape. ECStore object API,
|
||||
storage-api contracts, and read option types are concentrated in
|
||||
`crates/ecstore/src/bucket/replication/replication_storage_boundary.rs`.
|
||||
- `ReplicationMetadataStore`: replication config, target reset headers,
|
||||
MRF/resync state, and status persistence. Metadata sys access and replication
|
||||
metadata path constants are exposed through the contract type in
|
||||
`crates/ecstore/src/bucket/replication/replication_metadata_boundary.rs`.
|
||||
- `ReplicationConfigStore`: replication config persistence and config-derived
|
||||
labels used by target options. Config read/save helpers and storage class
|
||||
labels are exposed through the contract type in
|
||||
`crates/ecstore/src/bucket/replication/replication_config_store.rs`.
|
||||
- `ReplicationFileMeta`: replication status, decisions, MRF entries, resync
|
||||
decisions, and target reset helpers. ECStore concentrates filemeta-to-
|
||||
replication compatibility conversions in
|
||||
`crates/ecstore/src/bucket/replication/replication_filemeta_boundary.rs`,
|
||||
while `FileInfo` remains in the storage boundary for storage trait bindings
|
||||
and walk options.
|
||||
- `ReplicationCrateFileMetaIndependence`: filemeta wire contracts consumed by
|
||||
replication workers are owned in `crates/replication/src/filemeta.rs`, and
|
||||
`rustfs-replication` must not import or depend on `rustfs-filemeta`.
|
||||
- `ReplicationCrateStorageApiIndependence`: delete work DTOs consumed by
|
||||
replication delete/queue/operation helpers are owned in
|
||||
`crates/replication/src/storage_api.rs`, and `rustfs-replication` must not
|
||||
import or depend on `rustfs-storage-api`.
|
||||
- `ReplicationCrateUtilsIndependence`: replication-specific HTTP metadata,
|
||||
header, ETag, and prefix helper contracts are owned in
|
||||
`crates/replication/src/http.rs`, and `rustfs-replication` must not import or
|
||||
depend on `rustfs-utils`.
|
||||
- `EcstoreReplicationBoundaryImports`: ECStore-side imports from
|
||||
`rustfs-replication` are concentrated in replication `*_boundary.rs` modules.
|
||||
- `RuntimeReplicationFacadeConsumers`: scanner, admin, storage-owner, and app
|
||||
storage replication status/DTO/helper consumers import through
|
||||
`rustfs-ecstore`; runtime code under `rustfs/src` does not import
|
||||
`rustfs-replication` directly, and the RustFS runtime/scanner crates do not
|
||||
depend on it.
|
||||
- `StorageApiReplicationContracts`: owner-facing storage-api delete DTO
|
||||
replication state/status helpers remain concentrated in
|
||||
`crates/storage-api/src/replication.rs`, while replication worker DTOs live in
|
||||
`rustfs-replication`.
|
||||
- `ReplicationErrorBoundary`: ECStore error/result contracts and
|
||||
replication-specific error classifiers. `crate::error` imports are
|
||||
concentrated in
|
||||
`crates/ecstore/src/bucket/replication/replication_error_boundary.rs`.
|
||||
- `ReplicationTargetStore`: bucket target listing, target client lookup,
|
||||
target offline checks, target config types, and target operation option
|
||||
types. Bucket target sys access, `BucketTargets`, and target operation types
|
||||
are exposed through the contract type in
|
||||
`crates/ecstore/src/bucket/replication/replication_target_boundary.rs`.
|
||||
- `ReplicationRuntime`: pool, stats, worker admission, bucket monitor, local
|
||||
node identity, cancellation, and queue sizing. Concrete ECStore object store
|
||||
and bucket monitor types stay behind local storage/bandwidth boundaries.
|
||||
- `ReplicationBandwidthLimiter`: target reader wrapping for replication
|
||||
bandwidth accounting and throttling.
|
||||
- `ReplicationVersioningStore`, `ReplicationLockTiming`, `ReplicationMsgpCodec`,
|
||||
and `ReplicationTagFilter`: smaller state/codec/filter contracts that keep
|
||||
bucket versioning, SetDisks lock timing, MessagePack helpers, and bucket
|
||||
tagging helper access behind local replication boundary types.
|
||||
- `ReplicationEventSink`: notification/audit events for skipped, failed, and
|
||||
completed replication operations, including local event host selection.
|
||||
- `ReplicationLifecycleBridge`: lifecycle-originated delete and version-purge
|
||||
scheduling is exposed through the contract type in
|
||||
`crates/ecstore/src/bucket/replication/replication_lifecycle_bridge.rs`.
|
||||
- `ReplicationMigrationBridge`: persisted resync status decode/encode access
|
||||
for bucket metadata migration is exposed through the contract type in
|
||||
`crates/ecstore/src/bucket/replication/replication_migration_bridge.rs`.
|
||||
- `ReplicationResyncContracts`: resync options, target/bucket resync status,
|
||||
status labels, and persisted status encoding live in `crates/replication`.
|
||||
- `ReplicationObjectBridge`: app and SetDisks object write/delete replication
|
||||
decisions and scheduling are exposed through the contract type in
|
||||
`crates/ecstore/src/bucket/replication/replication_object_bridge.rs`.
|
||||
- `ObsReplicationStatsSnapshot`: observability reads replication bucket/site
|
||||
metrics through obs-local snapshot DTOs in
|
||||
`crates/obs/src/metrics/storage_api.rs` instead of carrying the ECStore
|
||||
replication stats handle through collectors.
|
||||
- `StorageReplicationPoolHandle` / `StorageReplicationStatsHandle`: RustFS app, admin,
|
||||
startup, and workload-admission code use storage-owner wrapper types from
|
||||
`rustfs/src/storage/storage_api.rs` for pool activity, resync, queue counts,
|
||||
proxy stats, and site metrics snapshots.
|
||||
- `ReplicationScannerBridge`: scanner-originated replication heal scheduling is
|
||||
exposed through the contract type in
|
||||
`crates/ecstore/src/bucket/replication/replication_scanner_bridge.rs`.
|
||||
Scanner consumers receive scanner-local replication config/admission/heal
|
||||
object DTOs from `crates/scanner/src/storage_api.rs` instead of constructing
|
||||
or inspecting replication queue DTOs directly.
|
||||
- `ReplicationTargetConfigBridge`: bucket target removal checks against
|
||||
replication target rules are exposed through the contract type in
|
||||
`crates/ecstore/src/bucket/replication/replication_target_config_bridge.rs`.
|
||||
- `ReplicationFacade`: the current `rustfs_ecstore::api::bucket::replication`
|
||||
compatibility surface is an explicit symbol list guarded against wildcard
|
||||
re-exports while downstream owners migrate to narrower contracts.
|
||||
|
||||
First safe PR:
|
||||
|
||||
- add a replication extraction inventory section or module-level README;
|
||||
- list current ECStore/runtime dependencies and the target contract owner for
|
||||
each dependency;
|
||||
- keep global pool/stat initialization and queue behavior unchanged.
|
||||
|
||||
The module-level inventory lives in
|
||||
`crates/ecstore/src/bucket/replication/README.md`.
|
||||
|
||||
Focused verification for the first code-bearing replication PR:
|
||||
|
||||
- `cargo test -p rustfs-ecstore replication --lib`
|
||||
- `cargo check -p rustfs-ecstore --tests`
|
||||
- `./scripts/check_architecture_migration_rules.sh`
|
||||
- `git diff --check`
|
||||
|
||||
## Facade Shrink Plan
|
||||
|
||||
The broad `rustfs_ecstore::api` facade remains a compatibility boundary, not a
|
||||
new architecture target. The current facade groups and external consumers are
|
||||
recorded in
|
||||
[`ecstore-api-facade-inventory.md`](ecstore-api-facade-inventory.md).
|
||||
Shrinking it must be monotonic:
|
||||
|
||||
1. Inventory every public facade group and consumer.
|
||||
2. Add compile-time coverage before removing or narrowing a facade item.
|
||||
3. Move outer consumers to storage-api or owner-local compatibility boundaries.
|
||||
4. Remove one facade group per PR only after downstream compatibility tests pass.
|
||||
|
||||
Do not delete facade groups only because the underlying module moved. Keep the
|
||||
facade stable until the replacement path is visible and tested.
|
||||
Facade groups, boundary files, and shrink rules are in [ecstore-api-facade-inventory.md](ecstore-api-facade-inventory.md). Shrinking is monotonic: inventory, add compile-time coverage, move consumers to storage-api or owner-local boundaries, then remove one group per change. Do not delete facade groups only because the underlying module moved.
|
||||
|
||||
## Ready-To-Split Checklist
|
||||
|
||||
A candidate split is ready for code movement only when all items below are true:
|
||||
A candidate is ready for code movement only when all of these hold:
|
||||
|
||||
- dependency graph shows no cycle with ECStore, storage-api, runtime sources, or
|
||||
owner-local compatibility modules;
|
||||
- the dependency graph shows no cycle with ECStore, storage-api, runtime sources, or owner-local compatibility modules;
|
||||
- contract traits compile without importing ECStore implementation modules;
|
||||
- old facade names have compatibility tests or explicit deprecation coverage;
|
||||
- focused tests cover the changed owner path before any full gate is attempted;
|
||||
- rollback preserves object IO, quorum, lifecycle/replication queues, scanner
|
||||
repair, notification/audit events, and metadata compatibility.
|
||||
- rollback preserves object IO, quorum, lifecycle/replication queues, scanner repair, notification/audit events, and metadata compatibility.
|
||||
|
||||
@@ -1,10 +1,13 @@
|
||||
# Erasure Coding — Normative Algorithm & On-Disk Compatibility Contract
|
||||
|
||||
**Use this when:** changing anything under `crates/ecstore/src/erasure/`, `crates/filemeta/`, `crates/ecstore/src/set_disk/`, storage-class or layout code, or any decode, quorum, or heal boundary; read §12 and §13 before editing.
|
||||
**Source of truth:** this document is normative for the algorithm and the on-disk / on-wire compatibility contract; the cited symbols are where the code enforces each rule.
|
||||
|
||||
Status: normative. This document is the source of truth for how RustFS erasure-codes, stores, reads, reconstructs, and heals user data, and for the on-disk / on-wire compatibility contract that every future change must preserve. It governs the highest-risk code in the system: a regression here can silently corrupt or lose all user data, or make existing (and MinIO-migrated) objects permanently unreadable.
|
||||
|
||||
Erasure coding, quorum/heal, and metadata/on-disk formats are **High-risk** per [AGENTS.md](../../AGENTS.md) ("Risk tiers"). Any behavior-affecting change to code this document governs requires the full seven-role adversarial validation and, for anything touching decode or the on-disk format, a regression test against real on-disk and MinIO-migrated samples before merge.
|
||||
Erasure coding, quorum/heal, and metadata/on-disk formats are **High-risk** per [AGENTS.md](../../AGENTS.md) ("Broad or High-Risk Changes"). Any behavior-affecting change to code this document governs requires adversarial review with the `adversarial-validation` skill and, for anything touching decode or the on-disk format, a regression test against real on-disk and MinIO-migrated samples before merge.
|
||||
|
||||
This document describes the baseline (`main`) algorithm. Where the baseline has a known defect that a specific change corrects, that is called out inline; the *invariant* stated is always the correct rule the code must converge to, never the defect.
|
||||
This document describes the algorithm as implemented on `main`. The *invariant* stated is always the rule the code must satisfy; where the code enforces it, the enforcing symbol is cited.
|
||||
|
||||
## How to use this document
|
||||
|
||||
@@ -34,7 +37,6 @@ This document describes the baseline (`main`) algorithm. Where the baseline has
|
||||
11. Compatibility contract and decode tolerance
|
||||
12. Invariants checklist (the frozen contract)
|
||||
13. Change procedure and guardrails
|
||||
14. References
|
||||
|
||||
---
|
||||
|
||||
@@ -80,7 +82,7 @@ Two storage classes: `STANDARD` (SC) and `REDUCED_REDUNDANCY` (RRS) ([storagecla
|
||||
|
||||
- **INVARIANT — parity bounds.** Parity must satisfy `parity ≤ N/2` for both classes, and `SC parity ≥ RRS parity` when both are non-zero ([storageclass.rs](../../crates/ecstore/src/config/storageclass.rs), `validate_parity` / `validate_parity_inner`). Enforcement nuance to be aware of: `validate_parity_inner` (the path a user-configured `EC:<parity>` storage class flows through) only applies the `parity ≤ N/2` check for `N > 2`, so degenerate small-set values (e.g. `EC:2` on `N = 2`, giving `data_blocks = 0`) are not caught there; the standalone `validate_parity` enforces the bound unconditionally but is applied only to the resolved default parity. A change that lets user-configured parity reach a write path must not assume the `≤ N/2` bound was enforced for `N ≤ 2`. Parity `0` is permitted (single-drive / capacity setups); there is no non-zero minimum.
|
||||
- **INVARIANT — per-pool validity.** Each pool's resolved parity must be valid for **that pool's own drive count**. A heterogeneous deployment (pools of different widths) must resolve parity per pool; applying one pool's parity to a narrower pool can drive `data_blocks = N − parity` to `0` and make encoding impossible.
|
||||
- Baseline defect: `main` computes `common_parity_drives` from the **first** pool only and applies it to every pool ([store/init.rs](../../crates/ecstore/src/store/init.rs), `ec_drives_no_config` at [store/init_format.rs](../../crates/ecstore/src/store/init_format.rs)); this is issue #4801 (a smaller later pool panics with `TooFewDataShards`). The correct rule is per-pool resolution.
|
||||
- Implemented by `resolve_write_layout` ([set_disk/mod.rs](../../crates/ecstore/src/set_disk/mod.rs)), which takes the pool index and resolves parity against that pool's own drive count; the default parity for a pool without explicit config comes from `ec_drives_no_config` ([store/init_format.rs](../../crates/ecstore/src/store/init_format.rs)).
|
||||
|
||||
Per-write layout (the numbers that go into `xl.meta`), from the storage class or `default_parity_count`, with `opts.max_parity` forcing `N/2` for internal writes ([set_disk/ops/object.rs](../../crates/ecstore/src/set_disk/ops/object.rs)):
|
||||
|
||||
@@ -219,7 +221,7 @@ Fields: `version_id`, `mod_time`, `signature: [u8;4]`, `version_type`, `flags: u
|
||||
|
||||
### 6.6 Inline data
|
||||
|
||||
Small objects store their payload inline after the container CRC ([filemeta_inline.rs](../../crates/filemeta/src/filemeta_inline.rs)): 1 version byte (`INLINE_DATA_VER = 1`) then a msgpack map of `version-key → bin`. **INVARIANT — the map key** is the version-id string, `"null"` (`NULL_VERSION_ID`) for the null/None version, else the lowercase hyphenated UUID. Presence is determined **on read** solely by the `meta_sys[inline-data]` body marker (`FileInfo::inline_data`); the read path gates inline extraction on that marker alone. The header `InlineData` flag is **written** (mirrored from the body on marshal) but is **not** consulted on read, and a disagreement is tolerated — MinIO may leave the header flag unset while inline data is present, so a reader must **not** require the flag and the marker to agree. The inline threshold is `should_inline` ([storageclass.rs](../../crates/ecstore/src/config/storageclass.rs)): inline if `shard_size ≤ inline_block/8` for versioned buckets, else `≤ inline_block`; `DEFAULT_INLINE_BLOCK = 128 KiB`.
|
||||
Small objects store their payload inline after the container CRC ([filemeta_inline.rs](../../crates/filemeta/src/filemeta_inline.rs)): 1 version byte (`INLINE_DATA_VER = 1`) then a msgpack map of `version-key → bin`. **INVARIANT — the map key** is the version-id string, `"null"` (`NULL_VERSION_ID`, [fileinfo.rs](../../crates/filemeta/src/fileinfo.rs)) for the null/None version, else the lowercase hyphenated UUID. Presence is determined **on read** solely by the `meta_sys[inline-data]` body marker (`FileInfo::inline_data`); the read path gates inline extraction on that marker alone. The header `InlineData` flag is **written** (mirrored from the body on marshal) but is **not** consulted on read, and a disagreement is tolerated — MinIO may leave the header flag unset while inline data is present, so a reader must **not** require the flag and the marker to agree. The inline threshold is `should_inline` ([storageclass.rs](../../crates/ecstore/src/config/storageclass.rs)): inline if `shard_size ≤ inline_block/8` for versioned buckets, else `≤ inline_block`; `DEFAULT_INLINE_BLOCK = 128 KiB`.
|
||||
|
||||
---
|
||||
|
||||
@@ -230,7 +232,7 @@ Small objects store their payload inline after the container CRC ([filemeta_inli
|
||||
- Encode-time gates: writable disks `< write_quorum` ⇒ `ErasureWriteQuorum`; committed shards `< write_quorum` after encode ⇒ error ([set_disk/ops/object.rs](../../crates/ecstore/src/set_disk/ops/object.rs)).
|
||||
- **INVARIANT — atomic commit with best-effort rollback.** Commit is `rename_data` (per-disk temp → final) fanned across all disks ([core/io_primitives.rs](../../crates/ecstore/src/set_disk/core/io_primitives.rs)). If write quorum is not met (`reduce_write_quorum_errs`), every successful disk is undone (`delete_version{undo_write:true}`) and the original quorum error is returned. **Baseline:** the rollback is **best-effort** — undo failures are counted and `warn!`-logged, never propagated or retried — so a write that both misses quorum *and* whose rollback partially fails can leave shards on some disks; that partial residue is reconciled later by heal/scanner, not by the commit path. The guarantee the commit path enforces is "never *reports* success below quorum", not "never leaves any bytes behind".
|
||||
- On success the newly committed dir is `fi.data_dir`. Separately, `reduce_common_data_dir` votes over each disk's **`old_data_dir`** (the *superseded* dir being dereferenced) and returns it when it reaches write_quorum, so the old dir can be reclaimed (`commit_rename_data_dir`) — it is a GC input, **not** the new `data_dir`. `classify_rename_convergence` classifies the commit (`PartialCommit` / `SignatureDivergent`), but **only the multipart-complete path consumes it** (`convergence.needs_heal()` → `send_heal_request`); the regular `put_object` path discards the convergence result and relies on the old-data-dir cleanup / `add_partial` heal enqueue instead.
|
||||
- The write layout (per-pool parity, storage class, `max_parity`) is computed **inline** in the write path ([set_disk/ops/object.rs](../../crates/ecstore/src/set_disk/ops/object.rs)); on `main` there is **no** `WriteLayout` type or `resolve_write_layout` function — do not cite either as if it exists (§13's symbol-citation rule). A future refactor may centralize this; add the symbol to the spec only once it lands in code.
|
||||
- The write layout (per-pool parity, storage class, `max_parity`) is resolved once by `resolve_write_layout` into a `WriteLayout` ([set_disk/mod.rs](../../crates/ecstore/src/set_disk/mod.rs)) and consumed by the object and multipart write paths ([set_disk/ops/object.rs](../../crates/ecstore/src/set_disk/ops/object.rs), [set_disk/ops/multipart.rs](../../crates/ecstore/src/set_disk/ops/multipart.rs)); resolution is per pool (§2.2).
|
||||
|
||||
---
|
||||
|
||||
@@ -238,7 +240,7 @@ Small objects store their payload inline after the container CRC ([filemeta_inli
|
||||
|
||||
- **INVARIANT — read quorum = `data_blocks`.** `object_quorum_from_meta` returns `(read_quorum = data_blocks, write_quorum)` ([set_disk/metadata.rs](../../crates/ecstore/src/set_disk/metadata.rs)); `parity_blocks = common_parity(...)` is the parity value held by the most disks that still reaches its own read quorum. When `default_parity_count == 0`, read = write = all shards.
|
||||
- Authoritative FileInfo selection — `find_file_info_in_quorum` ([set_disk/metadata.rs](../../crates/ecstore/src/set_disk/metadata.rs)) groups valid metas by a content-identity SHA-256 (`file_info_quorum_hash`) that hashes size/flags/mod_time/transition/version_id/data_dir/parts and, for real objects, data/parity/distribution — **excluding replication-status keys** so replication noise never splits quorum. A meta counts only if its mod_time equals the common mod_time (or etag matches when mod_time is absent). The winning hash must reach quorum, else `ErasureReadQuorum`. Latest-version reads may escalate to write_quorum to avoid resurrecting a partially-overwritten version.
|
||||
- **INVARIANT — decode needs ≥ `data_blocks` shards.** The stripe reader requires `available_shards ≥ data_shards`; below that the read fails closed with a read-quorum error (never silent truncation) ([set_disk/read.rs](../../crates/ecstore/src/set_disk/read.rs), [set_disk/shard_source.rs](../../crates/ecstore/src/set_disk/shard_source.rs)). Before any `block_size` / `data_shards` division, `has_valid_dimensions()` must hold (`block_size > 0 && data_shards > 0`) or the read fails instead of dividing by zero ([erasure.rs](../../crates/ecstore/src/erasure/coding/erasure.rs)); note this guard runs *after* codec construction and fully covers only `block_size == 0` — a `data_blocks == 0` geometry panics earlier in the constructor (§13).
|
||||
- **INVARIANT — decode needs ≥ `data_blocks` shards.** The stripe reader requires `available_shards ≥ data_shards`; below that the read fails closed with a read-quorum error (never silent truncation) ([set_disk/read.rs](../../crates/ecstore/src/set_disk/read.rs), [set_disk/shard_source.rs](../../crates/ecstore/src/set_disk/shard_source.rs)). Codec geometry is validated at construction: `Erasure::try_new` / `try_new_with_options` ([erasure.rs](../../crates/ecstore/src/erasure/coding/erasure.rs)) return `ErasureConstructionError` for `data_shards == 0`, `block_size == 0`, shard-count overflow, or an unsupported shard configuration, so a read never reaches a `block_size` / `data_shards` division with invalid geometry (§13).
|
||||
- If `available ≥ data_blocks` but some shards are missing, the read is served **and** a background read-repair heal is enqueued.
|
||||
- **INVARIANT — cross-stripe read verification.** When a data shard is missing and `available > data_blocks`, reconstruction regenerates parity and compares it to the surviving parity; a mismatch is `InvalidData "inconsistent read source shards"` (backlog#832), catching corruption that passed per-shard bitrot but disagrees across the stripe ([erasure.rs](../../crates/ecstore/src/erasure/coding/erasure.rs)).
|
||||
|
||||
@@ -324,22 +326,12 @@ Decode tolerance
|
||||
|
||||
## 13. Change procedure and guardrails
|
||||
|
||||
- **Risk tier.** All of the above is High-risk ([AGENTS.md](../../AGENTS.md)). Any behavior-affecting change requires the full seven-role adversarial validation.
|
||||
- **Risk tier.** All of the above is High-risk ([AGENTS.md](../../AGENTS.md) "Broad or High-Risk Changes"). Any behavior-affecting change requires adversarial review with the `adversarial-validation` skill.
|
||||
- **Keep this document in sync.** A change to any governed behavior, formula, format field, or invariant must update this spec in the same PR; renaming a cited symbol must update its reference here. The spec is normative and is the checklist the next change is reviewed against, so drift is a correctness defect. References are symbol-based (not line numbers) specifically so ordinary refactors do not invalidate them — but semantic changes still must.
|
||||
- **Adding an on-disk field** must be additive: new msgpack key or a `minor`/`meta_ver` bump with a read path for the old value; keep decoders skipping unknown keys; write both internal-key prefixes; never repurpose or reorder existing keys or header array positions.
|
||||
- **Never make a decode boundary stricter** than what §11 allows without (a) proving no legitimate older-RustFS or MinIO-migrated shape is rejected, and (b) a regression test against real on-disk and MinIO fixtures. New validation belongs at the trust boundary and must fail *open to a tolerant default*, not closed to `FileCorrupt`, for anything recoverable. (Concretely: rejecting a negative `actual_size`, or hard-failing a non-16-byte `transitioned-versionID`, breaks existing data — see §11.)
|
||||
- **Codec construction (baseline gap).** The baseline exposes panicking `Erasure::new` / `new_with_options`: the codec's shard-count validation surfaces as an `.expect` panic when `data_shards == 0 && parity_shards > 0` (`ReedSolomon::new` ⇒ `TooFewDataShards`). `has_valid_dimensions()` (`block_size > 0 && data_shards > 0`) is a **`&self`** method, so it can only run *after* construction — the read path builds the codec from on-disk geometry first and checks the guard second ([erasure.rs](../../crates/ecstore/src/erasure/coding/erasure.rs), [set_disk/read.rs](../../crates/ecstore/src/set_disk/read.rs)). It therefore reliably catches only the `block_size == 0` case (block size is never passed to `ReedSolomon::new`, so construction succeeds and the guard rejects it before any division); a `data_blocks == 0` xl.meta with `parity > 0` **panics in the constructor before the guard can run**. The correct fix is a **fallible constructor** (returning `Result`, not `.expect`) on any path reachable from untrusted metadata; until then `has_valid_dimensions()` is a partial preflight, not a complete guard.
|
||||
- **Codec construction is fallible.** Read, heal, multipart, and object write paths construct the codec through `Erasure::try_new` / `try_new_with_options` ([erasure.rs](../../crates/ecstore/src/erasure/coding/erasure.rs)) and surface `ErasureConstructionError` instead of panicking on geometry decoded from untrusted metadata (`data_shards == 0`, `block_size == 0`, shard-count overflow, unsupported shard counts). Do not reintroduce a panicking constructor on any path reachable from on-disk metadata; `has_valid_dimensions()` remains only a `&self` preflight for already-built codecs.
|
||||
- **Guardrail scripts** (part of `make pre-commit` / `make pre-pr`):
|
||||
- [check_architecture_migration_rules.sh](../../scripts/check_architecture_migration_rules.sh) keeps the erasure engine crate-private and under its owner module, and keeps erasure-cache / `GLOBAL_IS_ERASURE*` access behind ecstore helpers.
|
||||
- [check_doc_paths.sh](../../scripts/check_doc_paths.sh) validates that every repo path this document cites exists — keep citations to real paths.
|
||||
- **Tooling.** Inspect on-disk metadata with `dump_fileinfo` / `dump_versions` per [../operations/tier-ilm-debugging.md](../operations/tier-ilm-debugging.md) rather than guessing at bytes.
|
||||
|
||||
---
|
||||
|
||||
## 14. References
|
||||
|
||||
- Reed–Solomon codes; MDS property and GF(2⁸) byte-oriented coding — the standard basis for `rs-vandermonde` (Vandermonde generator matrix over GF(2⁸)).
|
||||
- `rustfs-erasure-codec` (RustFS fork of `reed-solomon-erasure`, GF(2⁸)) and `reed-solomon-simd` (GF(2¹⁶)) — declared in the workspace `Cargo.toml`.
|
||||
- HighwayHash-256 — the bitrot checksum family; π-derived default key.
|
||||
- MinIO `xl.meta` v1.3 format lineage — RustFS is byte-compatible for read + one-way migration; see [minio-file-format-compat.md](minio-file-format-compat.md) for the fixture-proven matrix and scope.
|
||||
- Related invariants: [placement-repair-invariants.md](placement-repair-invariants.md), [ecstore-layout-boundary.md](ecstore-layout-boundary.md), [decommission-compatibility.md](decommission-compatibility.md), [../operations/tier-ilm-debugging.md](../operations/tier-ilm-debugging.md), and [AGENTS.md](../../AGENTS.md) Cross-Cutting Domain Invariants.
|
||||
|
||||
@@ -1,202 +1,67 @@
|
||||
# Global State And Crate Split Plan
|
||||
|
||||
This document records the late global-state cleanup plan after the AppContext
|
||||
foundation, storage API contracts, ECStore layout, runtime lifecycle, and cluster
|
||||
control-plane boundaries are stable.
|
||||
**Use this when:** business logic needs runtime state (object store, endpoints, lock clients, lifecycle state, config) and you must pick the right boundary, or you are evaluating a new crate split out of ECStore.
|
||||
**Source of truth:** `crates/ecstore/src/runtime/global.rs` and `crates/ecstore/src/runtime/sources.rs` (ECStore-owned state and its adapter), `rustfs/src/app/context.rs` and the `runtime_sources.rs` owner modules under `rustfs/src` (RustFS resolvers), and the `rustfs_ecstore::api::global` boundary list in `scripts/check_architecture_migration_rules.sh`. The static inventory is [global-state-inventory.md](global-state-inventory.md).
|
||||
|
||||
As of the Phase 7 closeout, runtime resolver fallbacks have been pushed out of
|
||||
the root facade and into explicit owner-local boundaries. Future work should
|
||||
therefore treat broad fallback removal as complete and use this document for the
|
||||
remaining ECStore-owned bootstrap state and crate-split decisions.
|
||||
|
||||
The issue #730 global-state baseline and runtime migration target inventory are
|
||||
recorded in [`global-state-inventory.md`](global-state-inventory.md).
|
||||
Broad resolver-fallback removal is complete: runtime resolver fallbacks live in explicit owner-local boundaries, not in the root facade. What remains is ECStore-owned bootstrap state and crate-split decisions.
|
||||
|
||||
## Remaining Global Owners
|
||||
|
||||
| Owner | Current role | Migration stance |
|
||||
| Owner | Role | Stance |
|
||||
|---|---|---|
|
||||
| `rustfs/src/app/context.rs` | AppContext-first resolver facade. | Resolver helpers stay context-first and do not construct concrete no-AppContext defaults. |
|
||||
| `rustfs/src/app/context/runtime_sources.rs` | Default adapters for KMS, IAM, object store, endpoints, config, metrics, and notification state used by AppContext construction. | This is an allowed adapter boundary, not a business logic owner. |
|
||||
| `rustfs/src/*/runtime_sources.rs` | Root, admin, app, server, startup, and storage owner-local runtime-source boundaries. | Business modules use these boundaries instead of calling global state directly; owner facades own any remaining no-AppContext compatibility defaults. |
|
||||
| `rustfs/src/*/storage_api.rs` | Root, admin, app, and storage owner-local storage contract/facade boundaries. | Storage helper and ECStore facade access remains visible at local owner boundaries. |
|
||||
| `crates/*/storage_api.rs` | External crate-local storage facade boundaries for IAM, scanner, heal, notify, observability, Swift, and S3 Select. | External runtime crates consume ECStore runtime state through `rustfs_ecstore::api::runtime` instead of the direct global facade. |
|
||||
| `crates/ecstore/src/runtime/global.rs` | ECStore bootstrap/runtime state owner. | Keep internal until ECStore has explicit owner handles for all remaining bootstrap state. |
|
||||
| `crates/ecstore/src/runtime/sources.rs` | ECStore runtime-source adapter over global state. | Preferred ECStore-internal access path while shrinking direct `runtime::global` reads. |
|
||||
| `rustfs/src/app/context/runtime_sources.rs` | Default adapters for KMS, IAM, object store, endpoints, config, metrics, and notification state used by AppContext construction. | Allowed adapter boundary, not a business-logic owner. |
|
||||
| `rustfs/src/runtime_sources.rs`, `rustfs/src/admin/runtime_sources.rs`, `rustfs/src/app/runtime_sources.rs`, `rustfs/src/server/runtime_sources.rs`, `rustfs/src/storage/runtime_sources.rs` | Owner-local runtime-source boundaries. | Business modules use these instead of global state; owner facades decide when to apply no-AppContext compatibility defaults. |
|
||||
| `rustfs/src/storage_api.rs`, `rustfs/src/admin/storage_api.rs`, `rustfs/src/app/storage_api.rs`, `rustfs/src/storage/storage_api.rs` | Owner-local storage contract/facade boundaries. | Storage helper and ECStore facade access stays visible at local owner boundaries. |
|
||||
| `crates/*/storage_api.rs` | External crate-local storage facade boundaries (IAM, scanner, heal, notify, observability, Swift, S3 Select). | External runtime crates read ECStore runtime state through `rustfs_ecstore::api::runtime`, never the global facade. |
|
||||
| `crates/ecstore/src/runtime/global.rs` | ECStore bootstrap/runtime state owner. | Internal until ECStore has explicit owner handles for all remaining bootstrap state. |
|
||||
| `crates/ecstore/src/runtime/sources.rs` | ECStore runtime-source adapter over global state. | Preferred ECStore-internal access path while direct `runtime::global` reads shrink. |
|
||||
|
||||
## Runtime Source Boundaries
|
||||
|
||||
Runtime-source modules are the allowed compatibility layer between migrated
|
||||
consumers and process-global state. They must keep these properties:
|
||||
Runtime-source modules are the allowed compatibility layer between migrated consumers and process-global state. They keep these properties:
|
||||
|
||||
- context-first lookup when an `AppContext` handle exists;
|
||||
- explicit fallback to the existing global only where compatibility still
|
||||
requires it;
|
||||
- explicit fallback to the existing global only where compatibility still requires it, decided by the owner facade;
|
||||
- no hidden service construction in business logic;
|
||||
- no startup, readiness, IAM, KMS, lock, notification, or storage behavior
|
||||
change in inventory or guardrail PRs.
|
||||
- the root `rustfs/src/runtime_sources.rs` is an entrypoint only: it composes no concrete fallback defaults (`unwrap_or`, `unwrap_or_else`, direct `init_global` or `new_global` calls);
|
||||
- production callers outside runtime-source and `storage_api.rs` boundary modules do not import ECStore global state directly.
|
||||
|
||||
## Guarded Boundary List
|
||||
### Guarded Boundary List
|
||||
|
||||
The architecture guard snapshots the files currently allowed to reference
|
||||
`rustfs_ecstore::api::global` directly:
|
||||
The guard pins the production files allowed to reference `rustfs_ecstore::api::global` directly:
|
||||
|
||||
- `rustfs/src/storage/storage_api.rs`
|
||||
|
||||
That boundary now keeps only bootstrap writes and lifecycle controls in the
|
||||
global facade. Read-only runtime getters must be exported through
|
||||
`rustfs_ecstore::api::runtime` and consumed through the local storage facade.
|
||||
New direct uses must either move behind an existing owner-local boundary or
|
||||
update this plan and the guard in the same reviewed migration PR.
|
||||
That boundary keeps only bootstrap writes and lifecycle controls (`set_global_endpoints`, `set_global_region`, `set_global_rustfs_port`, `set_object_store_resolver`, `shutdown_background_services`, `update_erasure_type`). Read-only runtime getters are exported through `rustfs_ecstore::api::runtime` and consumed through the local storage facade. A new direct use either moves behind an existing owner-local boundary or updates this plan and the guard in the same reviewed change.
|
||||
|
||||
## Fallback Removal Plan
|
||||
|
||||
1. Keep AppContext-first lookup as the stable resolver contract.
|
||||
2. Keep concrete no-AppContext compatibility defaults only at owner-local
|
||||
runtime-source facades that consume them.
|
||||
3. Do not let business logic call `AppContext` or ECStore globals directly when
|
||||
an owner-local runtime-source boundary exists.
|
||||
4. Keep embedded startup and tests working before deleting any remaining owner
|
||||
fallback.
|
||||
5. Do not remove ECStore bootstrap globals until ownership handles exist for
|
||||
local disks, endpoint pools, lock clients, notification state, tier config,
|
||||
lifecycle state, and object-store publication.
|
||||
|
||||
## GLOB-007 Closeout Boundary
|
||||
|
||||
`GLOB-007` is complete when these invariants hold:
|
||||
|
||||
- root `rustfs/src/runtime_sources.rs` is an AppContext/root facade entrypoint
|
||||
and no longer composes concrete fallback defaults with `unwrap_or`,
|
||||
`unwrap_or_else`, direct `init_global`, or direct `new_global` calls;
|
||||
- private AppContext resolver helpers are context-first and do not hide fallback
|
||||
closure parameters;
|
||||
- admin, app, storage, server, startup, and config owner facades decide when to
|
||||
apply no-AppContext compatibility defaults;
|
||||
- production callers outside runtime-source and storage-api boundary modules do
|
||||
not import ECStore global state directly;
|
||||
- the architecture guard keeps the direct `rustfs_ecstore::api::global`
|
||||
boundary list explicit.
|
||||
|
||||
Allowed remaining fallbacks are owner compatibility decisions, not resolver
|
||||
fallback families. They are kept so embedded startup, tests, and no-context
|
||||
callers preserve the previous behavior while higher layers continue migrating
|
||||
to explicit AppContext ownership.
|
||||
1. AppContext-first lookup is the stable resolver contract.
|
||||
2. Concrete no-AppContext compatibility defaults exist only at the owner-local runtime-source facades that consume them.
|
||||
3. Business logic does not call `AppContext` or ECStore globals directly when an owner-local runtime-source boundary exists.
|
||||
4. Embedded startup and tests keep working before any remaining owner fallback is deleted.
|
||||
5. ECStore bootstrap globals stay until ownership handles exist for local disks, endpoint pools, lock clients, notification state, tier config, lifecycle state, and object-store publication.
|
||||
|
||||
## Crate Split Evaluation
|
||||
|
||||
`ecstore-erasure` and `storage-cluster` remain proposal-only until dependency
|
||||
cycles and hot-path risks are proven safe. The Phase 7 evaluation is complete
|
||||
for now: neither split is ready for code movement in this migration round.
|
||||
The follow-up ECStore module split plan is recorded in
|
||||
[`ecstore-module-split-plan.md`](ecstore-module-split-plan.md), including the
|
||||
remaining `SetDisks`, lifecycle, replication, and facade-shrink boundaries.
|
||||
`ecstore-erasure` and `storage-cluster` are proposal-only; neither is ready for code movement. Lifecycle and replication split status is tracked in [ecstore-module-split-plan.md](ecstore-module-split-plan.md).
|
||||
|
||||
### CRATE-001: `ecstore-erasure`
|
||||
### `ecstore-erasure`
|
||||
|
||||
Current coupling:
|
||||
Coupling: erasure decoding depends on disk errors, disk read timeouts, and set-disk shard sources; set-disk read/write/heal paths construct codecs in hot object I/O paths; bitrot readers/writers live in ECStore IO support and serve both erasure and set-disk code; `rustfs_ecstore::api::erasure` is still a public compatibility surface.
|
||||
|
||||
- erasure decoding depends on disk errors, disk read timeouts, and set-disk
|
||||
shard sources;
|
||||
- set-disk read/write/heal paths construct erasure codecs in hot object I/O
|
||||
paths;
|
||||
- bitrot readers/writers live in ECStore IO support and are used by both
|
||||
erasure and set-disk code;
|
||||
- public compatibility still exposes erasure symbols through
|
||||
`rustfs_ecstore::api::erasure`.
|
||||
Decision: do not split. The boundary becomes a candidate only after shard-source, disk-error, bitrot, and metrics contracts are explicit enough to avoid a dependency cycle back into ECStore, backed by encode/decode/reconstruction benchmarks and a rollback plan that keeps read/write quorum and old-version decode unchanged.
|
||||
|
||||
Decision: do not split in code yet. The erasure boundary is a candidate only
|
||||
after the shard-source, disk-error, bitrot, and metrics contracts are explicit
|
||||
enough to avoid a dependency cycle back into ECStore.
|
||||
### `storage-cluster`
|
||||
|
||||
Required evidence before proposing the split:
|
||||
Coupling: cluster RPC remote-disk code depends on disk stores, disk health tracking, set-disk buffer sizing, local disk scan guards, internode metrics, and runtime credential/signature sources; peer S3 and peer REST clients share bucket metadata, disk quorum reduction, endpoint layout, local disk initialization, and store helpers; control-plane snapshots are separate from data-plane RPC, but remote disk and peer clients still own data-movement side effects inside ECStore.
|
||||
|
||||
- `cargo tree -p rustfs-ecstore -e normal --depth 2` snapshot for dependency
|
||||
impact;
|
||||
- focused benchmarks for encode/decode, read reconstruction, bitrot verification,
|
||||
and large-object streaming;
|
||||
- contract sketch for shard sources, disk errors, bitrot IO, metrics, and file
|
||||
metadata without importing ECStore implementation modules;
|
||||
- compatibility plan for `rustfs_ecstore::api::erasure` and test harnesses;
|
||||
- rollback plan that keeps object read/write quorum and old-version file decode
|
||||
behavior unchanged.
|
||||
|
||||
### CRATE-002: `storage-cluster`
|
||||
|
||||
Current coupling:
|
||||
|
||||
- cluster RPC remote disk code depends on disk stores, disk health tracking,
|
||||
set-disk buffer sizing, local disk scan guards, internode metrics, and runtime
|
||||
credential/signature sources;
|
||||
- peer S3 and peer REST clients share bucket metadata, disk quorum reduction,
|
||||
endpoint layout, local disk initialization, and store helpers;
|
||||
- control-plane snapshots are separated from data-plane RPC, but remote disk and
|
||||
peer clients still own data movement side effects inside ECStore.
|
||||
|
||||
Decision: do not split in code yet. The storage-cluster boundary is a candidate
|
||||
only after remote disk, peer health, lock/quorum, runtime metrics, and endpoint
|
||||
layout contracts are explicit enough to stand below ECStore without circular
|
||||
dependencies.
|
||||
|
||||
Required evidence before proposing the split:
|
||||
|
||||
- dependency graph showing no cycle with ECStore, `rustfs-storage-api`, runtime
|
||||
source owners, or cluster control-plane owners;
|
||||
- RPC contract sketch for remote disk, peer S3, peer REST, auth/signature,
|
||||
internode metrics, and cancellation;
|
||||
- compatibility plan for `rustfs_ecstore::api::cluster`, `api::rpc`, and test
|
||||
fixtures that build local disks or endpoint pools;
|
||||
- focused tests for remote disk error classification, peer health recovery,
|
||||
per-pool quorum reduction, lock behavior, and data-stream request paths;
|
||||
- rollback plan that preserves quorum, remote disk IO, lock, peer health, and
|
||||
data movement behavior.
|
||||
|
||||
### CRATE-003: `bucket-lifecycle`
|
||||
|
||||
Decision: do not split in code yet. Lifecycle remains coupled to ECStore object
|
||||
operations, bucket metadata, `SetDisks` stale multipart cleanup, tier config,
|
||||
runtime lifecycle state, scanner metrics, notification/audit side effects, and
|
||||
replication delete scheduling.
|
||||
|
||||
Required evidence before proposing the split:
|
||||
|
||||
- contract sketch for lifecycle object operations, metadata access, runtime
|
||||
state, replication delete scheduling, and audit/notification sinks;
|
||||
- dependency graph showing the candidate crate can avoid importing ECStore
|
||||
implementation modules;
|
||||
- focused tests for lifecycle evaluation, expiry, transition, stale multipart
|
||||
cleanup, tier journal recovery, and lifecycle-originated replication deletes;
|
||||
- compatibility plan for `rustfs_ecstore::api::bucket::lifecycle` consumers;
|
||||
- rollback plan that preserves lifecycle queues, scanner repair accounting,
|
||||
tier transitions, object deletion behavior, and notification/audit events.
|
||||
|
||||
### CRATE-004: `bucket-replication`
|
||||
|
||||
Decision: do not split in code yet. Replication remains coupled to ECStore
|
||||
object APIs, bucket target clients, metadata systems, file metadata replication
|
||||
state, ECStore-owned runtime replication pool/stat handles, bucket monitor
|
||||
state, scanner repair classification, lifecycle-originated deletes, and
|
||||
notification events. RustFS-facing runtime consumers should use storage-owner
|
||||
wrapper handles while that state remains in ECStore.
|
||||
|
||||
Required evidence before proposing the split:
|
||||
|
||||
- contract sketch for replication storage operations, metadata/target access,
|
||||
runtime pool and stats, event sinks, and lifecycle/heal bridges;
|
||||
- dependency graph showing the candidate crate can avoid importing ECStore
|
||||
implementation modules;
|
||||
- focused tests for object replication, delete replication, resync state, heal
|
||||
repair queueing, target error handling, and queue admission;
|
||||
- compatibility plan for `rustfs_ecstore::api::bucket::replication` consumers;
|
||||
- rollback plan that preserves replication queues, MRF/resync state, target
|
||||
client behavior, scanner repair, and event emission.
|
||||
Decision: do not split. The boundary becomes a candidate only after remote disk, peer health, lock/quorum, runtime metrics, and endpoint layout contracts can stand below ECStore without cycles, with compatibility plans for `rustfs_ecstore::api::cluster` and `api::rpc` and focused tests for remote disk error classification, peer health recovery, per-pool quorum reduction, lock behavior, and data-stream request paths.
|
||||
|
||||
## Preservation Rules
|
||||
|
||||
- Do not reintroduce AppContext resolver fallback families in broad cleanup PRs.
|
||||
- Do not introduce direct global reads in admin, app, server, storage, scanner,
|
||||
heal, IAM, notify, observability, Swift, or S3 Select business logic.
|
||||
- Do not split crates in the same PR that moves runtime state.
|
||||
- Do not change startup order, readiness, KMS fatal boundaries, IAM recovery,
|
||||
lock quorum, object placement, reader behavior, or notification/audit
|
||||
lifecycle while shrinking global state.
|
||||
- Do not reintroduce AppContext resolver fallback families in broad cleanups.
|
||||
- Do not introduce direct global reads in admin, app, server, storage, scanner, heal, IAM, notify, observability, Swift, or S3 Select business logic.
|
||||
- Do not split crates in the same change that moves runtime state.
|
||||
- Do not change startup order, readiness, KMS fatal boundaries, IAM recovery, lock quorum, object placement, reader behavior, or notification/audit lifecycle while shrinking global state.
|
||||
|
||||
@@ -1,146 +1,64 @@
|
||||
# Global State Inventory
|
||||
|
||||
This inventory records the issue #730 baseline for global runtime state after
|
||||
the AppContext foundation and owner-local runtime-source boundaries were added.
|
||||
It is intentionally documentation-only: it classifies migration targets without
|
||||
changing startup, readiness, object IO, lifecycle, replication, or notification
|
||||
behavior.
|
||||
|
||||
## Counting Baseline
|
||||
|
||||
The audit uses the current workspace Rust sources and keeps broad static
|
||||
caches separate from runtime migration targets.
|
||||
|
||||
| Scope | Count | Command |
|
||||
|---|---:|---|
|
||||
| Rust source files | 1,252 | `rg --files -g '*.rs'` |
|
||||
| `OnceLock` references | 221 lines | `rg -n --glob '*.rs' 'OnceLock'` |
|
||||
| `GLOBAL_*` references | 273 lines | `rg -n --glob '*.rs' '\bGLOBAL_[A-Za-z0-9_]*\b'` |
|
||||
| `static NAME:` definitions | 621 lines | `rg -n --glob '*.rs' '^\s*(pub(\([^)]*\))?\s+)?static(\s+mut)?\s+[A-Za-z_][A-Za-z0-9_]*\s*:'` |
|
||||
| `lazy_static!` `static ref` definitions | 58 lines | `rg -n --glob '*.rs' '^\s*(pub\s+)?static\s+ref\s+[A-Za-z_][A-Za-z0-9_]*\s*:'` |
|
||||
| `static mut` definitions | 0 lines | `rg -n --glob '*.rs' '^\s*(pub(\([^)]*\))?\s+)?static\s+mut\s+'` |
|
||||
**Use this when:** you meet a `GLOBAL_*` static or an `OnceLock` and need to know whether it is a runtime ownership handle (reach it through a boundary), an owner-local static (leave it inside its module), or process-global by design.
|
||||
**Source of truth:** `crates/ecstore/src/api/mod.rs` (the `pub mod runtime` and `pub mod global` re-export lists), `crates/ecstore/src/runtime/global.rs`, `crates/ecstore/src/runtime/sources.rs`, and the statics themselves. Boundary rules are in [global-state-crate-split-plan.md](global-state-crate-split-plan.md).
|
||||
|
||||
## Global State Classification
|
||||
|
||||
| Category | Rule | Representative owners |
|
||||
|---|---|---|
|
||||
| Process-global | Process identity, metrics registries, lock manager, audit guard, TLS material, or other state that is intentionally one per process. | `crates/credentials`, `crates/common`, `crates/io-metrics`, `crates/lock`, `crates/obs`, `crates/tls-runtime` |
|
||||
| Runtime migration target | Mutable runtime state that describes the active object store, endpoints, local disks, lifecycle, replication, notification, config, or background controllers. | `crates/ecstore/src/runtime/global.rs`, `crates/ecstore/src/runtime/sources.rs`, `rustfs/src/app/context/*` |
|
||||
| Owner-local compatibility | Existing compatibility adapters that are allowed to read globals while callers migrate to AppContext-first or owner-local runtime-source APIs. | `rustfs/src/*/runtime_sources.rs`, `rustfs/src/*/storage_api.rs`, `crates/*/storage_api.rs` |
|
||||
| Test or fixture state | Static setup used by tests to amortize expensive ECStore setup or isolate compatibility harness state. | `rustfs/src/app/*_test.rs`, `crates/scanner/tests/*`, `crates/ecstore/src/**/tests` |
|
||||
| Cache or constant | Regexes, metrics descriptors, defaults, KVS registrations, headers, path constants, and small process caches that are not runtime ownership handles. | `crates/config`, `crates/obs/src/metrics`, `crates/utils`, `rustfs/src/server/readiness.rs` |
|
||||
| Legacy naming or review-needed | Old MinIO-port naming, stale comments, or names that need owner confirmation before code movement. | `GLOBAL_OBJECT_API` |
|
||||
| Process-global | Process identity, metrics registries, lock manager, audit guard, TLS material, or other state intentionally one per process. | `GLOBAL_LOCK_MANAGER` (`crates/lock`), `GLOBAL_CONN_MAP` (`crates/common`), `GLOBAL_RUSTFS_RPC_SECRET` (`crates/credentials`), `AUDIT_SYSTEM` (`crates/audit`), `crates/io-metrics`, `crates/obs`, `crates/tls-runtime` |
|
||||
| Runtime migration target | Mutable runtime state describing the active object store, endpoints, local disks, lifecycle, replication, notification, config, or background controllers. | `crates/ecstore/src/runtime/global.rs`, `crates/ecstore/src/runtime/sources.rs`, `rustfs/src/app/context/` |
|
||||
| Owner-local compatibility | Adapters allowed to read globals while callers migrate to AppContext-first or owner-local runtime-source APIs. | `rustfs/src/*/runtime_sources.rs`, `rustfs/src/*/storage_api.rs`, `crates/*/storage_api.rs` |
|
||||
| Owner-local static | A static private to one module and reached only through that module's functions: caches, single-run guards, admission locks, module toggles. | The RustFS inventory below |
|
||||
| Test or fixture state | Static setup that amortizes expensive ECStore setup or isolates harness state. | `rustfs/src/app/*_test.rs`, `crates/scanner/tests/`, `crates/test-utils/src/ecstore_test_compat.rs` |
|
||||
| Cache or constant | Regexes, metrics descriptors, defaults, KVS registrations, headers, path constants. | `crates/config`, `crates/obs/src/metrics`, `crates/utils` |
|
||||
|
||||
## Runtime Migration Inventory
|
||||
|
||||
These are the issue #730 targets that should remain visible until an owner
|
||||
migration PR removes or replaces each item.
|
||||
Runtime ownership handles that exist today. Reads go through `rustfs_ecstore::api::runtime`, bootstrap writes go through `rustfs_ecstore::api::global`, and RustFS code reaches both only from `rustfs/src/storage/storage_api.rs` and the AppContext resolvers.
|
||||
|
||||
| State | Current boundary | Category | Migration stance |
|
||||
|---|---|---|---|
|
||||
| `APP_CONTEXT_SINGLETON` | `rustfs/src/app/context/global.rs` | Owner-local compatibility | Keep as the context-first facade while no-context startup and embedded callers still exist. |
|
||||
| `GLOBAL_OBJECT_API`, `GLOBAL_OBJECT_STORE_RESOLVER` | `crates/ecstore/src/runtime/global.rs`, `rustfs/src/app/context/global.rs`, and storage compatibility APIs | Runtime migration target | Do not migrate first; it is tied to storage startup, IAM-after-storage AppContext publication, and data-plane resolver compatibility. The object-store resolver is now published from the AppContext owner path, no longer re-exported from the RustFS storage root, and RustFS AppContext tests no longer use the old `new_object_layer_fn` fallback chain. RustFS storage root no longer re-exports ECStore runtime/global facade symbols; callers must use storage/app/admin facades. |
|
||||
| `GLOBAL_ENDPOINTS`, `GLOBAL_IS_ERASURE`, `GLOBAL_IS_DIST_ERASURE`, `GLOBAL_IS_ERASURE_SD`, `GLOBAL_ROOT_DISK_THRESHOLD` | `crates/ecstore/src/runtime/global.rs` and `crates/ecstore/src/runtime/sources.rs` | Runtime migration target | Endpoint and setup-type reads now flow through ECStore `api::runtime` helpers at the RustFS storage facade boundary; root-disk-threshold access stays behind ECStore runtime helpers. Move endpoint ownership only after readiness and quorum behavior have explicit coverage. |
|
||||
| `GLOBAL_LOCAL_DISK_MAP`, `GLOBAL_LOCAL_DISK_ID_MAP`, `GLOBAL_LOCAL_DISK_SET_DRIVES` | `crates/ecstore/src/runtime/global.rs` and `crates/ecstore/src/runtime/sources.rs` | Runtime migration target | Local disk map, disk-id cache, and set-drive access now stay behind ECStore runtime-source helpers instead of direct global access; preserve disk lookup, remote/local classification, and test reset hooks in later ownership changes. |
|
||||
| `GLOBAL_EXPIRY_STATE`, `GLOBAL_TRANSITION_STATE`, `GLOBAL_LIFECYCLE_SYS` | `crates/ecstore/src/bucket/lifecycle/*`, `crates/ecstore/src/runtime/global.rs`, and `crates/ecstore/src/runtime/sources.rs` | Runtime migration target | Lifecycle state globals now stay behind ECStore lifecycle owner helpers and ECStore runtime-source helpers; RustFS AppContext has expiry/transition state interfaces and resolver coverage, and daily tier stats derive from the transition-state handle instead of a separate context boundary; scanner expiry-state access still uses the ECStore runtime `expiry_state_handle` boundary until scanner gets an injected provider. |
|
||||
| `GLOBAL_REPLICATION_POOL`, `GLOBAL_REPLICATION_STATS`, `GLOBAL_BUCKET_MONITOR` | `crates/ecstore/src/bucket/replication/*`, `crates/ecstore/src/runtime/global.rs` | Runtime migration target | Replication pool/stat access now stays behind replication owner and ECStore runtime-source helpers; bucket-monitor reads now flow through ECStore `api::runtime` at the RustFS storage facade boundary while AppContext/runtime-source resolvers remain the caller boundary. |
|
||||
| `GLOBAL_TIER_CONFIG_MGR`, `GLOBAL_STORAGE_CLASS`, `GLOBAL_CONFIG_SYS`, `GLOBAL_SERVER_CONFIG` | `crates/ecstore/src/config`, `crates/config`, `rustfs/src/app/context/runtime_sources.rs` | Runtime migration target | Tier config manager reads and reloads now use the ECStore runtime-source helper; move remaining config state through config/runtime-source owners only, without combining storage-class behavior or persistence changes. |
|
||||
| `GLOBAL_EVENT_NOTIFIER`, `GLOBAL_NOTIFICATION_SYS` | `crates/ecstore/src/runtime/global.rs`, `crates/ecstore/src/runtime/sources.rs`, and `crates/ecstore/src/services/*` | Runtime migration target | `GLOBAL_EVENT_NOTIFIER` and `GLOBAL_NOTIFICATION_SYS` access now stay behind ECStore runtime-source and notification owner helpers; move remaining notification ownership only through notify/runtime-source boundaries. |
|
||||
| `EVENT_DISPATCH_HOOK` | `crates/ecstore/src/services/event_notification.rs`, RustFS server event bridge, and storage compatibility APIs | Runtime migration target / owner helper | Direct hook storage stays inside the ECStore event-notification owner; RustFS registers the bridge through the storage compatibility facade until event dispatch ownership moves behind an injected notification sink. |
|
||||
| `GLOBAL_BUCKET_METADATA_SYS` | `crates/ecstore/src/bucket/metadata_sys.rs`, `crates/ecstore/src/runtime/sources.rs`, and RustFS storage compatibility APIs | Runtime migration target | Bucket metadata system direct access now stays inside the ECStore metadata owner; callers use metadata owner helpers or storage/runtime-source compatibility functions until metadata ownership moves behind an injected runtime context. |
|
||||
| `GLOBAL_BOOT_TIME`, `GLOBAL_BACKGROUND_SERVICES_CANCEL_TOKEN`, `GLOBAL_DEPLOYMENT_ID`, `GLOBAL_REGION`, `GLOBAL_RUSTFS_PORT`, `GLOBAL_LOCAL_NODE_NAME_FALLBACK`, `GLOBAL_LOCAL_NODE_NAME_HEX_FALLBACK` | `crates/ecstore/src/runtime/global.rs`, `crates/ecstore/src/runtime/sources.rs` | Runtime migration target | Boot time, background service cancellation token reads, ECStore local-node-name fallback reads, and deployment ID/region/port reads now stay behind the ECStore runtime-source API; scalar writes remain behind bootstrap owner helpers until ownership handles replace them. |
|
||||
| `WORKLOAD_ADMISSION_SNAPSHOT_PROVIDER` | `crates/ecstore/src/runtime/sources.rs`, RustFS startup background setup, and storage compatibility APIs | Runtime migration target / owner helper | Startup publishes the workload provider through the storage compatibility facade, and ECStore data movement reads it only through the runtime-source helper until workload admission ownership moves into an explicit runtime context. |
|
||||
| `GLOBAL_LOCAL_LOCK_CLIENT`, `GLOBAL_LOCK_CLIENTS`, `GLOBAL_LOCK_MANAGER` | `crates/ecstore/src/runtime/global.rs`, `crates/lock` | Runtime migration target / process-global split | ECStore lock client reads now flow through ECStore `api::runtime` helpers at the RustFS storage facade boundary; preserve lock quorum and lock client selection while keeping the process-level lock manager separate from endpoint-specific clients. |
|
||||
| `GLOBAL_CONN_MAP`, `GLOBAL_LOCAL_NODE_NAME`, `GLOBAL_RUSTFS_HOST`, `GLOBAL_RUSTFS_ADDR`, `GLOBAL_ROOT_CERT`, `GLOBAL_MTLS_IDENTITY`, `GLOBAL_OUTBOUND_TLS_GENERATION` | `crates/common`, `crates/tls-runtime`, `crates/ecstore/src/runtime/sources.rs` | Runtime migration target / process-global split | Internode connection cache, common local node name, RustFS host/address reads, and outbound TLS material reads are now owned behind `rustfs_common` helpers; migrate the remaining transport and TLS state only after internode transport and outbound TLS ownership are explicit, without changing cached channel reuse or TLS reload semantics. |
|
||||
| `GLOBAL_RUSTFS_RPC_SECRET` | `crates/credentials`, `crates/ecstore/src/runtime/sources.rs` | Runtime migration target / process-global split | RPC auth token writes now stay behind the `rustfs_credentials` helper boundary; migrate only if runtime secret ownership changes, preserving lazy environment and credential-derived token semantics. |
|
||||
| `GLOBAL_HEAL_MANAGER`, `GLOBAL_HEAL_CHANNEL_PROCESSOR`, `GLOBAL_AHM_SERVICES_CANCEL_TOKEN` | `crates/heal/src/lib.rs` | Runtime migration target / process-global split | Direct access now stays inside the heal owner; callers use heal helper functions until heal runtime ownership moves behind explicit owner handles. |
|
||||
| `AUDIT_SYSTEM` | `crates/audit/src/global.rs` | Runtime migration target / process-global split | Direct global access now stays inside the audit owner; callers use audit helper functions until audit lifecycle ownership moves behind AppContext or a runtime-source boundary. |
|
||||
| `GLOBAL_PROCESSORS` | `crates/ecstore/src/services/batch_processor.rs`, `crates/ecstore/src/runtime/sources.rs` | Runtime migration target / owner helper | Direct static access now stays inside the ECStore batch processor owner; callers use `get_global_processors` or the ECStore runtime-source helper until processor ownership moves into an injected runtime context. |
|
||||
| `INTERNODE_DATA_TRANSPORT` | `crates/ecstore/src/cluster/rpc/internode_data_transport.rs` | Runtime migration target / owner helper | Direct static access now stays inside the ECStore internode transport owner; callers use `build_internode_data_transport_from_env` until backend selection moves into an injected runtime context. |
|
||||
| `GLOBAL_KMS_SERVICE_MANAGER` | `crates/kms/src/service_manager.rs`, RustFS KMS runtime sources | Runtime migration target / owner helper | Direct static access now stays inside the `rustfs_kms` service manager owner; RustFS callers use KMS helpers or AppContext/runtime-source handles until KMS ownership fully moves into runtime context. |
|
||||
| `GLOBAL_CAPACITY_MANAGER` | `crates/object-capacity/src/capacity_manager.rs`, RustFS capacity service | Runtime migration target / owner helper | Direct static access now stays inside the object-capacity owner; callers use `get_capacity_manager` or isolated manager factories until capacity ownership moves into an injected runtime context. |
|
||||
| `GLOBAL_BUCKET_TARGET_SYS` | `crates/ecstore/src/bucket/bucket_target_sys.rs`, admin/app/scanner/replication target paths | Runtime migration target / owner helper | Direct static access now stays inside the ECStore bucket target owner; callers still use `BucketTargetSys::get()` until bucket target ownership moves behind a runtime-source or replication target boundary. |
|
||||
| `USAGE_MEMORY_CACHE`, `USAGE_CACHE_UPDATING` | `crates/ecstore/src/data_usage/mod.rs` | Runtime migration target / owner-local cache | Data-usage memory overlay and singleflight state stay private to the ECStore data-usage owner; callers use data-usage functions until scanner/data-usage ownership moves behind an injected runtime context. |
|
||||
| Handle (`rustfs_ecstore::api::runtime`) | Backing state | Stance |
|
||||
|---|---|---|
|
||||
| `object_store_handle` | `GLOBAL_OBJECT_API`, `GLOBAL_OBJECT_STORE_RESOLVER` (`crates/ecstore/src/runtime/global.rs`); the resolver is published from the AppContext owner path | Do not migrate first: tied to storage startup, IAM-after-storage AppContext publication, and data-plane resolver compatibility. |
|
||||
| `endpoint_pools`, `setup_is_erasure`, `setup_is_dist_erasure`, `setup_is_erasure_sd`, `first_cluster_node_is_local` | `GLOBAL_ENDPOINTS` and setup-type state (`crates/ecstore/src/runtime/global.rs`) | Move endpoint ownership only after readiness and quorum behavior have explicit coverage. |
|
||||
| `local_disk_map_read` | Local disk map and set-drive state (`crates/ecstore/src/runtime/sources.rs`) | Preserve disk lookup, remote/local classification, and test reset hooks. |
|
||||
| `expiry_state_handle`, `transition_state_handle` | Lifecycle expiry and transition state, `GLOBAL_LIFECYCLE_SYS` (`crates/ecstore/src/runtime/global.rs`) | Lifecycle owner helpers and the AppContext `ExpiryStateInterface` (`rustfs/src/app/context/interfaces.rs`) are the caller boundary; the scanner still reads `expiry_state_handle` until it gets an injected provider. |
|
||||
| `global_tier_config_mgr` | Tier config manager | Reads and reloads stay behind this helper. |
|
||||
| `bucket_monitor` | Replication bandwidth monitor | Replication pool/stat handles are projected into RustFS wrapper types at the storage boundary. |
|
||||
| `global_lock_client`, `global_lock_clients` | `GLOBAL_LOCAL_LOCK_CLIENT`, `GLOBAL_LOCK_CLIENTS` (`crates/ecstore/src/runtime/global.rs`) | Preserve lock quorum and client selection; the process-level `GLOBAL_LOCK_MANAGER` stays separate. |
|
||||
| `boot_time`, `deployment_id`, `region`, `rustfs_port` | `GLOBAL_BOOT_TIME`, deployment id, region, and port state (`crates/ecstore/src/runtime/global.rs`) | Scalar writes remain behind the `api::global` setters (`set_global_endpoints`, `set_global_region`, `set_global_rustfs_port`, `set_object_store_resolver`, `shutdown_background_services`, `update_erasure_type`). |
|
||||
|
||||
## Owner-Local Cache Inventory
|
||||
Owner-helper handles outside the runtime-source list stay inside their owner and are reached through owner functions: `GLOBAL_EVENT_NOTIFIER` (`crates/ecstore/src/runtime/global.rs`); `GLOBAL_NOTIFICATION_SYS`, `EVENT_DISPATCH_HOOK`, `GLOBAL_PROCESSORS`, `INTERNODE_DATA_TRANSPORT`, `GLOBAL_BUCKET_TARGET_SYS`, `GLOBAL_CONFIG_SYS`, `GLOBAL_STORAGE_CLASS`, `WORKLOAD_ADMISSION_SNAPSHOT_PROVIDER` (ECStore owner modules); `GLOBAL_SERVER_CONFIG` (`crates/config/src/server_config.rs`); `GLOBAL_HEAL_RUNTIME`, `GLOBAL_AHM_SERVICES_CANCEL_TOKEN` (`crates/heal/src/lib.rs`); `GLOBAL_KMS_SERVICE_MANAGER` (`crates/kms/src/service_manager.rs`); `GLOBAL_CAPACITY_MANAGER` (`crates/object-capacity/src/capacity_manager.rs`); `APP_CONTEXT_SINGLETON` (`rustfs/src/app/context/global.rs`).
|
||||
|
||||
These owner-local caches and static guards are part of the broad issue #730
|
||||
`OnceLock` audit, but they are not runtime ownership handles. They stay private
|
||||
to the defining owner module; callers must use the existing owner APIs instead
|
||||
of reaching across module boundaries.
|
||||
Regenerate:
|
||||
|
||||
| State | Owner boundary | Category | Migration stance |
|
||||
|---|---|---|---|
|
||||
| `READ_REPAIR_HEAL_CACHE` | `crates/ecstore/src/set_disk/read.rs` | Cache or constant / owner-local cache | Read-repair heal suppression stays local to set-disk read handling. |
|
||||
| `DISK_COMPRESSION_CONFIG` | `crates/ecstore/src/io_support/compress.rs` | Cache or constant / owner-local cache | Disk compression environment parsing stays local to IO support compression helpers. |
|
||||
| `CACHED_MAX_INFLIGHT_BYTES`, `CACHED_BATCH_BLOCKS`, `CACHED_BYTESMUT_INGEST` | `crates/ecstore/src/erasure/coding/encode.rs` | Cache or constant / owner-local cache | Erasure encode tuning caches stay local to the coding owner. |
|
||||
| `CACHED_PUT_LARGE_BATCH_MIN_SIZE_BYTES`, `CACHED_MULTIPART_PUT_LARGE_BATCH_MIN_SIZE_BYTES`, `OBJECT_LOCK_DIAG_ENABLED` | `crates/ecstore/src/set_disk/mod.rs` | Cache or constant / owner-local cache | Set-disk batching and diagnostics caches stay local to the set-disk owner. |
|
||||
| `DRIVE_TIMEOUT_PROFILE_CACHE`, `DRIVE_TIMEOUT_HEALTH_POLICY_CACHE` | `crates/ecstore/src/disk/disk_store.rs` | Cache or constant / owner-local cache | Drive timeout environment caches stay local to the disk-store owner. |
|
||||
| `TIER_FREE_VERSION_RECOVERY_STARTED`, `TIER_DELETE_JOURNAL_RECOVERY_STARTED` | `crates/ecstore/src/bucket/lifecycle/bucket_lifecycle_ops.rs` | Cache or constant / owner-local static guard | Lifecycle recovery single-run guards stay local to lifecycle operations. |
|
||||
| `REMOTE_DELETE_INFLIGHT`, `REMOTE_DELETE_LIMITER`, `REMOTE_DELETE_BREAKER`, `REMOTE_TIER_DELETE_TEST_HOOK` | `crates/ecstore/src/bucket/lifecycle/tier_sweeper.rs` | Cache or constant / owner-local static guard | Remote tier delete concurrency, breaker, and test hook state stay local to the tier sweeper owner. |
|
||||
| `ACTIVE_REGISTRY`, `BackendCapacity` | `crates/kms/src/policy.rs` | Process-global owner-local admission capacity registry | KMS policy generations share only active semaphore capacity by backend identity; each generation owns fresh bounded queues and circuit breakers. Callers access this state only through `RetryPolicy`. |
|
||||
```bash
|
||||
rg -n -A4 'pub use crate::runtime::(sources|global)::' crates/ecstore/src/api/mod.rs
|
||||
rg -n --glob '*.rs' 'static (ref )?GLOBAL_[A-Z_]+' crates rustfs/src
|
||||
```
|
||||
|
||||
## RustFS Owner-Local Static Inventory
|
||||
|
||||
These RustFS-side lazy, atomic, and `OnceLock` statics are also part of the
|
||||
issue #730 process-static audit. They are private implementation details for
|
||||
their owner modules, not shared runtime ownership handles. This section excludes
|
||||
allocator statics, public contract/error references, route handler constants,
|
||||
and `APP_CONTEXT_SINGLETON`, which is classified in the runtime migration
|
||||
inventory. Generic function-local names such as `CACHE`, `LOCK`, `INIT`, and
|
||||
`ENABLED` are documented by owner row instead of name-regex guarded.
|
||||
RustFS-side statics that matter architecturally because other modules are tempted to reach them. They stay private to their owner module; callers use the owner's functions.
|
||||
|
||||
| State | Owner boundary | Category | Migration stance |
|
||||
|---|---|---|---|
|
||||
| `KEYSTONE_AUTH`, `KEYSTONE_MAPPER`, `KEYSTONE_CONFIG` | `rustfs/src/auth_keystone.rs` | Process-global owner-local state | Keystone authentication provider, identity mapper, and config stay private to the Keystone auth owner. |
|
||||
| `LICENSE_STATE`, `LICENSE_VERIFIER` | `rustfs/src/license.rs` | Process-global owner-local state | License state and verifier selection stay private to the license owner; callers use license helper functions. |
|
||||
| `CPU_CONT_GUARD`, `PROFILING_CANCEL_TOKEN` | `rustfs/src/profiling.rs` | Process-global owner-local guard | CPU profiling guard and cancellation state stay private to the profiling owner. |
|
||||
| `MEMORY_SYSTEM` | `rustfs/src/memory_observability.rs` | Process-global owner-local cache | Memory sampling keeps the `sysinfo::System` cache private to the memory observability owner. |
|
||||
| `DISPLAY_CONFIG_SNAPSHOT`, `GLOBAL_CONFIG_SNAPSHOT` | `rustfs/src/config/snapshot.rs` | Process-global owner-local state | Config snapshots stay private to the config snapshot owner. |
|
||||
| `BUFFER_CONFIG_SINGLETON`, `BUFFER_PROFILE_ENABLED` | `rustfs/src/config/workload_profiles.rs` | Process-global owner-local state | Workload buffer profile configuration stays private to workload profile helpers. |
|
||||
| `LEGACY_CREDENTIAL_WARNED_KEYS` | `rustfs/src/config/config_struct.rs` | Process-global owner-local cache | Legacy credential warning de-duplication stays private to config parsing. |
|
||||
| `CONSOLE_CONFIG` | `rustfs/src/admin/console.rs` | Process-global owner-local state | Console bootstrap config stays private to the admin console owner. |
|
||||
| `ACTIVE_HTTP_REQUESTS` | `rustfs/src/server/http.rs` | Process-global owner-local counter | HTTP request inflight accounting stays private to the HTTP server owner. |
|
||||
| Function-local `CACHE` and `LOCK` statics | `rustfs/src/server/readiness.rs` | Cache or constant / owner-local cache | Readiness and cluster-health caches stay function-local to readiness probes. |
|
||||
| `USE_STARSHARD_CACHE`, `BUCKET_CACHE_SMALL`, `BUCKET_CACHE_LARGE` | `rustfs/src/storage/ecfs_extend.rs` | Cache or constant / owner-local cache | Bucket validation cache backend selection and cache storage stay private to the ECFS extension owner. |
|
||||
| `GLOBAL_SSE_DEK_PROVIDER`, `SSE_TEST_LOCK` | `rustfs/src/storage/sse.rs` | Owner-local cache / test state | SSE DEK provider cache and test serialization lock stay private to the SSE owner. |
|
||||
| `AUTH_FS` | `rustfs/src/storage/access.rs` | Cache or constant / owner-local cache | Authorization tag-condition lookup keeps its filesystem helper private to the access owner. |
|
||||
| `DEADLOCK_DETECTOR` | `rustfs/src/storage/deadlock_detector.rs` | Process-global owner-local state | Deadlock detector lifecycle state stays private to the storage deadlock detector owner. |
|
||||
| `CONCURRENCY_MANAGER`, `ACTIVE_GET_REQUESTS`, `ACTIVE_PUT_REQUESTS` | `rustfs/src/storage/concurrency/*` | Process-global owner-local scheduler state | Storage concurrency manager and request counters remain inside the storage concurrency owner boundary. |
|
||||
| `GET_OBJECT_BUFFER_THRESHOLD_WARNED`, `GET_READER_STREAM_BUFFER_SIZE_OVERRIDE`, function-local `ENABLED`, `OBJECT_SEEK_SUPPORT_THRESHOLD`, `OBJECT_SEEK_SUPPORT_CONCURRENCY_THRESHOLDS` | `rustfs/src/app/object/get.rs` | Cache or constant / owner-local cache | Object GET/seek tuning caches and warning guards stay private to object usecase helpers. |
|
||||
| `SUPPORTED_HEADERS` | `rustfs/src/storage/options.rs` | Cache or constant / owner-local constant | Supported-header lookup state stays private to storage option parsing. |
|
||||
| `AUDIT_TARGET_SPECS`, `NOTIFICATION_TARGET_SPECS` | `rustfs/src/admin/handlers/audit.rs`, `rustfs/src/admin/handlers/event.rs`, `rustfs/src/admin/handlers/plugins_instances.rs` | Cache or constant / owner-local constant | Admin target descriptor tables stay private to their handler owners. |
|
||||
| `SITE_REPLICATION_PEER_CLIENT` | `rustfs/src/site_replication/transport.rs` | Process-global owner-local cache | Site-replication peer client cache stays private to the site-replication transport module. The state RMW transaction holds no process-local mutex — see `rustfs/src/site_replication/state_lock.rs`. |
|
||||
| `AUDIT_MODULE_ENABLED`, `NOTIFY_MODULE_ENABLED`, `PERSISTED_NOTIFY_MODULE_ENABLED`, `PERSISTED_AUDIT_MODULE_ENABLED`, `PERSISTED_MODULE_SWITCH_CONFIGURED` | `rustfs/src/server/audit.rs`, `rustfs/src/server/event.rs`, `rustfs/src/server/module_switch.rs` | Process-global owner-local toggles | Audit/notify module snapshots stay private to the server module switch owners. |
|
||||
| `DELETE_TAIL_TOTAL`, `DELETE_CLEANUP_TOTAL`, `DELETE_REPLICATION_TOTAL`, `DELETE_NOTIFY_TOTAL` | `rustfs/src/delete_tail_activity.rs` | Process-global owner-local counters | Delete-tail activity counters stay private behind delete-tail activity helpers. |
|
||||
| `EMBEDDED_SERVER_STARTED` | `rustfs/src/startup_lifecycle.rs` | Process-global owner-local guard | Embedded startup single-start protection stays private to startup lifecycle. |
|
||||
| `TEST_OUTBOUND_TLS_GENERATION` | `rustfs/src/site_replication/mod.rs` | Test or fixture state | Outbound TLS generation test hook state stays private to site-replication transport tests. |
|
||||
| `TEST_REMAINING_FAILURES` | `rustfs/src/startup_iam.rs` | Test or fixture state | IAM startup retry injection state stays private to debug/test startup code. |
|
||||
| `CAPACITY_DIRTY_SCOPE_ENV`, `CAPACITY_DIRTY_SCOPE_INIT`, `GLOBAL_ENV`, function-local `INIT` | `rustfs/src/app/*_test.rs` | Test or fixture state | App integration test fixture state stays private to the owning test modules. |
|
||||
| Static | Owner | Stance |
|
||||
|---|---|---|
|
||||
| `KEYSTONE_AUTH`, `KEYSTONE_MAPPER`, `KEYSTONE_CONFIG` | `rustfs/src/auth_keystone.rs` | Keystone provider, mapper, and config stay private to the Keystone owner. |
|
||||
| `DEADLOCK_DETECTOR` | `rustfs/src/storage/deadlock_detector.rs` | Detector lifecycle stays private to the storage deadlock detector. |
|
||||
| `CONCURRENCY_MANAGER` | `rustfs/src/storage/concurrency/manager.rs` | Storage concurrency scheduler state stays inside the concurrency owner. |
|
||||
| `GLOBAL_KMS_DEK_PROVIDER`, `GLOBAL_SSE_DEK_PROVIDER` | `rustfs/src/storage/sse.rs` | DEK provider caches stay private to the SSE owner. |
|
||||
| `ECSTORE_EVENT_DISPATCH_HOOK` | `rustfs/src/server/event.rs` | Event bridge registration goes through the storage facade. |
|
||||
| `AUDIT_MODULE_ENABLED`, `NOTIFY_MODULE_ENABLED` | `rustfs/src/module_switches.rs` | Module toggles are read through module-switch helpers; `MODULE_SWITCH_RMW_LOCK` (`rustfs/src/server/module_switch.rs`) serializes persisted updates. |
|
||||
| `RUNTIME_CONFIG_RELOAD_MUTEX` | `rustfs/src/admin/service/config.rs` | Serializes dynamic config reload fanout. |
|
||||
| `EMBEDDED_RUNTIME_OWNERS` | `rustfs/src/startup_shutdown.rs` | Embedded runtime owner handles used for shutdown ordering. |
|
||||
| `SERVICE_FROZEN` | `rustfs/src/admin/handlers/system.rs` | Service freeze flag stays behind the system admin handler. |
|
||||
| `RECONCILER` | `rustfs/src/site_replication_reconcile.rs` | Site-replication reconciler singleton. |
|
||||
| `CONSOLE_CONFIG` | `rustfs/src/admin/console.rs` | Console bootstrap config. |
|
||||
| `LICENSE_STATE`, `LICENSE_VERIFIER` | `rustfs/src/license.rs` | License state and verifier stay behind license helpers. |
|
||||
|
||||
## First Code-Bearing Candidate
|
||||
Regenerate the full list (long, mostly caches and test hooks):
|
||||
|
||||
`GLOBAL_EXPIRY_STATE` is the safest first runtime migration candidate:
|
||||
|
||||
- AppContext already exposes `ExpiryStateInterface` and resolver coverage in
|
||||
`rustfs/src/app/context.rs`.
|
||||
- ECStore access is already concentrated in
|
||||
`crates/ecstore/src/runtime/sources.rs`.
|
||||
- The main external readers can be moved through storage/observability facades
|
||||
before changing lifecycle queue ownership.
|
||||
|
||||
Do not migrate `GLOBAL_OBJECT_API` first. It is coupled to storage startup,
|
||||
object-store resolver publication, IAM-after-storage AppContext initialization,
|
||||
and broad data-plane compatibility.
|
||||
|
||||
## Verification
|
||||
|
||||
Inventory and guardrail PRs should run:
|
||||
|
||||
- `bash -n scripts/check_architecture_migration_rules.sh`
|
||||
- `./scripts/check_architecture_migration_rules.sh`
|
||||
- `cargo fmt --all --check`
|
||||
- `git diff --check`
|
||||
|
||||
Code-bearing migration PRs must add focused tests for the owner being moved
|
||||
before running broader gates.
|
||||
```bash
|
||||
rg -n '^\s*(pub(\(crate\))? )?static [A-Z_]+' rustfs/src
|
||||
```
|
||||
|
||||
@@ -0,0 +1,77 @@
|
||||
# Heal concurrency model
|
||||
|
||||
**Use this when:** changing heal, PUT/multipart commit, delete, lifecycle expiry, or data-movement code that touches the same `(bucket, object)` commit surface; or evaluating whether RustFS needs a persistent per-object healing marker like MinIO's `x-minio-healing`.
|
||||
**Source of truth:** `crates/ecstore/src/set_disk/ops/heal.rs` (`heal_object_with_explicit_version_regen`, `HealObjectLockKind`, `HEAL_RENAME_INCOMPLETE`), `crates/ecstore/src/set_disk/ops/object.rs` (PUT/DELETE lock sections, `reconcile_old_data_cleanup_receipts`), `crates/ecstore/src/set_disk/core/io_primitives.rs` (`commit_rename_data_dir`, `report_old_data_dir_cleanup`, `reclaim_orphan_data_dirs`), `crates/filemeta/src/fileinfo.rs` (`FileInfo::set_healing`), `crates/heal/src/heal/manager/queue.rs` (dedup keys).
|
||||
|
||||
## Model
|
||||
|
||||
Heal and every foreground or background write path serialize on the same object-level namespace write lock (a quorum lock RPC in distributed mode, the in-process lock manager on a single node; granularity is the object, the version component is always `None`), and heal holds its guard across the whole rename commit. MinIO's `x-minio-healing` marker is an out-of-lock defence against version-cleanup logic inside `RenameData` interleaving with a heal commit; RustFS's commit model has no such interleaving, so no persistent marker exists (`x-minio-healing` does not occur in `crates/` or `rustfs/`) and none is needed. Three layers replace it:
|
||||
|
||||
| Layer | Mechanism | Owner |
|
||||
| --- | --- | --- |
|
||||
| In-lock mutual exclusion | Heal and all write-path commit points take the `(bucket, object)` namespace write lock. | `acquire_heal_object_lock` in `crates/ecstore/src/set_disk/ops/heal.rs`; lock sections in `ops/object.rs` and `ops/multipart.rs` |
|
||||
| Commit-model isolation | `rename_data` contains no version cleanup that could interleave with heal. Physical deletion of a replaced old `data_dir` runs after the object lock is released (the commit tail) and only for unshared directories already superseded by the new commit. | `commit_rename_data_dir` in `crates/ecstore/src/set_disk/core/io_primitives.rs` |
|
||||
| Transient healing flag | `FileInfo::set_healing` sets the internal `SUFFIX_HEALING` key on the in-memory `FileInfo` of a heal commit; `rename_data` reads it through `is_healing` to clear a stale non-empty target `data_dir` before the rename (in-place repair reuses the `data_dir`, and `rename(2)` cannot replace a non-empty directory). The key is never persisted (`is_skip_meta_key` in `crates/filemeta/src/filemeta.rs`). A non-heal commit that meets a non-empty target fails explicitly; tests lock both directions. | `crates/filemeta/src/fileinfo.rs`, `crates/ecstore/src/disk/local.rs` |
|
||||
|
||||
## Heal lock scope
|
||||
|
||||
`heal_object` delegates to `heal_object_with_explicit_version_regen`, which takes the namespace write lock at entry unless `opts.no_lock` is set and binds the guard to the function scope. The guard covers the quorum metadata read, EC reconstruction, per-disk rename commit, tmp cleanup, the `HEAL_RENAME_INCOMPLETE` partial-commit return, and orphan `data_dir` reclamation (`reclaim_orphan_data_dirs`).
|
||||
|
||||
Read-repair heals (`opts.read_repair`) hold a shared lock (`HealObjectLockKind::Read`) during reconstruction so readers keep flowing, then `acquire_revalidated_read_repair_commit_lock` takes the write lock and re-reads a commit fingerprint; a changed fingerprint aborts the commit (`read_repair_commit_stale`).
|
||||
|
||||
## Lock-intersection matrix
|
||||
|
||||
| # | Concurrent path | Lock held by that path | Outcome | Where |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| 1 | PUT commit | object write lock; `rename_data` inside it | serialized | `ops/object.rs` put commit |
|
||||
| 2 | PUT old `data_dir` tail cleanup | none (runs after the lock is dropped) | unlocked, semantically safe ([commit tail](#commit-tail-cleanup)) | `commit_rename_data_dir` in `core/io_primitives.rs` |
|
||||
| 3 | DELETE object or version | object write lock; `delete_version` inside it | serialized | `ops/object.rs` `delete_object` |
|
||||
| 4 | Batch DELETE | per-object write locks (batch lock RPC in distributed mode) | serialized | `ops/object.rs` `delete_objects` |
|
||||
| 5 | CompleteMultipartUpload | object write lock plus upload-path lock; rename inside | serialized | `ops/multipart.rs` |
|
||||
| 6 | CompleteMultipart tail cleanup | none (after lock drop) | unlocked, semantically safe ([commit tail](#commit-tail-cleanup)) | `ops/multipart.rs` |
|
||||
| 7 | AbortMultipartUpload | upload-path lock in the multipart bucket only | disjoint resources: abort never touches the object `data_dir` or `xl.meta` | `ops/multipart.rs` |
|
||||
| 8 | ILM expiry including DeleteAllVersions | `delete_prefix_object=true` keeps the object lock; `FreeVersionTask` locks explicitly; noncurrent batches use batch locks | serialized | `crates/ecstore/src/bucket/lifecycle/bucket_lifecycle_ops.rs` |
|
||||
| 9 | Pure prefix delete | `delete_prefix` without `delete_prefix_object` takes no child-object lock | unlocked; no production caller ([prefix delete](#pure-prefix-delete)) | `ops/object.rs` lock condition in `delete_object` |
|
||||
| 10 | Orphan `data_dir` reclamation | none inside the function; its only production caller runs inside the heal lock | serialized within heal | `reclaim_orphan_data_dirs` in `core/io_primitives.rs` |
|
||||
| 11 | Old-cleanup receipt reconciliation | none inside the function; caller runs inside the heal lock and an epoch fence rejects stale receipts | serialized | `reconcile_old_data_cleanup_receipts` in `ops/object.rs` |
|
||||
| 12 | Replication | data plane writes to the remote over HTTP; local metadata write-back takes the object lock | serialized or disjoint | `crates/ecstore/src/bucket/replication/replication_resyncer.rs` |
|
||||
| 13 | Data movement, rebalance, decommission source cleanup | explicit object lock plus version-unchanged recheck; `no_lock` only reuses an already-held guard | serialized | `crates/ecstore/src/data_movement/mod.rs` |
|
||||
| 14 | CopyObject | destination object lock through the PUT chain | serialized | `ops/object.rs` `copy_object` |
|
||||
| 15 | Another heal task (different `HealType`, or `force_start`) | dedup keys are per `HealType` and `force_start` skips dedup, so tasks may coexist | serialized on the namespace write lock | `make_dedup_key_for_type` in `crates/heal/src/heal/manager/queue.rs` |
|
||||
| 16 | Admin heal with `nolock=true` | caller bypasses the lock | unlocked by operator choice ([no_lock](#no_lock-and-force_start)) | `rustfs/src/admin/handlers/heal.rs` |
|
||||
| 17 | Stale multipart cleanup | upload-path lock in the multipart bucket | disjoint resources | `crates/ecstore/src/bucket/lifecycle/bucket_lifecycle_ops.rs` |
|
||||
|
||||
## Residual windows
|
||||
|
||||
### Commit tail cleanup
|
||||
|
||||
Rows 2 and 6. After a write path commits and releases the object lock, it best-effort deletes the replaced old `data_dir`; the code deliberately does not block the next operation on this. The deletion can race a concurrent heal reading or rebuilding that same old `data_dir`, but the race is semantically safe:
|
||||
|
||||
- The target is an unshared `data_dir` already replaced by the new commit. Heal's canonical metadata comes from quorum arbitration (ETag, mod time), and quorum already points at the new version, so heal cannot resurrect the replaced version as canonical.
|
||||
- The worst outcome is one transient failure or no-op for the heal round on the old version; the next round converges. Cleanup residue is reported and re-queued for heal via `report_old_data_dir_cleanup`.
|
||||
- Long heals such as drive replacement request explicit versions and read quorum metadata inside the lock, so the tail does not affect them.
|
||||
|
||||
### Pure prefix delete
|
||||
|
||||
Row 9. `delete_prefix && !delete_prefix_object` takes no child-object locks (an object namespace lock cannot protect a recursive prefix delete), so a heal running during the prefix delete could theoretically rebuild a version from stale quorum metadata. Every production `delete_prefix: true` call site also sets `delete_prefix_object: true` (and therefore takes the object lock); the remaining `delete_prefix`-only call sites are in test modules. A future caller that needs a pure prefix delete must prove isolation from heal and scanner at the call site (for example a bucket-level scan fence).
|
||||
|
||||
### `no_lock` and `force_start`
|
||||
|
||||
Row 16. Admin heal requests pass the client's `nolock` parameter through (`rustfs/src/admin/handlers/heal.rs`), matching the MinIO madmin option. Setting it is an explicit operator choice that accepts races with concurrent writes; it is documented, not restricted.
|
||||
|
||||
Heal-side invariants that hold regardless of the caller:
|
||||
|
||||
- Dedup keys are disjoint across `HealType` (object, metadata, MRF, EC decode, prefix), and admin `force_start` skips dedup. Several heal tasks for one object can therefore exist at once, but every production entry calls `heal_object` with `no_lock=false`, so their execution bodies serialize on the namespace write lock.
|
||||
- Read-repair's local TTL reservation dedups only its own source and does not block heals from other sources; the namespace lock is the backstop.
|
||||
- The healing flag is never persisted, so there is no reverse risk of a leftover marker making a later commit yield incorrectly.
|
||||
|
||||
## Regression tests
|
||||
|
||||
Both live in the test module of `crates/ecstore/src/set_disk/ops/heal.rs`:
|
||||
|
||||
| Test | Invariant |
|
||||
| --- | --- |
|
||||
| `heal_racing_version_delete_never_resurrects_the_deleted_version` | With a doomed version's shards corrupted, a versioned DELETE and a deep heal contend on the same lock; the deleted version is not resurrected and the surviving version is intact. |
|
||||
| `heal_racing_unversioned_overwrites_preserves_the_last_commit` | Unversioned overwrite commits (exercising the commit-tail old `data_dir` deletion) race a deep-heal loop; the final current version is exactly the last commit (ETag-level equality). |
|
||||
|
||||
Related: the atomic-commit and best-effort-rollback invariants for the write path are in [erasure-coding.md](erasure-coding.md).
|
||||
@@ -1,184 +1,143 @@
|
||||
# KMS Bulk Rekey Job Contract
|
||||
|
||||
This document defines the contract for the object-side bulk rekey job: a long-running administrative job that re-wraps stored data-key envelopes under the current key-encryption key (KEK) without rewriting object bodies. A first execution engine has shipped: the sweep in `rustfs/src/kms_rekey.rs`, driven by the admin endpoints in `rustfs/src/admin/handlers/kms_rekey.rs`. The contract remains the acceptance bar; where the shipped v1 sweep deliberately narrows it, the [Implementation Status](#implementation-status-v1-sweep) section records the deviation so the document and the tree cannot drift apart silently.
|
||||
**Use this when:** changing the bulk envelope re-wrap sweep (`rustfs/src/kms_rekey.rs`), its admin endpoints (`rustfs/src/admin/handlers/kms_rekey.rs`), the re-wrap primitive, or anything that decides which objects a rekey may touch.
|
||||
**Source of truth:** `rustfs/src/kms_rekey.rs`, `rustfs/src/admin/handlers/kms_rekey.rs`, `rewrap_object_encryption_metadata` in `rustfs/src/storage/sse.rs`, `KmsManager::rewrap_data_key` / `KmsManager::describe_data_key_wrapping` in `crates/kms/src/manager.rs`, `put_object_metadata` in `crates/ecstore/src/set_disk/ops/object.rs`.
|
||||
|
||||
It tracks [`rustfs/backlog#1642`](https://github.com/rustfs/backlog/issues/1642), which lands the `bulk migrate/rekey` line of [`rustfs/backlog#1562`](https://github.com/rustfs/backlog/issues/1562).
|
||||
The bulk rekey job re-wraps stored data-key envelopes under the current key-encryption key (KEK) without rewriting object bodies. This document is the acceptance bar; where the shipped v1 sweep deliberately narrows it, [Implementation Status](#implementation-status-v1-sweep) records the deviation.
|
||||
|
||||
## Scope
|
||||
|
||||
- Applies to: the job lifecycle, ownership, idempotency, failure semantics, exclusion rules, and completion evidence for bulk envelope re-wrap.
|
||||
- Out of scope, and deliberately so: the cryptographic definition of a single-object re-wrap (owned by the re-wrap primitive), master key material migration between KMS backends, a pause state, multi-node parallel execution, and destruction of superseded key versions.
|
||||
|
||||
### Why master key material migration is not this job
|
||||
|
||||
Vault Transit, AWS KMS, and HSM backends are designed so that key material cannot be exported. There is no path that moves a Local master key into Transit, and the reverse direction would export production key material from an HSM onto local disk, which is a security regression. The one case that is both possible and useful, Local to Local, is already served by the KMS backup and restore bundle in `crates/kms/src/backup/local_export.rs` and `crates/kms/src/backup/local_restore.rs`. Nothing in this contract creates a second, weaker copy of that capability.
|
||||
- Applies to: job lifecycle, ownership, idempotency, failure semantics, exclusion rules, and completion evidence for bulk envelope re-wrap.
|
||||
- Out of scope: the cryptographic definition of a single-object re-wrap (owned by the primitive), master key material migration between backends, a pause state, multi-node parallel execution, destruction of superseded key versions.
|
||||
- Master key material migration is not this job: Vault Transit, AWS KMS, and HSM backends do not export key material, and the one useful case (Local to Local) is already served by `crates/kms/src/backup/local_export.rs` and `crates/kms/src/backup/local_restore.rs`.
|
||||
|
||||
## Implementation Status (v1 Sweep)
|
||||
|
||||
The shipped sweep (`rustfs/src/kms_rekey.rs`, admin surface `POST /rustfs/admin/v3/kms/keys/rekey` plus `/status` and `/cancel`, all gated on the cluster-scoped `kms:Rekey` action) implements the contract with these deliberate narrowings:
|
||||
The shipped sweep (`POST /rustfs/admin/v3/kms/keys/rekey` plus `/status` and `/cancel`, gated on the cluster-scoped `kms:Rekey` action) narrows the contract as follows:
|
||||
|
||||
- **One sweep per process, not scope-scoped admission.** A single in-memory slot serializes sweeps cluster-wide on the node that received the request; a second start request is refused with the running job id. This is narrower than the scope-scoped ownership below — two disjoint-scope jobs cannot run concurrently — which is the safe direction: concurrent sweeps would double every KMS round-trip for zero extra coverage. The persisted CAS job record, lease, and crash-recovered ownership described under [Skeleton, Ownership, And Admission](#skeleton-ownership-and-admission) are not implemented; job state and counters are process-local and reset on restart. Correctness does not depend on them: the envelope itself is the resume state.
|
||||
- **Cursor-free convergence.** No checkpoint exists at all. The contract already declared the cursor a performance optimization; v1 takes that to its limit — recovery from a crash, cancel, or partial failure is re-running the sweep, and every already-current envelope costs one describe-shaped KMS call and no write.
|
||||
- **Backend gate at start.** The start endpoint refuses with `501` when the configured backend does not advertise `BackendCapabilities::rewrap`. Vault KV2 and Vault Transit pass; Local, Static, and AWS are refused. This is the "refused at admission" behavior the contract requires for AWS, and it is also what disarms the Local blocker below: a sweep can only run where superseded key versions demonstrably remain decryptable.
|
||||
- **Collapsed exclusion counting.** Plaintext objects, SSE-C objects, and MinIO-sealed envelopes are counted together as `not_applicable` rather than per-class; delete markers and directory entries are skipped without counting. Per-class exclusion counts remain future work.
|
||||
- **No dry run.** The dry-run report model below is not implemented; the closest present capability is reading `/status` counters from a completed sweep.
|
||||
- **Admission posture.** The sweep processes exactly one object at a time — each iteration awaits a KMS round-trip and, on rewrap, one metadata write — so its foreground contention is bounded by strict serialization, the KMS policy layer's shared concurrency cap, and the storage layer's own namespace locks and quorum rules. It does not integrate with a workload-admission mechanism, because [workload-admission-contracts.md](workload-admission-contracts.md) currently defines an observation-only snapshot surface repo-wide, with no runtime admission API for any background job to join. When such a mechanism exists, this job joins it alongside the scanner, heal, and decommission; until then, the requirement is bounded contention, which serialization provides.
|
||||
| Contract item | v1 behavior |
|
||||
|---|---|
|
||||
| Ownership / admission | One in-memory slot per process serializes sweeps; a second start request is refused with the running job id. No persisted CAS job record, lease, or crash-recovered ownership; counters are process-local and reset on restart. |
|
||||
| Resume cursor | None. Recovery from crash, cancel, or partial failure is re-running the sweep; every already-current envelope costs one describe-shaped KMS call and no write. |
|
||||
| Backend gate | Start refuses with `501` when the backend does not advertise `BackendCapabilities::rewrap` (`crates/kms/src/backends/mod.rs`). Vault KV2 and Vault Transit pass; Local, Static, and AWS are refused. |
|
||||
| Exclusion counting | Plaintext, SSE-C, and MinIO-sealed envelopes are counted together as `not_applicable`; delete markers and directory entries are skipped without counting. |
|
||||
| Dry run | Not implemented; the closest capability is `/status` counters from a completed sweep. |
|
||||
| Admission posture | Exactly one object at a time (one KMS round-trip, then at most one metadata write). No workload-admission integration: [workload-admission-contracts.md](workload-admission-contracts.md) defines an observation-only snapshot surface with no runtime admission API for a background job to join. |
|
||||
|
||||
What v1 keeps exactly as contracted: work units are `(bucket, object, versionId)` with `latest_only: false`; `mod_time` is never set on the rewrap write; object-lock retention is inherited from `put_object_metadata`; the rewrap replaces every stored envelope copy by value match across the RustFS-internal and MinIO-compatible slots, and treats "no replaceable copy found" as an error rather than a silent success — the stale-branch hazard rule from [Metadata Write Contract](#metadata-write-contract); failures are counted and logged per object and never abort the sweep; cancellation is cooperative and terminal.
|
||||
Kept exactly as contracted: work units are `(bucket, object, versionId)` with `latest_only: false`; `mod_time` is never set on the rewrap write; object-lock retention is inherited from `put_object_metadata`; every stored envelope copy is replaced by value match across the RustFS-internal and MinIO-compatible slots, and "no replaceable copy found" is an error, not a silent success; failures are counted and logged per object and never abort the sweep; cancellation is cooperative and terminal.
|
||||
|
||||
## Terms
|
||||
|
||||
| Term | Meaning |
|
||||
|---|---|
|
||||
| Envelope | The sealed data key (DEK) stored on an object version's metadata, together with the identifiers needed to unseal it. |
|
||||
| Re-wrap primitive | A single-object operation that unseals one envelope and re-seals it under the target KEK, changing metadata only. Implemented as `rewrap_object_encryption_metadata` in `rustfs/src/storage/sse.rs`, over `KmsManager::rewrap_data_key`. |
|
||||
| Rekey job | The scan-and-drive layer defined by this document, which applies the re-wrap primitive across a scope. |
|
||||
| Envelope | The sealed data key (DEK) stored on an object version's metadata, with the identifiers needed to unseal it. |
|
||||
| Re-wrap primitive | Single-object operation that unseals one envelope and re-seals it under the target KEK, changing metadata only: `rewrap_object_encryption_metadata` over `KmsManager::rewrap_data_key`. |
|
||||
| Rekey job | The scan-and-drive layer defined here, applying the primitive across a scope. |
|
||||
| Work unit | One `(bucket, object, versionId)` triple. Never `(bucket, object)`: each version carries its own envelope. |
|
||||
| Scope | The bucket and prefix selector that bounds one job, and the unit of admission exclusion. |
|
||||
| Target state | The envelope state the job is driving toward: sealed under the intended key id at the current KEK version. |
|
||||
| Scope | The bucket and prefix selector that bounds one job; the unit of admission exclusion. |
|
||||
| Target state | Envelope sealed under the intended key id at the current KEK version. |
|
||||
|
||||
## What the Job Does And Does Not Do
|
||||
## What The Job Does And Does Not Do
|
||||
|
||||
The job re-wraps envelopes. It never rewrites object bodies. Erasure-coded shards, part layout, ETag, and storage usage must be unchanged across a rekey; only encryption metadata keys may differ. Metadata-only rewrite is supported by the storage layer: `put_object_metadata` is declared on `ObjectStore` in `crates/ecstore/src/store/mod.rs`, dispatched in `crates/ecstore/src/core/sets.rs`, and implemented in `crates/ecstore/src/set_disk/ops/object.rs`, where it takes a namespace write lock, selects the version named by `opts.version_id`, and merges `opts.eval_metadata` into the existing `FileInfo` metadata under read and write quorum.
|
||||
|
||||
**The job never destroys a superseded key version.** This is the hardest constraint in this contract, and every other guarantee rests on it. A job that fails halfway leaves some objects wrapped under the new KEK version and some under the old one. That state is fully serviceable — reads and writes both succeed — precisely and only because the old version can still decrypt. Destroying old versions from inside the job would convert a resumable operational action into irreversible data loss on partial failure. Destruction stays a separate, human-initiated operation gated on usage evidence.
|
||||
|
||||
A job must therefore refuse to start when the target key's retention policy would allow the superseded version to leave the retention window while the job runs.
|
||||
- Re-wraps envelopes only. Erasure-coded shards, part layout, ETag, and storage usage are unchanged; only encryption metadata keys may differ. The metadata-only write is `put_object_metadata` (declared on `ObjectStore` in `crates/ecstore/src/store/mod.rs`, dispatched in `crates/ecstore/src/core/sets.rs`, implemented in `crates/ecstore/src/set_disk/ops/object.rs`).
|
||||
- **Never destroys a superseded key version.** A half-finished job leaves some envelopes under the new KEK version and some under the old; that state is serviceable only because the old version still decrypts. Destruction stays a separate, human-initiated operation gated on usage evidence.
|
||||
- Must refuse to start when the target key's retention policy would let the superseded version leave the retention window while the job runs.
|
||||
|
||||
## Idempotency Model
|
||||
|
||||
Re-running the job must be safe and must converge. The intended source of idempotency is the object metadata itself: the envelope's own state is the target state, so a re-run reads what is already correct and skips it. No separate idempotency table is required, and the job identity is only a `job_id: Uuid` for reporting and ownership, following the ILM manual transition job record in `crates/ecstore/src/bucket/lifecycle/manual_transition_job.rs`.
|
||||
Idempotency comes from object metadata itself: the envelope's state is the target state, so a re-run reads what is already correct and skips it. No idempotency table; the job identity is a `job_id: Uuid` for reporting and ownership, following `ManualTransitionJobRecord` in `crates/ecstore/src/bucket/lifecycle/manual_transition_job.rs`.
|
||||
|
||||
Two consequences follow, and both are contract requirements:
|
||||
|
||||
- **The resume cursor is a performance optimization, not a correctness dependency.** Losing a checkpoint may cause a rescan and a higher skip count, never a wrong result. This is what makes crash recovery cheap: checkpoints may be throttled rather than written per object, following the `PersistThrottle` policy in `crates/heal/src/heal/resume.rs`, which flushes after a bounded number of buffered mutations or a bounded interval, whichever comes first. That module states the same reasoning for heal: because the operation is idempotent, a crash re-does at most one throttle window.
|
||||
- **The job is at-least-once with target-state idempotency, never exactly-once.** No design may introduce exactly-once machinery for work units.
|
||||
- **The resume cursor is a performance optimization, not a correctness dependency.** Losing a checkpoint may cause a rescan and a higher skip count, never a wrong result. Checkpoints may therefore be throttled (`PersistThrottle` in `crates/heal/src/heal/resume.rs`).
|
||||
- **At-least-once with target-state idempotency, never exactly-once.** No design may introduce exactly-once machinery for work units.
|
||||
|
||||
### Reading the wrapping KEK version
|
||||
|
||||
The self-evidencing property above holds only when the wrapping KEK version is observable. It is, for every backend that actually rotates, but not from a dedicated metadata field and not by the same mechanism on each backend.
|
||||
|
||||
There is no key-version metadata key: object metadata carries the key **id** (`x-rustfs-encryption-key-id` in `rustfs/src/storage/sse.rs`, defaulting to `default`) and the sealed blob under `x-rustfs-encryption-key`, and nothing else names a version. `DecryptResponse` in `crates/kms/src/types.rs` does not report one either, though `EncryptResponse` does.
|
||||
|
||||
The version is nonetheless recoverable, because the sealed blob is structured. `x-rustfs-encryption-key` stores the base64 of the backend ciphertext, and for every backend that builds one that ciphertext is the JSON of `DataKeyEnvelope` (`crates/kms/src/encryption/dek.rs`). Reading it needs no new metadata: base64-decode the value, then parse the JSON. The read path in `rustfs/src/storage/sse.rs` already does exactly this discrimination, calling `is_data_key_envelope` on the decoded blob to pick a provider, so this is an established in-tree pattern rather than a new capability.
|
||||
|
||||
Where the version sits inside that structure is backend-specific:
|
||||
There is no key-version metadata key. Object metadata carries the key id (`x-rustfs-encryption-key-id`) and the sealed blob under `x-rustfs-encryption-key`; `DecryptResponse` in `crates/kms/src/types.rs` does not report a version either. The version is recoverable because the sealed blob is structured: for every backend that builds one, the ciphertext is the JSON of `DataKeyEnvelope` (`crates/kms/src/encryption/dek.rs`), and the read path already discriminates on it via `is_data_key_envelope` in `rustfs/src/storage/sse.rs`.
|
||||
|
||||
| Backend | Rotates | Where the wrapping version lives | Recoverable by a scan |
|
||||
|---|---|---|---|
|
||||
| Vault KV2 (`crates/kms/src/backends/vault.rs`) | Yes | `DataKeyEnvelope::master_key_version`, populated from the key record's version | Yes, from the envelope JSON |
|
||||
| Vault Transit (`crates/kms/src/backends/vault_transit.rs`) | Yes | The `vault:vN:` prefix of the ciphertext held in the envelope's `encrypted_key`; the envelope's own version field is deliberately `None` because Transit ciphertext self-describes | Yes, by parsing that prefix |
|
||||
| Local (`crates/kms/src/backends/local.rs`) | No — rotation is rejected | Nowhere; the version field is hardcoded `None` because a key has exactly one material | Moot while rotation is rejected |
|
||||
| Static (`crates/kms/src/backends/static_kms.rs`) | No — single fixed key | Nowhere; hardcoded `None` | Moot |
|
||||
| AWS (`crates/kms/src/backends/aws.rs`) | AWS-managed | Inside the opaque `CiphertextBlob`; no `DataKeyEnvelope` is built at all | **No** |
|
||||
| Vault KV2 (`crates/kms/src/backends/vault.rs`) | Yes | `DataKeyEnvelope::master_key_version` | Yes, from the envelope JSON |
|
||||
| Vault Transit (`crates/kms/src/backends/vault_transit.rs`) | Yes | `vault:vN:` prefix of the ciphertext in `encrypted_key`; the envelope's version field is deliberately `None` | Yes, by parsing that prefix |
|
||||
| Local (`crates/kms/src/backends/local.rs`) | No, rotation is rejected | Nowhere; hardcoded `None` | Moot while rotation is rejected |
|
||||
| Static (`crates/kms/src/backends/static_kms.rs`) | No | Nowhere; hardcoded `None` | Moot |
|
||||
| AWS (`crates/kms/src/backends/aws.rs`) | AWS-managed | Inside the opaque `CiphertextBlob`; no `DataKeyEnvelope` | **No** |
|
||||
|
||||
Two traps follow, and both are contract rules.
|
||||
Contract rules that follow:
|
||||
|
||||
**`None` does not mean one thing.** On Vault KV2 it means a pre-versioning envelope, and `resolve_envelope_master_key_version` resolves it to the key's recorded baseline version, or to the current version for a key that was never rotated — never implicitly to whatever is current now. On Transit it is permanent and expected, and the version must be read from the ciphertext prefix instead. On Local and Static it is unconditional. A scan that reads `None` as a single condition will misclassify three different situations, so version extraction must be dispatched by backend, never inferred from the field alone.
|
||||
|
||||
**Local's `None` is coupled to the blocker below.** The Local backend omits the version specifically because rotation is rejected there. When [`rustfs/backlog#1565`](https://github.com/rustfs/backlog/issues/1565) gives Local a rotation history, that construction must begin recording the wrapping version in the same change, or Local silently becomes a second unreadable backend and loses idempotent skip along with it. This coupling is not obvious from either issue and must not be discovered later.
|
||||
|
||||
The requirement this places on the re-wrap primitive is therefore narrower than "record a version", most of which the tree already satisfies:
|
||||
|
||||
- The primitive must expose the wrapping version through **one backend-dispatched accessor** — satisfied by `KmsManager::describe_data_key_wrapping`, which dispatches per backend so callers never reimplement envelope-field or ciphertext-prefix parsing, which would also put KMS format knowledge on the wrong side of the crate boundary.
|
||||
- The primitive must report **"already at target state" as an outcome distinct from "re-wrapped"**, so the job counts a skip instead of inferring one.
|
||||
- For AWS, neither is achievable by inspection, and the contract must say so rather than pretend otherwise (see below).
|
||||
|
||||
### The cost of recognizing the target state
|
||||
|
||||
Skipping already-current objects is achievable, and it is not free. Every scanned work unit costs a base64 decode plus a JSON parse of its envelope, and on Transit an additional prefix parse. That is CPU and allocation per object version, not extra I/O: the metadata is already being read by the scan, and no KMS round trip is involved. Envelopes are small, so the cost is bounded per object, but at bulk scale it is the dominant cost of a dry run and of the skip check in a re-run, and it belongs in the rate and admission budget rather than being treated as free.
|
||||
|
||||
This cost buys three things, all of which the contract requires and none of which are available without it: a re-run that skips completed work and performs zero metadata writes, a dry run that reports which KEK versions are actually in scope, and the per-object half of completion evidence.
|
||||
|
||||
**AWS is the exception, and it is a scoping exception rather than a cost.** Its ciphertext is opaque to RustFS, so no inspection can tell a current envelope from a stale one. A rekey scope on an AWS-backed key therefore cannot skip, cannot report version composition in a dry run, and cannot self-evidence completion; a re-run would re-wrap every object again. AWS also rotates backing key material transparently on decrypt, so the operational need that motivates this job is weaker there to begin with. Until there is a reason to do otherwise, AWS-backed keys are out of scope for bulk rekey, and a job must refuse such a scope at admission rather than start one whose re-runs silently rewrite everything.
|
||||
- **`None` does not mean one thing.** KV2: pre-versioning envelope, resolved by `resolve_envelope_master_key_version` to the key's recorded baseline, never implicitly to "current". Transit: permanent and expected; read the ciphertext prefix. Local/Static: unconditional. Version extraction must be dispatched by backend, never inferred from the field alone.
|
||||
- **Local's `None` is coupled to the Local blocker.** If Local gains rotation history (`rustfs/backlog#1565`), envelope version recording must land in the same change, or Local becomes a second unreadable backend.
|
||||
- The primitive exposes the wrapping version through **one backend-dispatched accessor** (`KmsManager::describe_data_key_wrapping`) and reports **"already at target state" as an outcome distinct from "re-wrapped"**.
|
||||
- **AWS is a scoping exception.** Its ciphertext is opaque, so no scan can skip, report version composition, or self-evidence completion; a re-run would rewrap everything. AWS-backed keys are out of scope and must be refused at admission.
|
||||
- Skip detection costs a base64 decode plus JSON parse (plus a prefix parse on Transit) per work unit: CPU, not I/O, and part of the rate budget rather than free.
|
||||
|
||||
## Failure Semantics
|
||||
|
||||
A partially complete rekey is a valid, serviceable state, not a damaged one. It requires no emergency handling, no fail-closed startup guard, and no rollback. This is the sharpest difference from KMS backup restore, whose intermediate state genuinely is unserviceable and which therefore fails closed on startup when its commit marker is present.
|
||||
|
||||
The precondition is that superseded key versions remain decryptable. Where that precondition does not hold, the whole model collapses (see Blockers).
|
||||
|
||||
Cancellation is cooperative and terminal. A canceled job reaches a terminal state with already-processed objects left in the target state; restarting on the same scope skips them.
|
||||
|
||||
## Pause Is Not Provided
|
||||
|
||||
The originating requirement asked for pause, resume, and idempotent retry. This contract provides cancel, cursor restart, and rate control instead, and does not provide a pause state.
|
||||
|
||||
Seven long-running job frameworks exist in the tree — ILM manual transition, heal resume (`crates/heal/src/heal/resume.rs`), tier mutation intent (`crates/ecstore/src/services/tier/tier_mutation_intent.rs`), decommission and rebalance (`crates/ecstore/src/core/pools.rs`), the scanner (`crates/scanner/src/scanner.rs`), and KMS backup restore. None of them has a pause state; each has cancel or stop only. That consistency is a design position, not an oversight. A paused job has to answer what it still holds: whether its lease is renewed, whether it keeps its scope admission slot, and how long it may stay paused before it is abandoned. Each answer adds state and a failure mode.
|
||||
|
||||
The two things pause is actually asked for are that the job must not overwhelm the data path, and that stopping it must not throw away progress. Rate and admission control delivers the first; cancel plus cursor restart delivers the second. Both are existing patterns.
|
||||
- A partially complete rekey is a valid, serviceable state: no emergency handling, no fail-closed startup guard, no rollback. This is the sharpest difference from KMS backup restore, whose intermediate state is unserviceable and fails closed on startup.
|
||||
- Precondition: superseded key versions remain decryptable (see Blockers).
|
||||
- Cancellation is cooperative and terminal; restarting on the same scope skips already-processed objects.
|
||||
- Pause is not provided. None of the tree's long-running job frameworks (ILM manual transition, heal resume, tier mutation intent, decommission/rebalance, scanner, KMS restore) has a pause state; rate control plus cancel-and-restart deliver what pause is asked for without lease/slot/abandonment state.
|
||||
|
||||
## Objects That Cannot Be Rekeyed
|
||||
|
||||
These must be enumerated during the scan and excluded with a counted reason. Encountering one is never a job failure, and the execution phase must not touch them.
|
||||
Enumerated during the scan and excluded with a counted reason; never a job failure; the execution phase must not touch them.
|
||||
|
||||
| Class | Disposition | Reason |
|
||||
|---|---|---|
|
||||
| SSE-C objects | Exclude and count | The server never holds the customer key, so it can neither unseal nor re-seal the envelope. |
|
||||
| Objects transitioned to a remote tier | Exclude and count | The body lives remotely; the relationship between local metadata and the remote object's encryption needs its own analysis first. See [tier-ilm-debugging.md](../operations/tier-ilm-debugging.md). |
|
||||
| In-progress multipart uploads | Exclude and count | Each part carries its own envelope and an incomplete upload is not a stable work unit. `crates/kms/src/key_impact.rs` already models this as a distinct reference scope. |
|
||||
| Unencrypted objects | Exclude and count | No envelope to re-wrap. |
|
||||
| Objects under object-lock retention | Governed by the storage layer, see below | |
|
||||
| SSE-C objects | Exclude and count | The server never holds the customer key. |
|
||||
| Objects transitioned to a remote tier | Exclude and count | Body lives remotely; see [tier-ilm-debugging.md](../operations/tier-ilm-debugging.md). |
|
||||
| In-progress multipart uploads | Exclude and count | Each part carries its own envelope; `crates/kms/src/key_impact.rs` models this as a distinct reference scope. |
|
||||
| Unencrypted objects | Exclude and count | No envelope. |
|
||||
| Objects under object-lock retention | Governed by the storage layer (see Metadata Write Contract) | |
|
||||
|
||||
Replication destinations are unresolved: whether an envelope metadata rewrite must propagate to a replica depends on [`rustfs/backlog#1619`](https://github.com/rustfs/backlog/issues/1619). Until that closes, this contract does not authorize propagation.
|
||||
Replication destinations are unresolved: propagation depends on `rustfs/backlog#1619`. Until it closes, a rewrap never propagates to a replica and each site runs its own sweep.
|
||||
|
||||
## Metadata Write Contract
|
||||
|
||||
Three properties of `put_object_metadata` constrain the re-wrap write, all confirmed in `crates/ecstore/src/set_disk/ops/object.rs`.
|
||||
Three properties of `put_object_metadata` (`crates/ecstore/src/set_disk/ops/object.rs`) constrain the re-wrap write:
|
||||
|
||||
**The merge is additive; it cannot remove keys.** `opts.eval_metadata` entries are inserted into the existing metadata map. There is no removal path. Overwriting a key that keeps its name is therefore safe, but a re-wrap that changes *which* metadata keys describe the envelope leaves the old keys behind permanently.
|
||||
|
||||
That is a structural hazard, not a theoretical one. `rustfs/src/storage/sse.rs` selects its decrypt branch on the mere presence of the MinIO-compatible seal-algorithm header: `parse_minio_managed_sealed_key` returns a sealed key whenever that header is present with the expected value, and the caller then takes the MinIO branch in preference to the RustFS-native one. A re-wrap that writes a RustFS-native envelope onto an object carrying MinIO-compatible headers, without clearing them, steers subsequent reads down the stale branch. Any re-wrap that changes envelope shape must neutralize the superseded keys in the same write, and cannot rely on deletion to do it.
|
||||
|
||||
**Object-lock retention is enforced before the merge.** `check_object_lock_retention_update`, defined in `crates/ecstore/src/set_disk/mod.rs`, runs before `eval_metadata` is applied. Rekey inherits that decision rather than restating it: whatever that check permits for a metadata update, rekey permits; whatever it refuses, rekey counts as an exclusion. Rekey must not acquire a bypass.
|
||||
|
||||
**`mod_time` is preserved unless the caller sets it.** The implementation assigns `fi.mod_time` only when `opts.mod_time` is `Some`. The re-wrap path must leave it unset, so that a rekey does not perturb lifecycle rule evaluation — an age-based expiry or transition rule reading a refreshed `mod_time` across a whole bucket would be a cross-feature regression.
|
||||
- **The merge is additive; it cannot remove keys.** Overwriting a key that keeps its name is safe; a re-wrap that changes *which* keys describe the envelope leaves the old keys behind. This is a live hazard: `parse_minio_managed_sealed_key` in `rustfs/src/storage/sse.rs` selects the MinIO decrypt branch on the mere presence of the MinIO seal-algorithm header, so a RustFS-native envelope written onto MinIO-compatible headers without neutralizing them steers reads down the stale branch. Any envelope-shape change must neutralize superseded keys in the same write.
|
||||
- **Object-lock retention is enforced before the merge.** `check_object_lock_retention_update` (`crates/ecstore/src/set_disk/mod.rs`) runs first; rekey inherits its decision and must not acquire a bypass.
|
||||
- **`mod_time` is preserved unless the caller sets it.** The re-wrap path leaves it unset so age-based lifecycle rules are not perturbed.
|
||||
|
||||
## Skeleton, Ownership, And Admission
|
||||
|
||||
The ILM manual transition job is the structural template. `ManualTransitionJobRecord` in `crates/ecstore/src/bucket/lifecycle/manual_transition_job.rs` already carries `job_id`, `scope_key`, `owner_id`, `lease_id` with an expiry, a state machine including an explicit `Unknown` state for a corrupt journal, `cancel_requested`, a report, and a queue snapshot. Records are persisted under dedicated metadata-bucket prefixes with a schema string and checksum, and mutated with S3 conditional writes (`if_match` for updates, `if_none_match` for creates) so that ownership transitions are compare-and-swap rather than last-write-wins. Crash recovery, cooperative cancel via `request_manual_transition_job_cancel`, and capability advertisement through `ManualTransitionJobCapabilities` in `rustfs/src/admin/handlers/system.rs` all follow from that shape.
|
||||
|
||||
Ownership is scope-scoped, not cluster-scoped. Two jobs on disjoint scopes may run concurrently; two jobs on the same scope must be refused by admission. The scanner's leader lock with epoch fencing in `crates/scanner/src/scanner.rs` is the wrong granularity here because it enforces exactly one worker per cluster; it stays a reference for fencing technique only.
|
||||
|
||||
The first implementation is single-node: one owner plus a lease plus recovery is sufficient for correctness. Multi-node parallel execution is a throughput optimization and is out of scope until correctness and its acceptance evidence are both in place.
|
||||
|
||||
Because the job runs online, it must not contend its way into the foreground data path. The v1 posture — strict serialization plus the KMS policy layer's shared cap and the storage layer's own locks — and the reason no workload-admission mechanism is joined yet are recorded under [Implementation Status](#implementation-status-v1-sweep); when a runtime admission mechanism exists per [workload-admission-contracts.md](workload-admission-contracts.md), this job joins it alongside the scanner, heal, and decommission.
|
||||
|
||||
## What Is Taken From KMS Backup, And What Is Not
|
||||
|
||||
Four things transfer:
|
||||
|
||||
- The durable file commit protocol in `crates/kms/src/backends/local.rs` — write, fsync the file, publish by rename or hard link, fsync the parent directory — together with its injectable `CommitStep` failpoints.
|
||||
- The write-receipt ownership proof in `crates/kms/src/backup/vault_restore.rs`. Its distinction is the reusable idea: the list of intended targets proves nothing about ownership, and only a receipt recording the version a write actually landed at may authorize touching that record later; everything else is reported as never-written or not-at-written-version. Bulk rekey faces the identical problem when a concurrent writer modifies an object between the job's read and its write-back. Such an object must be counted as a conflict and skipped, never overwritten.
|
||||
- The sequence guard `VaultRestoreSequence` in the same module: a small, domain-free state machine that makes phase order structural. Rekey's phases are scan, plan, apply, verify.
|
||||
- The three-part dry-run report model in `crates/kms/src/backup/dry_run.rs` — blockers, conflicts, and external mismatches, with a permission predicate that requires all three to be empty — and its zero-write contract: the report is pure data with no handles and no drop-time side effects.
|
||||
|
||||
The lifecycle model does not transfer, and must not be adapted. Backup restore is synchronous, one-shot, single-node, requires an empty target, has no progress surface, and requires the KMS service to be out of `Running` state; its admin layer says as much in `rustfs/src/admin/handlers/kms_backup.rs`. Its commit marker enumerates every file up front, which does not scale to object counts. Its publish primitive is no-clobber, whereas rekey rewrites existing state by definition. Forcing rekey into that four-phase protocol produces an all-or-nothing transaction over the whole scope, which is not operable at this scale.
|
||||
|
||||
The job also does not belong in the KMS crate. `crates/kms/Cargo.toml` does not depend on `rustfs-ecstore` and must not: the job body is object scanning and metadata rewriting, which is ecstore and admin territory. The KMS crate supplies the re-wrap primitive only.
|
||||
- Structural template: `ManualTransitionJobRecord` (`job_id`, `scope_key`, `owner_id`, `lease_id` with expiry, state machine with explicit `Unknown`, `cancel_requested`, report, queue snapshot), persisted with S3 conditional writes so ownership transitions are compare-and-swap; capability advertisement via `ManualTransitionJobCapabilities` in `rustfs/src/admin/handlers/system.rs`.
|
||||
- Ownership is scope-scoped: disjoint scopes may run concurrently; same-scope jobs are refused by admission. The scanner leader lock in `crates/scanner/src/scanner.rs` is the wrong granularity (one worker per cluster) and is a fencing reference only.
|
||||
- First implementation is single-node; multi-node parallelism is a throughput optimization deferred until correctness evidence exists.
|
||||
- Taken from KMS backup: the durable file commit protocol in `crates/kms/src/backends/local.rs` (`CommitStep` failpoints); the write-receipt ownership proof and `VaultRestoreSequence` phase guard in `crates/kms/src/backup/vault_restore.rs` (a concurrent writer between read and write-back is a conflict-and-skip, never an overwrite); the three-part zero-write dry-run report model in `crates/kms/src/backup/dry_run.rs`. Not taken: the synchronous, empty-target, all-or-nothing restore lifecycle. The job does not belong in `crates/kms` (which must not depend on `rustfs-ecstore`); KMS supplies the primitive only.
|
||||
|
||||
## API Surface
|
||||
|
||||
This section originally required reusing the MinIO-compatible batch-job endpoints in `rustfs/src/admin/handlers/batch_job.rs` and forbade a second REST surface. The shipped v1 superseded that rule: the sweep landed on RustFS-specific endpoints (`/v3/kms/keys/rekey`, `/status`, `/cancel`), reviewed and merged with the engine. The batch-job surface parses MinIO's full job-definition format, whose semantics (per-job flags, retries, notifications) the v1 sweep does not implement — and accepting a job definition whose semantics cannot be executed is exactly what this section forbids.
|
||||
|
||||
The rule that survives is about live semantics, not endpoint shape: **one operation must never have two live semantics.** Today there is one live surface (the RustFS endpoints) and one refusing stub — `KNOWN_JOB_TYPES` in `batch_job.rs` still lists `keyrotate`, and `start-job` still returns a deliberate `NotImplemented`, unknown types get `InvalidRequest`, `list-jobs` returns an empty list, and status, describe, and cancel return a no-such-job error. That `NotImplemented` remains an external promise: the batch-job `keyrotate` type must keep refusing until it either proxies to this same engine with full batch-job semantics or is removed. It must never report success while it executes nothing, and it must never grow a second, divergent rekey implementation.
|
||||
- Live surface: the RustFS endpoints above. The MinIO-compatible batch-job surface (`rustfs/src/admin/handlers/batch_job.rs`) still lists `keyrotate` in `KNOWN_JOB_TYPES` and returns a deliberate `NotImplemented` from `start-job`.
|
||||
- Rule: **one operation must never have two live semantics.** The batch-job `keyrotate` type must keep refusing until it proxies to this engine with full batch-job semantics or is removed; it must never report success while executing nothing.
|
||||
|
||||
## Completion Evidence
|
||||
|
||||
A job that reports success has not proven anything until no object in the scope still references the superseded key version. That evidence surface is the key usage inventory, whose typed foundation already exists in `crates/kms/src/key_impact.rs`. That module is deliberately built so a report can never claim a key is unused: it has no `in_use`, no `unreferenced`, and no `safe_to_delete` field, and instead reports which sources were consulted and how completely they could be read. It lists object envelopes and in-progress multipart uploads among its reference scopes and currently marks both as not scanned.
|
||||
|
||||
Rekey must inherit that discipline. An empty result means nothing was found in the sources that were scanned, never that nothing references the key. A report that cannot state its own coverage is not completion evidence, and must not be used to authorize destroying anything.
|
||||
Completion is proven only when no object in scope still references the superseded key version. The evidence surface is the key usage inventory in `crates/kms/src/key_impact.rs`, which deliberately has no `in_use` / `unreferenced` / `safe_to_delete` field and instead reports which sources were consulted and how completely. Rekey inherits that discipline: an empty result means nothing was found in the sources scanned, never that nothing references the key.
|
||||
|
||||
## Blockers
|
||||
|
||||
**Resolved by capability gating — Local rotation history.** [`rustfs/backlog#1565`](https://github.com/rustfs/backlog/issues/1565) (no rotation history in the Local backend) was a hard blocker while a sweep could run against Local: without retained superseded versions, a rekey interrupted halfway would leave every unprocessed object permanently unreadable after rotation, falsifying the partial-completion guarantee this contract is built on. The shipped resolution is not rotation history but scope: the Local backend is positioned as non-production, rotation stays rejected there, and the sweep's start endpoint refuses any backend that does not advertise `BackendCapabilities::rewrap` — so a sweep can only run where the retained-versions invariant holds by construction (Vault KV2 and Vault Transit). If Local ever gains rotation, the coupling recorded under [Reading the wrapping KEK version](#reading-the-wrapping-kek-version) still applies: rotation history and envelope version recording must land in the same change before Local may advertise `rewrap`.
|
||||
|
||||
**Resolved — the execution chain is complete.** The envelope-level primitive (`KmsManager::rewrap_data_key`, `KmsManager::describe_data_key_wrapping` in `crates/kms/src/manager.rs`), the object-level adapter (`rewrap_object_encryption_metadata` in `rustfs/src/storage/sse.rs`, which reads a version's envelope, reconstructs its encryption context, re-wraps, and returns the metadata overrides), and the sweep that drives the adapter and persists through `put_object_metadata` (`rustfs/src/kms_rekey.rs`) all exist.
|
||||
|
||||
**Affects acceptance, not start — still open.** Key usage inventory coverage over object envelopes: `crates/kms/src/key_impact.rs` still reports `ObjectEnvelopes` and `InProgressMultipartUploads` as not scanned, so a completed sweep's counters are evidence from that run only, not inventory-grade completion proof. KMS key list pagination, which a job enumerating keys would hit. And [`rustfs/backlog#1619`](https://github.com/rustfs/backlog/issues/1619), which decides replica propagation — until it closes, a rewrap never propagates to a replica site and each site runs its own sweep.
|
||||
| Item | Status |
|
||||
|---|---|
|
||||
| Local rotation history (`rustfs/backlog#1565`) | Resolved by capability gating: Local stays non-production, rotation stays rejected, and the start endpoint refuses any backend without `BackendCapabilities::rewrap`. If Local ever gains rotation, envelope version recording must land in the same change. |
|
||||
| Execution chain | Resolved: primitive (`rewrap_data_key`, `describe_data_key_wrapping`), object adapter (`rewrap_object_encryption_metadata`), sweep (`rustfs/src/kms_rekey.rs`). |
|
||||
| Key usage inventory coverage | Open: `key_impact.rs` still reports `ObjectEnvelopes` and `InProgressMultipartUploads` as not scanned, so sweep counters are evidence from that run only. |
|
||||
| KMS key list pagination | Open; a job enumerating keys would hit it. |
|
||||
| Replica propagation (`rustfs/backlog#1619`) | Open; no propagation until it closes. |
|
||||
|
||||
## Verification Expectations
|
||||
|
||||
This list is the acceptance bar for the full contract, not a claim about what the v1 sweep has already demonstrated: the dry-run and checkpoint items await the features themselves (a cursor-free sweep satisfies the checkpoint-deletion clause vacuously), and per-class exclusion counting is narrowed as recorded under [Implementation Status](#implementation-status-v1-sweep).
|
||||
Acceptance bar for the full contract (dry-run and checkpoint items await those features; a cursor-free sweep satisfies the checkpoint clause vacuously):
|
||||
|
||||
Implementation work under this contract must be able to demonstrate, at minimum: that dry run performs zero storage writes; that non-rekeyable objects are excluded and counted rather than failing the job; that an immediate second run skips every object and writes no metadata, on both a KV2-backed and a Transit-backed scope, since the two recover the wrapping version by different mechanisms; that a scope on an AWS-backed key is refused at admission rather than accepted as a job whose re-runs rewrite everything; that an envelope with no recorded version is classified by backend rather than by the bare `None`; that deleting the checkpoint changes only the skip count, not the outcome; that a killed and recovered job reaches a terminal state while every object remains readable throughout; that a concurrent writer causes a conflict-and-skip rather than an overwrite; that ETag, part layout, and storage usage are unchanged at the `xl.meta` level; that each version of a multi-version object is processed independently with its `versionId` intact; that superseded key versions still exist and still decrypt afterward; and that success, skip, exclusion, conflict, and failure counts sum to the number of work units scanned.
|
||||
1. Dry run performs zero storage writes.
|
||||
2. Non-rekeyable objects are excluded and counted rather than failing the job.
|
||||
3. An immediate second run skips every object and writes no metadata, on both a KV2-backed and a Transit-backed scope.
|
||||
4. A scope on an AWS-backed key is refused at admission.
|
||||
5. An envelope with no recorded version is classified by backend, not by the bare `None`.
|
||||
6. Deleting the checkpoint changes only the skip count, not the outcome.
|
||||
7. A killed and recovered job reaches a terminal state while every object stays readable throughout.
|
||||
8. A concurrent writer causes conflict-and-skip, not an overwrite.
|
||||
9. ETag, part layout, and storage usage are unchanged at the `xl.meta` level.
|
||||
10. Each version of a multi-version object is processed independently with its `versionId` intact.
|
||||
11. Superseded key versions still exist and still decrypt afterward.
|
||||
12. Success, skip, exclusion, conflict, and failure counts sum to the number of work units scanned.
|
||||
|
||||
@@ -1,416 +1,148 @@
|
||||
# MinIO File-Format Interoperability — Gap Analysis & Phased Plan
|
||||
# MinIO On-Disk Format Interoperability
|
||||
|
||||
Assesses how closely the RustFS on-disk format matches MinIO's, so that a
|
||||
MinIO drive set can be read (and eventually served) by RustFS and vice versa.
|
||||
This is a **plan and analysis document**. It changes no storage code. Every
|
||||
claim below cites the code that backs it.
|
||||
**Use this when:** deciding whether a MinIO drive set, bucket-metadata blob, or SSE object can be read or imported by a given RustFS build, or before touching any constant or codec listed under Version Anchors.
|
||||
**Source of truth:** `crates/filemeta/src/filemeta.rs`, `crates/filemeta/src/filemeta/codec.rs`, `crates/ecstore/src/bucket/metadata.rs`, `crates/ecstore/src/bucket/migration.rs`, `rustfs/src/storage/sse.rs`, `rustfs/Cargo.toml` `[features]`, `.github/workflows/ci.yml`, `.github/workflows/minio-interop.yml`.
|
||||
|
||||
Scope: the two on-disk artifacts that matter for interop are the per-object
|
||||
`xl.meta` (object metadata + inline data) and the per-bucket `.metadata.bin`
|
||||
(bucket configuration blob). IAM/config layout is noted where it affects
|
||||
bucket-metadata migration.
|
||||
This is an interop contract, not a plan. Migration is one-way (MinIO to RustFS). Erasure-coding internals are owned by [erasure-coding.md](erasure-coding.md); this document owns the interop claim, the fixture evidence, and the out-of-scope list.
|
||||
|
||||
Refs rustfs/backlog#580.
|
||||
## Scope Matrix By Build Variant
|
||||
|
||||
## Executive Summary
|
||||
Build variants are the `rustfs` crate features in `rustfs/Cargo.toml`: `default`, `full`, and `rio-v2` (which enables `rustfs-ecstore/rio-v2` and pulls in `crates/rio-v2`). `rio-v2` is absent from both `default` and `full`.
|
||||
|
||||
- **`xl.meta`**: RustFS writes `XL_META_VERSION = 3` and reads meta_ver ≤ 3,
|
||||
including legacy meta_ver 2 objects with legacy checksums. Magic `XL2 `,
|
||||
erasure algorithm `rs-vandermonde` (Reed-Solomon), and HighwayHash256 bitrot
|
||||
all match MinIO. `xl.meta` interop is the **strong** part of the story.
|
||||
- **`.metadata.bin`**: RustFS uses the same filename, the same 4-byte
|
||||
`format|version` header, the same MessagePack blob layout, and the same
|
||||
per-config field encodings (XML/JSON) as MinIO's `bucketMetadata`. The
|
||||
divergence is a small set of RustFS-only fields (table-bucket support,
|
||||
bucket-targets meta) — not a format mismatch.
|
||||
- **Migration**: RustFS already ships a one-way importer that reads a legacy
|
||||
meta bucket and rewrites bucket-metadata + IAM config into the RustFS meta
|
||||
bucket (`crates/ecstore/src/bucket/migration.rs`).
|
||||
- **Server-side encryption**: not covered by the above. Objects MinIO wrote with SSE-S3, SSE-KMS, or SSE-C are **not readable by RustFS** in any shipped build. See [Part C](#part-c--server-side-encryption-sse) before planning a migration that includes encrypted objects.
|
||||
| MinIO artifact | `default` / `full` build | `rio-v2` build | Notes |
|
||||
|---|:--:|:--:|---|
|
||||
| Unencrypted `xl.meta` (meta_ver 1-3, inline, multipart, versioned, delete marker) | Read | Read | Part A. Normalized to meta_ver 3 on rewrite. |
|
||||
| Transitioned (tiered) `xl.meta` | Not fixture-proven | Not fixture-proven | Out of scope; see erasure-coding.md for the tolerant `transitioned-versionID` read rule. |
|
||||
| `.metadata.bin` bucket config | Read and imported | Read and imported | Part B. Importer reads a `.minio.sys` layout end to end. |
|
||||
| IAM config under `config/iam/` | Imported | Imported | `try_migrate_iam_config`; legacy field aliases normalized. |
|
||||
| SSE-S3 / SSE-KMS objects, MinIO builtin static KMS | Fail closed, diagnosed | Read | Part C. Requires the shared master key. |
|
||||
| SSE-C objects | Fail closed, diagnosed | Read | Part C. Customer key supplied per request. |
|
||||
| Any SSE object, MinIO backed by KES / KMS plugin / MinKMS | Fail closed | Fail closed | Not planned; the DEK is sealed by the KES service. |
|
||||
| RustFS-written drive set read by a live MinIO binary | Unsupported | Unsupported | Set-level divergence: MinIO looks for `.minio.sys`, RustFS writes `.rustfs.sys`. |
|
||||
| RustFS-written SSE objects read by MinIO | Unsupported | Unsupported | Part C, reverse direction. |
|
||||
|
||||
For unencrypted objects the remaining work is verification breadth and closing
|
||||
per-config parsing gaps, not a format rewrite. Encrypted objects are a separate,
|
||||
unsolved axis (rustfs/backlog#1638).
|
||||
## Version Anchors
|
||||
|
||||
---
|
||||
These constants are compatibility anchors. Bumping any of them requires a read-compat path for the prior value and a migration story, exactly as the meta_ver 2 to 3 read path provides. Values live in code; do not copy them elsewhere.
|
||||
|
||||
| Anchor | Symbol | File | Rule |
|
||||
|---|---|---|---|
|
||||
| `xl.meta` magic | `XL_FILE_HEADER` | `crates/filemeta/src/filemeta.rs` | Must equal MinIO's XL2 magic. |
|
||||
| Container major / minor | `XL_FILE_VERSION_MAJOR`, `XL_FILE_VERSION_MINOR` | `crates/filemeta/src/filemeta.rs` | `check_xl2_v1` (`crates/filemeta/src/filemeta/codec.rs`) rejects `major > XL_FILE_VERSION_MAJOR`. |
|
||||
| Header version | `XL_HEADER_VERSION` | `crates/filemeta/src/filemeta.rs` | `decode_xl_headers` rejects `header_ver > XL_HEADER_VERSION`. |
|
||||
| Metadata version | `XL_META_VERSION` | `crates/filemeta/src/filemeta.rs` | Written by `FileMeta::new`; `decode_xl_headers` rejects `meta_ver > XL_META_VERSION` (accept-older, reject-newer). |
|
||||
| Bucket metadata header | `BUCKET_METADATA_FORMAT`, `BUCKET_METADATA_VERSION` | `crates/ecstore/src/bucket/metadata.rs` | Checked by `check_header`; both match MinIO's `bucketMetadataFormat` / `bucketMetadataVersion`. |
|
||||
| Erasure algorithm string | `ERASURE_ALGORITHM` | `crates/ecstore/src/object_api/mod.rs` | `rs-vandermonde`; enum `ErasureAlgo` in `crates/filemeta/src/fileinfo.rs`. |
|
||||
| Meta bucket names | `RUSTFS_META_BUCKET`, `MIGRATING_META_BUCKET`, `BUCKET_META_PREFIX` | `crates/ecstore/src/disk/mod.rs` | `.rustfs.sys` is the live meta bucket; `.minio.sys` is the importer source. |
|
||||
|
||||
## Part A — `xl.meta` Object Format
|
||||
|
||||
### Version support
|
||||
|
||||
| Aspect | Value | Evidence |
|
||||
| Aspect | Contract | Where |
|
||||
|---|---|---|
|
||||
| Write version (`meta_ver`) | 3 | `crates/filemeta/src/filemeta.rs:54` (`XL_META_VERSION = 3`), written in `FileMeta::new` at `crates/filemeta/src/filemeta.rs:121` |
|
||||
| Read versions accepted | ≤ 3 (1, 2, 3) | Decode rejects only `meta_ver > XL_META_VERSION` — see `crates/filemeta/src/filemeta/codec.rs` (`decode_xl_headers`); `load_or_convert` doc at `crates/filemeta/src/filemeta.rs:864` |
|
||||
| Legacy meta_ver 2 read | Supported (with legacy checksum) | Regression fixtures `test_issue_2265_legacy_meta_v2_object_compatibility` / `test_issue_2288_legacy_xlmeta_compatibility` at `crates/filemeta/src/filemeta.rs:1130`, `:1152`; `uses_legacy_checksum` asserted at `:1174` |
|
||||
| Version probe | `read_format_versions` returns `(major, minor, header_ver, meta_ver)` without a full parse | `crates/filemeta/src/filemeta/codec.rs` |
|
||||
| Read compatibility | Accepts meta_ver 1-3 including legacy meta_ver 2 with legacy checksums (`uses_legacy_checksum`); `load_or_convert` normalizes on rewrite | `crates/filemeta/src/filemeta.rs`, `crates/ecstore/src/set_disk/read.rs` |
|
||||
| Container layout | 8-byte header, bin-length-prefixed msgpack header block, CRC trailer, optional inline data (MinIO XL2 v1 shape) | `crates/filemeta/src/filemeta/codec.rs` |
|
||||
| Erasure coding | Reed-Solomon Vandermonde, `rs-vandermonde` identifier, codec crate `rustfs-erasure-codec` | `Cargo.toml`, [erasure-coding.md](erasure-coding.md) |
|
||||
| Bitrot | `HighwayHash256S` default; `HighwayHash256SLegacy` (fixed key) for older shards | `crates/ecstore/src/io_support/bitrot.rs`, `crates/ecstore/tests/legacy_bitrot_read_test.rs` |
|
||||
| Inline data | Inline block after the CRC trailer; `null` / version-id keying via `data_key_for_version`; `physical_data_dir` accounting | `crates/filemeta/src/filemeta.rs`, `crates/filemeta/src/filemeta/inline_data.rs` |
|
||||
|
||||
RustFS is a **read-forward-compatible** consumer of MinIO's `xl.meta`: it can
|
||||
parse older MinIO objects and normalizes them to meta_ver 3 on rewrite. It does
|
||||
not write MinIO's older versions.
|
||||
|
||||
### Container header
|
||||
|
||||
| Field | RustFS value | Evidence |
|
||||
|---|---|---|
|
||||
| Magic | `XL2 ` (`[b'X', b'L', b'2', b' ']`) | `crates/filemeta/src/filemeta.rs:46` |
|
||||
| File version major / minor | 1 / 3 | `crates/filemeta/src/filemeta.rs:51-52` |
|
||||
| Header version | 3 | `crates/filemeta/src/filemeta.rs:53` |
|
||||
| Magic + version check (decode entry) | `check_xl2_v1` validates magic and rejects `major > 1` | `crates/filemeta/src/filemeta/codec.rs:45-61` |
|
||||
| Version-only probe (no full parse) | `read_format_versions` returns `(major, minor, header_ver, meta_ver)` | `crates/filemeta/src/filemeta/codec.rs:30-43` |
|
||||
|
||||
The layout after the 8-byte header is `bin-length-prefixed msgpack header block`
|
||||
followed by a CRC trailer and optional inline data — matching MinIO's XL2 v1
|
||||
container.
|
||||
|
||||
### Erasure coding
|
||||
|
||||
| Aspect | Value | Evidence |
|
||||
|---|---|---|
|
||||
| Algorithm enum | `ErasureAlgo::ReedSolomon = 1` | `crates/filemeta/src/fileinfo.rs:83-106` |
|
||||
| Algorithm string | `rs-vandermonde` | `crates/filemeta/src/fileinfo.rs:31` (`ERASURE_ALGORITHM`); also `crates/ecstore/src/object_api/mod.rs:52` |
|
||||
| Codec crate | `rustfs-erasure-codec` (Reed-Solomon, SIMD) | `Cargo.toml:277` |
|
||||
|
||||
Same Reed-Solomon Vandermonde scheme and identifier string as MinIO.
|
||||
|
||||
### Bitrot / shard integrity
|
||||
|
||||
| Aspect | Value | Evidence |
|
||||
|---|---|---|
|
||||
| Default hash | `HashAlgorithm::HighwayHash256S` | Bitrot read/write paths in `crates/ecstore/src/io_support/bitrot.rs` (e.g. `:564`, `:767`) |
|
||||
| Legacy variant | `HighwayHash256SLegacy` (fixed key) for old objects | referenced from `rustfs_utils::HashAlgorithm` (imported at `crates/ecstore/src/io_support/bitrot.rs:26`) |
|
||||
| HighwayHash crate | `highway` 1.3.0 | `Cargo.toml:252` |
|
||||
| Legacy bitrot read coverage | dedicated test | `crates/ecstore/tests/legacy_bitrot_read_test.rs` |
|
||||
|
||||
MinIO uses HighwayHash256 for bitrot; RustFS's default `HighwayHash256S` is
|
||||
compatible, with a legacy-key variant retained for older shards.
|
||||
|
||||
### Inline data
|
||||
|
||||
Small objects are inlined into the `xl.meta` container after the CRC trailer
|
||||
rather than written as a separate `part.1`. Handling lives in
|
||||
`crates/filemeta/src/filemeta/inline_data.rs` (e.g. `physical_data_dir` and the
|
||||
shared-data-dir accounting), and the inline block is appended/consumed by the
|
||||
codec in `crates/filemeta/src/filemeta/codec.rs`. This mirrors MinIO's inline
|
||||
data feature and the `null`/version-id keying used for the inline map
|
||||
(`data_key_for_version` at `crates/filemeta/src/filemeta.rs:69`, legacy key at
|
||||
`:77`).
|
||||
|
||||
### `xl.meta` interop verdict
|
||||
|
||||
| Item | Done | Partial | Todo |
|
||||
|---|:--:|:--:|:--:|
|
||||
| Read MinIO meta_ver ≤ 3 | ✅ | | |
|
||||
| Legacy meta_ver 2 + legacy checksum read | ✅ | | |
|
||||
| XL2 container magic/version parity | ✅ | | |
|
||||
| Reed-Solomon `rs-vandermonde` parity | ✅ | | |
|
||||
| HighwayHash256 bitrot parity | ✅ | | |
|
||||
| Inline data parity | ✅ | | |
|
||||
| Broad fixture corpus from real MinIO writers | | ⚠️ | |
|
||||
| Write-back parity for round-trip (RustFS→MinIO read) | | ⚠️ | |
|
||||
|
||||
The two ⚠️ items are verification breadth, not known incompatibilities: the
|
||||
current fixtures are targeted regressions (issues #2265, #2288), and there is no
|
||||
CI job proving a MinIO binary can re-read a RustFS-written `xl.meta`.
|
||||
|
||||
---
|
||||
MinIO stores an inlined object body as `[HighwayHash256 (32 B)][body]`. Feeding the raw inline shard through RustFS's `BitrotReader` with `HighwayHash256S` verifies the checksum and yields the exact payload; the bitrot prefix is not a format incompatibility.
|
||||
|
||||
## Part B — Bucket Metadata (`.metadata.bin`)
|
||||
|
||||
### On-disk layout
|
||||
|
||||
| Aspect | RustFS value | Evidence |
|
||||
| Aspect | Contract | Where |
|
||||
|---|---|---|
|
||||
| Meta bucket | `.rustfs.sys` | `crates/ecstore/src/disk/mod.rs:29` (`RUSTFS_META_BUCKET`) |
|
||||
| Bucket-config prefix | `buckets` | `crates/ecstore/src/disk/mod.rs:34` (`BUCKET_META_PREFIX`) |
|
||||
| Blob file | `.metadata.bin` | `crates/ecstore/src/bucket/metadata.rs:227` (`BUCKET_METADATA_FILE`) |
|
||||
| Full path | `buckets/{bucket}/.metadata.bin` | `crates/ecstore/src/bucket/metadata.rs:415-416` (`save_file_path`) |
|
||||
| Header | `format: u16 LE` + `version: u16 LE`, both `= 1` | `crates/ecstore/src/bucket/metadata.rs:228-229`, checked in `check_header` at `:595-614` |
|
||||
| Body | MessagePack-encoded `BucketMetadata` | `marshal_msg`/`unmarshal` at `crates/ecstore/src/bucket/metadata.rs:582-593`; read strips the 4-byte header (`unmarshal(&data[4..])` at `:1079`) |
|
||||
| Path | `buckets/{bucket}/.metadata.bin` under the meta bucket (`BUCKET_METADATA_FILE`, `save_file_path`) | `crates/ecstore/src/bucket/metadata.rs` |
|
||||
| Header | 4 bytes: `format: u16 LE` + `version: u16 LE`, stripped before `unmarshal` | `check_header` in `crates/ecstore/src/bucket/metadata.rs` |
|
||||
| Body | MessagePack-encoded `BucketMetadata`; field names map one-to-one onto MinIO's `bucketMetadata` (PascalCase on the wire) | `BucketMetadata` in `crates/ecstore/src/bucket/metadata.rs` |
|
||||
| Per-config encoding | XML for S3-XML configs, JSON for policy / quota / targets / ACL; the per-config filename constants (`policy.json`, `lifecycle.xml`, ...) are `update_config` field-selector keys, not separate files | `update_config`, `parse_all_configs` in `crates/ecstore/src/bucket/metadata.rs` |
|
||||
| RustFS-only fields | `bucket_targets_config_meta_json`, `table_bucket_config_json`; a MinIO reader ignores unknown msgpack fields | `crates/ecstore/src/bucket/metadata.rs` |
|
||||
| Partial interop | `bucket_targets` meta side-channel is RustFS-specific; `bucket_acl` round-trips as a blob but only canned ACLs are enforced (see [minio-rustfs-router-compatibility.md](minio-rustfs-router-compatibility.md)) | |
|
||||
|
||||
This is the same design as MinIO's bucket metadata: a single
|
||||
`.minio.sys/buckets/<bucket>/.metadata.bin` blob with a 4-byte
|
||||
`bucketMetadataFormat|bucketMetadataVersion` header and a msgpack body. The
|
||||
filename, header shape, and format/version values (`1`/`1`) all match. The
|
||||
`BucketMetadata` field names correspond one-to-one to MinIO's `bucketMetadata`
|
||||
struct (`policyConfigJSON`, `lifecycleConfigXML`, `objectLockConfigXML`, …).
|
||||
### Importer
|
||||
|
||||
> Correction to a common misconception: modern MinIO does **not** store each
|
||||
> bucket config as a separate loose `versioning.json` / `lifecycle.json` file —
|
||||
> it embeds them in the same `.metadata.bin` blob, with XML for the S3-XML
|
||||
> configs and JSON for policy/quota/targets. The per-config filename constants
|
||||
> in RustFS (`policy.json`, `lifecycle.xml`, …) are the **keys used by
|
||||
> `update_config`** to select a field, not separate on-disk files.
|
||||
`crates/ecstore/src/bucket/migration.rs` is a one-way, idempotent importer from a `MIGRATING_META_BUCKET` (`.minio.sys`) layout into `.rustfs.sys`, run at startup from `rustfs/src/startup_bucket_metadata.rs`:
|
||||
|
||||
### Interop matrix (backlog#580 items)
|
||||
|
||||
Field/constant references are in `crates/ecstore/src/bucket/metadata.rs`.
|
||||
"Encoding" is the payload RustFS stores in that field and must match MinIO's for
|
||||
byte-level interop. Getter functions live in
|
||||
`crates/ecstore/src/bucket/metadata_sys.rs`.
|
||||
|
||||
| Config item | RustFS field / constant | Encoding | MinIO field | Status |
|
||||
|---|---|---|---|---|
|
||||
| versioning | `versioning_config_xml` / `BUCKET_VERSIONING_CONFIG` = `versioning.xml` | XML | versioningConfigXML | Done |
|
||||
| quota | `quota_config_json` / `BUCKET_QUOTA_CONFIG_FILE` = `quota.json` | JSON | quotaConfigJSON | Done |
|
||||
| object_lock | `object_lock_config_xml` / `OBJECT_LOCK_CONFIG` = `object-lock.xml` | XML | objectLockConfigXML | Done |
|
||||
| replication | `replication_config_xml` / `BUCKET_REPLICATION_CONFIG` = `replication.xml` | XML | replicationConfigXML | Done |
|
||||
| policy | `policy_config_json` / `BUCKET_POLICY_CONFIG` = `policy.json` | JSON | policyConfigJSON | Done |
|
||||
| lifecycle | `lifecycle_config_xml` / `BUCKET_LIFECYCLE_CONFIG` = `lifecycle.xml` | XML | lifecycleConfigXML | Done |
|
||||
| tagging | `tagging_config_xml` / `BUCKET_TAGGING_CONFIG` = `tagging.xml` | XML | taggingConfigXML | Done |
|
||||
| bucket_targets | `bucket_targets_config_json` + `bucket_targets_config_meta_json` / `BUCKET_TARGETS_FILE` = `bucket-targets.json` | JSON | bucketTargetsConfigJSON (+ meta variant) | Partial |
|
||||
| notification | `notification_config_xml` / `BUCKET_NOTIFICATION_CONFIG` = `notification.xml` | XML | notificationConfigXML | Done |
|
||||
| encryption | `encryption_config_xml` / `BUCKET_SSECONFIG` = `bucket-encryption.xml` | XML | encryptionConfigXML | Done |
|
||||
| cors | `cors_config_xml` / `BUCKET_CORS_CONFIG` = `cors.xml` | XML | corsConfigXML | Done |
|
||||
| public_access | `public_access_block_config_xml` / `BUCKET_PUBLIC_ACCESS_BLOCK_CONFIG` = `public-access-block.xml` | XML | publicAccessBlockConfigXML | Done |
|
||||
| bucket_acl | `bucket_acl_config_json` / `BUCKET_ACL_CONFIG` = `bucket-acl.json` | JSON | bucketACLConfigJSON | Partial |
|
||||
|
||||
Field definitions: `crates/ecstore/src/bucket/metadata.rs:274-336`. Constants:
|
||||
`:227-247`. `update_config` field routing: `:678-761`. `parse_all_configs` is
|
||||
invoked on load (`load_bucket_metadata_parse` at `:1043`).
|
||||
|
||||
Notes on the two "Partial" rows:
|
||||
|
||||
- **bucket_targets** — RustFS carries an extra `bucket_targets_config_meta_json`
|
||||
field (`:288`) beyond MinIO's single targets blob. The primary
|
||||
`bucket-targets.json` payload is interoperable; the meta side-channel is
|
||||
RustFS-specific and a MinIO reader would ignore it. ACL enforcement itself is
|
||||
bounded (S3 `PutBucketAcl`/`PutObjectAcl` accept canned ACLs only — see
|
||||
[minio-rustfs-router-compatibility.md](minio-rustfs-router-compatibility.md)).
|
||||
- **bucket_acl** — stored and round-tripped in the blob, but ACL grant
|
||||
semantics are intentionally limited at the S3 layer.
|
||||
|
||||
RustFS also defines fields with no interop requirement from backlog#580 but
|
||||
worth noting so a migration tool does not choke on them: `logging_config_xml`,
|
||||
`website_config_xml`, `accelerate_config_xml`, `request_payment_config_xml`
|
||||
(`:242-245`), and the RustFS-only `table_bucket_config_json`
|
||||
(`BUCKET_TABLE_CONFIG` = `table-bucket.json`, `:248`). A MinIO reader that does
|
||||
not know `table_bucket_config_json` will ignore the unknown msgpack field.
|
||||
|
||||
### Old-RustFS → new-RustFS migration
|
||||
|
||||
RustFS ships a one-way importer that reads a legacy meta bucket
|
||||
(`MIGRATING_META_BUCKET`) and rewrites both bucket metadata and IAM config into
|
||||
the current RustFS meta bucket, skipping entries that already exist
|
||||
(idempotent). See `crates/ecstore/src/bucket/migration.rs`:
|
||||
|
||||
- `try_migrate_bucket_metadata` copies `buckets/{bucket}/.metadata.bin` and the
|
||||
replication resync blob for each bucket (`crates/ecstore/src/bucket/migration.rs:193`).
|
||||
- `try_migrate_iam_config` walks `config/iam/` and normalizes legacy IAM
|
||||
records — legacy timestamp fields (`update_at` → `updatedAt`) and legacy
|
||||
policy-mapping field aliases (`policies` → `policy`) are rewritten
|
||||
(`normalize_iam_config_blob` at `:97`; regression test at `:428`).
|
||||
- Bucket resync metadata is re-encoded through `ReplicationMigrationBridge`
|
||||
(`normalize_bucket_meta_blob` at `:178`).
|
||||
|
||||
This importer is the practical basis for a MinIO → RustFS bucket-metadata
|
||||
migration: because the blob layout and field encodings already match, the
|
||||
missing piece is a source adapter that points the importer at a MinIO
|
||||
`.minio.sys` layout rather than the RustFS legacy layout.
|
||||
|
||||
### Bucket-metadata interop verdict
|
||||
|
||||
| Item | Done | Partial | Todo |
|
||||
|---|:--:|:--:|:--:|
|
||||
| `.metadata.bin` filename + header + msgpack layout parity | ✅ | | |
|
||||
| Per-config field encodings (XML/JSON) match MinIO | ✅ | | |
|
||||
| versioning/quota/object_lock/replication/policy/lifecycle/tagging/notification/encryption/cors/public_access round-trip | ✅ | | |
|
||||
| bucket_targets primary blob | ✅ | | |
|
||||
| bucket_targets meta side-channel + ACL grant semantics | | ⚠️ | |
|
||||
| Old-RustFS → new-RustFS importer | ✅ | | |
|
||||
| MinIO `.minio.sys` source adapter for the importer | | | ❌ |
|
||||
| CI proof a MinIO-written `.metadata.bin` loads unchanged | | | ❌ |
|
||||
|
||||
---
|
||||
| Function | Imports |
|
||||
|---|---|
|
||||
| `try_migrate_bucket_metadata` | `buckets/{bucket}/.metadata.bin` plus the replication resync blob (`normalize_bucket_meta_blob` via `ReplicationMigrationBridge`) |
|
||||
| `try_migrate_iam_config` | `config/iam/` records; `normalize_iam_config_blob` rewrites legacy timestamp and policy-mapping aliases |
|
||||
|
||||
## Part C — Server-Side Encryption (SSE)
|
||||
|
||||
Reading MinIO-written SSE objects is implemented, with a deliberate build boundary. The read path lives behind the `rio-v2` feature and is a **special-purpose migration capability**: it is not compiled into released binaries or container images, and there is no short-term plan to promote it into default builds. A default build fails such reads closed with a diagnosed error (see "How default builds fail" below); a `rio-v2` build reads them, within the scenario matrix below. The read-path work was tracked in rustfs/backlog#1638 (landed across rustfs/rustfs#6191, #6784, #6785).
|
||||
The `xl.meta` around a MinIO SSE object parses in every build, so such objects list, HEAD, and report plausible sizes; only payload readability depends on the build. KMS wire protocols (AWS `awsJson1_1` client in `crates/kms/src/backends/aws.rs`, MinIO KES) are non-targets.
|
||||
|
||||
### Scope boundary: KMS wire protocols and the production gate
|
||||
|
||||
This document covers MinIO on-disk metadata and object-encryption seams only. The **AWS KMS wire protocol** and the **MinIO KES wire protocol** are explicit non-targets: RustFS's AWS backend uses the AWS SDK's `awsJson1_1` client path (`crates/kms/src/backends/aws.rs:830`), while KES compatibility is outside this interop work. Those ecosystem evaluations remain separate work in the [#1562 Production Ready exit gate](https://github.com/rustfs/backlog/issues/1562), whose compatibility criterion covers MinIO/RustFS SSE data and rolling upgrades. Closing #1638 does not by itself close that gate.
|
||||
|
||||
Note the asymmetry with Parts A and B: the `xl.meta` around a MinIO SSE object parses fine, so such objects list, HEAD, and report plausible sizes. Only payload readability depends on the build and the scenario.
|
||||
|
||||
### What can and cannot be migrated
|
||||
|
||||
| Object class | Default build | `rio-v2` build | Notes |
|
||||
| Object class | `default` / `full` | `rio-v2` | Requirement |
|
||||
|---|:--:|:--:|---|
|
||||
| Unencrypted objects | ✅ | ✅ | Parts A and B apply. |
|
||||
| Bucket metadata, IAM config | ✅ | ✅ | Via the importer, once a `.minio.sys` source adapter exists (see Part B). |
|
||||
| Bucket-level default-encryption *configuration* | ✅ | ✅ | The `encryption` config blob round-trips as a blob; it does not make existing ciphertext readable. |
|
||||
| SSE-S3 / SSE-KMS, MinIO builtin static KMS (`MINIO_KMS_SECRET_KEY`), single- and multipart | ❌ diagnosed | ✅ | Requires `RUSTFS_SSE_S3_MASTER_KEY` set to the same 32-byte key material as MinIO's static secret. Proven against real MinIO fixtures (rustfs/rustfs#6191). |
|
||||
| SSE-C, MinIO-written | ❌ diagnosed | ✅ | Detection via MinIO's sealed-key slot; the customer key is proven by the AEAD unseal, since MinIO stores no key MD5 (rustfs/rustfs#6785). |
|
||||
| Any SSE, MinIO backed by KES / KMS plugin / MinKMS | ❌ | ❌ **not planned** | The wrapped DEK is sealed by the KES service itself; it is not a Vault/Transit ciphertext RustFS could be pointed at. Re-encrypt on the MinIO side before migrating. |
|
||||
| Objects sealed with legacy `DARE-SHA256` (`InsecureSealAlgorithm`) | ❌ | ❌ out of scope | Pre-DAREv2-HMAC MinIO; `parse_minio_managed_sealed_key` rejects the algorithm and the read fails closed. |
|
||||
| RustFS-written SSE objects read back by MinIO | ❌ | ❌ | See "Reverse direction". |
|
||||
| SSE-S3 / SSE-KMS, MinIO builtin static KMS, single- and multipart | Fail closed, diagnosed | Read | `RUSTFS_SSE_S3_MASTER_KEY` (base64, 32 bytes) equal to the source MinIO's static secret. |
|
||||
| SSE-C, MinIO-written | Fail closed, diagnosed | Read | Client supplies the customer key per request; MinIO stores no key MD5, so the AEAD unseal is the key proof. |
|
||||
| Any SSE, MinIO backed by KES / KMS plugin / MinKMS | Fail closed | Fail closed | Not planned. Re-encrypt or decrypt on the MinIO side first. |
|
||||
| Bucket default-encryption *configuration* | Round-trips | Round-trips | A config blob; it does not make existing ciphertext readable. |
|
||||
|
||||
### The seams, and where they closed
|
||||
### Seams
|
||||
|
||||
The cryptographic primitives were never the gap — RustFS implements the same DARE V2 stream format, object-key derivation, and sealing. Three seams above the cryptography rejected MinIO-written objects; all three are closed in `rio-v2` builds.
|
||||
The cryptography (DARE v2 stream format, object-key derivation, sealing) was never the gap. Three metadata seams above it rejected MinIO objects; all are closed in `rio-v2` builds. Symbols are in `rustfs/src/storage/sse.rs` unless noted.
|
||||
|
||||
| # | Seam | Resolution |
|
||||
|---|---|---|
|
||||
| 1 | Managed-SSE detection required the *persisted* public `x-amz-server-side-encryption` key, which MinIO synthesizes at response time and never stores. | Closed by rustfs/rustfs#6191: `infer_minio_managed_sse_type` infers the scheme from which MinIO sealed-key slot is present (the slot also selects the sealing-key domain, so a wrong inference cannot silently derive a wrong key). Inference from the KMS key id would misclassify — MinIO writes `-S3-Kms-Key-Id` on SSE-S3 objects too. |
|
||||
| 2 | MinIO's wrapped-DEK ciphertext was not accepted by any envelope parser. | Closed by rustfs/rustfs#6191: `decrypt_minio_kms_data_key` implements MinIO's builtin-KMS sealing (`sealingKey = HMAC-SHA256(master, iv)`), accepting both the raw `sealed‖iv‖nonce` layout and the legacy `{"aead": ...}` JSON. Routing is by the data key's own byte shape — RustFS's strict JSON envelopes are recognized positively, everything else goes to the MinIO decoder — because slot names cannot distinguish the writer. `LocalSseDekEnvelope` keeps `deny_unknown_fields`. |
|
||||
| 3 | SSE-C detection keyed on the stored customer-algorithm header, which MinIO also never persists, and the early key check demanded a stored key MD5 MinIO does not write. | Closed by rustfs/rustfs#6785: `stored_ssec_metadata` also accepts MinIO's SSE-C sealed-key slot (rio-v2 builds only), and `verify_ssec_key_match` tolerates a missing stored MD5 for exactly that shape — the AEAD unseal remains the key proof, and a wrong key still fails there. |
|
||||
|
||||
Two further single-part defects were fixed on the way (both rustfs/rustfs#6191 follow-ups): multipart classification now trusts MinIO's own `X-Minio-Internal-Encrypted-Multipart` marker instead of an ETag-length heuristic (MinIO stores *encrypted* ETags, so every single-part SSE object mis-classified as multipart), and single-part plaintext sizes are recovered by DARE reverse-size arithmetic (`dare_v2_decrypted_size`) since MinIO records an explicit size only for multipart uploads.
|
||||
| Seam | Resolution |
|
||||
|---|---|
|
||||
| Managed-SSE detection required the persisted public `x-amz-server-side-encryption` key, which MinIO synthesizes at response time | `infer_minio_managed_sse_type` infers the scheme from which MinIO sealed-key slot is present; the slot also selects the sealing-key domain, so a wrong inference cannot derive a wrong key. |
|
||||
| MinIO's wrapped-DEK ciphertext was accepted by no envelope parser | `decrypt_minio_kms_data_key` implements MinIO's builtin-KMS sealing for both the raw `sealed‖iv‖nonce` layout and the legacy `{"aead": ...}` JSON. Routing is by byte shape: `LocalSseDekEnvelope` (`deny_unknown_fields`) is recognized positively, everything else goes to the MinIO decoder. |
|
||||
| SSE-C detection keyed on the stored customer-algorithm header, which MinIO also never persists, and demanded a stored key MD5 | `stored_ssec_metadata` accepts MinIO's SSE-C sealed-key slot (`rio-v2` only); `verify_ssec_key_match` tolerates a missing stored MD5 for exactly that shape. |
|
||||
| Multipart classification used an ETag-length heuristic (MinIO stores encrypted ETags) | Trusts MinIO's own `X-Minio-Internal-Encrypted-Multipart` marker (`crates/utils/src/http/header_compat.rs`). |
|
||||
|
||||
### How default builds fail
|
||||
|
||||
The read fails closed: ciphertext is never served as plaintext. `is_object_encryption_marker` matches the whole `x-minio-internal-server-side-encryption-` prefix, so `ObjectInfo::is_encrypted()` is true for these objects, and the read plan refuses to construct a reader without decryption material. Since rustfs/rustfs#6784 the refusal is diagnosed: the resolver raises a typed error naming the condition — in default builds it points at the MinIO-compatible sealed format and the `rio-v2` read path it would require — and it surfaces as S3 `InvalidObjectState` (non-retryable) instead of the former undiagnosed 500 `InternalError`. List and HEAD still succeed, because `xl.meta` parses normally.
|
||||
|
||||
### What a `rio-v2` migration build needs
|
||||
|
||||
- A binary built with `--features rio-v2`. The feature is deliberately absent from `default` and `full` in `rustfs/Cargo.toml`; released binaries and images never include it.
|
||||
- For SSE-S3/SSE-KMS objects: `RUSTFS_SSE_S3_MASTER_KEY` (base64, 32 bytes) set to the same key material as the source MinIO's `MINIO_KMS_SECRET_KEY`. For SSE-C objects: nothing server-side — the client supplies the customer key per request, as on MinIO.
|
||||
- The interop harness is the evidence chain: `rustfs/src/storage/minio_generated_read_test.rs` (`#[ignore]` reader tests over real MinIO-generated fixtures, run with `--features rio-v2`), the fixture lab under `crates/rio-v2/tests/minio_fixture_lab/`, and the `minio-interop` workflow. The SSE-C lane of that harness (customer-key handout from a fixture capture to the reader test) is not wired yet; SSE-C coverage currently lives in the unit suite, which builds the MinIO shape with the same sealing primitives the fixture suite proved byte-compatible.
|
||||
|
||||
Known unverified edge: MinIO seals ETags on SSE objects (`SealETag`); RustFS does not unseal them, so ETag display and `If-Match` semantics on migrated SSE objects are not guaranteed to match MinIO's.
|
||||
`is_object_encryption_marker` (`crates/utils/src/http/header_compat.rs`) matches the whole `x-minio-internal-server-side-encryption-` prefix, so `ObjectInfo::is_encrypted()` is true and the read plan refuses to construct a reader without decryption material. The refusal is a typed error that names the MinIO-compatible sealed format and the `rio-v2` read path it requires, surfaced as S3 `InvalidObjectState` (non-retryable). Ciphertext is never served as plaintext.
|
||||
|
||||
### Reverse direction
|
||||
|
||||
Migrating back is also unsupported. Under `rio-v2` RustFS writes its own DEK envelope into MinIO's sealed-key metadata slots and labels it with MinIO's seal algorithm (`rustfs/src/storage/sse.rs:1830-1852`), so the metadata is MinIO-shaped while the key bytes are not MinIO-openable. Default builds do not populate those slots at all (`rustfs/src/storage/sse.rs:1796-1798`). Treat RustFS-written SSE objects as readable only by RustFS.
|
||||
Under `rio-v2` RustFS writes its own DEK envelope into MinIO's sealed-key metadata slots labelled with MinIO's seal algorithm, so the metadata is MinIO-shaped while the key bytes are not MinIO-openable. Default builds do not populate those slots. Treat RustFS-written SSE objects as readable only by RustFS. Known unverified edge: MinIO seals ETags on SSE objects; RustFS does not unseal them, so ETag display and `If-Match` on migrated SSE objects are not guaranteed to match MinIO.
|
||||
|
||||
### Migration options
|
||||
### Migration options for encrypted objects
|
||||
|
||||
- For static-KMS MinIO sources: run the migration through a `rio-v2` build with the shared master key (see above), either serving reads in place or copying objects out into a default-build cluster (the copy re-encrypts under RustFS's own KMS).
|
||||
- For KES/MinKMS-backed sources, or when a special-purpose build is not wanted: decrypt on the MinIO side first — rewrite the affected objects as plaintext, or copy them out through MinIO's S3 endpoint, which decrypts on read — and let RustFS apply its own encryption on ingest.
|
||||
- Leave encrypted objects on MinIO and migrate only unencrypted data.
|
||||
1. Static-KMS source: run the migration through a `rio-v2` build with the shared master key, serving in place or copying into a default-build cluster (the copy re-encrypts under RustFS's own KMS).
|
||||
2. KES / MinKMS source, or no special-purpose build wanted: decrypt on the MinIO side (rewrite as plaintext, or copy out through MinIO's S3 endpoint) and let RustFS encrypt on ingest.
|
||||
3. Leave encrypted objects on MinIO and migrate only unencrypted data.
|
||||
|
||||
Inventory the source first — bucket default-encryption settings mean objects can be encrypted without the uploader having asked for it, so "we never set SSE headers" is not sufficient evidence that a bucket has no encrypted objects.
|
||||
Inventory the source first: bucket default encryption means objects can be encrypted without the uploader asking, so "we never set SSE headers" is not evidence that a bucket has no encrypted objects.
|
||||
|
||||
### SSE interop verdict
|
||||
## rio-v2 variant lifecycle
|
||||
|
||||
| Item | Done | Partial | Todo |
|
||||
|---|:--:|:--:|:--:|
|
||||
| DARE V2 stream format parity | ✅ | | |
|
||||
| Object-key derivation / sealing parity | ✅ | | |
|
||||
| Managed-SSE detection accepts MinIO-written metadata (`rio-v2`) | ✅ | | |
|
||||
| MinIO builtin-KMS wrapped-DEK parser (raw + legacy JSON) | ✅ | | |
|
||||
| SSE-C detection accepts MinIO-written metadata (`rio-v2`) | ✅ | | |
|
||||
| Read MinIO-written SSE-S3 / SSE-KMS end to end, single- and multipart | ✅ | | |
|
||||
| Read MinIO-written SSE-C end to end | | ⚠️ unit-proven; fixture-lab lane unwired | |
|
||||
| Migrated-object sealed-ETag semantics | | | ❌ unverified |
|
||||
| KES / MinKMS / legacy `DARE-SHA256` sources | | | ❌ not planned |
|
||||
| RustFS-written SSE objects readable by MinIO | | | ❌ |
|
||||
| CI proof of SSE read parity | | ⚠️ `minio-interop` workflow; nightly once re-enabled | |
|
||||
`rio-v2` is a dormant, special-purpose migration variant tracked under `rustfs/backlog#1835`.
|
||||
|
||||
---
|
||||
| Fact | Value |
|
||||
|---|---|
|
||||
| Shipping status | Ships in no default build: absent from `default` and `full` in `rustfs/Cargo.toml`; released binaries and container images never include it. Enable with `--features rio-v2`. |
|
||||
| Pull-request coverage | `test-and-lint-rio-v2` in `.github/workflows/ci.yml`: clippy plus `cargo nextest` for `rustfs` and `rustfs-ecstore` with `--features rio-v2`. This is the cfg-seam guard; it keeps the feature compiling and its unit suite green on every pull request. |
|
||||
| Full-suite lane | `build-rustfs-debug-binary-rio-v2` and `e2e-tests-rio-v2` in `ci.yml` run only on the weekly schedule and manual dispatch (`cache-warm.yml` keeps the `ci-feat-rio` cache warm so the scheduled build fits its timeout). |
|
||||
| Interop evidence | `.github/workflows/minio-interop.yml` (nightly plus manual) regenerates real MinIO backend trees via `crates/rio-v2/tests/minio_fixture_lab/` and runs the `#[ignore]` reader tests in `rustfs/src/storage/minio_generated_read_test.rs` with `--features rio-v2`. Its freshness is tracked in `.github/scheduled-validations.json`. |
|
||||
| Promote-or-delete condition | The variant stays dormant until one of two things happens. **Promote**: a release commits to shipping MinIO SSE migration as a supported capability; then `rio-v2` joins `default`/`full`, the scheduled lanes run on every pull request, and this section is rewritten. **Delete**: no release commits to it and the scheduled lanes are not kept green; then the feature flag, `crates/rio-v2`, the cfg seams in `rustfs/src/storage/sse.rs`, the three `ci.yml` jobs, the `cache-warm.yml` warm step, `minio-interop.yml`, and its `scheduled-validations.json` entry are removed in one change. Either outcome must update `ARCHITECTURE.md` and the `ci.yml` job comments that cite this section. |
|
||||
|
||||
## Phased Plan
|
||||
## Fixture Evidence
|
||||
|
||||
The format is already close; the plan is verification, a source adapter, and
|
||||
closing the two partial encodings — not a rewrite.
|
||||
Fixtures were captured from a real MinIO single-drive instance and live under `crates/filemeta/tests/fixtures/minio/` and `crates/ecstore/tests/fixtures/minio/`. These tests run in the normal `cargo test` / nextest lanes.
|
||||
|
||||
### Phase 1 — Read parity, proven (verification)
|
||||
| Test | File | Proves |
|
||||
|---|---|---|
|
||||
| `parses_real_minio_object_xlmeta` | `crates/filemeta/src/filemeta.rs` | Inline, two-version plus delete-marker, and multipart `xl.meta` parse to the expected `FileInfo`. |
|
||||
| `parses_real_minio_bucket_metadata_blob_without_loss` | `crates/ecstore/src/bucket/metadata.rs` | The msgpack blob decodes via MinIO's field names and `parse_all_configs` loads every config in the corpus, including MinIO's lifecycle `<ExpiryUpdatedAt>` and replication `DeleteMarkerReplication` / `ExistingObjectReplication` extensions. |
|
||||
| `reads_minio_inline_bucket_metadata_via_bitrot` | `crates/ecstore/src/bucket/metadata.rs` | The inline shard's HighwayHash prefix verifies under `HighwayHash256S` and yields the exact `.metadata.bin` blob. |
|
||||
| `migrates_real_minio_bucket_metadata_end_to_end` | `crates/ecstore/src/bucket/migration.rs` | A real `.metadata.bin` seeded under a `.minio.sys` layout is imported by `try_migrate_bucket_metadata` into `.rustfs.sys` byte-identical, through the object layer on a 4-drive `ECStore`. |
|
||||
| `test_issue_2265_legacy_meta_v2_object_compatibility`, `test_issue_2288_legacy_xlmeta_compatibility` | `crates/filemeta/src/filemeta.rs` | Legacy meta_ver 2 objects with legacy checksums still read. |
|
||||
| `minio_generated_read_test.rs` (`#[ignore]`, `rio-v2`) | `rustfs/src/storage/minio_generated_read_test.rs` | Byte-identical plaintext reconstruction of MinIO SSE-S3 / SSE-KMS fixtures; driven by `minio-interop.yml`. |
|
||||
|
||||
- Add a MinIO-writer fixture corpus for `xl.meta` (inline + multipart +
|
||||
versioned + delete-marker + transitioned) and assert RustFS parses each to a
|
||||
`FileInfo` equivalent to MinIO's, alongside the existing issue #2265 / #2288
|
||||
fixtures in `crates/filemeta/src/filemeta.rs`.
|
||||
- Add a fixture `.metadata.bin` written by MinIO and assert
|
||||
`BucketMetadata::unmarshal` + `parse_all_configs` load every field without
|
||||
loss (`crates/ecstore/src/bucket/metadata.rs`).
|
||||
- Exit criterion: a CI job that fails if a real MinIO-written object or bucket
|
||||
blob cannot be read.
|
||||
Not fixture-proven: transitioned `xl.meta`; CORS, public-access-block, and bucket-ACL configs (the SNSD corpus did not exercise them); bucket-targets credentials (MinIO stores them KMS-encrypted); the SSE-C fixture-lab lane (customer-key handout is not wired; SSE-C coverage is unit-level).
|
||||
|
||||
#### Phase 1 status — first fixtures landed (verified 2026-07-07)
|
||||
## Out Of Scope
|
||||
|
||||
A real MinIO `RELEASE.2025-07-23` single-drive instance wrote a bucket with
|
||||
versioning, object-lock (GOVERNANCE default), lifecycle, tagging, quota, and a
|
||||
public-download policy, plus inline / versioned / multipart objects. The on-disk
|
||||
`xl.meta` blobs are captured as hex fixtures
|
||||
(`crates/filemeta/tests/fixtures/minio/`, `crates/ecstore/tests/fixtures/minio/`).
|
||||
|
||||
Proven by regression tests:
|
||||
|
||||
- **Object `xl.meta` read parity** — `parses_real_minio_object_xlmeta`
|
||||
(`crates/filemeta/src/filemeta.rs`): small inline, two-object-version + delete
|
||||
marker, and multipart objects all parse to the expected `FileInfo`.
|
||||
- **Bucket-metadata parse parity** — `parses_real_minio_bucket_metadata_blob_without_loss`
|
||||
(`crates/ecstore/src/bucket/metadata.rs`): the msgpack blob decodes via the
|
||||
PascalCase MinIO field names, and `parse_all_configs` loads **all ten** config
|
||||
types present in the corpus without loss — policy, lifecycle (**including
|
||||
MinIO's `<ExpiryUpdatedAt>` extension**), object-lock, versioning, tagging,
|
||||
quota, notification, encryption (SSE-S3), and replication (**including the
|
||||
`DeleteMarkerReplication` / `ExistingObjectReplication` MinIO extensions**).
|
||||
- **Inline bucket-metadata read parity** — `reads_minio_inline_bucket_metadata_via_bitrot`
|
||||
(`crates/ecstore/src/bucket/metadata.rs`): MinIO stores an inlined object body
|
||||
as `[HighwayHash256 (32B)][body]`. The "`inline_data` 前缀不同" that weisd
|
||||
raised on 2026-03-06 is exactly that bitrot prefix — **not** a format
|
||||
incompatibility. Feeding the raw inline shard through RustFS's `BitrotReader`
|
||||
with the default `HighwayHash256S` verifies the checksum (confirming RustFS's
|
||||
hash matches MinIO's) and yields the exact `.metadata.bin` blob, which then
|
||||
parses. So the object-layer inline read is compatible; the earlier "extract
|
||||
`fi.data` directly" concern was reading the shard before the bitrot layer
|
||||
strips its prefix.
|
||||
- **End-to-end migration** — `migrates_real_minio_bucket_metadata_end_to_end`
|
||||
(`crates/ecstore/src/bucket/migration.rs`): on a throwaway 4-drive local
|
||||
`ECStore`, a real MinIO `.metadata.bin` seeded under a `.minio.sys` layout is
|
||||
migrated by `try_migrate_bucket_metadata` into `.rustfs.sys`, and the migrated
|
||||
blob carries every config (policy / lifecycle / object-lock / versioning /
|
||||
tagging / quota / notification / encryption / replication) byte-identical to
|
||||
the source. This exercises the Phase 2 source adapter
|
||||
(`MIGRATING_META_BUCKET = ".minio.sys"`) end-to-end through the object layer —
|
||||
proven, not just present.
|
||||
|
||||
Still to broaden: transitioned `xl.meta`; CORS, public-access-block, and bucket
|
||||
ACL configs (the SNSD test binary/`mc` did not expose these); and bucket-targets
|
||||
credentials, which MinIO stores KMS-encrypted (a documented partial). These run
|
||||
as ordinary crate tests, so they already execute in the normal `cargo
|
||||
test`/nextest CI jobs.
|
||||
|
||||
### Phase 2 — MinIO source adapter for migration
|
||||
|
||||
- Generalize the importer in `crates/ecstore/src/bucket/migration.rs` so the
|
||||
source can be a MinIO `.minio.sys/buckets/<bucket>/.metadata.bin` layout, not
|
||||
only the RustFS legacy meta bucket. Because the blob format matches, this is
|
||||
mostly source-path plumbing plus IAM record normalization reuse.
|
||||
- Exit criterion: importing a MinIO backup reproduces all backlog#580
|
||||
bucket-config items with byte-identical config payloads.
|
||||
|
||||
### Phase 3 — Close the two partial encodings
|
||||
|
||||
- bucket_targets: document/normalize the RustFS-only
|
||||
`bucket_targets_config_meta_json` so a round-trip through MinIO and back does
|
||||
not silently drop it; or fold its content into a MinIO-compatible
|
||||
representation.
|
||||
- bucket_acl: decide whether ACL grant semantics beyond canned ACLs are in
|
||||
scope; if not, keep the blob round-trippable but document the enforcement
|
||||
limit (already reflected in the router compatibility matrix).
|
||||
|
||||
### Phase 4 — Round-trip / write-back parity (non-goal for migration)
|
||||
|
||||
Proving a MinIO binary can re-read a *RustFS-written drive set* (the reverse
|
||||
direction) is **out of scope for the migration use case**, which is one-way
|
||||
MinIO → RustFS:
|
||||
|
||||
- RustFS's meta bucket is `.rustfs.sys` (`crates/ecstore/src/disk/mod.rs:29`);
|
||||
MinIO looks for `.minio.sys`. A MinIO binary pointed at a RustFS drive set
|
||||
does not find `format.json` or bucket configs and refuses the set — this is a
|
||||
set-level divergence, not an object-format one.
|
||||
- The object-level `xl.meta` format *does* match (proven above), so the reverse
|
||||
direction is limited by drive-set discovery, not by per-object encoding.
|
||||
- The supported flow is one-way: `try_migrate_bucket_metadata` /
|
||||
`try_migrate_iam_config` / `format.json` migration import a MinIO layout into
|
||||
RustFS. There is no requirement to keep a live MinIO able to serve
|
||||
RustFS-written drives.
|
||||
|
||||
If a true bidirectional round-trip is ever needed, it would require RustFS to
|
||||
optionally write the `.minio.sys` set layout — a separate feature, not part of
|
||||
the interop/migration story tracked here.
|
||||
|
||||
---
|
||||
- A live MinIO binary serving a RustFS-written drive set (set-level `.minio.sys` vs `.rustfs.sys` divergence). A bidirectional round-trip would require RustFS to optionally write the `.minio.sys` set layout, which is a separate feature.
|
||||
- RustFS-written SSE objects readable by MinIO.
|
||||
- KES / MinKMS / KMS-plugin-sealed MinIO objects.
|
||||
- Objects sealed with pre-DARE-v2-HMAC MinIO seal algorithms; `parse_minio_managed_sealed_key` rejects unknown algorithms and the read fails closed.
|
||||
- AWS KMS and KES wire-protocol compatibility.
|
||||
|
||||
## Guardrails
|
||||
|
||||
- This document is analysis only. Any change to `crates/filemeta` or
|
||||
`crates/ecstore/src/bucket` metadata encoding is a storage-format change and
|
||||
must follow the migration and readiness contracts in
|
||||
[README.md](README.md) and the ecstore layout boundary rules.
|
||||
- The version constants (`XL_META_VERSION`,
|
||||
`BUCKET_METADATA_FORMAT`/`BUCKET_METADATA_VERSION`) are compatibility anchors.
|
||||
Bumping any of them requires a read-compat path for the prior value and a
|
||||
migration story, exactly as the current meta_ver 2 → 3 read path provides.
|
||||
- Any change to `crates/filemeta` or `crates/ecstore/src/bucket` metadata encoding is a storage-format change and follows the migration and readiness contracts in [README.md](README.md) and the ecstore layout boundary rules.
|
||||
- Do not bump a Version Anchor without a read path for the prior value; see [erasure-coding.md](erasure-coding.md) for the accept-older, reject-newer rule.
|
||||
- `.github/workflows/ci.yml`, `.github/workflows/cache-warm.yml`, and `ARCHITECTURE.md` cite the [rio-v2 variant lifecycle](#rio-v2-variant-lifecycle) heading; keep it when editing this file.
|
||||
|
||||
@@ -1,213 +1,50 @@
|
||||
# MinIO ↔ RustFS Router Compatibility Matrix
|
||||
# MinIO ↔ RustFS Router Compatibility (Exceptions Only)
|
||||
|
||||
Tracks how RustFS covers the MinIO HTTP router surface, split into the S3
|
||||
data-plane router (`cmd/api-router.go` in MinIO: object + bucket APIs) and the
|
||||
admin control-plane router (`cmd/admin-router.go`: admin `/v3/` and `/v4/`
|
||||
APIs). Each row records the current RustFS implementation status and the
|
||||
landing point in the code so the matrix can be re-verified after refactors.
|
||||
**Use this when:** a client or `mc` call that works against MinIO fails against RustFS and you need to know whether the endpoint is missing, stubbed, or deliberately different.
|
||||
**Source of truth:** S3 plane: the `s3s::S3` trait impl in `rustfs/src/storage/ecfs.rs`. Admin plane: `make_admin_route` in `rustfs/src/admin/mod.rs`, the registration inventory `rustfs/src/admin/route_registration_test.rs`, and the route/action guardrail [admin-route-action-snapshot.md](admin-route-action-snapshot.md).
|
||||
|
||||
This complements two neighbouring documents and does not duplicate them:
|
||||
|
||||
- [s3-compatibility-matrix.md](s3-compatibility-matrix.md) — the release-facing
|
||||
S3 compatibility claim and the Ceph s3tests lists that gate it.
|
||||
- [admin-route-action-snapshot.md](admin-route-action-snapshot.md) — the admin
|
||||
route/handler/authorization-action migration guardrail (the source of truth
|
||||
for exact route patterns and auth contracts).
|
||||
|
||||
Refs rustfs/backlog#596 rustfs/backlog#603.
|
||||
This document lists only exceptions. Anything not listed here is implemented with MinIO-equivalent behavior. For the s3tests-level claim see [s3-compatibility-matrix.md](s3-compatibility-matrix.md).
|
||||
|
||||
## Status Legend
|
||||
|
||||
| Status | Meaning |
|
||||
|---|---|
|
||||
| 已实现 (implemented) | Handler is registered and performs the real operation. |
|
||||
| 部分兼容 (partial) | Registered and functional, but a documented subset of the MinIO behavior is rejected or unsupported. |
|
||||
| 已注册未完成 (registered, incomplete) | Route is registered but the handler returns `NotImplemented` (a behavior contract, not a real implementation). |
|
||||
| 缺失 (missing) | No RustFS route/handler for the MinIO endpoint. |
|
||||
| 行为不一致 (behavior differs) | Implemented but intentionally diverges from MinIO's response contract. |
|
||||
| 缺失 (missing) | No RustFS route or handler for the MinIO endpoint. |
|
||||
| 部分兼容 (partial) | Registered and functional, but a documented subset of MinIO behavior is rejected. |
|
||||
| 已注册未完成 (registered, incomplete) | Route is registered; the handler returns `NotImplemented` as a behavior contract. |
|
||||
| 行为不一致 (behavior differs) | Implemented, but intentionally diverges from MinIO's response contract. |
|
||||
|
||||
Prefixes: RustFS registers admin routes under the canonical `/rustfs/admin`
|
||||
prefix and accepts `/minio/admin` as a compatibility alias via router
|
||||
canonicalization (see
|
||||
[admin-route-action-snapshot.md](admin-route-action-snapshot.md)). Admin paths
|
||||
below are shown relative to that prefix (e.g. `/v3/info`).
|
||||
Admin paths are relative to the canonical `/rustfs/admin` prefix; `/minio/admin` is accepted as an alias via router canonicalization.
|
||||
|
||||
---
|
||||
## S3 Data Plane
|
||||
|
||||
## Part 1 — S3 Data Plane (MinIO `cmd/api-router.go`)
|
||||
All `s3s::S3` trait methods are implemented in `rustfs/src/storage/ecfs.rs` except the following.
|
||||
|
||||
RustFS implements the S3 surface through the `s3s` service trait in
|
||||
`rustfs/src/storage/ecfs.rs`, delegating to use-case layers under
|
||||
`rustfs/src/app/`. Line numbers are indicative landing points on the branch
|
||||
this matrix was written against and may drift; the file paths are stable.
|
||||
|
||||
### Bucket-level operations
|
||||
|
||||
| MinIO / S3 operation | Status | RustFS landing point |
|
||||
| S3 operation | Status | Detail |
|
||||
|---|---|---|
|
||||
| CreateBucket | 已实现 | `rustfs/src/storage/ecfs.rs` (`create_bucket`) |
|
||||
| DeleteBucket | 已实现 | `rustfs/src/storage/ecfs.rs` (`delete_bucket`) |
|
||||
| HeadBucket | 已实现 | `rustfs/src/storage/ecfs.rs` (`head_bucket`) |
|
||||
| ListBuckets | 已实现 | `rustfs/src/storage/ecfs.rs` (`list_buckets`) |
|
||||
| GetBucketLocation | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_location`) |
|
||||
| ListObjects (v1) | 已实现 | `rustfs/src/storage/ecfs.rs` (`list_objects`) |
|
||||
| ListObjectsV2 | 已实现 | `rustfs/src/storage/ecfs.rs` (`list_objects_v2`) |
|
||||
| ListObjectVersions | 已实现 | `rustfs/src/storage/ecfs.rs` (`list_object_versions`) |
|
||||
| ListMultipartUploads | 已实现 | `rustfs/src/storage/ecfs.rs` (`list_multipart_uploads`) |
|
||||
| Get/PutBucketVersioning | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_versioning`, `put_bucket_versioning`) |
|
||||
| Get/Put/DeleteBucketPolicy | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_policy`, `put_bucket_policy`, `delete_bucket_policy`) |
|
||||
| GetBucketPolicyStatus | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_policy_status`) |
|
||||
| Get/Put/DeleteBucketTagging | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_tagging`, `put_bucket_tagging`, `delete_bucket_tagging`) |
|
||||
| Get/Put/DeleteBucketLifecycle | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_lifecycle_configuration`, `put_bucket_lifecycle_configuration`, `delete_bucket_lifecycle`) |
|
||||
| Get/Put/DeleteBucketReplication | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_replication`, `put_bucket_replication`, `delete_bucket_replication`) |
|
||||
| Get/Put/DeleteBucketEncryption | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_encryption`, `put_bucket_encryption`, `delete_bucket_encryption`) |
|
||||
| Get/PutObjectLockConfiguration | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_object_lock_configuration`, `put_object_lock_configuration`) |
|
||||
| Get/Put/DeletePublicAccessBlock | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_public_access_block`, `put_public_access_block`, `delete_public_access_block`) |
|
||||
| Get/Put/DeleteBucketCors | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_cors`, `put_bucket_cors`, `delete_bucket_cors`) |
|
||||
| GetBucketNotificationConfiguration | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_notification_configuration`) |
|
||||
| PutBucketNotificationConfiguration | 已实现 | `rustfs/src/storage/ecfs.rs` (`put_bucket_notification_configuration`) |
|
||||
| Get/PutBucketRequestPayment | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_request_payment`, `put_bucket_request_payment`) |
|
||||
| Get/PutBucketLogging | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_logging`, `put_bucket_logging`) |
|
||||
| Get/Put/DeleteBucketWebsite | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_website`, `put_bucket_website`, `delete_bucket_website`) |
|
||||
| Get/PutBucketAccelerateConfiguration | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_accelerate_configuration`, `put_bucket_accelerate_configuration`) |
|
||||
| GetBucketAcl | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_bucket_acl`) |
|
||||
| PutBucketAcl | 部分兼容 | `rustfs/src/storage/ecfs.rs` (`put_bucket_acl`) — canned-ACL headers only; XML grant policies return `NotImplemented`. |
|
||||
| GetBucketReplicationMetrics | 缺失 | No S3-path handler; replication metrics are exposed via the admin API `/v3/replicationmetrics` instead. |
|
||||
| GetBucketOwnershipControls | 缺失 | Not implemented (matches the "bucket ownership controls: planned" note in `s3-compatibility-matrix.md`). |
|
||||
| Put/DeleteBucketOwnershipControls | 缺失 | Not implemented. |
|
||||
| GetBucketReplicationMetrics | 缺失 | No `get_bucket_replication_metrics`; replication metrics are exposed via admin `/v3/replicationmetrics`. |
|
||||
| GetBucketOwnershipControls | 缺失 | No handler; s3tests entries remain in `scripts/s3-tests/unimplemented_tests.txt`. |
|
||||
| PutBucketOwnershipControls, DeleteBucketOwnershipControls | 缺失 | No handler. |
|
||||
| DeleteBucketNotification, DeleteBucketLogging, DeleteBucketRequestPayment, DeleteBucketAccelerate | 部分兼容 | No distinct DELETE handlers; clear the config by writing an empty configuration through the PUT path. |
|
||||
| PutBucketAcl, PutObjectAcl | 部分兼容 | Canned-ACL headers only; XML grant bodies return `NotImplemented` (`put_bucket_acl`, `put_object_acl`). |
|
||||
| GetObjectTorrent | 行为不一致 | `get_object_torrent` returns `404 NoSuchKey` by design, not `501 NotImplemented`, so clients degrade gracefully. |
|
||||
|
||||
Note on delete verbs: several S3 sub-resource DELETE operations
|
||||
(DeleteBucketNotification, DeleteBucketLogging, DeleteBucketRequestPayment,
|
||||
DeleteBucketAccelerate) are not exposed as distinct handlers; the corresponding
|
||||
config is cleared by writing an empty configuration through the PUT path. Treat
|
||||
these as 部分兼容 at the client level.
|
||||
## Admin Control Plane
|
||||
|
||||
### Object-level operations
|
||||
Every route asserted in `rustfs/src/admin/route_registration_test.rs` is registered. Exceptions:
|
||||
|
||||
| MinIO / S3 operation | Status | RustFS landing point |
|
||||
| MinIO admin family | Status | Detail |
|
||||
|---|---|---|
|
||||
| GetObject | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_object`) |
|
||||
| PutObject | 已实现 | `rustfs/src/storage/ecfs.rs` (`put_object`) |
|
||||
| DeleteObject | 已实现 | `rustfs/src/storage/ecfs.rs` (`delete_object`) |
|
||||
| DeleteObjects (multi-delete) | 已实现 | `rustfs/src/storage/ecfs.rs` (`delete_objects`) |
|
||||
| HeadObject | 已实现 | `rustfs/src/storage/ecfs.rs` (`head_object`) |
|
||||
| CopyObject | 已实现 | `rustfs/src/storage/ecfs.rs` (`copy_object`) |
|
||||
| GetObjectAcl | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_object_acl`) |
|
||||
| PutObjectAcl | 部分兼容 | `rustfs/src/storage/ecfs.rs` (`put_object_acl`) — canned-ACL headers only; XML grants return `NotImplemented`. |
|
||||
| Get/Put/DeleteObjectTagging | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_object_tagging`, `put_object_tagging`, `delete_object_tagging`) |
|
||||
| GetObjectAttributes | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_object_attributes`) |
|
||||
| Get/PutObjectLegalHold | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_object_legal_hold`, `put_object_legal_hold`) |
|
||||
| Get/PutObjectRetention | 已实现 | `rustfs/src/storage/ecfs.rs` (`get_object_retention`, `put_object_retention`) |
|
||||
| RestoreObject (POST restore) | 已实现 | `rustfs/src/storage/ecfs.rs` (`restore_object`) |
|
||||
| SelectObjectContent | 已实现 | `rustfs/src/storage/ecfs.rs` (`select_object_content`) → `rustfs/src/app/select_object.rs` |
|
||||
| CreateMultipartUpload | 已实现 | `rustfs/src/storage/ecfs.rs` (`create_multipart_upload`) |
|
||||
| UploadPart / UploadPartCopy | 已实现 | `rustfs/src/storage/ecfs.rs` (`upload_part`, `upload_part_copy`) |
|
||||
| CompleteMultipartUpload | 已实现 | `rustfs/src/storage/ecfs.rs` (`complete_multipart_upload`) |
|
||||
| AbortMultipartUpload | 已实现 | `rustfs/src/storage/ecfs.rs` (`abort_multipart_upload`) |
|
||||
| ListParts | 已实现 | `rustfs/src/storage/ecfs.rs` (`list_parts`) |
|
||||
| PostObject (POST form upload) | 已实现 | Routed via the POST-object marker into the put-object path (`rustfs/src/app/object/put.rs`). See the "POST Object form upload checksum handling: planned" note in `s3-compatibility-matrix.md`. |
|
||||
| GetObjectTorrent | 行为不一致 | `rustfs/src/storage/ecfs.rs` (`get_object_torrent`) — returns `404 NoSuchKey` by design (not `501 NotImplemented`) so clients degrade gracefully. |
|
||||
| Batch jobs (`/v3/start-job`, `/v3/list-jobs`, `/v3/status-job`, `/v3/describe-job`, `/v3/cancel-job`) | 已注册未完成 | `rustfs/src/admin/handlers/batch_job.rs`: `start-job` returns `NotImplemented` for known job types (`KNOWN_JOB_TYPES`) and `InvalidRequest` for unknown ones; `list-jobs` returns an empty list; status/describe/cancel return no-such-job. See [kms-bulk-rekey-contract.md](kms-bulk-rekey-contract.md) for why `keyrotate` must keep refusing. |
|
||||
| Service control (`POST /v3/service`) | 行为不一致 | `ServiceHandle` in `rustfs/src/admin/handlers/system.rs`: `restart` and `stop` both initiate graceful shutdown (the process manager must relaunch; no in-process restart); `freeze` / `unfreeze` toggle a global freeze flag under `ServiceFreezeAdminAction`. `rustfs/src/admin/route_policy.rs` still classifies the route as deferred `NotImplemented`. |
|
||||
| Inspect data (`GET|POST /v3/inspect-data`) | 行为不一致 | `InspectDataHandler` in `system.rs` returns the raw bytes of one exact `volume` + `file`, size-capped, instead of MinIO's encrypted raw-drive-file archive. The bounded archive lives at `POST /v4/inspect/archive` (`rustfs/src/admin/handlers/inspect_archive.rs`). `route_policy.rs` still classifies the v3 route as deferred `NotImplemented`. |
|
||||
| Pools decommission / cancel / clear | 部分兼容 | `rustfs/src/admin/handlers/pools.rs` returns `NotImplemented` when endpoints are not initialized (single-pool or uninitialized clusters). |
|
||||
| `/v3/top/drives`, `/v3/top/net` | 缺失 | Only `/v3/top/locks` is registered (`rustfs/src/admin/handlers/diagnostics.rs`). |
|
||||
| Bucket / site replication per-object diff | 缺失 | `/v3/replicationmetrics` and site-replication status exist; no diff endpoint. |
|
||||
| MRF (most-recent-failures) replication metrics breakdown | 缺失 | Only the generic `/v3/metrics` stream and replication metrics wire (`rustfs/src/admin/replication_metrics_wire.rs`). |
|
||||
|
||||
For the gate-level view of which of these are covered by executable s3tests,
|
||||
defer to [s3-compatibility-matrix.md](s3-compatibility-matrix.md); this table is
|
||||
the router/handler view, not the test-list view.
|
||||
Formerly-missing families that are now registered and therefore not exceptions: `/v3/healthinfo`, `/v3/obdinfo`, `/v3/force-unlock`, `/v3/top/locks`, `/v3/speedtest*`, `/v3/log`, `/v3/trace`, `/v3/profile`, `/v3/profiling/*`, `/v3/idp/{ldap|openid}/*`, `/v3/idp-config/*`.
|
||||
|
||||
---
|
||||
## Update Rule
|
||||
|
||||
## Part 2 — Admin Control Plane (MinIO `cmd/admin-router.go`)
|
||||
|
||||
Router assembly is `rustfs/src/admin/mod.rs::register_admin_routes`; the exact
|
||||
route patterns, handler ownership, and authorization actions are the guardrail
|
||||
in [admin-route-action-snapshot.md](admin-route-action-snapshot.md). This table
|
||||
maps MinIO admin route families to RustFS status.
|
||||
|
||||
### Implemented / registered families
|
||||
|
||||
| MinIO admin family | Status | RustFS landing point |
|
||||
|---|---|---|
|
||||
| STS / is-admin probe | 已实现 | `rustfs/src/admin/handlers/sts.rs`, `is_admin.rs` |
|
||||
| User lifecycle (list/add/info/remove/status) | 已实现 | `rustfs/src/admin/handlers/user_lifecycle.rs`, `user.rs` |
|
||||
| Groups | 已实现 | `rustfs/src/admin/handlers/group.rs` |
|
||||
| Service accounts / access keys | 已实现 | `rustfs/src/admin/handlers/service_account.rs` |
|
||||
| Canned policies + builtin policy attach/detach + policy-entities | 已实现 | `rustfs/src/admin/handlers/policies.rs` |
|
||||
| IAM import/export | 已实现 | `rustfs/src/admin/handlers/user_iam.rs`, `user.rs` |
|
||||
| Account info | 已实现 | `rustfs/src/admin/handlers/account_info.rs` |
|
||||
| Config KV (get/set/del/help/history/restore + `/v3/config`) | 已实现 | `rustfs/src/admin/handlers/config_admin.rs` |
|
||||
| Server info / storageinfo / datausageinfo | 已实现 | `rustfs/src/admin/handlers/system.rs` |
|
||||
| Metrics stream (`/v3/metrics`) | 已实现 | `rustfs/src/admin/handlers/metrics.rs` via `system.rs` |
|
||||
| Runtime capabilities (`/v4/runtime/capabilities`) | 已实现 | `rustfs/src/admin/handlers/system.rs` |
|
||||
| Pools list/status | 已实现 | `rustfs/src/admin/handlers/pools.rs` |
|
||||
| Pools decommission/cancel/clear | 部分兼容 | `rustfs/src/admin/handlers/pools.rs` — returns `NotImplemented` when endpoints are not initialized (single-pool / uninitialized clusters). |
|
||||
| Rebalance start/status/stop | 已实现 | `rustfs/src/admin/handlers/rebalance.rs` |
|
||||
| Heal + background-heal status | 已实现 | `rustfs/src/admin/handlers/heal.rs` |
|
||||
| Tier (list/stats/verify/add/edit/remove/clear) | 已实现 | `rustfs/src/admin/handlers/tier.rs` |
|
||||
| Quota (legacy + bucket-scoped + stats/check) | 已实现 | `rustfs/src/admin/handlers/quota.rs` |
|
||||
| Bucket metadata export/import | 已实现 | `rustfs/src/admin/handlers/bucket_meta.rs` |
|
||||
| Scanner status | 已实现 | `rustfs/src/admin/handlers/scanner.rs` |
|
||||
| Notification targets (list/arns/put/reset) | 已实现 | `rustfs/src/admin/handlers/event.rs` |
|
||||
| Audit targets (list/put/reset) | 已实现 | `rustfs/src/admin/handlers/audit.rs` |
|
||||
| Module switches | 已实现 | `rustfs/src/admin/handlers/module_switch.rs` |
|
||||
| Plugin catalog + instances (`/v4/plugins/*`) | 已实现 | `rustfs/src/admin/handlers/plugins_catalog.rs`, `plugins_instances.rs` |
|
||||
| Extension catalog + instances (`/v4/extensions/*`) | 已实现 | `rustfs/src/admin/handlers/extensions.rs` |
|
||||
| Object ZIP download (`/v3/zip-downloads`) | 已实现 | `rustfs/src/admin/handlers/object_zip_download.rs` |
|
||||
| Cluster snapshot (`/v4/cluster/snapshot`) | 已实现 | `rustfs/src/admin/handlers/cluster_snapshot.rs` |
|
||||
| Bucket-level remote targets (list/metrics/set/remove) | 已实现 | `rustfs/src/admin/handlers/replication.rs` |
|
||||
| Site replication (add/remove/info/status/peer/resync + devnull/netperf) | 已实现 | `rustfs/src/admin/handlers/site_replication.rs` |
|
||||
| Admin profiling (`/debug/pprof/profile`, `/debug/pprof/status`) | 已实现 | `rustfs/src/admin/handlers/profile_admin.rs`, `profile.rs` |
|
||||
| TLS debug (`/debug/tls/status`) | 已实现 | `rustfs/src/admin/handlers/tls_debug.rs`, `profile.rs` |
|
||||
| KMS management / dynamic / keys | 已实现 | `rustfs/src/admin/handlers/kms_management.rs`, `kms_dynamic.rs`, `kms_keys.rs` |
|
||||
| OIDC public + config | 已实现 | `rustfs/src/admin/handlers/oidc.rs` |
|
||||
| Table catalog (Iceberg) | 已实现 | `rustfs/src/admin/handlers/table_catalog/mod.rs` |
|
||||
|
||||
### Registered-but-incomplete
|
||||
|
||||
| MinIO admin family | Status | RustFS landing point |
|
||||
|---|---|---|
|
||||
| Service restart/stop (`POST /v3/service`) | 已注册未完成 | `rustfs/src/admin/handlers/system.rs` — handler returns `NotImplemented`. |
|
||||
| Inspect data (`GET|POST /v3/inspect-data`) | 已注册未完成 | `rustfs/src/admin/handlers/system.rs` — handler returns `NotImplemented`. |
|
||||
|
||||
These registered-but-`NotImplemented` routes are behavior contracts; per the
|
||||
migration rules in [admin-route-action-snapshot.md](admin-route-action-snapshot.md),
|
||||
implementing or removing them is a behavior-change PR.
|
||||
|
||||
---
|
||||
|
||||
## Gaps Only — Missing Admin Endpoints (follow-up checklist)
|
||||
|
||||
The following MinIO admin `/v3/` route families have **no** RustFS registration
|
||||
today. This is the actionable checklist for closing admin-API parity. Verified
|
||||
against `rustfs/src/admin/mod.rs` and `rustfs/src/admin/handlers/` on the branch
|
||||
this doc was written on.
|
||||
|
||||
- [ ] **Server profiling start/stop** — MinIO `/v3/profile` (bulk profiling
|
||||
session). RustFS only exposes `/debug/pprof/profile` and
|
||||
`/debug/pprof/status`, which are a different, single-shot pprof surface.
|
||||
- [ ] **Health info** — MinIO `/v3/healthinfo` (cluster health report / subnet
|
||||
diagnostics). No RustFS route.
|
||||
- [ ] **LDAP / generic IDP config CRUD** — MinIO `/v3/idp/{ldap|openid}/...`
|
||||
config management. RustFS exposes OIDC config under `/v3/oidc/*` only; there
|
||||
is no LDAP IDP config route.
|
||||
- [ ] **Bucket / site replication diff** — MinIO replication-diff endpoints.
|
||||
RustFS exposes `/v3/replicationmetrics` (metrics) and site-replication
|
||||
status, but no per-object diff.
|
||||
- [ ] **MRF metrics** — MinIO's most-recent-failures replication metrics
|
||||
breakdown. RustFS has only the generic `/v3/metrics` stream.
|
||||
- [ ] **Batch jobs** — MinIO `/v3/batch`, `/v3/list-batch-jobs`, job
|
||||
describe/cancel. No RustFS batch API.
|
||||
- [ ] **Distributed locks introspection** — MinIO `/v3/force-unlock` and
|
||||
`/v3/top/locks`. No RustFS locks-management API.
|
||||
- [ ] **Speedtest / perf** — MinIO `/v3/speedtest` (object/drive/net perf).
|
||||
RustFS has `netperf`/`devnull` **only** inside the site-replication family,
|
||||
not as standalone admin speedtest endpoints.
|
||||
- [ ] **Console log stream** — MinIO `/v3/log` (kstream / log search). No RustFS
|
||||
route.
|
||||
- [ ] **Top introspection** — MinIO `/v3/top/locks`, `/v3/top/drives`,
|
||||
`/v3/top/net`. No RustFS unified `top` family.
|
||||
- [ ] **Trace stream** — MinIO `/v3/trace`. A `trace.rs` handler skeleton
|
||||
exists under `rustfs/src/admin/handlers/` but its registration function is
|
||||
**not** called from `register_admin_routes`, so no route is live.
|
||||
|
||||
When one of these lands, register it in `rustfs/src/admin/mod.rs`, extend
|
||||
`rustfs/src/admin/route_registration_test.rs`, update
|
||||
[admin-route-action-snapshot.md](admin-route-action-snapshot.md) with the
|
||||
route/handler/action rows, and move the item out of this checklist.
|
||||
When an exception above changes state, edit its row here in the same PR that changes the handler, and extend `rustfs/src/admin/route_registration_test.rs` and [admin-route-action-snapshot.md](admin-route-action-snapshot.md) for admin routes. Do not add "implemented" rows to this document; absence from this list is the implemented claim.
|
||||
|
||||
@@ -1,65 +1,35 @@
|
||||
# Observability ECStore Dependency Inventory
|
||||
|
||||
This inventory closes the first `rustfs/backlog#735` step: make every
|
||||
observability dependency on ECStore visible before introducing traits or moving
|
||||
dependency direction.
|
||||
**Use this when:** adding, removing, or moving any `rustfs_ecstore` or `rustfs_storage_api` reference inside `crates/obs`.
|
||||
**Source of truth:** the `use` block at the top of `crates/obs/src/metrics/storage_api.rs`; the guard in `scripts/check_architecture_migration_rules.sh`.
|
||||
|
||||
No behavior or crate movement is planned in inventory PRs. The current boundary
|
||||
is `crates/obs/src/metrics/storage_api.rs`; all direct `rustfs_ecstore` and
|
||||
`rustfs_storage_api` source references in `rustfs-obs` must stay in that file
|
||||
until the contracts below are extracted.
|
||||
`rustfs-obs` still depends on `rustfs-ecstore` (`crates/obs/Cargo.toml`). Every direct reference is confined to one boundary file so the dependency can later be replaced by provider traits without touching collectors.
|
||||
|
||||
## Dependency Inventory
|
||||
|
||||
| Current symbol in `crates/obs/src/metrics/storage_api.rs` | Consumed by | Classification | Purpose |
|
||||
|---|---|---|---|
|
||||
| `rustfs_ecstore::api::storage::ECStore` as `ObsStore` | `stats_collector.rs` | Type dependency | Concrete object-store handle used to call storage admin methods and data-usage loaders. |
|
||||
| `rustfs_storage_api::{BucketOperations, BucketOptions, StorageAdminApi}` | `stats_collector.rs` | Type and trait dependency | Method-resolution and associated type contracts for bucket listing, backend info, and storage info. |
|
||||
| `rustfs_ecstore::api::runtime::object_store_handle` | `stats_collector.rs` | Runtime dependency | Resolves the currently published object-store handle for metric collection. |
|
||||
| `rustfs_ecstore::api::data_usage::load_data_usage_from_backend` | `stats_collector.rs` | Behavior dependency | Loads bucket/object usage and is projected into obs-local DTOs before collectors consume it. |
|
||||
| `rustfs_ecstore::api::capacity::{get_total_usable_capacity, get_total_usable_capacity_free}` | `stats_collector.rs` | Behavior dependency | Computes usable and free capacity from ECStore storage info. |
|
||||
| `rustfs_ecstore::api::bucket::metadata_sys::get_quota_config` | `stats_collector.rs` | Behavior dependency | Reads per-bucket quota limits used in bucket usage metrics. |
|
||||
| `rustfs_ecstore::api::bucket::bandwidth::monitor::Monitor` | `runtime_sources.rs`, `stats_collector.rs` | Type and runtime dependency | Reads replication bandwidth reports from the global bucket monitor handle. |
|
||||
| `rustfs_ecstore::api::runtime::bucket_monitor` | `runtime_sources.rs` | Runtime dependency | Resolves the global bucket bandwidth monitor for metric collection. |
|
||||
| `rustfs_ecstore::api::bucket::replication::get_global_replication_stats` | `storage_api.rs` snapshot helpers | Runtime dependency | Reads replication status, transfer, failure, and site-replication stats, then projects them into obs-local snapshot DTOs. |
|
||||
| `rustfs_ecstore::api::bucket::lifecycle::bucket_lifecycle_ops::{GLOBAL_ExpiryState, GLOBAL_TransitionState}` | `runtime_sources.rs`, `stats_collector.rs` | Runtime dependency | Reads lifecycle expiry and transition queue counters. |
|
||||
| `rustfs_ecstore::api::error::Result` as `ObsEcstoreResult` | `stats_collector.rs` | Type dependency | Preserves ECStore error propagation while data-usage behavior remains ECStore-owned. |
|
||||
The authoritative list is the `pub(crate) use rustfs_ecstore::api::...` block in `crates/obs/src/metrics/storage_api.rs`; it is not copied here. Each import belongs to one of three coupling categories:
|
||||
|
||||
## Classification
|
||||
| Category | Covers | Examples (aliases defined in the boundary file) |
|
||||
|---|---|---|
|
||||
| Type coupling | Concrete ECStore types and storage-api traits used for method resolution | `ObsStore`, `ObsEcstoreResult`, `ObsBucketBandwidthMonitor`, the `rustfs_storage_api` trait imports |
|
||||
| Runtime handle coupling | Resolving process-wide handles for metric collection | object-store handle, bucket monitor, expiry and transition state handles (`rustfs_ecstore::api::runtime::*`), replication stats read inside the snapshot helpers |
|
||||
| Behavior coupling | ECStore-owned computations whose output is projected into obs-local DTOs | data-usage loading, compression totals, quota lookup, usable-capacity math |
|
||||
|
||||
The remaining coupling is not just a dependency declaration problem:
|
||||
|
||||
- type coupling: `ObsStore`, `ObsEcstoreResult`, `ObsBucketBandwidthMonitor`,
|
||||
`StorageAdminApi`, `BucketOperations`, and `BucketOptions`;
|
||||
- runtime handle coupling: object-store handle, bucket monitor, replication
|
||||
stats inside snapshot helpers, expiry state, and transition state;
|
||||
- behavior coupling: data-usage loading, quota lookup, and capacity math.
|
||||
|
||||
Removing `rustfs-ecstore` from `crates/obs/Cargo.toml` is unsafe until those
|
||||
three categories have replacement contracts and compile coverage.
|
||||
Collectors consume only the aliases and the obs-local DTOs. Removing `rustfs-ecstore` from `crates/obs/Cargo.toml` is unsafe until all three categories have replacement contracts and compile coverage.
|
||||
|
||||
## Extraction Plan
|
||||
|
||||
1. Keep all direct ECStore and storage-api imports centralized in
|
||||
`crates/obs/src/metrics/storage_api.rs`.
|
||||
2. Keep projecting ECStore data-usage and replication stats output into
|
||||
obs-local DTOs before collectors consume it.
|
||||
3. Introduce obs-owned provider traits for storage info, bucket info, quota,
|
||||
data usage, replication, bandwidth, and lifecycle queue snapshots.
|
||||
4. Implement those traits in ECStore or an ECStore-owned adapter crate after the
|
||||
trait shapes are covered by focused tests.
|
||||
5. Remove the `rustfs-ecstore` dependency from `rustfs-obs` only after metrics
|
||||
behavior is unchanged through the provider traits.
|
||||
1. Keep all direct ECStore and storage-api imports centralized in `crates/obs/src/metrics/storage_api.rs`.
|
||||
2. Keep projecting ECStore data-usage and replication stats into obs-local DTOs before collectors consume them.
|
||||
3. Introduce obs-owned provider traits for storage info, bucket info, quota, data usage, replication, bandwidth, and lifecycle queue snapshots.
|
||||
4. Implement those traits in ECStore or an ECStore-owned adapter crate once the trait shapes are covered by focused tests.
|
||||
5. Remove the `rustfs-ecstore` dependency from `rustfs-obs` only after metrics behavior is unchanged through the provider traits.
|
||||
|
||||
## Guardrails
|
||||
|
||||
The architecture guard enforces this inventory boundary:
|
||||
Enforced by `scripts/check_architecture_migration_rules.sh`:
|
||||
|
||||
- `crates/obs/src/metrics/storage_api.rs` is the only `rustfs-obs` source file
|
||||
allowed to reference `rustfs_ecstore` or `rustfs_storage_api`;
|
||||
- raw replication stats handles and ECStore replication stat methods must stay
|
||||
behind the snapshot helpers in `crates/obs/src/metrics/storage_api.rs`;
|
||||
- `rustfs-obs` must not add `storage_compat.rs` or `ecstore_compat.rs`
|
||||
passthrough bridges;
|
||||
- future extraction PRs must update this inventory and the guard in the same
|
||||
reviewed change when a dependency category is removed.
|
||||
- `crates/obs/src/metrics/storage_api.rs` is the only `rustfs-obs` source file allowed to reference `rustfs_ecstore` or `rustfs_storage_api`.
|
||||
- Raw replication stats handles and ECStore replication stat methods stay behind the snapshot helpers in that file.
|
||||
- `rustfs-obs` must not add passthrough bridge modules (a second `storage_api.rs`, an `ecstore_compat.rs`, or similar) that re-export ECStore items to other crates.
|
||||
- An extraction PR that removes a dependency category updates this inventory and the guard in the same change.
|
||||
|
||||
@@ -1,60 +1,20 @@
|
||||
# RustFS Architecture Evolution
|
||||
|
||||
This document set tracks the architecture migration from
|
||||
[`rustfs/backlog#660`](https://github.com/rustfs/backlog/issues/660).
|
||||
**Use this when:** you need the historical framing of the architecture-migration program or the phase order that the migration contracts assume.
|
||||
**Source of truth:** [README.md](README.md) is the index of architecture documents; the per-topic contracts it lists are authoritative.
|
||||
|
||||
## Baseline
|
||||
|
||||
- Baseline branch: `upstream/main`
|
||||
- Baseline commit: `61f0dfbc40f748be313be84d834d8259cf3e19c9`
|
||||
- Baseline title: `fix(ecstore): invalidate wiped disk id cache (#3251)`
|
||||
- First migration PR type: `docs-only`
|
||||
The architecture-migration program (`rustfs/backlog#660`) closed in 2026-07. Its original baseline commit predates the current `main` lineage and is no longer reachable from `main`; treat it as historical. The guardrails the program introduced remain enforced by `scripts/check_architecture_migration_rules.sh`.
|
||||
|
||||
## Core Principle
|
||||
|
||||
Cut wrong dependency directions with directories and contracts first, migrate global
|
||||
state in small steps next, and split crates only after boundaries are stable. Storage
|
||||
hot-path behavior must not drift during this migration.
|
||||
|
||||
## Architecture Documents
|
||||
|
||||
- [`runtime-lifecycle.md`](runtime-lifecycle.md): runtime, AppContext,
|
||||
startup/readiness, and shutdown contracts.
|
||||
- [`readiness-matrix.md`](readiness-matrix.md): request-surface behavior,
|
||||
runtime dependency readiness, probe semantics, and preservation rules.
|
||||
- [`s3-tables-support-matrix.md`](s3-tables-support-matrix.md): supported,
|
||||
preview, reference-only, and not-claimed S3 Tables and Iceberg REST Catalog
|
||||
surfaces.
|
||||
- [`storage-control-data-plane.md`](storage-control-data-plane.md): boundaries
|
||||
between StorageCore, ECStore, ClusterControlPlane, and BackgroundControllers.
|
||||
- [`background-services-inventory.md`](background-services-inventory.md): current
|
||||
scanner, heal, lifecycle, replication, config reload, metrics, and shutdown
|
||||
surface before BackgroundController work.
|
||||
- [`background-controller-contract.md`](background-controller-contract.md):
|
||||
desired/current/status/reconcile vocabulary and lifecycle boundaries for
|
||||
future read-only BackgroundController work.
|
||||
- [`crate-boundaries.md`](crate-boundaries.md): PR types, crate direction,
|
||||
compatibility rules, and migration guardrails.
|
||||
- [`global-state-crate-split-plan.md`](global-state-crate-split-plan.md): late
|
||||
global-state cleanup, runtime-source boundaries, fallback removal rules, and
|
||||
crate-split evaluation criteria.
|
||||
- [`obs-ecstore-dependency-inventory.md`](obs-ecstore-dependency-inventory.md):
|
||||
observability-to-ECStore dependency inventory, classification, and extraction
|
||||
guardrails.
|
||||
- [`ecstore-config-consumer-inventory.md`](ecstore-config-consumer-inventory.md):
|
||||
current `ecstore::config::{Config, KV, KVS}` definitions, consumers,
|
||||
migration risks, and do-not-change contract.
|
||||
- [`ecstore-api-facade-inventory.md`](ecstore-api-facade-inventory.md): current
|
||||
`rustfs_ecstore::api` facade groups, external consumer boundaries, shrink
|
||||
rules, and split dependency inventory.
|
||||
- [`config-model-boundary-adr.md`](config-model-boundary-adr.md): target crate,
|
||||
module path, dependency rules, and verification gates for moving the pure
|
||||
server-config model.
|
||||
- [`compat-cleanup-register.md`](compat-cleanup-register.md): temporary
|
||||
compatibility code that must be removed later.
|
||||
Cut wrong dependency directions with directories and contracts first, migrate global state in small steps next, and split crates only after boundaries are stable. Storage hot-path behavior must not drift during this migration.
|
||||
|
||||
## Phase Order
|
||||
|
||||
Historical sequencing of the migration phases. All phases are closed; the diagram is kept because later documents refer to phase names.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
G["Phase 0: Baseline and guardrails"]
|
||||
@@ -80,10 +40,4 @@ flowchart LR
|
||||
GS --> CR
|
||||
```
|
||||
|
||||
The first implementation sequence is conservative:
|
||||
|
||||
1. Record baseline and migration context.
|
||||
2. Establish PR and compatibility rules.
|
||||
3. Add dependency and loss-prevention checks in a separate `ci-gate` PR.
|
||||
4. Inventory `ecstore::config::{Config, KV, KVS}` before moving any code.
|
||||
5. Decide the config model boundary before extracting or migrating consumers.
|
||||
The document index is [README.md](README.md). The ECStore facade boundary that the storage phases converged on is described in [ecstore-api-facade-inventory.md](ecstore-api-facade-inventory.md).
|
||||
|
||||
@@ -1,53 +1,38 @@
|
||||
# Placement And Repair Invariants
|
||||
|
||||
This inventory covers `G-012` for `rustfs/backlog#666`. It records the current
|
||||
object placement, readiness, lock quorum, scanner, and repair boundaries that
|
||||
later scheduler or topology work must preserve.
|
||||
**Use this when:** changing anything that resolves an object to a pool, set, or disk, or that admits scanner or heal work; these are the behaviors later scheduler or topology work must preserve.
|
||||
**Source of truth:** `Sets::get_disks_by_key` / `get_hashed_set_index` in `crates/ecstore/src/core/sets.rs`; `DistributionAlgoVersion` in `crates/ecstore/src/layout/format.rs`; `crc_hash` / `sip_hash` in `crates/utils/src/hash.rs`; `ScannerCycleBudget` in `crates/scanner/src/scanner_budget.rs`; `crates/scanner/src/scanner_heal_admission_baseline.rs`.
|
||||
|
||||
## Object To Set Hash Rule
|
||||
|
||||
Objects reach a set through `Sets::get_disks_by_key`, which calls
|
||||
`get_hashed_set_index` on the object key:
|
||||
Objects reach a set through `Sets::get_disks_by_key`, which calls `get_hashed_set_index` on the object key:
|
||||
|
||||
- `DistributionAlgoVersion::V1` uses `crc_hash(input, set_count)`.
|
||||
- `DistributionAlgoVersion::V2` and `V3` use
|
||||
`sip_hash(input, set_count, format_id_bytes)`.
|
||||
- The format ID is part of the V2/V3 distribution seed, so changing the seed,
|
||||
object key, set count, or algorithm changes placement.
|
||||
- `DistributionAlgoVersion::V2` and `V3` use `sip_hash(input, set_count, format_id_bytes)`.
|
||||
- The format ID is part of the V2/V3 distribution seed, so changing the seed, object key, set count, or algorithm changes placement.
|
||||
|
||||
Preservation rule: every object read, write, list, heal, repair, and
|
||||
decommission path that resolves a set for an existing object must preserve the
|
||||
same object key and format distribution algorithm.
|
||||
Preservation rule: every object read, write, list, heal, repair, and decommission path that resolves a set for an existing object must preserve the same object key and format distribution algorithm.
|
||||
|
||||
## Pool, Set, And Disk Assignment Boundary
|
||||
|
||||
Pool selection is separate from set hashing:
|
||||
|
||||
- Existing objects are discovered across pools and resolved to the best current
|
||||
pool candidate before reads or updates continue.
|
||||
- New object writes select an available pool from current per-pool free-space
|
||||
inputs after suspended or rebalancing pools are skipped.
|
||||
- Set selection inside a pool still uses the object-to-set hash rule above.
|
||||
- Disk index assignment comes from endpoint and format metadata, not from a
|
||||
scheduler decision.
|
||||
- Existing objects are discovered across pools and resolved to the best current pool candidate before reads or updates continue.
|
||||
- New object writes select an available pool from current per-pool free-space inputs after suspended or rebalancing pools are skipped.
|
||||
- Set selection inside a pool uses the object-to-set hash rule above.
|
||||
- Disk index assignment comes from endpoint and format metadata, not from a scheduler decision.
|
||||
|
||||
Boundary rule: schedulers may influence admission, worker concurrency, or
|
||||
buffer sizing, but they must not rewrite pool, set, or disk indexes.
|
||||
Boundary rule: schedulers may influence admission, worker concurrency, or buffer sizing, but they must not rewrite pool, set, or disk indexes.
|
||||
|
||||
## Readiness And Lock Quorum Boundary
|
||||
|
||||
Runtime readiness currently checks storage and lock health independently:
|
||||
Runtime readiness checks storage and lock health independently:
|
||||
|
||||
- Storage readiness requires every observed set to meet write quorum based on
|
||||
the set drive count and storage class data/parity shape.
|
||||
- Lock readiness aggregates per-set lock-client host quorum and fails fast if
|
||||
any set loses quorum.
|
||||
- Object and bucket mutations acquire namespace locks through the existing
|
||||
storage lock wrappers before changing object or bucket state.
|
||||
- Storage readiness requires every observed set to meet write quorum based on the set drive count and storage class data/parity shape.
|
||||
- Lock readiness aggregates per-set lock-client host quorum and fails fast if any set loses quorum.
|
||||
- Object and bucket mutations acquire namespace locks through the existing storage lock wrappers before changing object or bucket state.
|
||||
|
||||
Boundary rule: readiness and lock quorum must stay set-aware. A global healthy
|
||||
disk count or global connected-host count is not sufficient when any individual
|
||||
set is below quorum.
|
||||
Boundary rule: readiness and lock quorum must stay set-aware. A global healthy disk count or global connected-host count is not sufficient when any individual set is below quorum.
|
||||
|
||||
## Scanner Budget Preservation
|
||||
|
||||
@@ -55,43 +40,45 @@ Scanner cycles are bounded by `ScannerCycleBudget`:
|
||||
|
||||
- Runtime budget cancels the child token after the configured duration.
|
||||
- Object budget cancels after the configured object count.
|
||||
- Directory budget rejects additional directories and cancels with the
|
||||
directories reason.
|
||||
- Directory budget rejects additional directories and cancels with the directories reason.
|
||||
- Partial-cycle metrics and checkpoints use the budget reason.
|
||||
|
||||
Preservation rule: later scheduler work can change how scan cycles are admitted
|
||||
only if it preserves the budget reason, checkpoint reason, and child-token
|
||||
cancellation behavior.
|
||||
Preservation rule: later scheduler work can change how scan cycles are admitted only if it preserves the budget reason, checkpoint reason, and child-token cancellation behavior.
|
||||
|
||||
## Heal Admission Preservation
|
||||
|
||||
Scanner and background repair work enter the heal manager through explicit
|
||||
admission:
|
||||
Scanner and background repair work enter the heal manager through explicit admission:
|
||||
|
||||
- Scanner object heal requests are low priority and may be accepted, merged,
|
||||
rejected as full, or dropped.
|
||||
- Required/high-priority heal candidates escalate on non-admission instead of
|
||||
silently disappearing.
|
||||
- Heal queue admission deduplicates queued and active work unless the request
|
||||
explicitly forces admission.
|
||||
- Full queues can drop low-priority work or displace lower-priority work for a
|
||||
higher-priority request according to current manager rules.
|
||||
- Scanner object heal requests are low priority and may be accepted, merged, rejected as full, or dropped.
|
||||
- Required/high-priority heal candidates escalate on non-admission instead of silently disappearing.
|
||||
- Heal queue admission deduplicates queued and active work unless the request explicitly forces admission.
|
||||
- Full queues can drop low-priority work or displace lower-priority work for a higher-priority request according to current manager rules.
|
||||
|
||||
Preservation rule: repair scheduling changes must keep admission outcomes
|
||||
observable and must not convert rejected or dropped repair work into silent
|
||||
success.
|
||||
Preservation rule: repair scheduling changes must keep admission outcomes observable and must not convert rejected or dropped repair work into silent success.
|
||||
|
||||
### Scanner/heal admission entry points
|
||||
|
||||
No cluster-wide coordinator or second generation token exists for scanner/heal admission; each entry point keeps its own guard. `crates/scanner/src/scanner_heal_admission_baseline.rs` `include_str!`s the scanner sources and asserts the named guards are still present, so a rename or guard removal fails that test instead of silently leaving this table stale. It also encodes the investigation matrix (scanner read and heal read may overlap; heal write conflicts with scanner reads; data-movement write conflicts with all work; independent sets stay concurrent) without claiming production enforces it.
|
||||
|
||||
| Work | Entry point | Current guard | Fallback / namespace semantics |
|
||||
|---|---|---|---|
|
||||
| Scanner read/list | `nsscanner_disk` in `crates/scanner/src/scanner_io/io_disk.rs` | Per-disk `start_scan()` guard; bucket lifecycle/replication/object-lock reads precede `scan_data_folder` | Scanner keeps its local disk and durable cursor; no HealManager set-level admission is consulted |
|
||||
| Scanner metadata read | Object-size and metadata branches in `crates/scanner/src/scanner_folder.rs` | Scanner cycle budget and per-disk scan marker | Corrupt metadata records the pending scanner ledger; MRF is a hint, not the durable owner |
|
||||
| Scanner heal admission | `send_required_scanner_heal_request` in `crates/scanner/src/scanner_folder.rs` | Manager queue dedup and pending ledger (`update_pending_scanner_heal_after_admission`) | MRF `Enqueued` / `Coalesced` is ledger-only; rejected MRF keeps immediate heal plus ledger |
|
||||
| Heal auto scan | `start_auto_disk_scanner` in `crates/heal/src/heal/manager/auto_scan.rs` | Queue-first then active-task check; replacement recovery blocklist | Scanning disks remain candidates when degraded quorum needs them; they are not globally excluded |
|
||||
| Heal object read | `heal_object` in `crates/ecstore/src/set_disk/ops/heal.rs` | Namespace write lock (`get_write_lock`) unless `no_lock`; reads file info before commit | The namespace lock is object-scoped and does not claim scanner cycle ownership |
|
||||
| Disk selection | `get_online_disks_with_healing_and_info` in `crates/ecstore/src/set_disk/ops/locking.rs` | Healing disks are ordered after new disks; scanning disks may remain candidates | Degraded/quorum fallback is preserved |
|
||||
| Data movement | `wait_for_data_movement_admission` in `crates/ecstore/src/data_movement/backpressure.rs` | Storage-owned backpressure on foreground pressure; no second coordinator | Any future admission token must be validated at the final metadata/format/delete commit (see [unified-object-generation.md](unified-object-generation.md)) |
|
||||
|
||||
Rules: cancellation or a local lease alone is not a fence; if a fixture ever demonstrates a stale destructive write, the fix extends the storage-owned generation/admission primitive and validates the token at the final commit rather than adding a coordinator.
|
||||
|
||||
## Behavior Change Gates
|
||||
|
||||
Any later placement or repair PR must use the following gates:
|
||||
|
||||
- Placement gate: prove object-to-set hashing is unchanged for existing object
|
||||
keys and format algorithms.
|
||||
- Pool gate: prove pool selection does not choose suspended or rebalancing
|
||||
pools unless the existing path already allows it.
|
||||
- Placement gate: prove object-to-set hashing is unchanged for existing object keys and format algorithms.
|
||||
- Pool gate: prove pool selection does not choose suspended or rebalancing pools unless the existing path already allows it.
|
||||
- Quorum gate: prove storage readiness and lock readiness remain per-set.
|
||||
- Scanner gate: prove scan budget reason and checkpoint mapping remain stable.
|
||||
- Heal gate: prove low-priority scanner heal, forced heal, duplicate merge, and
|
||||
queue-full outcomes remain distinct.
|
||||
- Rollback gate: if a new scheduler sidecar is disabled, placement and repair
|
||||
must fall back to the current direct ECStore/scanner/heal behavior.
|
||||
- Heal gate: prove low-priority scanner heal, forced heal, duplicate merge, and queue-full outcomes remain distinct.
|
||||
- Rollback gate: if a new scheduler sidecar is disabled, placement and repair must fall back to the current direct ECStore/scanner/heal behavior.
|
||||
|
||||
@@ -1,8 +1,7 @@
|
||||
# Readiness Matrix
|
||||
|
||||
This document records the current request and dependency behavior around
|
||||
startup readiness. It is a behavior-preservation baseline for architecture
|
||||
migration work, not a new readiness policy.
|
||||
**Use this when:** changing what a request surface does before storage or IAM is ready, changing probe semantics, or adding a runtime dependency that readiness must wait for.
|
||||
**Source of truth:** `rustfs/src/server/readiness.rs` (probe paths, `Retry-After`), `crates/common/src/readiness.rs` (`StorageReady`, `IamReady`, `FullReady`), `crates/config/src/constants/health.rs` (`RUSTFS_HEALTH_*` gates). This matrix is a behavior-preservation baseline, not a new readiness policy.
|
||||
|
||||
## Request Behavior Matrix
|
||||
|
||||
@@ -23,7 +22,7 @@ migration work, not a new readiness policy.
|
||||
| IamReady | Inline IAM bootstrap or deferred IAM recovery publication. | Yes. | Deferred recovery can publish IAM readiness after HTTP has already started. |
|
||||
| Lock quorum | Per-set write quorum readiness. | Yes. | Do not replace the distributed lock quorum check with node count or endpoint count. |
|
||||
| Peer health | `peer_health_ready` runtime status. | Only when `RUSTFS_HEALTH_PEER_READY_CHECK_ENABLE` is enabled. | The gate is disabled by default; unknown peer health degrades readiness only when enabled. |
|
||||
| KMS compatibility | KMS health compatibility readiness. | Only when the KMS compatibility readiness check is enabled. | KMS startup fatality and health reporting remain separate from pure docs work. |
|
||||
| KMS compatibility | KMS health compatibility readiness. | Only when `RUSTFS_HEALTH_COMPAT_KMS_READY_CHECK_ENABLE` is enabled (default off; `crates/config/src/constants/health.rs`). | When enabled, `/health/ready` additionally requires the KMS service to be running if a global KMS manager exists. KMS startup fatality and health reporting remain separate from pure docs work. |
|
||||
|
||||
Effective `FullReady` is:
|
||||
|
||||
|
||||
@@ -1,70 +1,28 @@
|
||||
# Runtime Capability Contracts
|
||||
|
||||
This document records the `rustfs/backlog#660` PR-08 and PR-09 contract slice.
|
||||
It adds read-only observability and topology snapshot shapes to
|
||||
`rustfs-storage-api` without coupling the contract crate to runtime, ECStore,
|
||||
admin routes, profiling, or observability implementation crates.
|
||||
**Use this when:** changing the read-only observability or topology snapshot contracts in `rustfs-storage-api`, their RustFS providers, or the `storage_classes` payload of `GET /rustfs/admin/v4/runtime/capabilities`.
|
||||
**Source of truth:** `ObservabilitySnapshot` in `crates/storage-api/src/observability.rs`; `TopologySnapshot` in `crates/storage-api/src/topology.rs`; `CapabilityState` and `CapabilitySnapshotError` in `crates/storage-api/src/capability.rs`; providers in `rustfs/src/runtime_capabilities.rs`; storage-class constants in `crates/ecstore/src/config/storageclass.rs`.
|
||||
|
||||
## Observability Snapshot Contract
|
||||
## Snapshot Contracts
|
||||
|
||||
`ObservabilitySnapshot` records:
|
||||
Field lists live on the defining types and are not repeated here.
|
||||
|
||||
- Runtime telemetry capability state.
|
||||
- Userspace CPU and memory profiling capability state.
|
||||
- Process, system, and cgroup memory sampling state.
|
||||
- Platform support for target triple, OS, architecture, allocator, eBPF, and
|
||||
NUMA capability.
|
||||
| Contract | Defining type | RustFS provider | Rule |
|
||||
|---|---|---|---|
|
||||
| Observability | `ObservabilitySnapshot` | `RustFsObservabilitySnapshotProvider` | Reports runtime telemetry, profiling, memory-sampling, platform, allocator, eBPF, and NUMA capability as `CapabilityState` values without starting telemetry, profiling, allocator reclaim, or memory-observability workers. |
|
||||
| Topology | `TopologySnapshot` | `EndpointTopologySnapshotProvider` | Maps `EndpointServerPools` into pool/set/disk indexes, optional stable IDs, and optional zone/rack/node/media/NUMA labels without changing endpoint construction, placement, readiness, locks, or ECStore metadata. Local file endpoint paths are never used as disk IDs or labels; extra labels go in the `additional` map so future inventory labels need no ECStore type leakage. |
|
||||
|
||||
Unsupported, disabled, and unknown states are represented by `CapabilityState`
|
||||
instead of failing snapshot construction. The contract is intentionally read
|
||||
only and does not replace existing profiling routes, telemetry APIs, exporter
|
||||
pipelines, or startup behavior.
|
||||
|
||||
## Topology Snapshot Contract
|
||||
|
||||
`TopologySnapshot` records:
|
||||
|
||||
- Pool, set, and disk identity indexes plus optional stable IDs.
|
||||
- Optional zone, rack, node, media, NUMA, and additional labels.
|
||||
- Topology-wide profiling, NUMA, failure-domain label, and media-label
|
||||
capability states.
|
||||
- Per-disk media, failure-domain, NUMA, and profiling capability states.
|
||||
|
||||
Missing labels are represented as absent `Option` values. Extra topology labels
|
||||
belong in the `additional` label map, so future inventory labels do not require
|
||||
ECStore type leakage.
|
||||
Unsupported, disabled, and unknown states are values of `CapabilityState`, not construction failures. Missing labels are `None`. Providers map implementation failures into `CapabilitySnapshotError` before crossing the contract boundary. Neither contract replaces existing profiling routes, telemetry APIs, exporter pipelines, or startup behavior.
|
||||
|
||||
## Boundary Rules
|
||||
|
||||
- No `rustfs-ecstore`, `rustfs-obs`, Axum, KMS, admin route, OTEL, eBPF, or
|
||||
profiling implementation dependency is added to `rustfs-storage-api`.
|
||||
- No placement, membership, NUMA pinning, profiling, startup, admin route, or
|
||||
exporter behavior changes are part of this contract slice.
|
||||
- Providers must map implementation failures into `CapabilitySnapshotError`
|
||||
before crossing the contract boundary.
|
||||
|
||||
## RustFS Provider Slice
|
||||
|
||||
`rustfs/src/runtime_capabilities.rs` wires the contracts to RustFS runtime
|
||||
owners through read-only providers:
|
||||
|
||||
- `RustFsObservabilitySnapshotProvider` maps current dial9, profiling,
|
||||
memory-sampling, platform, allocator, eBPF, and NUMA capability state without
|
||||
starting telemetry, profiling, allocator reclaim, or memory-observability
|
||||
workers.
|
||||
- `EndpointTopologySnapshotProvider` maps `EndpointServerPools` into pool, set,
|
||||
and disk topology snapshots without changing endpoint construction,
|
||||
placement, readiness, locks, or ECStore metadata. Local file endpoint paths are
|
||||
intentionally not used as disk IDs or labels.
|
||||
|
||||
Unsupported or unavailable runtime capabilities are reported as `unsupported`
|
||||
or `unknown` contract states instead of activating fallback behavior.
|
||||
- `rustfs-storage-api` gains no dependency on `rustfs-ecstore`, `rustfs-obs`, Axum, KMS, admin routes, OTEL, eBPF, or profiling implementation crates.
|
||||
- Providers are read-only. Adding or changing a provider changes no placement, membership, NUMA pinning, profiling, startup, admin-route, or exporter behavior.
|
||||
- Unsupported or unavailable runtime capabilities are reported as `unsupported` or `unknown`; they never activate fallback behavior.
|
||||
|
||||
## Storage-Class Write Contract
|
||||
|
||||
Authenticated clients discover the storage-class write contract from
|
||||
`GET /rustfs/admin/v4/runtime/capabilities`. The additive
|
||||
`storage_classes` object is versioned independently from the route:
|
||||
Authenticated clients discover the storage-class write contract from `GET /rustfs/admin/v4/runtime/capabilities`. The additive `storage_classes` object is versioned independently from the route:
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -77,15 +35,11 @@ Authenticated clients discover the storage-class write contract from
|
||||
}
|
||||
```
|
||||
|
||||
`supported_write_classes` is the complete client-selectable write allowlist.
|
||||
Any other value fails before object or multipart mutation with the stable S3
|
||||
error named by `unsupported_write_error`. `legacy_label_behavior` means
|
||||
non-transitioned historical label-only metadata is reported as its effective
|
||||
local class; actual lifecycle transition tier names remain unchanged.
|
||||
| Field | Meaning |
|
||||
|---|---|
|
||||
| `supported_write_classes` | The complete client-selectable write allowlist. Any other value fails before object or multipart mutation with the S3 error named by `unsupported_write_error`. |
|
||||
| `unsupported_write_error` | Stable S3 error code (`UNSUPPORTED_WRITE_ERROR` in `crates/ecstore/src/config/storageclass.rs`). |
|
||||
| `legacy_label_behavior` | Non-transitioned historical label-only metadata is reported as its effective local class; lifecycle transition tier names are unchanged. |
|
||||
| `contract_version` | Consumers must branch on it before assigning meaning to future fields. |
|
||||
|
||||
The values are sourced from
|
||||
[`crates/ecstore/src/config/storageclass.rs`](../../crates/ecstore/src/config/storageclass.rs),
|
||||
which also owns write validation and response normalization. Consumers must
|
||||
branch on `contract_version` before assigning meaning to future fields. The
|
||||
admin route continues to require `ServerInfoAdminAction`; capability discovery
|
||||
does not weaken authentication or authorization.
|
||||
Values, write validation, and response normalization are owned by `crates/ecstore/src/config/storageclass.rs`. The route continues to require `ServerInfoAdminAction`; capability discovery does not weaken authentication or authorization.
|
||||
|
||||
@@ -1,120 +1,57 @@
|
||||
# Runtime And Lifecycle Contracts
|
||||
|
||||
Runtime and lifecycle work must preserve startup ordering, readiness behavior, and
|
||||
shutdown semantics.
|
||||
**Use this when:** moving or reordering anything in `rustfs/src/startup_*.rs`, changing readiness publication, or touching shutdown ordering.
|
||||
**Source of truth:** the `rustfs/src/startup_*.rs` modules listed below; readiness semantics in [readiness-matrix.md](readiness-matrix.md); global-state targets in [global-state-inventory.md](global-state-inventory.md).
|
||||
|
||||
Runtime and lifecycle work must preserve startup ordering, readiness behavior, and shutdown semantics.
|
||||
|
||||
## Startup And Readiness
|
||||
|
||||
- HTTP can listen early, but normal requests must remain behind readiness gates.
|
||||
- Effective `FullReady = storage_ready && iam_ready && lock_quorum_ready &&
|
||||
peer_health_ready`; `peer_health_ready` is true by default unless
|
||||
`RUSTFS_HEALTH_PEER_READY_CHECK_ENABLE` is enabled.
|
||||
- Boot phases must keep the old fatal and non-fatal boundaries.
|
||||
- AppContext migration keeps context-first lookup with global fallback until the
|
||||
global path is proven unused.
|
||||
- HTTP can listen early, but normal requests stay behind the readiness gate.
|
||||
- The `FullReady` formula, its dependencies, and the `RUSTFS_HEALTH_PEER_READY_CHECK_ENABLE` gate are defined once in [readiness-matrix.md](readiness-matrix.md); do not restate them elsewhere.
|
||||
- Boot phases keep the existing fatal and non-fatal boundaries.
|
||||
- AppContext migration keeps context-first lookup with global fallback until the global path is proven unused.
|
||||
- Notify and audit lifecycle behavior must not drift during lifecycle movement.
|
||||
- IAM and KMS startup, deferred recovery, and fatal boundary behavior must not be
|
||||
changed by pure movement PRs.
|
||||
- Request-surface and dependency details are tracked in
|
||||
[`readiness-matrix.md`](readiness-matrix.md).
|
||||
- IAM and KMS startup, deferred recovery, and fatal-boundary behavior must not be changed by pure movement PRs.
|
||||
|
||||
## Service Registry Scope
|
||||
## Startup Module Ownership
|
||||
|
||||
`ServiceRegistry` is only for lifecycle and shutdown ordering. It must not become a
|
||||
general dependency injection container.
|
||||
Each module owns one concern; orchestration order is owned by `startup_services`, `startup_lifecycle`, and `startup_shutdown`. Movement PRs may pass handles between modules but must not reorder the steps a module owns.
|
||||
|
||||
Allowed responsibilities:
|
||||
|
||||
- Register start and stop order.
|
||||
- Expose read-only status snapshots.
|
||||
- Coordinate graceful shutdown.
|
||||
|
||||
Disallowed responsibilities:
|
||||
|
||||
- Construct arbitrary dependencies for business logic.
|
||||
- Hide globals behind a service-locator API.
|
||||
- Change startup side effects while moving code.
|
||||
| Module (`rustfs/src/`) | Owns |
|
||||
|---|---|
|
||||
| `startup_entrypoint.rs` | CLI command dispatch into preflight and the runtime lifecycle. |
|
||||
| `startup_preflight.rs` | License init, external env compatibility, runtime foundation bootstrap. |
|
||||
| `startup_runtime.rs` | Runtime foundation orchestration; outbound TLS fatal boundary when configured material fails to load. |
|
||||
| `startup_runtime_hooks.rs` | Startup diagnostics, profiling hook dispatch, default crypto provider installation. |
|
||||
| `startup_tls_material.rs` | Outbound TLS material loading, global publication, generation recording, TLS metrics init. |
|
||||
| `startup_runtime_sources.rs` | Process-local runtime source publication (port, buffer profile, KMS manager, TLS generation). |
|
||||
| `startup_fs_guard.rs` | Unsupported-filesystem policy enforcement for endpoint paths. |
|
||||
| `startup_deadlock.rs` | Deadlock detector state logging. |
|
||||
| `startup_server.rs` | HTTP listener start and `ServiceStateManager` publication. |
|
||||
| `startup_storage.rs` | Endpoints, local disks, ECStore, lock clients, global config, background replication init; `StorageReady`. |
|
||||
| `startup_bucket_metadata.rs` | Bucket metadata system init, legacy meta-bucket import (`try_migrate_bucket_metadata`, `try_migrate_iam_config`), resync intents. |
|
||||
| `startup_iam.rs` | IAM init, deferred recovery, `IamReady` publication. |
|
||||
| `startup_auth.rs` | OIDC and federated identity setup. |
|
||||
| `startup_notification.rs` | Notification system and bucket notification configuration. |
|
||||
| `startup_audit.rs` | Event notifier and audit system start. |
|
||||
| `startup_observability.rs` | Auto-tuner, update check, server info, compression totals. |
|
||||
| `startup_background.rs` | Scanner, heal, bitrot self-test, workload-admission provider publication. |
|
||||
| `startup_protocols.rs` | FTP/FTPS/SFTP/WebDAV sidecar start and shutdown senders. |
|
||||
| `startup_optional_runtime_sidecars.rs` | Handles, shutdown planning, and shutdown execution for optional sidecars that are not readiness boundaries (currently protocol servers only). New sidecars enter here with explicit shutdown handles and status snapshots, not ad hoc work in `startup_services`. |
|
||||
| `startup_services.rs` | Orchestration order of runtime service startup: KMS, optional runtimes, audit, metadata, IAM, auth, notification, background services, observability. |
|
||||
| `startup_lifecycle.rs` | Ready publication, global init-time publication, scanner startup, shutdown-signal wait, shutdown delegation, final stopped-state log. |
|
||||
| `startup_shutdown.rs` | The shutdown sequence (see below). |
|
||||
| `startup_embedded.rs`, `startup_embedded_optional.rs` | Embedded-mode reuse of the phase owners above (see below). |
|
||||
|
||||
## Shutdown Lifecycle Boundary
|
||||
|
||||
`startup_shutdown` owns the main shutdown sequence after the process receives a
|
||||
shutdown signal. Startup modules may pass handles into this boundary, but they
|
||||
must not reorder runtime-token cancellation, background service shutdown,
|
||||
optional runtime shutdown planning, notifier/audit/profiling shutdown, HTTP
|
||||
shutdown, optional runtime waits, or final service-state publication.
|
||||
`startup_shutdown` owns the main shutdown sequence after the process receives a shutdown signal. Startup modules may pass handles into this boundary, but they must not reorder runtime-token cancellation, background service shutdown, optional runtime shutdown planning, notifier/audit/profiling shutdown, HTTP shutdown, optional runtime waits, or final service-state publication.
|
||||
|
||||
## Startup Lifecycle Boundary
|
||||
## Embedded Startup Reuse
|
||||
|
||||
`startup_lifecycle` owns the ready-to-shutdown orchestration after runtime
|
||||
services initialize. Service modules may return initialized handles into this
|
||||
boundary, but they must not reorder ready publication, global init-time
|
||||
publication, scanner startup, shutdown-signal wait, shutdown delegation, or the
|
||||
final stopped-state log.
|
||||
|
||||
## Startup Service Component Boundary
|
||||
|
||||
`startup_service_components` owns individual runtime service startup component
|
||||
helpers while `startup_services` preserves their orchestration order. Migration
|
||||
PRs must not change KMS, optional runtime, audit, metadata, IAM, auth,
|
||||
notification, background service, or observability startup ordering while moving
|
||||
these helpers.
|
||||
|
||||
## Optional Runtime Boundary
|
||||
|
||||
`startup_optional_runtime_sidecars` owns startup handles, shutdown planning, and
|
||||
shutdown execution for optional runtime services that are not readiness
|
||||
boundaries. `startup_optional_runtimes` remains a compatibility handoff for the
|
||||
old module path. The current owner set is protocol servers only. Future optional
|
||||
sidecars must enter this boundary with explicit shutdown handles and status
|
||||
snapshots instead of adding ad hoc startup or shutdown work to
|
||||
`startup_services`.
|
||||
|
||||
## Startup Runtime Hook Boundary
|
||||
|
||||
`startup_runtime_hooks` owns runtime hook side effects that wrap the startup
|
||||
foundation but are not TLS material loading: startup diagnostics, profiling
|
||||
hook dispatch, and default crypto provider installation. `startup_runtime`
|
||||
preserves BOOT-006 orchestration and outbound TLS fatal behavior, while
|
||||
`startup_profiling` remains a compatibility handoff for the old profiling hook
|
||||
path.
|
||||
|
||||
## Startup TLS Material Boundary
|
||||
|
||||
`startup_tls_material` owns configured outbound TLS material loading, global
|
||||
outbound TLS publication, generation recording, and TLS metrics initialization.
|
||||
`startup_runtime` still owns BOOT-006 ordering and must preserve the fatal
|
||||
boundary when configured TLS material fails to load.
|
||||
|
||||
## Embedded Startup Phase Reuse
|
||||
|
||||
Embedded startup should use the same startup server and storage phase owners for
|
||||
listen context, endpoint/local disk setup, storage runtime setup, readiness
|
||||
publication, and replication startup. Embedded-specific behavior still owns its
|
||||
stable-port requirement, one-shot global initialization guard placement, S3-only
|
||||
HTTP listener, and non-fatal KMS/audit/notification policy.
|
||||
|
||||
## Embedded Runtime Service Reuse
|
||||
|
||||
Embedded runtime service setup should share startup service helpers for optional
|
||||
service initialization, bucket metadata/IAM setup, notification setup, and
|
||||
shutdown cleanup. Embedded-specific behavior still owns warning-only
|
||||
KMS/audit/notification failures, no binary-only background sidecars, no state
|
||||
manager, and the one-shot server handle cleanup used by embedded shutdown.
|
||||
|
||||
## Embedded Lifecycle Publication Reuse
|
||||
|
||||
Embedded ready publication should share startup lifecycle helpers for IAM
|
||||
readiness publication, global init-time publication, and ready-state logging.
|
||||
Embedded-specific behavior still owns server handle construction, endpoint
|
||||
address normalization, and process-local shutdown cleanup.
|
||||
Embedded startup reuses the same phase owners as the binary: server and storage phases for listen context, endpoint/local disk setup, storage runtime setup, readiness publication, and replication startup; service helpers for optional service init, bucket metadata/IAM setup, notification setup, and shutdown cleanup; lifecycle helpers for IAM readiness publication, global init-time publication, and ready-state logging. Embedded-specific behavior that stays in `startup_embedded*.rs`: stable-port requirement, one-shot global initialization guard placement, S3-only HTTP listener, warning-only KMS/audit/notification failures, no binary-only background sidecars, no state manager, server handle construction, endpoint address normalization, and process-local one-shot shutdown cleanup.
|
||||
|
||||
## AppContext Foundation
|
||||
|
||||
Early AppContext work should split resolver files and add compatibility tests before
|
||||
boot extraction or consumer migration. This keeps the migration context-first while
|
||||
preserving the old global fallback path during transition.
|
||||
|
||||
AppContext remains a context-first facade, not a full replacement for every
|
||||
process global. New migration work must keep fallback reads inside owner-local
|
||||
runtime-source boundaries and follow the global-state target inventory in
|
||||
[`global-state-inventory.md`](global-state-inventory.md).
|
||||
AppContext is a context-first facade, not a full replacement for every process global. Resolver files are split and covered by compatibility tests before boot extraction or consumer migration, so the old global fallback path keeps working during transition. New migration work keeps fallback reads inside owner-local runtime-source boundaries and follows the target inventory in [global-state-inventory.md](global-state-inventory.md).
|
||||
|
||||
@@ -1,44 +1,30 @@
|
||||
# S3 Compatibility Matrix
|
||||
|
||||
This matrix records the user-facing S3 compatibility claim for RustFS and ties
|
||||
it to the executable Ceph s3tests lists under `scripts/s3-tests/`.
|
||||
**Use this when:** writing or checking a user-facing S3 compatibility claim, or moving a Ceph s3tests case between lists.
|
||||
**Source of truth:** the test lists under `scripts/s3-tests/` and the runner `scripts/s3-tests/run.sh`; counts are derived from those files and are not recorded here.
|
||||
|
||||
## Current Claim
|
||||
|
||||
RustFS provides broad S3 API compatibility for supported features. It does not
|
||||
claim complete coverage of every standard or vendor-specific S3 behavior.
|
||||
|
||||
The root README should use the same wording: supported S3-compatible clients and
|
||||
features are covered by the compatibility matrix and test lists.
|
||||
RustFS provides broad S3 API compatibility for supported features. It does not claim complete coverage of every standard or vendor-specific S3 behavior. The root README uses the same wording: supported S3-compatible clients and features are covered by the compatibility matrix and test lists.
|
||||
|
||||
## Test List Sources
|
||||
|
||||
| List | Purpose | Current count | Source |
|
||||
|---|---:|---:|---|
|
||||
| Implemented tests | Standard S3 tests expected to pass and used by the default local s3tests run. | 452 | `scripts/s3-tests/implemented_tests.txt` |
|
||||
| Lifecycle behavior tests | Expiration behavior cases gated by the dedicated `s3-lifecycle-behavior-tests` lane (debug-accelerated day + scanner enabled). | 5 | `scripts/s3-tests/lifecycle_behavior_tests.txt` |
|
||||
| Unimplemented tests | Standard S3 features planned but not yet implemented. | 17 | `scripts/s3-tests/unimplemented_tests.txt` |
|
||||
| Excluded tests | Vendor-specific or intentionally unsupported behavior excluded from RustFS compatibility gating. | 273 | `scripts/s3-tests/excluded_tests.txt` |
|
||||
| List | Purpose | Source |
|
||||
|---|---|---|
|
||||
| Implemented tests | Standard S3 tests expected to pass; the default local s3tests run. | `scripts/s3-tests/implemented_tests.txt` |
|
||||
| Lifecycle behavior tests | Days-based expiration cases gated by the `s3-lifecycle-behavior-tests` lane in `.github/workflows/ci.yml`. | `scripts/s3-tests/lifecycle_behavior_tests.txt` |
|
||||
| Unimplemented tests | Standard S3 features not yet passing. | `scripts/s3-tests/unimplemented_tests.txt` |
|
||||
| Excluded tests | Vendor-specific or intentionally unsupported behavior excluded from RustFS gating. | `scripts/s3-tests/excluded_tests.txt` |
|
||||
|
||||
Counts ignore blank lines and comments.
|
||||
|
||||
The lifecycle behavior lane runs real Days-based expiration cases that need
|
||||
`RUSTFS_ILM_DEBUG_DAY_SECS` (Ceph `lc_debug_interval` equivalent) and an enabled
|
||||
background scanner; it cannot share the default single-server gate because a
|
||||
global debug day would also shrink the `x-amz-expiration` header asserted by the
|
||||
`test_lifecycle_expiration_header_*` cases. See `scripts/s3-tests/run.sh`
|
||||
(`IMPLEMENTED_TESTS_FILE` override) and the `s3-lifecycle-behavior-tests` job in
|
||||
`.github/workflows/ci.yml`.
|
||||
Counts ignore blank lines and comments; compute them from the files. The lifecycle lane runs separately because its cases need `RUSTFS_ILM_DEBUG_DAY_SECS` and an enabled scanner, and a global debug day would also shrink the `x-amz-expiration` header asserted by `test_lifecycle_expiration_header_*`; see `IMPLEMENTED_TESTS_FILE` in `scripts/s3-tests/run.sh`.
|
||||
|
||||
## Supported Coverage
|
||||
|
||||
The implemented test list currently covers the common object-storage surface:
|
||||
|
||||
| Area | Status | Evidence |
|
||||
|---|---|---|
|
||||
| Bucket create/delete/list/head | Supported | `implemented_tests.txt` |
|
||||
| Object put/get/delete/copy/head | Supported | `implemented_tests.txt` |
|
||||
| CopyObject checksums (CRC32, CRC32C, CRC64NVME, SHA1, SHA256, MD5, SHA512, XXHASH3, XXHASH64, XXHASH128), including source preservation and explicit override | Supported in the first RustFS release containing this change | `crates/e2e_test/src/copy_object_checksum_test.rs` |
|
||||
| CopyObject checksums (CRC32, CRC32C, CRC64NVME, SHA1, SHA256, MD5, SHA512, XXHASH3, XXHASH64, XXHASH128), including source preservation and explicit override | Supported | `crates/e2e_test/src/copy_object_checksum_test.rs` |
|
||||
| ListObjects/ListObjectsV2 prefix, delimiter, marker, max-keys | Supported | `implemented_tests.txt` |
|
||||
| Multipart upload create/upload/complete/abort and selected multipart copy/checksum/object-attribute behavior | Supported | `implemented_tests.txt` |
|
||||
| Bucket and object tagging | Supported | `implemented_tests.txt` |
|
||||
@@ -50,33 +36,25 @@ The implemented test list currently covers the common object-storage surface:
|
||||
| SSE-C and selected SSE-KMS edge cases | Supported | `implemented_tests.txt` |
|
||||
| Selected versioning, object-lock, checksum, CORS, raw request, and conditional write behavior | Supported | `implemented_tests.txt` |
|
||||
|
||||
"Supported" for the SSE row means RustFS encrypts and decrypts its own objects. It does not mean RustFS can read objects another implementation encrypted: objects MinIO wrote with SSE-S3, SSE-KMS, or SSE-C are not readable by RustFS today, which matters when migrating. See [MinIO file-format interoperability, Part C](minio-file-format-compat.md#part-c--server-side-encryption-sse) and rustfs/backlog#1638.
|
||||
"Supported" for the SSE row means RustFS encrypts and decrypts its own objects. MinIO SSE objects (SSE-S3, SSE-KMS, SSE-C) are not readable in default builds; see [minio-file-format-compat.md Part C](minio-file-format-compat.md#part-c--server-side-encryption-sse) for the `rio-v2` migration build.
|
||||
|
||||
## Planned Standard Coverage
|
||||
## Not Yet Passing
|
||||
|
||||
These are standard S3 areas that remain planned work and must not be described
|
||||
as already complete:
|
||||
Standard S3 areas that must not be described as complete:
|
||||
|
||||
| Area | Status | Evidence |
|
||||
|---|---|---|
|
||||
| Bucket access logging | Planned | `unimplemented_tests.txt` |
|
||||
| POST Object form upload checksum handling | Planned | `unimplemented_tests.txt` |
|
||||
| Bucket ownership controls | Planned | `unimplemented_tests.txt` |
|
||||
| Bucket access logging | Handlers exist (`get_bucket_logging`, `put_bucket_logging` in `rustfs/src/storage/ecfs.rs`); the `test_*bucket_logging*` s3tests cases are still listed as unimplemented | `unimplemented_tests.txt` |
|
||||
| POST Object form upload checksum handling | Not yet passing | `unimplemented_tests.txt` |
|
||||
| Bucket ownership controls | No handler | `unimplemented_tests.txt` |
|
||||
| Multipart upload listing and part lookup compatibility edge cases | Not part of default gate | `excluded_tests.txt` |
|
||||
| IAM-account or multi-storage-class dependent cases | Not part of default gate | `unimplemented_tests.txt` |
|
||||
| Tenanted bucket policy edge cases | Needs investigation | `unimplemented_tests.txt` |
|
||||
|
||||
## Intentional Exclusions
|
||||
|
||||
`excluded_tests.txt` contains tests that should not block the RustFS
|
||||
compatibility gate. They fall into two classes:
|
||||
|
||||
- vendor-specific or non-portable behavior not required for RustFS S3
|
||||
compatibility;
|
||||
- intentionally unsupported product behavior, such as ACL authorization.
|
||||
`excluded_tests.txt` holds tests that must not block the compatibility gate: vendor-specific or non-portable behavior, and intentionally unsupported product behavior such as ACL authorization.
|
||||
|
||||
## Update Rule
|
||||
|
||||
When a planned S3 feature is implemented, move its passing test entries from
|
||||
`unimplemented_tests.txt` to `implemented_tests.txt`, update this matrix, and
|
||||
avoid changing README wording beyond the supported coverage.
|
||||
When a feature starts passing, move its test entries from `unimplemented_tests.txt` to `implemented_tests.txt` and update the row here in the same PR. Do not change README wording beyond the supported coverage. Handler-level status (missing, stubbed, or diverging endpoints) is tracked in [minio-rustfs-router-compatibility.md](minio-rustfs-router-compatibility.md).
|
||||
|
||||
@@ -1,327 +1,152 @@
|
||||
# S3 Tables Support Matrix
|
||||
|
||||
This matrix records the RustFS S3 Tables surfaces that are supported,
|
||||
previewed, referenced, or intentionally not claimed. It is the release-facing
|
||||
boundary for the Iceberg REST Catalog work in RustFS.
|
||||
**Use this when:** writing a release note, README claim, or client-compatibility statement about RustFS S3 Tables / Iceberg REST Catalog, or deciding whether a feature is supported, preview, or not claimed.
|
||||
**Source of truth:** the table-catalog handlers under `rustfs/src/admin/handlers/table_catalog/`; the conformance scripts and their README in `scripts/table-catalog/`; the durable-backing cutover procedure in [docs/operations/s3-tables-cutover-runbook.md](../operations/s3-tables-cutover-runbook.md).
|
||||
|
||||
RustFS S3 Tables is an Iceberg REST Catalog and table-bucket implementation on
|
||||
top of the RustFS S3 data plane. This document does not claim full parity with
|
||||
the AWS S3 Tables control-plane API or with every vendor-specific Iceberg
|
||||
catalog extension.
|
||||
RustFS S3 Tables is an Iceberg REST Catalog and table-bucket implementation on top of the RustFS S3 data plane. It does not claim parity with the AWS S3 Tables control-plane API or with vendor-specific Iceberg catalog extensions.
|
||||
|
||||
## Status Labels
|
||||
|
||||
| Label | Meaning |
|
||||
|---|---|
|
||||
| Automated | Covered by a runnable RustFS script or server test. |
|
||||
| Manual/live harness | RustFS can generate pinned client package inputs, commands, expected outputs, and CI opt-in gates for a live endpoint, but the live run is not enabled by default in CI. |
|
||||
| Generated harness | RustFS can generate client configuration or probe input, but live execution is not automated in CI. |
|
||||
| Manual/live harness | RustFS generates pinned client inputs, commands, expected outputs, and CI opt-in gates for a live endpoint; the live run is not enabled by default in CI. |
|
||||
| Generated harness | RustFS generates client configuration or probe input; live execution is not automated. |
|
||||
| Supported | Implemented server-side and covered by focused RustFS tests. |
|
||||
| Preview / controlled | Implemented behind explicit operator action or a run-once endpoint. No automatic background claim is made. |
|
||||
| Documented, not automated | Configuration or behavior is documented, but the live client run is not automated. |
|
||||
| Reference only | Kept as a compatibility reference. RustFS does not claim live interoperability yet. |
|
||||
| Not claimed | Out of scope for the current S3 Tables implementation. |
|
||||
| Preview / controlled | Implemented behind explicit operator action or a run-once endpoint; no automatic background claim. |
|
||||
| Documented, not automated | Behavior is documented; the live client run is not automated. |
|
||||
| Reference only | Compatibility reference; live interoperability is not claimed. |
|
||||
| Not claimed | Out of scope for the current implementation. |
|
||||
|
||||
## Endpoint And Profile Matrix
|
||||
|
||||
| Surface | Status | Notes |
|
||||
|---|---|---|
|
||||
| `/iceberg/v1` | Supported | Canonical RustFS Iceberg REST Catalog prefix. Default REST signing name is `s3`. |
|
||||
| `/_iceberg/v1` | Supported compatibility alias | MinIO AIStor-style alias. The smoke profile defaults to REST signing name `s3tables`. |
|
||||
| S3 object data plane | Supported | Data, metadata, manifest, and delete files remain ordinary S3 objects, with table-aware policy checks for table warehouse paths. |
|
||||
| Table bucket enablement | Supported | A regular RustFS bucket can be enabled for table catalog use and then addressed as the REST catalog warehouse. |
|
||||
| Catalog-vended table credentials | Automated when enabled | Disabled by default. When enabled, LoadTable vends credentials only when `X-Iceberg-Access-Delegation` contains the exact `vended-credentials` token; the dedicated credentials endpoint uses the same issuer path. |
|
||||
| AWS S3 Tables endpoint shape | Profile generator | Generates the AWS catalog URI and S3 Tables warehouse ARN shape for migration docs. Full AWS S3 Tables API parity is not claimed. |
|
||||
| MinIO AIStor Tables profile | Profile generator plus RustFS alias smoke | RustFS exposes the alias shape, but does not claim all AIStor private extensions. |
|
||||
| Cloudflare R2 Data Catalog profile | Profile generator | Generates the catalog URI and warehouse-name shape for migration docs. Live RustFS interoperability is not claimed. |
|
||||
| Alibaba OSS Tables profile | Profile generator | Generates provider endpoint, `acs:osstables` warehouse ARN, `osstables` signing-name, and `https://oss-{region}.aliyuncs.com` S3FileIO endpoint shapes for migration docs. Live RustFS interoperability is not claimed. |
|
||||
| `/iceberg/v1` | Supported | Canonical REST Catalog prefix; default REST signing name `s3`. |
|
||||
| `/_iceberg/v1` | Supported compatibility alias | MinIO AIStor-style alias; smoke profile defaults to signing name `s3tables`. |
|
||||
| S3 object data plane | Supported | Data, metadata, manifest, and delete files are ordinary S3 objects with table-aware policy checks on warehouse paths. |
|
||||
| Table bucket enablement | Supported | A regular bucket is enabled for catalog use and addressed as the REST catalog warehouse. |
|
||||
| Catalog-vended table credentials | Automated when enabled | Disabled by default. LoadTable vends credentials only when `X-Iceberg-Access-Delegation` contains the exact `vended-credentials` token; the dedicated credentials endpoint uses the same issuer path. |
|
||||
| AWS S3 Tables endpoint shape | Profile generator | Generates the AWS catalog URI and warehouse ARN shape for migration docs. API parity not claimed. |
|
||||
| MinIO AIStor Tables profile | Profile generator plus alias smoke | Alias shape only; AIStor private extensions not claimed. |
|
||||
| Cloudflare R2 Data Catalog profile | Profile generator | Catalog URI and warehouse-name shape only; live interop not claimed. |
|
||||
| Alibaba OSS Tables profile | Profile generator | Endpoint, `acs:osstables` warehouse ARN, `osstables` signing name, and `https://oss-{region}.aliyuncs.com` S3FileIO endpoint shapes only; live interop not claimed. |
|
||||
|
||||
## Client And Engine Matrix
|
||||
|
||||
| Client or engine | Status | Current RustFS claim |
|
||||
| Client or engine | Status | Claim |
|
||||
|---|---|---|
|
||||
| PyIceberg | Automated | Creates namespace and table, appends rows, reloads, scans, probes metadata-location, refs, views, maintenance, diagnostics, and optional catalog-vended table credentials with an exact-prefix data-plane scope check. |
|
||||
| Spark Iceberg REST catalog | Manual/live harness | RustFS can generate pinned Spark/Iceberg package inputs, REST catalog properties, SQL, run commands, expected `row_count=2`, and a CI opt-in gate for namespace creation, table creation, append, refresh, count, and cleanup. Live Spark execution and commit-conflict probing are still manual validation items unless explicitly enabled in the runner. |
|
||||
| Trino Iceberg REST catalog | Manual/live harness | RustFS can generate catalog properties and a read-only `SELECT COUNT(*)` command for a table created by PyIceberg or Spark. Write compatibility is not claimed. |
|
||||
| DuckDB Iceberg 1.5.5 | Automated | `duckdb_smoke.py` verifies the metadata-location read path and generic REST Catalog single-table create, insert, update, delete, merge, schema evolution, snapshots, concurrent writers, normal drop, PyIceberg cross-read, `/iceberg` with `s3` signing, and `/_iceberg` with `s3tables` signing. Staged create, purge-on-drop, and format v3 are verified as fail-closed boundaries. DuckDB's endpoint-disabled two-table mode is exercised without claiming cross-table atomicity. AWS `ENDPOINT_TYPE S3_TABLES` and catalog-vended credential integration are not claimed. |
|
||||
| StarRocks Iceberg REST catalog | Documented, not automated | External catalog read-path reference only. Write compatibility is not claimed. |
|
||||
| Databend | Manual/live harness | RustFS can generate an S3 stage read probe for table data files. RustFS does not claim Databend Iceberg REST Catalog integration yet. |
|
||||
| Snowflake Open Catalog / Iceberg integrations | Generated harness | RustFS can generate an operator-adapted external volume/catalog SQL template. Live RustFS interoperability is not claimed. |
|
||||
| PyIceberg | Automated | Namespace and table create, append, reload, scan, metadata-location, refs, views, maintenance, diagnostics, optional vended credentials with an exact-prefix data-plane scope check. |
|
||||
| Spark Iceberg REST catalog | Manual/live harness | Pinned package inputs, catalog properties, SQL, expected `row_count=2`, and a CI opt-in gate for create/append/refresh/count/cleanup. Live execution and commit-conflict probing remain manual unless enabled in the runner. |
|
||||
| Trino Iceberg REST catalog | Manual/live harness | Catalog properties and a read-only `SELECT COUNT(*)` against a PyIceberg- or Spark-created table. Write compatibility not claimed. |
|
||||
| DuckDB Iceberg 1.5.5 | Automated | `duckdb_smoke.py` covers metadata-location read, single-table create/insert/update/delete/merge, schema evolution, snapshots, concurrent writers, drop, PyIceberg cross-read, and both signing profiles. Staged create, purge-on-drop, and format v3 are verified fail-closed. Two-table mode runs without claiming cross-table atomicity. AWS `ENDPOINT_TYPE S3_TABLES` and vended-credential integration not claimed. |
|
||||
| StarRocks Iceberg REST catalog | Documented, not automated | External catalog read-path reference only. |
|
||||
| Databend | Manual/live harness | S3 stage read probe for table data files only; Iceberg REST integration not claimed. |
|
||||
| Snowflake Open Catalog / Iceberg integrations | Generated harness | Operator-adapted external volume/catalog SQL template only; live interop not claimed. |
|
||||
|
||||
## Live Evidence And Operations Matrix
|
||||
|
||||
| Area | Status | Current RustFS claim |
|
||||
| Area | Status | Claim |
|
||||
|---|---|---|
|
||||
| Live conformance evidence | Automated for PyIceberg and DuckDB | `engine_compatibility.py --print-live-evidence-schema` defines the required evidence schema and claim promotion boundaries. `pyiceberg_smoke.py --live-evidence-output` and `duckdb_smoke.py --live-evidence-output` write validated client evidence records after successful live smoke runs. |
|
||||
| Production operations guide | Generated harness | `engine_compatibility.py --print-operations-guide` records command, evidence, pass criteria, and fail-closed signals for live conformance, durable backing cutover, maintenance, recovery, permissions, credential vending, and unsupported-claim governance. |
|
||||
| Vendor compatibility gap audit | Generated harness | `engine_compatibility.py --print-vendor-audit` records provider source URLs, catalog path and warehouse shapes, signing/auth models, error/permission/maintenance validation categories, and not-claimed boundaries for AWS S3 Tables, MinIO AIStor Tables, Cloudflare R2 Data Catalog, and Alibaba OSS Tables. |
|
||||
| Client claim promotion | Automated for scoped clients | PyIceberg and DuckDB claims remain bounded by their repeatable smoke entrypoints and recorded versions. Spark can be promoted only with recorded manual/live evidence; Trino remains read-only; Snowflake and vendor profiles remain reference-only without repeatable live evidence. |
|
||||
| Live conformance evidence | Automated for PyIceberg and DuckDB | `engine_compatibility.py --print-live-evidence-schema` defines the evidence schema and promotion boundaries; both smoke scripts write validated evidence records via `--live-evidence-output`. |
|
||||
| Production operations guide | Generated harness | `engine_compatibility.py --print-operations-guide` records commands, evidence, pass criteria, and fail-closed signals for conformance, cutover, maintenance, recovery, permissions, credential vending, and claim governance. |
|
||||
| Vendor compatibility gap audit | Generated harness | `engine_compatibility.py --print-vendor-audit` records provider URLs, path and warehouse shapes, auth models, validation categories, and not-claimed boundaries for the four vendor profiles. |
|
||||
| Client claim promotion | Automated for scoped clients | PyIceberg and DuckDB claims are bounded by their smoke entrypoints and recorded versions; Spark needs recorded live evidence; Trino stays read-only; Snowflake and vendor profiles stay reference-only. |
|
||||
|
||||
## Catalog API Matrix
|
||||
|
||||
| Area | Status | Covered behavior |
|
||||
|---|---|---|
|
||||
| Catalog config | Supported | `GET /v1/config` advertises RustFS catalog defaults and only the supported OpenAPI REST paths in `endpoints`. RustFS administration, maintenance, migration, diagnostics, refs, and metadata-location extensions remain available but are not presented as standard Iceberg REST endpoints. |
|
||||
| Table bucket discovery | Supported | `PUT` and `GET /v1/buckets/{warehouse}` enable and inspect table bucket state. |
|
||||
| Namespaces | Supported | Create, list, load, existence check, and drop namespace routes are registered on both catalog prefixes. List responses support Iceberg REST `pageSize`/`pageToken` pagination with context-bound tokens and bounded catalog-store reads. Namespace identifiers are limited to 512 ASCII characters so persisted paths and stateless continuation tokens remain bounded. |
|
||||
| Tables | Supported | Create, register, list, load, existence check, rename, commit, metadata-location get/update, and drop table routes are registered on both catalog prefixes. Object-backed rename uses a bucket-scoped persistent fence, recoverable intent, and conditional publication of the destination, source tombstone, and warehouse index; the source identifier is reusable only through an ETag-conditional tombstone replacement. Table and view listings support Iceberg REST `pageSize`/`pageToken` pagination with context-bound tokens and bounded catalog-store reads. Commit identifiers must match the URL resource; unknown requirements, updates, and snapshot operations fail as bad requests; staged create, register overwrite, purge-on-drop, and v3-only encryption-key updates return an explicit unsupported-operation response. Standard statistics, partition statistics, and schema/spec cleanup updates are accepted. |
|
||||
| Commit CAS | Supported | Single-table commits validate base metadata, expected version token, referenced object existence, warehouse scope, and Iceberg commit requirements before advancing the current metadata pointer. Externally supplied metadata transitions preserve monotonic column, partition, and sequence assignment watermarks and immutable definitions for retained schemas, partition specs, sort orders, and snapshots. Standard commits preserve the normal commit-token file name and use an immutable-table-scoped fallback when rename followed by source-name reuse would otherwise collide at the same generation and commit ID. The catalog does not advertise `idempotency-key-lifetime`; clients must treat standard mutation-wide `Idempotency-Key` semantics as unsupported. |
|
||||
| Commit recovery | Supported | Commit log, idempotency lookup, diagnostics, and recovery routes expose staged/finalization gaps and repair safe idempotency gaps without moving the table pointer. |
|
||||
| Snapshot refs | Supported | Refs can be listed, created or replaced, and deleted through catalog commits. `main` is protected and refs with explicit retention require forced delete. |
|
||||
| Iceberg views | Supported | Basic create, list, load, replace, existence check, and drop routes persist view metadata with view-scoped authorization. Replace identifiers must match the URL resource, `schema-id: -1` resolves to the last added schema, one commit timestamp is used consistently, and only Iceberg view format version 1 is accepted. |
|
||||
| LoadTable and table credentials endpoint | Supported | LoadTable keeps the client-provided mode unless the request negotiates `vended-credentials`. Successful vending returns one temporary session for both the table warehouse prefix and the exact current metadata location. Missing credential permission falls back to metadata-only LoadTable with an explicit reason; issuer failures remain errors. Negotiated and dedicated credential responses set `Cache-Control: no-store, private`, `Pragma: no-cache`, and `Expires: 0`. |
|
||||
| Catalog diagnostics and export | Supported | Exposes recovery state, consistency state, backing manifest, recoverable commit-log WAL state, strong backing migration target, single-active-writer policy, and scale validation matrix. |
|
||||
| Catalog import and rollback | Supported | Import/register and online rollback use catalog validation and commit paths rather than direct pointer mutation. Online rollback accepts only a forward-safe metadata target that preserves assignment watermarks and retained definitions. Restoring an older target that lowers those watermarks is an offline disaster-recovery operation and requires every writer to be stopped. |
|
||||
| External catalog bridge | Supported operator path | Operator-supplied metadata pointer sync/import is supported for external catalog identity boundaries. Online vendor SDK polling and policy mirroring are not claimed. |
|
||||
| Multi-table transactions | Not claimed | RustFS currently claims single-table commit atomicity only. |
|
||||
| Catalog config | Supported | `GET /v1/config` advertises defaults and only the supported OpenAPI REST paths in `endpoints`; RustFS extensions (administration, maintenance, migration, diagnostics, refs, metadata-location) are not presented as standard endpoints. |
|
||||
| Table bucket discovery | Supported | `PUT` / `GET /v1/buckets/{warehouse}` enable and inspect table bucket state. |
|
||||
| Namespaces | Supported | Create, list, load, exists, drop on both prefixes; `pageSize`/`pageToken` pagination with context-bound tokens; identifiers limited to 512 ASCII characters. |
|
||||
| Tables | Supported | Create, register, list, load, exists, rename, commit, metadata-location get/update, drop on both prefixes. Rename uses a bucket-scoped persistent fence, recoverable intent, and conditional publication; the source name is reusable only via an ETag-conditional tombstone replacement. Commit identifiers must match the URL; unknown requirements/updates fail as bad requests; staged create, register overwrite, purge-on-drop, and v3-only encryption-key updates return an explicit unsupported-operation response. |
|
||||
| Commit CAS | Supported | Single-table commits validate base metadata, version token, referenced object existence, warehouse scope, and Iceberg requirements before advancing the pointer; external metadata transitions preserve monotonic assignment watermarks and immutable retained definitions. `idempotency-key-lifetime` is not advertised; mutation-wide `Idempotency-Key` semantics are unsupported. |
|
||||
| Commit recovery | Supported | Commit log, idempotency lookup, diagnostics, and recovery routes expose and repair finalization gaps without moving the pointer. |
|
||||
| Snapshot refs | Supported | List, create/replace, delete via commits; `main` is protected; refs with explicit retention need forced delete. |
|
||||
| Iceberg views | Supported | Create, list, load, replace, exists, drop with view-scoped authorization; only view format version 1. |
|
||||
| LoadTable and table credentials endpoint | Supported | Vending only on negotiated `vended-credentials`; one temporary session scoped to the warehouse prefix and current metadata location; missing credential permission falls back to metadata-only with an explicit reason; responses carry `Cache-Control: no-store, private`. |
|
||||
| Catalog diagnostics and export | Supported | Recovery state, consistency, backing manifest, WAL state, migration target, single-active-writer policy, scale validation matrix. |
|
||||
| Catalog import and rollback | Supported | Import/register and online rollback go through validation and commit paths. Online rollback accepts only forward-safe targets; restoring an older target that lowers watermarks is an offline disaster-recovery operation with all writers stopped. |
|
||||
| External catalog bridge | Supported operator path | Operator-supplied metadata pointer sync/import. Vendor SDK polling and policy mirroring not claimed. |
|
||||
| Multi-table transactions | Not claimed | Single-table commit atomicity only. |
|
||||
|
||||
## Data Plane And Credential Matrix
|
||||
|
||||
| Area | Status | Covered behavior |
|
||||
|---|---|---|
|
||||
| Table-aware S3 policy bridge | Supported | Ordinary S3 actions against table warehouse paths are checked through the table data-plane bridge so table policy cannot be bypassed by direct object access. |
|
||||
| Reserved catalog protection | Supported | Catalog-reserved internal prefixes are protected from ordinary object mutation. |
|
||||
| Static S3 credentials | Automated | The default PyIceberg smoke path uses configured S3 credentials for REST signing and object data-plane access. |
|
||||
| Catalog-vended credentials | Automated when enabled | `rustfs-vended-credentials` verifies the returned table prefix, then checks `PutObject`, `HeadObject`, `GetObject`, and `DeleteObject` inside the prefix and denies access outside the prefix. |
|
||||
| Credential lifetime | Supported | Vended credential TTL is server-side and clamped to a short-lived range. |
|
||||
| No-long-term-data-credential bootstrap | Not claimed | The current credential-vending flow still uses the configured principal for catalog setup before table-scoped credentials are requested. |
|
||||
| Table-aware S3 policy bridge | Supported | Ordinary S3 actions on warehouse paths are checked through the table bridge; table policy cannot be bypassed by direct object access. |
|
||||
| Reserved catalog protection | Supported | Catalog-reserved prefixes are protected from ordinary object mutation. |
|
||||
| Static S3 credentials | Automated | Default PyIceberg smoke path. |
|
||||
| Catalog-vended credentials | Automated when enabled | `rustfs-vended-credentials` verifies the returned prefix, then checks Put/Head/Get/DeleteObject inside it and denies access outside it. |
|
||||
| Credential lifetime | Supported | Server-side TTL clamped to a short-lived range. |
|
||||
| No-long-term-data-credential bootstrap | Not claimed | Catalog setup still uses the configured principal before table-scoped credentials are requested. |
|
||||
|
||||
## Maintenance Matrix
|
||||
|
||||
| Capability | Status | Current RustFS claim |
|
||||
| Capability | Status | Claim |
|
||||
|---|---|---|
|
||||
| Metadata retention dry-run | Supported | Reports retained metadata and deletion candidates without moving the table pointer. |
|
||||
| Metadata cleanup delete | Supported | Deletes only candidates that pass the safety window and current-pointer checks. |
|
||||
| Ordinary bucket lifecycle expiry | Disabled for table buckets | Table bucket objects are excluded from ordinary lifecycle expiration, including already queued expiry work. Snapshot expiration and orphan cleanup remain catalog maintenance operations so referenced Iceberg files cannot be deleted outside publication fencing. |
|
||||
| Snapshot expiration planning | Supported | Produces expiration plans with retained and candidate snapshots. |
|
||||
| Snapshot expiration commit | Preview / controlled | Can manually commit safe snapshot expiration through the catalog. Stale plans fail closed. |
|
||||
| Manifest/data/delete reachability cleanup | Supported | Reads manifest-list and manifest Avro references, reports reachable objects, and deletes only unreferenced table objects that pass the safety window. |
|
||||
| Maintenance scheduler run endpoint | Preview / controlled | Lets an external scheduler durably queue one maintenance job per table, reuse an active queued job, and recover expired queued leases before requeuing. |
|
||||
| Maintenance worker run endpoint | Preview / controlled | Supports queued-job claim, run-once execution, current-job backpressure, retry deferral, lease expiry recovery, and heartbeat updates. |
|
||||
| Maintenance scheduler guardrails | Preview / controlled | Exposes disabled, paused, ready, queued-job handoff, active-job backpressure, retry deferral, quarantine boundary, recommended actions, and recent maintenance job audit timeline state for external schedulers and operators. |
|
||||
| Maintenance audit events | Preview / controlled | Job reports and scheduler job summaries include structured audit events for planning, worker transitions, heartbeats, lease expiry recovery, and mutating quarantine operations. |
|
||||
| Maintenance quarantine operations | Preview / controlled | Lets operators inspect, release, retry, or abandon the current quarantined maintenance job without moving the table pointer. |
|
||||
| Compaction planning | Preview / controlled | Plans partition-local and sort-order-local binpack candidates for Parquet files and does not mix data files from different partition directories or sort orders in one rewrite group. |
|
||||
| Delete-file or row-level compaction planning | Preview / controlled | Manifests with position or equality delete files produce machine-readable row-level planning and force the compaction report into manual review before any rewrite can run. |
|
||||
| Compaction commit | Preview / controlled | Can commit a safe partition-local Parquet rewrite through the catalog while preserving Iceberg data file sort order IDs in the rewritten manifest. |
|
||||
| Built-in periodic scheduler | Not claimed | Operators can trigger scheduler and worker ticks, but continuous in-process scheduling is not claimed. |
|
||||
| Delete-file or row-level compaction execution | Not claimed | RustFS does not rewrite delete files or execute row-level compaction; those cases remain manual-review maintenance items. |
|
||||
| Metadata retention dry-run | Supported | Reports retained metadata and deletion candidates without moving the pointer. |
|
||||
| Metadata cleanup delete | Supported | Deletes only candidates passing the safety window and current-pointer checks. |
|
||||
| Ordinary bucket lifecycle expiry | Disabled for table buckets | Table bucket objects are excluded from lifecycle expiration, including queued work; snapshot expiration and orphan cleanup stay inside catalog maintenance. |
|
||||
| Snapshot expiration planning | Supported | Plans with retained and candidate snapshots. |
|
||||
| Snapshot expiration commit | Preview / controlled | Manual commit through the catalog; stale plans fail closed. |
|
||||
| Manifest/data/delete reachability cleanup | Supported | Reads manifest-list and manifest Avro references; deletes only unreferenced objects passing the safety window. |
|
||||
| Maintenance scheduler run endpoint | Preview / controlled | Durably queues one job per table, reuses an active queued job, recovers expired leases. |
|
||||
| Maintenance worker run endpoint | Preview / controlled | Claim, run-once, backpressure, retry deferral, lease expiry recovery, heartbeats. |
|
||||
| Maintenance scheduler guardrails | Preview / controlled | Disabled/paused/ready state, handoff, backpressure, quarantine boundary, recommended actions, recent job audit timeline. |
|
||||
| Maintenance audit events | Preview / controlled | Structured events for planning, worker transitions, heartbeats, lease recovery, quarantine mutations. |
|
||||
| Maintenance quarantine operations | Preview / controlled | Inspect, release, retry, or abandon the quarantined job without moving the pointer. |
|
||||
| Compaction planning | Preview / controlled | Partition-local and sort-order-local binpack candidates for Parquet; never mixes partitions or sort orders in one group. |
|
||||
| Delete-file or row-level compaction planning | Preview / controlled | Position or equality delete files force machine-readable planning into manual review. |
|
||||
| Compaction commit | Preview / controlled | Commits a safe partition-local Parquet rewrite preserving sort order IDs. |
|
||||
| Built-in periodic scheduler | Not claimed | Ticks are operator-triggered; no continuous in-process scheduling. |
|
||||
| Delete-file or row-level compaction execution | Not claimed | Manual-review maintenance item. |
|
||||
|
||||
## Recovery And Strong Backing Matrix
|
||||
|
||||
| Area | Status | Current RustFS claim |
|
||||
| Area | Status | Claim |
|
||||
|---|---|---|
|
||||
| Single-table CAS | Supported | The table pointer advances only through expected-token and expected-metadata-location validation. |
|
||||
| Idempotent retry | Supported | Repeated commit IDs can return the already finalized result or surface recoverable finalization gaps. |
|
||||
| Commit publication fencing | Supported with rolling-upgrade gate | Existing deployments retain exact object guards so older writers cannot mutate referenced files during publication. Set `RUSTFS_TABLE_CATALOG_PUBLICATION_FENCE_FLEET_CONFIRMED=true` only after every serving node supports table and table-bucket publication fences. In scalable mode, active table warehouse prefixes must not overlap, ordinary lifecycle expiry remains disabled for table buckets, and first enablement, first publication, drop, and warehouse relocation are serialized by the table-bucket fence. |
|
||||
| Post-CAS finalization recovery | Supported | Diagnostics and recovery can repair stale or missing idempotency indexes without changing the current table pointer. |
|
||||
| Catalog export | Supported | Exposes table state, commit recovery state, and backing migration information for operator inspection. |
|
||||
| Strong backing state transfer | Supported | Object-backed table bucket, namespace, table, view, commit-log, and idempotency state can be materialized into the durable strong snapshot. The transfer is deterministic, ETag-CAS protected, idempotent after an interrupted finalization, validates candidate state through the restart decoder before publication, preserves resource-backed implicit namespaces, and fails closed when an inactive explicit namespace conflicts with active descendants or resources. Snapshot hydration requires a stable non-empty ETag, caps the encoded snapshot at 64 MiB, shares state and reload serialization across requests in one server context, and rejects disappearance or format-version rollback after observation. Configured durable-strong mode rejects a missing snapshot on its first catalog access after startup; only object-backed migration may initialize an empty target. |
|
||||
| Durable backing migration preflight | Supported | `GET /iceberg/v1/{warehouse}/catalog/migration` and the `/_iceberg/v1` alias inspect object-backed catalog inventory, recovery blockers, warehouse prefix index readiness, active table/view identifier collisions, persistent write-fence state, target snapshot agreement, and whether every table bucket is ready for cutover. |
|
||||
| Durable backing migration execution | Preview / controlled | `POST /iceberg/v1/{warehouse}/catalog/migration` fences table-bucket registry changes, acquires a persistent per-bucket write fence, records whether a global strong snapshot existed before publication, drains in-flight catalog mutations, materializes the target snapshot, and reports `ready_to_enable_durable_strong`. Retries and `DELETE` may restore a known-absent initial target after an ambiguous first write, but fail closed if a previously existing or materialized global snapshot disappears. `DELETE` releases the bucket fence only while its target state has not advanced, and releases the registry fence after the last bucket is cancelled. Both mutations require `admin:MigrateTableCatalog`. |
|
||||
| Strong snapshot rolling compatibility | Supported | Durable strong control-plane reads snapshot versions 1 and 2, writes version 1 by default, and writes version 2 only after both the requested and fleet-confirmed gates are enabled. A running process rejects any lower-format snapshot after observing a higher format. Once version 2 is fleet-confirmed, table data-plane resolution fails closed until the persisted snapshot is version 2; a missing table-bucket entry also fails closed instead of bypassing table-aware authorization. |
|
||||
| Disaster recovery rehearsal | Manual/live harness | `failure_coverage.py --print-disaster-recovery-rehearsal` generates an operator runbook covering catalog export, diagnostics, safe recovery repair, rollback/import, durable backing migration dry-run, post-recovery loadTable, and table data-plane policy probes. |
|
||||
| Scale and fault rehearsal | Manual/live harness | `failure_coverage.py --print-scale-fault-rehearsal` generates an opt-in runbook for concurrent writer stress, maintenance scheduler lease recovery, durable backing cutover preflight, recovery/rollback/import under load, and post-run evidence capture. |
|
||||
| Durable strong snapshot backing cutover | Preview / controlled | Operators can select the ETag-CAS snapshot backing with `RUSTFS_TABLE_CATALOG_BACKING=durable-strong` only after every table bucket reports `SNAPSHOT_MATERIALIZED` and `ready_to_enable_durable_strong: true`. This mode does not claim a separate external KV/WAL service, and object-only advanced operations fail closed. Version 1 backing manifests retain the legacy `STRONG_KV_WAL` and `CUT_OVER_LINEARIZABLE_READS` wire labels for client compatibility; those labels do not expand the implementation claim. |
|
||||
| Single-table CAS | Supported | Pointer advances only through expected-token and expected-metadata-location validation. |
|
||||
| Idempotent retry | Supported | Repeated commit IDs return the finalized result or surface recoverable finalization gaps. |
|
||||
| Commit publication fencing | Supported with rolling-upgrade gate | Exact object guards protect referenced files during publication. Set `RUSTFS_TABLE_CATALOG_PUBLICATION_FENCE_FLEET_CONFIRMED=true` only after every serving node supports table and table-bucket fences. In scalable mode, warehouse prefixes must not overlap and enablement, first publication, drop, and relocation are serialized by the table-bucket fence. |
|
||||
| Post-CAS finalization recovery | Supported | Repairs stale or missing idempotency indexes without changing the pointer. |
|
||||
| Catalog export | Supported | Table state, commit recovery state, and backing migration information. |
|
||||
| Strong backing state transfer | Supported | Object-backed catalog state is materialized into the durable strong snapshot deterministically, ETag-CAS protected, idempotent after interrupted finalization, and validated through the restart decoder before publication; conflicts between inactive explicit namespaces and active descendants fail closed. Hydration requires a stable non-empty ETag, caps the snapshot at 64 MiB, and rejects disappearance or format-version rollback after observation. Configured durable-strong mode rejects a missing snapshot on first access; only object-backed migration may initialize an empty target. |
|
||||
| Durable backing migration preflight | Supported | `GET /iceberg/v1/{warehouse}/catalog/migration` (and the alias) reports inventory, recovery blockers, prefix-index readiness, identifier collisions, fence state, target agreement, and per-bucket cutover readiness. |
|
||||
| Durable backing migration execution | Preview / controlled | `POST /iceberg/v1/{warehouse}/catalog/migration` fences registry changes, acquires a persistent per-bucket write fence, drains in-flight mutations, materializes the snapshot, and reports `ready_to_enable_durable_strong`. `DELETE` cancels only while the target has not advanced. Both mutations require `admin:MigrateTableCatalog`. Procedure: [s3-tables-cutover-runbook.md](../operations/s3-tables-cutover-runbook.md). |
|
||||
| Strong snapshot rolling compatibility | Supported | Reads snapshot versions 1 and 2; writes version 1 by default and version 2 only after both `RUSTFS_TABLE_CATALOG_STRONG_SNAPSHOT_V2` and `RUSTFS_TABLE_CATALOG_STRONG_SNAPSHOT_V2_FLEET_CONFIRMED` are set. A process rejects a lower format after observing a higher one; once v2 is fleet-confirmed, data-plane resolution fails closed until the persisted snapshot is v2. |
|
||||
| Disaster recovery rehearsal | Manual/live harness | `failure_coverage.py --print-disaster-recovery-rehearsal`. |
|
||||
| Scale and fault rehearsal | Manual/live harness | `failure_coverage.py --print-scale-fault-rehearsal`. |
|
||||
| Durable strong snapshot backing cutover | Preview / controlled | `RUSTFS_TABLE_CATALOG_BACKING=durable-strong` only after every table bucket reports `SNAPSHOT_MATERIALIZED` and `ready_to_enable_durable_strong: true`. No separate external KV/WAL service is claimed; the legacy `STRONG_KV_WAL` and `CUT_OVER_LINEARIZABLE_READS` labels in v1 manifests are wire-compatibility labels, not claims. |
|
||||
| Single active writer region | Supported policy | Diagnostics publish single-active-writer semantics and read-only replica limits. |
|
||||
| Active-active multi-region writes | Not claimed | A table must not accept independent concurrent writers in multiple active regions. |
|
||||
|
||||
## Durable Backing Cutover Runbook
|
||||
|
||||
Use the migration dry-run before changing the table catalog backing for a
|
||||
warehouse:
|
||||
|
||||
1. Take an object-backed catalog backup and record the current metadata pointer
|
||||
and version token for representative tables.
|
||||
2. Run `GET /iceberg/v1/{warehouse}/catalog/migration` with a principal that has
|
||||
`GetTableCatalogAction` on each table bucket. Treat every `blockers` entry as
|
||||
fail-closed; repair commit recovery state and backfill the warehouse prefix
|
||||
index before continuing.
|
||||
3. Before the migration `POST`, drain every catalog writer that predates the
|
||||
durable-backing migration fence and restart it on a fence-aware release. An
|
||||
older writer does not recognize the persisted fence and can otherwise
|
||||
mutate the object-backed source after the snapshot inventory is captured.
|
||||
Keep all catalog writers on the fence-aware release until cutover completes.
|
||||
4. Inventory object-only advanced operations, including maintenance workers,
|
||||
catalog recovery, export, diagnostics, and external catalog bridge writes.
|
||||
Quiesce mutating operations before cutover and confirm that each required
|
||||
operation is supported by durable-strong mode; unsupported operations fail
|
||||
closed after cutover rather than continuing against object-backed state.
|
||||
5. Run `POST /iceberg/v1/{warehouse}/catalog/migration` with
|
||||
`admin:MigrateTableCatalog`. This acquires the exclusive migration fence to
|
||||
drain in-flight fence-aware mutations, persists the source fence while
|
||||
exclusivity is held, and then copies the catalog state.
|
||||
6. Repeat the preflight and materialization for every table bucket. Do not set
|
||||
`RUSTFS_TABLE_CATALOG_BACKING=durable-strong` until the preflight reports
|
||||
`SNAPSHOT_MATERIALIZED`, no blockers, and
|
||||
`ready_to_enable_durable_strong: true`.
|
||||
7. Restart with durable strong backing enabled, then verify catalog config,
|
||||
table and view loads, commit idempotency, and table data-plane policy
|
||||
resolution before admitting writers.
|
||||
8. Before restarting into durable-strong mode, `DELETE` on the migration
|
||||
endpoint can remove a migration-created target bucket snapshot and release
|
||||
the source fence. After the durable-strong state advances, cancellation
|
||||
fails closed; recovery requires an operator-selected restore or reverse
|
||||
migration instead of restarting against the stale object-backed pointer.
|
||||
9. Preserve the object-backed catalog backup until durable strong backing has
|
||||
passed the operator's retention window.
|
||||
10. Keep strong snapshot writes on version 1 during a rolling binary upgrade.
|
||||
After every catalog writer can read version 2, set both
|
||||
`RUSTFS_TABLE_CATALOG_STRONG_SNAPSHOT_V2=true` and
|
||||
`RUSTFS_TABLE_CATALOG_STRONG_SNAPSHOT_V2_FLEET_CONFIRMED=true`, then restart
|
||||
the catalog writers. Perform a controlled catalog write or migration
|
||||
materialization and confirm that the persisted snapshot is version 2 before
|
||||
serving table data-plane traffic. Setting only one gate does not change the
|
||||
write format.
|
||||
11. After any version 2 snapshot is persisted, do not roll catalog writers back
|
||||
to a binary that only reads version 1. Current binaries preserve version 2
|
||||
even when the gates are later disabled. A running process rejects restored
|
||||
version 1 content after observing version 2, but cannot distinguish an older
|
||||
snapshot with the same format version from a deliberate restore. The format
|
||||
high-water mark is process-local: restoring any older snapshot and restarting
|
||||
every writer is a privileged disaster-recovery rollback that cannot be
|
||||
inferred from the restored object alone. Recovery must restore a compatible
|
||||
binary and a snapshot selected through the operator recovery procedure.
|
||||
12. Migration preflight rejects an active table/view identifier collision before
|
||||
it writes a migration fence. A pre-existing version 1 strong snapshot with
|
||||
such a collision is loaded in cleanup-only quarantine. Ambiguous reads fail
|
||||
closed; each cleanup mutation must reduce the collision set, and unrelated
|
||||
writes remain blocked until all collisions are removed. Drain catalog
|
||||
writers that predate cleanup quarantine before starting this repair, and
|
||||
complete cleanup before the first version 2 write. Restoring any version 1
|
||||
snapshot after a writer has observed version 2 fails closed instead of
|
||||
replacing the in-process catalog state.
|
||||
| Active-active multi-region writes | Not claimed | A table must not accept independent concurrent writers in multiple regions. |
|
||||
|
||||
## Production Failure Coverage
|
||||
|
||||
Positive client smoke proves a client can use a table. Production failure probes
|
||||
prove RustFS does not silently advance table state when a failure happens.
|
||||
Failure probes prove RustFS does not silently advance table state on failure. Tracked cases: stale commit token or base metadata returns a conflict without advancing the pointer; missing metadata, manifest, data, or delete objects fail closed before commit or maintenance; concurrent writers produce one CAS winner and retryable conflicts; catalog and S3 permission denials prevent data-plane bypass; stale maintenance plans fail closed before deletion or commit; post-CAS finalization gaps are visible and safely recoverable; external catalog sync conflicts leave pointer, token, and generation unchanged; backing migration stays blocked until WAL and recovery replay are clean.
|
||||
|
||||
The tracked failure cases are:
|
||||
|
||||
- stale commit token or stale base metadata returns a conflict without advancing
|
||||
the table pointer
|
||||
- missing metadata, manifest, data, or delete objects fail closed before commit
|
||||
or maintenance can advance state
|
||||
- concurrent writers produce a single winning CAS and retryable conflicts for
|
||||
stale writers
|
||||
- table catalog and ordinary S3 permission denials prevent data-plane bypass
|
||||
- stale maintenance plans fail closed before object deletion or catalog commit
|
||||
- post-CAS finalization gaps are visible through diagnostics and safe recovery
|
||||
- external catalog sync conflicts leave pointer, token, and generation unchanged
|
||||
- backing migration remains blocked until WAL and recovery replay are clean
|
||||
|
||||
Do not promote a failure case from a required live probe or load test to an
|
||||
automated claim until the exact RustFS build, client version, and expected
|
||||
response shape are recorded.
|
||||
Do not promote a failure case from live probe or load test to an automated claim until the exact RustFS build, client version, and expected response shape are recorded.
|
||||
|
||||
## Unsupported Or Not Claimed
|
||||
|
||||
RustFS does not currently claim:
|
||||
Full AWS S3 Tables control-plane parity; full MinIO AIStor private extensions; full Cloudflare R2 Data Catalog or Alibaba OSS Tables interoperability; built-in periodic maintenance scheduling; active-active multi-region writes; multi-table transactions; no-long-term-data-credential bootstrap; online vendor SDK polling; external catalog policy mirroring; delete-file rewrite or row-level compaction execution; built-in SQL execution; Delta Lake or Hudi; end-to-end SQL row-level DML validation through Spark, Trino, or another engine.
|
||||
|
||||
- full AWS S3 Tables control-plane API parity
|
||||
- full MinIO AIStor Tables private extension parity
|
||||
- full Cloudflare R2 Data Catalog interoperability
|
||||
- full Alibaba OSS Tables interoperability
|
||||
- built-in periodic maintenance scheduling; external schedulers can queue maintenance jobs and workers can claim them, but RustFS does not claim a continuous in-process scheduler
|
||||
- active-active multi-region table writes
|
||||
- multi-table transactions
|
||||
- no-long-term-data-credential table bootstrap
|
||||
- online external catalog vendor SDK polling
|
||||
- external catalog policy mirroring
|
||||
- delete-file rewrite or row-level compaction execution
|
||||
- built-in SQL query execution
|
||||
- Delta Lake or Hudi table format support
|
||||
- end-to-end SQL row-level DML validation through Spark, Trino, or another SQL engine
|
||||
## Verification
|
||||
|
||||
## Verification Commands
|
||||
|
||||
Use these commands when updating this matrix, release notes, or client
|
||||
compatibility claims:
|
||||
|
||||
```bash
|
||||
python3 scripts/table-catalog/test_pyiceberg_smoke.py
|
||||
python3 scripts/table-catalog/test_engine_compatibility.py
|
||||
python3 scripts/table-catalog/test_duckdb_smoke.py
|
||||
python3 scripts/table-catalog/test_failure_coverage.py
|
||||
python3 scripts/table-catalog/pyiceberg_smoke.py --print-client-matrix
|
||||
python3 scripts/table-catalog/pyiceberg_smoke.py --print-engine-compatibility
|
||||
python3 scripts/table-catalog/pyiceberg_smoke.py --print-production-failure-coverage
|
||||
python3 scripts/table-catalog/pyiceberg_smoke.py --print-vendor-profiles
|
||||
python3 scripts/table-catalog/pyiceberg_smoke.py --print-production-readiness
|
||||
python3 scripts/table-catalog/engine_compatibility.py --print-vendor-audit
|
||||
python3 scripts/table-catalog/engine_compatibility.py --print-spark-config
|
||||
python3 scripts/table-catalog/engine_compatibility.py --print-duckdb-rest-sql
|
||||
python3 scripts/table-catalog/engine_compatibility.py \
|
||||
--profile aws-s3tables \
|
||||
--region us-east-1 \
|
||||
--account-id 123456789012 \
|
||||
--table-bucket analytics \
|
||||
--print-spark-config
|
||||
python3 scripts/table-catalog/engine_compatibility.py \
|
||||
--metadata-location s3://rustfs-s3table-smoke/tables/table-id/metadata/v1.metadata.json \
|
||||
--print-live-conformance \
|
||||
--cleanup
|
||||
python3 scripts/table-catalog/engine_compatibility.py --print-live-evidence-schema
|
||||
python3 scripts/table-catalog/pyiceberg_smoke.py \
|
||||
--endpoint http://127.0.0.1:9000 \
|
||||
--bucket rustfs-s3table-smoke \
|
||||
--replace \
|
||||
--cleanup \
|
||||
--rustfs-build rustfs-v1.0.0-beta.8 \
|
||||
--git-sha "$(git rev-parse HEAD)" \
|
||||
--catalog-backing durable-strong \
|
||||
--live-evidence-output /tmp/rustfs-pyiceberg-live-evidence.json
|
||||
python3 scripts/table-catalog/engine_compatibility.py \
|
||||
--warehouse rustfs-s3table-smoke \
|
||||
--namespace smoke \
|
||||
--table events \
|
||||
--print-operations-guide
|
||||
python3 scripts/table-catalog/failure_coverage.py \
|
||||
--warehouse rustfs-s3table-smoke \
|
||||
--namespace smoke \
|
||||
--table events \
|
||||
--print-failure-probes
|
||||
python3 scripts/table-catalog/failure_coverage.py \
|
||||
--warehouse rustfs-s3table-smoke \
|
||||
--namespace smoke \
|
||||
--table events \
|
||||
--table-warehouse-location s3://rustfs-s3table-smoke/tables/table-id \
|
||||
--print-disaster-recovery-rehearsal
|
||||
python3 scripts/table-catalog/failure_coverage.py \
|
||||
--warehouse rustfs-s3table-smoke \
|
||||
--namespace smoke \
|
||||
--table events \
|
||||
--table-warehouse-location s3://rustfs-s3table-smoke/tables/table-id \
|
||||
--writer-count 8 \
|
||||
--maintenance-worker-count 2 \
|
||||
--iteration-count 50 \
|
||||
--print-scale-fault-rehearsal
|
||||
```
|
||||
Commands for updating this matrix, release notes, or client claims are maintained in [scripts/table-catalog/README.md](../../scripts/table-catalog/README.md); the unit tests are `scripts/table-catalog/test_*.py`.
|
||||
|
||||
## Release Claim Guidance
|
||||
|
||||
Use conservative release wording that matches the matrix.
|
||||
Acceptable: "RustFS includes a core Iceberg REST Catalog-based S3 Tables implementation with PyIceberg and DuckDB smoke coverage, table-aware S3 data-plane policy checks, controlled maintenance, catalog recovery diagnostics, manual conformance input for Spark, Trino, Databend, and Snowflake, production-failure probe harnesses, disaster-recovery and scale/fault rehearsal probes, and a machine-readable operations evidence guide."
|
||||
|
||||
Acceptable wording:
|
||||
Do not claim: "RustFS is fully compatible with AWS S3 Tables."
|
||||
|
||||
> RustFS includes a core Iceberg REST Catalog-based S3 Tables implementation
|
||||
> with PyIceberg and DuckDB smoke coverage, table-aware S3 data-plane policy checks,
|
||||
> controlled maintenance, catalog recovery diagnostics, manual conformance
|
||||
> input for Spark, Trino, Databend, and Snowflake, production-failure
|
||||
> probe harnesses, disaster-recovery and scale/fault rehearsal probes, and a
|
||||
> machine-readable production operations evidence guide.
|
||||
|
||||
Do not claim:
|
||||
|
||||
> RustFS is fully compatible with AWS S3 Tables.
|
||||
|
||||
Any stronger vendor or engine claim needs a repeatable live validation harness,
|
||||
the exact client versions used, and the expected response shapes recorded in the
|
||||
table-catalog inventories.
|
||||
Any stronger vendor or engine claim needs a repeatable live harness, the exact client versions, and the expected response shapes recorded in the table-catalog inventories.
|
||||
|
||||
## Related
|
||||
|
||||
- [Table catalog conformance scripts](../../scripts/table-catalog/README.md)
|
||||
- [Durable backing cutover runbook](../operations/s3-tables-cutover-runbook.md)
|
||||
- [Admin route action snapshot](admin-route-action-snapshot.md)
|
||||
- [Runtime capability contracts](runtime-capability-contracts.md)
|
||||
|
||||
@@ -1,29 +0,0 @@
|
||||
# Scanner/Heal admission Phase 0 baseline
|
||||
|
||||
This document records the current entry points and safety boundaries for backlog #1939. It is an inventory and test contract, not a lease design. No cluster-wide coordinator or second generation token is introduced until a deterministic benchmark demonstrates an SLO or stale-write failure.
|
||||
|
||||
## Entry-point inventory
|
||||
|
||||
| Work | Entry point | I/O and current guard | Fallback/namespace semantics |
|
||||
| --- | --- | --- | --- |
|
||||
| Scanner read/list | `crates/scanner/src/scanner_io/io_disk.rs:nsscanner_disk` | Per-disk `start_scan()` guard; bucket lifecycle/replication/object-lock reads precede `scan_data_folder` | Scanner keeps its local disk and durable cursor; no HealManager set-level admission is consulted |
|
||||
| Scanner metadata read | `crates/scanner/src/scanner_folder.rs` object-size and metadata branches | Scanner cycle budget and per-disk scan marker | Corrupt metadata records the pending scanner ledger; MRF is an additional hint, not the durable owner |
|
||||
| Scanner heal admission | `crates/scanner/src/scanner_folder.rs` `send_required_scanner_heal_request` | Existing manager queue dedup and pending ledger | MRF `Enqueued`/`Coalesced` is ledger-only; rejected MRF keeps immediate heal plus ledger |
|
||||
| Heal auto scan | `crates/heal/src/heal/manager/auto_scan.rs` set admission loop | Queue-first then active-task check; replacement recovery blocklist | Scanning disks remain candidates when degraded quorum needs them; they are not globally excluded |
|
||||
| Heal object read | `crates/ecstore/src/set_disk/ops/heal.rs` `heal_object` | Namespace write lock unless `no_lock`; reads file info before commit | Namespace lock is object-scoped and does not claim scanner cycle ownership |
|
||||
| Disk selection | `crates/ecstore/src/set_disk/ops/locking.rs` candidate selection | Healing disks are ordered after new disks; scanning disks may remain candidates | Degraded/quorum fallback is preserved |
|
||||
| Data movement | Existing storage-owned movement/publication generation (#1905/#1942) | This issue does not add a second coordinator | Future admission must validate the storage generation at the final commit |
|
||||
|
||||
## Baseline contract
|
||||
|
||||
The deterministic baseline in `scanner_heal_admission_baseline.rs` encodes the investigation matrix only: ScannerRead+HealRead may overlap, HealWrite conflicts with scanner reads, DataMovementWrite conflicts with all work, and independent set identities remain concurrent. It does not claim that production currently enforces the matrix.
|
||||
|
||||
The production facts that must be measured before Phase 1 are scanner p99, heal p99, cursor/checkpoint delay, queue and pending-ledger depth, and starvation by set. The benchmark matrix must include restart recovery, degraded quorum/scanning-disk fallback, urgent replacement heal, and at least two independent sets.
|
||||
|
||||
The executable fixture uses a fixed eight-sample restart/degraded sequence so the baseline is reproducible without wall-clock noise: two sets each receive ScannerRead, HealRead, HealWrite and a follow-up ScannerRead. Its expected synthetic p99 is 420 microseconds, maximum modeled backlog is 2, two HealWrite samples are deferred, and the independent second set still services three reads. These are fixture values, not production SLO claims; production benchmark output must replace them with measured p99, backlog and per-set wait distributions.
|
||||
|
||||
The inventory test reads the current source files and asserts the named guards/fallback branches are still present (`start_scan`, pending-ledger admission, Heal queue/active checks, namespace `get_write_lock`, and scanning-disk re-append). A source rename or guard removal therefore fails the baseline instead of silently leaving stale documentation.
|
||||
|
||||
Commit-time generation-fencing, lease-expiry, and lock-order tests are intentionally deferred until a Phase-0 fixture demonstrates a stale write or an SLO violation; arithmetic-only placeholders would stay green if production paths regressed.
|
||||
|
||||
If a future fixture demonstrates stale destructive writes, the fix must extend the storage-owned generation/admission primitive and validate the token at the final metadata/format/delete commit. Cancellation or a local lease alone is not a fence.
|
||||
@@ -1,7 +1,7 @@
|
||||
# Storage, Control Plane, And Background Controllers
|
||||
|
||||
This document defines migration boundaries for the storage hot path and adjacent
|
||||
control-plane responsibilities.
|
||||
**Use this when:** adding a storage API surface, a cluster read model, or a background-service status/reconcile surface, and you need to know which layer owns it and what must not drift.
|
||||
**Source of truth:** `crates/storage-api` (trait contracts), `crates/ecstore/src/api/mod.rs` (facade groups, `api::cluster`), `crates/ecstore/src/cluster/` (control plane), [background-controller-contract.md](background-controller-contract.md) (controller vocabulary).
|
||||
|
||||
## Storage API Contracts
|
||||
|
||||
@@ -42,9 +42,11 @@ It maps existing endpoint pools into the shared storage-api topology contract an
|
||||
an ECStore-owned static membership snapshot. It must not expose local disk paths,
|
||||
start health checks, mutate endpoint ownership, or change placement/readiness.
|
||||
The same facade also owns static pool-state, local-node storage, and peer-health
|
||||
status projections. Peer health remains explicitly unknown until a later slice
|
||||
wires real health signals; this document does not authorize background probes or
|
||||
RPC-based health checks.
|
||||
status projections. `peer_health_snapshot` in
|
||||
`crates/ecstore/src/cluster/control_plane.rs` projects the internode health
|
||||
tracker's per-node reachability (`PEER_HEALTH_REACHABLE` /
|
||||
`PEER_HEALTH_UNREACHABLE`, or not-reported); the facade itself starts no probes
|
||||
and issues no RPC-based health checks.
|
||||
Readiness impact for storage, lock quorum, peer health, probes, admin routes,
|
||||
RPC, and the S3 data plane is recorded in
|
||||
[`readiness-matrix.md`](readiness-matrix.md).
|
||||
|
||||
@@ -1,384 +1,112 @@
|
||||
# Unified Per-Object Generation Authority
|
||||
# Object Transaction UUID And Generation-Fencing Contract
|
||||
|
||||
Establishes a **single per-object generation authority** that spans object
|
||||
commit, GET snapshots, garbage collection, and quota accounting, and pins the
|
||||
transport, encoding, proto-evolution, and mixed-version contracts that every
|
||||
consumer must obey.
|
||||
**Use this when:** adding or changing anything that fences a commit, scopes a read lease, gates old-directory cleanup, binds prepared pool reads, or settles quota against "the current version of an object", or when adding a field that rides internode RPC or `xl.meta`.
|
||||
**Source of truth:** `assign_object_transaction_epoch` in `crates/ecstore/src/set_disk/ops/object.rs` and `crates/ecstore/src/set_disk/ops/multipart.rs`; `FileInfo::set_object_transaction_epoch` in `crates/filemeta/src/fileinfo.rs`; `commit_rename_data_dir` and `RenameConvergence` in `crates/ecstore/src/set_disk/core/io_primitives.rs`; `PreparedPoolReadFallbackBarrier` in `crates/ecstore/src/store/rebalance.rs`; `crates/protos/src/node.proto`; env constants in `crates/config/src/constants/object.rs` and `crates/config/src/constants/internode.rs`.
|
||||
|
||||
This is a **design and contract document**. It changes no storage code. It is
|
||||
the shared prerequisite for five implementation sub-issues under the
|
||||
[#1307](https://github.com/rustfs/backlog/issues/1307) adversarial-review
|
||||
program:
|
||||
[#1312](https://github.com/rustfs/backlog/issues/1312) (commit fencing),
|
||||
[#1313](https://github.com/rustfs/backlog/issues/1313) (read lease),
|
||||
[#1314](https://github.com/rustfs/backlog/issues/1314) (prepared pool read),
|
||||
[#1318](https://github.com/rustfs/backlog/issues/1318) (quota reservation), and
|
||||
[#1323](https://github.com/rustfs/backlog/issues/1323) (old-dir GC).
|
||||
Design tracking lives in `rustfs/backlog#1326`. This document holds only the invariants.
|
||||
|
||||
Tracks [rustfs/backlog#1326](https://github.com/rustfs/backlog/issues/1326).
|
||||
## Authority
|
||||
|
||||
## Why one authority
|
||||
The target contract requires **one per-object commit identity** consumed by commit fencing, read leases, cleanup, prepared reads, and quota settlement. No consumer may mint a second value and call it the same generation.
|
||||
|
||||
The #1307 adversarial-review verdict (issuecomment-4992565957) found that the
|
||||
five sub-issues each reach for their own generation / fencing / lease token to
|
||||
solve the same underlying problem — **commit mutual-exclusion plus snapshot
|
||||
lifetime**. Left independent, they diverge and punch through one another:
|
||||
What exists today is an **object transaction UUID**, not the target authority:
|
||||
|
||||
- #1323 old-dir GC can reclaim a directory still referenced by a #1313 lease if
|
||||
the two disagree on what "current generation" means.
|
||||
- #1312 fence epoch and #1318 quota reservation token, if derived from two
|
||||
different monotonic sources, cannot be compared — a late commit fenced on one
|
||||
plane can still settle quota on the other.
|
||||
| Property | Current implementation |
|
||||
|---|---|
|
||||
| Minting | `assign_object_transaction_epoch` mints a random non-nil UUID for PUT and CompleteMultipartUpload when the object-transaction gate is active. |
|
||||
| Persistence | Written through `FileInfo::set_object_transaction_epoch` into the version's internal metadata map under the dual-key contract (`x-rustfs-internal-*` / `x-minio-internal-*`). |
|
||||
| Fence check | The coordinator reads the current UUID (or `Absent`) and revalidates exact equality immediately before `rename_data`. |
|
||||
| Cleanup | Old-data cleanup receipts carry the committed UUID; reconciliation deletes only when the receipt UUID still equals the current object UUID. |
|
||||
|
||||
The fix is a single authority with one selected comparison rule, one persistence
|
||||
semantics, and one transport binding, that every consumer references rather than
|
||||
re-derives.
|
||||
This is an equality-CAS fence and cleanup identity. It is not a monotonic epoch, is not minted by the distributed lock grant, and is not compared atomically at each disk's `xl.meta` commit point. Documents and issues must call it the *object transaction UUID*, not proof that the generation authority exists.
|
||||
|
||||
## Target authority and the current bounded token
|
||||
### Authority modes (one must be selected)
|
||||
|
||||
The target contract still requires **one per-object commit identity** consumed
|
||||
by commit fencing, read leases, cleanup, prepared reads, and quota settlement.
|
||||
No consumer may mint a second value and call it the same generation.
|
||||
|
||||
The concrete ordering semantics are not settled, however. The original #1326
|
||||
proposal requires a total-ordered, monotonic lock-grant epoch. Current main does
|
||||
not implement that proposal. PR #6077 instead implements an opaque transaction
|
||||
identity:
|
||||
|
||||
- `assign_object_transaction_epoch` mints a random non-nil UUID for PUT and
|
||||
CompleteMultipartUpload when the object-transaction gate is active.
|
||||
- The UUID is written through `FileInfo::set_object_transaction_epoch` into the
|
||||
dual internal metadata map.
|
||||
- The coordinator reads the current UUID (or `Absent`) and revalidates exact
|
||||
equality immediately before `rename_data`.
|
||||
- Old-data cleanup receipts carry the committed UUID and reconciliation deletes
|
||||
only when the receipt UUID still equals the current object UUID.
|
||||
|
||||
This is a useful **equality-CAS fence and cleanup identity**. It is not a
|
||||
monotonic epoch, is not minted by the distributed lock grant, and is not
|
||||
compared atomically at each disk's `xl.meta` commit point. Until the decision
|
||||
below is made, documents and issue checklists must call it the *object
|
||||
transaction UUID* rather than use it as proof that the target generation
|
||||
authority exists.
|
||||
|
||||
### Ordering decision required
|
||||
|
||||
Before #1313, #1314, or a unified quota binding can consume the authority, one
|
||||
of these contracts must be selected and tested:
|
||||
|
||||
1. **Total-ordered fencing epoch.** A lock grant returns a durable per-object
|
||||
`(term, counter)` (or another specified total-order type). Every disk rejects
|
||||
a lower epoch at the atomic metadata commit point. The value never regresses
|
||||
across lock-plane restart, failover, or minority recovery.
|
||||
2. **Opaque commit-generation identity.** Consumers compare only exact identity;
|
||||
no `<` / `>` semantics are permitted. The authoritative commit must perform
|
||||
an atomic expected-generation CAS, and all lease, cleanup, prepared-read, and
|
||||
quota contracts must be rewritten in terms of “references this exact
|
||||
generation,” not “lower/newer generation.”
|
||||
|
||||
The current UUID implementation proves neither a durable total order nor a
|
||||
per-disk atomic expected-generation CAS, so it does not by itself decide between
|
||||
these options.
|
||||
|
||||
### Persistence semantics if total order is selected
|
||||
|
||||
A total-ordered epoch must be **monotonic across lock-plane restart and
|
||||
failover**. The distributed lock entry remains in-memory; deriving a counter
|
||||
from that entry alone would reset it after restart. The chosen source therefore
|
||||
must be either quorum-persisted before grant or derived from a durable term whose
|
||||
full `(term, counter)` comparison cannot regress. This requirement does not
|
||||
apply to an opaque UUID as an ordering rule; the opaque alternative instead
|
||||
requires atomic expected-identity comparison and durable crash recovery.
|
||||
|
||||
## Consumer binding contracts
|
||||
|
||||
### Current implementation snapshot (2026-08-31, main@9ee7b1221)
|
||||
|
||||
This table separates code that exists on current main from the target contract.
|
||||
Closing an implementation issue does not imply that its token is already the
|
||||
unified authority.
|
||||
|
||||
| Surface | Current main | Gap against this contract |
|
||||
| Mode | Contract | Persistence requirement |
|
||||
|---|---|---|
|
||||
| PUT / CompleteMultipartUpload (#1312, PR #6077) | Owned commit tasks retain the relevant guards; an opt-in gate persists a random object transaction UUID and performs a quorum metadata equality recheck before rename | no lock-grant monotonic source; no per-disk atomic epoch/CAS comparison; the live proof is the reused remote-version-state fleet proof, not a dedicated generation capability |
|
||||
| Old-data cleanup (#1323, PR #6077) | JSON receipt carries transaction UUID, old dir, and committed dir; reconciliation is gated and requires UUID equality | no generation-bound read lease is consulted, so this is crash cleanup fencing rather than the full #1313/#1323 lease lifetime contract |
|
||||
| Read lease (#1313) | short-term streaming/multipart path holds the namespace read lock through EOF/drop; deterministic part-boundary coverage is tracked by PR #6887 | no cross-node generation-bound lease registry, TTL reclamation, or crash recovery |
|
||||
| Prepared pool read (#1314) | PR #6889 tracks a pool-local prepared identity and fails closed/refetches when pool state changes | not merged on this snapshot; pool-local identity is not a cross-pool generation authority; black-box mixed-version/rebalance coverage remains open |
|
||||
| Quota reservation (#1318) | durable per-bucket ledger plus independent snapshot-lease mutation-fence tokens; issue closed after PR #6058 | reservation and settle are not bound to the object transaction UUID; the independent fence must be reconciled with the selected authority or explicitly proven to be a separate, non-generation arbitration domain |
|
||||
| Internode integrity (#1327, #1541, #1542) | v2/v3 HMAC binds audience, exact method, timestamp, nonce, canonical body digest, and receiver boot epoch; body-bound RPC policy has exact-set coverage | signature/body/replay strict switches remain default-off rollout gates; generation enforcement cannot treat an unrelated fleet-version proof as proof that these strict contracts converged |
|
||||
| Total-ordered fencing epoch | A lock grant returns a durable per-object `(term, counter)`; every disk rejects a lower epoch at the atomic metadata commit point; the value never regresses across lock-plane restart, failover, or minority recovery. | Quorum-persisted before grant, or derived from a durable term whose full comparison cannot regress. The in-memory distributed lock entry alone is insufficient. |
|
||||
| Opaque commit-generation identity | Consumers compare exact identity only; no `<` / `>` semantics. The authoritative commit performs an atomic expected-generation CAS; lease, cleanup, prepared-read, and quota contracts are phrased as "references this exact generation". | Atomic expected-identity comparison plus durable crash recovery. |
|
||||
|
||||
| Consumer | How it binds generation | Key invariant |
|
||||
|---|---|---|
|
||||
| #1312 commit fence | selected generation is checked at `rename`, rollback restore/delete, and cleanup mutation points using the chosen ordered or exact-CAS rule | a stale writer is rejected on **all** disks; an already-ACK'd write is never rolled back |
|
||||
| #1313 read lease | lease binds the exact generation observed at read time; GC runs only after every lease referencing that generation is released | lease is visible across nodes; a crashed reader's lease is reclaimed by TTL |
|
||||
| #1323 old-dir GC | cleanup job carries the committed generation; before deleting `old_dir` it confirms that no lease for the generation owning that directory remains | `old_dir != committed_dir`; a still-referenced directory is never deleted |
|
||||
| #1314 prepared pool read | the `PreparedPoolRead` bundle carries the generation resolved during pool lookup; the chosen pool's reader setup reuses it only after a match | generation mismatch forces a fallback to full metadata fanout |
|
||||
| #1318 quota reservation | reservation / settle record binds the exact object generation (and an ordered epoch too, if that option is selected) | a late commit cannot settle quota for a different committed generation |
|
||||
The current UUID proves neither a durable total order nor a per-disk atomic CAS, so it does not decide between the modes.
|
||||
|
||||
### Fence coverage is three disk-write points, not one (#1312 B2)
|
||||
## Consumer Binding
|
||||
|
||||
Checking the generation only before the `rename` fanout is insufficient. The
|
||||
authoritative commit sequence is `tmp sync → data-dir rename → xl.meta commit →
|
||||
directory sync` in `crates/ecstore/src/disk/local.rs`, and there are two further
|
||||
detachable disk-write points in
|
||||
`crates/ecstore/src/set_disk/core/io_primitives.rs`:
|
||||
| Consumer | Binds generation how | Key invariant | Current state |
|
||||
|---|---|---|---|
|
||||
| Commit fence (PUT / CompleteMultipartUpload) | Checked at `rename`, rollback restore/delete, and cleanup mutation points using the selected rule | A stale writer is rejected on **all** disks; an already-ACK'd write is never rolled back | Opt-in UUID equality recheck before rename; no per-disk atomic comparison |
|
||||
| Read lease | Lease binds the exact generation observed at read time; GC runs only after every lease on that generation is released | Lease visible across nodes; crashed reader's lease reclaimed by TTL | Streaming/multipart GET holds the namespace read lock through EOF/drop (part-boundary coverage: `#6887`); no cross-node generation-bound registry |
|
||||
| Old-dir GC | Cleanup job carries the committed generation and confirms no lease owns `old_dir` before deleting | `old_dir != committed_dir`; a still-referenced directory is never deleted | UUID receipt equality (`#6077`); no lease consultation |
|
||||
| Prepared pool read | The prepared bundle carries the generation resolved during pool lookup; the chosen pool reuses it only after a match | Mismatch forces fallback to full metadata fanout | `PreparedPoolReadFallbackBarrier` (`#6889`) is a pool-local identity that fails closed / refetches on pool state change; it is not a cross-pool authority |
|
||||
| Quota reservation | Reserve / settle record binds the exact object generation (and the ordered epoch too, if selected) | A late commit cannot settle quota for a different committed generation | Durable per-bucket ledger with independent snapshot-lease fence tokens (`#6058`); not bound to the transaction UUID |
|
||||
|
||||
- **Rollback restore/delete** — on quorum failure each disk can restore backup
|
||||
metadata or delete the failed version. A stale writer's rollback must compare
|
||||
the expected generation, otherwise it can overwrite or delete the winner's
|
||||
already-committed metadata.
|
||||
- **`commit_rename_data_dir`** — a cancel-then-detach disk-write point; the
|
||||
coordinator's "reap all child tasks" must explicitly include it so a cancelled
|
||||
writer cannot bypass fence/lease and keep deleting directories.
|
||||
## Fence Coverage: Three Disk-Write Points
|
||||
|
||||
If generation is validated only after data-dir rename, a fenced writer
|
||||
may already have renamed its data-dir into the object path, leaving a staged
|
||||
orphan. Either move the fence ahead of the data-dir rename, or declare that
|
||||
orphan an acceptable residue accounted for by GC metrics — the white-box
|
||||
acceptance "no background disk write after release" must be rewritten
|
||||
accordingly.
|
||||
Checking generation only before the `rename` fanout is insufficient. The commit sequence is `tmp sync → data-dir rename → xl.meta commit → directory sync` in `crates/ecstore/src/disk/local.rs`, and `crates/ecstore/src/set_disk/core/io_primitives.rs` has two further detachable disk-write points:
|
||||
|
||||
Current PR #6077 performs a quorum metadata equality recheck before rename and
|
||||
reaps owned commit work. That closes important cancellation windows, but it is
|
||||
not evidence that every disk mutation above performs the selected generation
|
||||
comparison atomically. The writer inventory and per-point CAS/ordering proof
|
||||
remain acceptance work for #1326 even though #1312 is closed.
|
||||
1. **Rollback restore/delete.** On quorum failure each disk can restore backup metadata or delete the failed version. A stale writer's rollback must compare the expected generation, or it can overwrite or delete the winner's committed metadata. Panic, cancel, and timeout outcomes must be reaped into coordinator convergence rather than skip rollback through an early return.
|
||||
2. **`commit_rename_data_dir`.** A cancel-then-detach disk-write point; the coordinator's "reap all child tasks" must include it so a cancelled writer cannot bypass fence or lease and keep deleting directories.
|
||||
|
||||
### Post-commit convergence is orthogonal to the fence (#1321)
|
||||
If generation is validated only after the data-dir rename, a fenced writer may already have renamed its data-dir into the object path, leaving a staged orphan. Either move the fence ahead of the data-dir rename, or declare that orphan an accepted residue accounted for by GC metrics.
|
||||
|
||||
The same `SetDisks::rename_data` path already returns a post-commit
|
||||
convergence classification (`RenameConvergence`, rustfs/backlog#1321) that
|
||||
tells the caller whether the *committed* replicas need heal to converge —
|
||||
`AllSuccessIdentical` (no heal), `PartialCommit` (a replica failed/offline),
|
||||
`SignatureDivergent` (committed replicas' version signatures differ), or
|
||||
`Unknown` (no signature was produced, e.g. >10 versions — scanner-backstopped).
|
||||
This replaced an earlier `Option<Vec<u8>>` heuristic under which any
|
||||
version signature looked like "needs heal", so every healthy multipart
|
||||
completion self-enqueued.
|
||||
`RenameConvergence` (`AllSuccessIdentical` / `PartialCommit` / `SignatureDivergent` / `Unknown`) is a *post-commit* heal signal on the same `rename_data` path; the fence is a *commit* gate. They compose: the fence decides whether a convergence is produced, `RenameConvergence` classifies it. A fence-aware convergence variant would be an additive enum change.
|
||||
|
||||
Convergence is a *post-commit* signal (the write landed; do the replicas need
|
||||
reconciliation), whereas the #1312 fence is a *commit* gate (a stale epoch is
|
||||
rejected before the write lands, surfaced through the existing `Result::Err`
|
||||
channel). They compose on the one `rename_data` path rather than competing:
|
||||
the fence decides whether a convergence is produced at all, and
|
||||
`RenameConvergence` classifies it once produced. A future fence-aware
|
||||
convergence variant, if ever needed, is an additive change to that enum and
|
||||
does not disturb the epoch comparison at the disk-write points above.
|
||||
## Transport And Security
|
||||
|
||||
## Transport and security contract
|
||||
Generation and derived tokens (lease, reservation) cross node boundaries in internode RPC bodies; every such flow must be signature-bound.
|
||||
|
||||
Generation and all derived tokens (lease, reservation) cross node boundaries in
|
||||
internode RPC bodies. Every such flow must be signature-bound.
|
||||
| Rule | Detail |
|
||||
|---|---|
|
||||
| HMAC scope | Target audience, exact service/method, timestamp, nonce, canonical body digest, receiver replay (boot) epoch. The receiver consumes the nonce in a bounded replay cache; a transmitted-but-unconsumed nonce is not replay protection. |
|
||||
| Current substrate | RPC v2/v3 in `crates/ecstore/src/cluster/rpc/http_auth.rs` binds all of the above. Body-bound policy covers mutating disk RPCs including `RenameData`, whose versioned canonical body includes every `RenameDataRequest` field, so the `FileInfo` metadata map carrying the UUID is authenticated. |
|
||||
| Strict switches | `RUSTFS_INTERNODE_RPC_SIGNATURE_STRICT`, `RUSTFS_INTERNODE_RPC_BODY_DIGEST_STRICT`, `RUSTFS_INTERNODE_RPC_REPLAY_SCOPE_STRICT` (`crates/config/src/constants/internode.rs`) are default-off rollout gates governed by [compat-cleanup-register.md](compat-cleanup-register.md). A generation capability may claim strong transport binding only after the relevant strict modes have converged fleet-wide. |
|
||||
| Acceptance tests per consumer | Method substitution, canonical body tamper, nonce replay, receiver restart, stripped-strict-metadata negatives. |
|
||||
|
||||
### RPC signature binding (#1312 B3, #1313, #1318)
|
||||
## Encoding Rules
|
||||
|
||||
**Requirement.** The canonical body carrying a generation or derived token must
|
||||
be folded into the internode HMAC. The authenticated scope binds the target
|
||||
audience, exact service/method, timestamp, nonce, canonical body digest, and
|
||||
receiver replay epoch. The receiver must consume the nonce in a bounded replay
|
||||
cache; transmitting a nonce without receiver-side consumption is not replay
|
||||
protection.
|
||||
| Rule | Reason |
|
||||
|---|---|
|
||||
| **Do not bump `XL_META_VERSION` or `XL_HEADER_VERSION`** (`crates/filemeta/src/filemeta.rs`). | `decode_xl_headers` in `crates/filemeta/src/filemeta/codec.rs` rejects newer values outright; a bump makes every new `xl.meta` unreadable by rolling-upgrade old nodes and by MinIO. See [minio-file-format-compat.md](minio-file-format-compat.md). |
|
||||
| **Do not add generation as a `FileInfo` struct field.** | Internode RPC serializes `FileInfo` with two msgpack encoders: positional-array encoding for the `read_version` family (a new positional field breaks mixed-version decode) and `encode_msgpack_named` (named-map) for `rename_data` in `rustfs/src/storage/rpc/node_service/disk.rs`. A field would have to be correct under both plus the JSON compatibility twin. Use the metadata map, which rides every encoder unchanged. |
|
||||
| **Metadata-map dual key.** | The UUID lives under `x-rustfs-internal-*` / `x-minio-internal-*`; missing, malformed, nil, or conflicting dual values fail closed when fencing is active. |
|
||||
| **No sidecar unless atomic.** | An epoch sidecar outside `xl.meta` is admissible only if it commits at the same atomic/CAS point as `xl.meta` with a specified crash-recovery protocol. None is implemented. |
|
||||
| **Regression guard.** | The real-MinIO `xl.meta` interop fixtures in `crates/filemeta/src/filemeta.rs` must keep passing: objects written by a new node stay readable by old RustFS nodes and by MinIO in both upgrade directions. |
|
||||
|
||||
**Current substrate (verified on main).** The original legacy-only description
|
||||
is obsolete:
|
||||
### Wire-encoding window (JSON and msgpack)
|
||||
|
||||
- RPC v2 binds target audience, exact method, POST, timestamp, nonce, and body
|
||||
digest.
|
||||
- Body-bound policy covers mutating disk RPCs including `RenameData`; its
|
||||
versioned canonical body includes every `RenameDataRequest` field, so the
|
||||
`FileInfo` metadata map carrying the transaction UUID is authenticated.
|
||||
- PR #5425 extended canonical-body enforcement to implemented non-disk mutating
|
||||
unary RPCs and added an exact policy/handler coverage partition.
|
||||
- PR #5455 added the receiver boot epoch and rotating replay scope so signatures
|
||||
captured before a receiver restart are rejected after capability convergence.
|
||||
|
||||
The rollout switches
|
||||
`RUSTFS_INTERNODE_RPC_SIGNATURE_STRICT`,
|
||||
`RUSTFS_INTERNODE_RPC_BODY_DIGEST_STRICT`, and
|
||||
`RUSTFS_INTERNODE_RPC_REPLAY_SCOPE_STRICT` remain default-off for rolling
|
||||
compatibility. The compatibility register and fallback/overflow metrics govern
|
||||
their fleet convergence. Therefore a generation capability may claim strong
|
||||
transport binding only when the relevant strict modes have converged; the
|
||||
object-transaction gate's current remote-version-state fleet proof is not, by
|
||||
itself, proof of RPC signature/body/replay strictness.
|
||||
|
||||
Acceptance for each generation consumer includes method substitution, canonical
|
||||
body tamper, nonce replay, receiver restart, and stripped-strict-metadata
|
||||
negative tests. Generation rollout must also record which strict-mode evidence
|
||||
authorized enforcement.
|
||||
|
||||
### Encoding contract (#1312 B1)
|
||||
|
||||
The on-disk persistence of generation must not perturb the file format:
|
||||
|
||||
- **Do not bump `XL_META_VERSION` / `XL_HEADER_VERSION`.**
|
||||
`crates/filemeta/src/filemeta/codec.rs` rejects `meta_ver > 3` and
|
||||
`header_ver > 3` outright (`decode_xl_headers`), and both constants are `3`
|
||||
(`crates/filemeta/src/filemeta.rs:53-54`). Bumping either makes every new
|
||||
`xl.meta` unreadable by rolling-upgrade old RustFS nodes and by MinIO — a
|
||||
total read failure, not a graceful downgrade.
|
||||
- **Do not add generation as a `FileInfo` struct field.** The internode RPC layer serializes `FileInfo` with two different msgpack encoders depending on the call site: `encode_msgpack` uses rmp_serde's default **array** (positional) encoding for the `read_version` family, where a new positional field breaks decode across mixed-version nodes; `encode_msgpack_named` uses `.with_struct_map()` (named-map) encoding for `rename_data` (`crates/ecstore/src/cluster/rpc/remote_disk.rs`), which is more tolerant but still requires `#[serde(default)]` and MinIO-side agreement. Because a `FileInfo` field would have to be correct under *both* encoders and under the JSON compatibility twin (see "Wire-encoding migration" below), do not add one — use the metadata map, which rides through every encoder unchanged.
|
||||
- **Where it lives today.** The object transaction UUID uses the version's
|
||||
internal metadata map under the dual-key contract
|
||||
(`x-rustfs-internal-*` / `x-minio-internal-*`) via
|
||||
`set_object_transaction_epoch`. Missing, malformed, nil, or conflicting dual
|
||||
values fail closed when fencing is active.
|
||||
- **Sidecars are not an equivalent alternative.** A future sidecar is admissible
|
||||
only if it commits atomically with `xl.meta` and has a specified crash-recovery
|
||||
protocol. No such protocol is implemented, so a sidecar cannot be selected by
|
||||
an implementation issue merely because this document mentions one.
|
||||
- **Regression guard.** Preserve the #4377 real-MinIO `xl.meta` interop
|
||||
regression (the fixture family around `crates/filemeta/src/filemeta.rs`):
|
||||
objects written by a new node must still be readable by old RustFS nodes and
|
||||
by MinIO, in both upgrade and downgrade directions.
|
||||
|
||||
### Wire-encoding migration (JSON → msgpack) interaction
|
||||
|
||||
The internode RPC layer retains a JSON/msgpack rolling-compatibility window, and
|
||||
generation-bearing fields must respect it.
|
||||
|
||||
- **Dual-field transport.** Each dual-encoded RPC field exists twice in `crates/protos/src/node.proto`: a JSON `string` field and a msgpack `bytes _bin` field (e.g. `file_info` #4 alongside `file_info_bin` #7 on `RenameDataRequest`). Senders emit both; receivers `decode_msgpack_or_json` prefer the `_bin` form and fall back to the JSON string only when `_bin` is empty (`crates/ecstore/src/cluster/rpc/remote_disk.rs`).
|
||||
- **Capability flags, default off.** `rustfs_protos::internode_rpc_msgpack_only()` only drops the redundant JSON copy when both `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true` and `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED=true` are deliberately enabled after the JSON-fallback metric reads zero fleet-wide and the convergence runbook is followed. Generation follows the same default-off, fleet-confirmed, metric-reads-zero rollout discipline, but a msgpack proof is not itself a generation capability proof.
|
||||
- **Generation must ride both encodings during the window.** If epoch lives in the version's internal metadata map, that map is carried inside `FileInfo`, so it is present in both the msgpack `_bin` and JSON copies automatically — good. But any new *top-level* generation datum must be added to **both** the msgpack and JSON representations (and, for msgpack, be safe under both the array and named-map encoders). A field added to only one encoding is silently lost the moment a peer falls back to the other — exactly the failure the JSON-fallback metric exists to catch.
|
||||
- **Signature binds a canonical form.** `RenameDataRequest` now has a versioned,
|
||||
injective canonical-body encoder that covers both compatibility fields and is
|
||||
authenticated independently of whichever JSON/msgpack decoder branch a peer
|
||||
consumes. A generation-capable strict request must reject missing or
|
||||
mismatched canonical-body metadata; it must not silently downgrade to an
|
||||
unauthenticated JSON twin.
|
||||
- Dual-encoded RPC fields exist twice in `crates/protos/src/node.proto`: a JSON `string` field and a msgpack `bytes *_bin` field (e.g. `file_info` and `file_info_bin` on `RenameDataRequest`). Senders emit both; receivers (`decode_msgpack_or_json` in `crates/ecstore/src/cluster/rpc/remote_disk.rs`) prefer `_bin` and fall back to JSON only when `_bin` is empty.
|
||||
- `rustfs_protos::internode_rpc_msgpack_only()` drops the JSON copy only when both `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY` and `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED` are set after the JSON-fallback metric reads zero fleet-wide.
|
||||
- Generation inside the `FileInfo` metadata map is carried in both copies automatically. Any new *top-level* generation datum must be added to both encodings and be safe under both msgpack encoders; a field in only one encoding is silently lost when a peer falls back.
|
||||
- `RenameDataRequest` has a versioned, injective canonical-body encoder covering both compatibility fields; a strict generation-capable request must reject missing or mismatched canonical-body metadata rather than downgrade to the unauthenticated JSON twin.
|
||||
|
||||
### Proto evolution
|
||||
|
||||
No top-level proto field is required by the current metadata-map UUID. If a
|
||||
future ordered epoch or explicit expected-generation is added to proto, it uses
|
||||
**proto3 `optional`** (explicit presence). A non-optional scalar is forbidden:
|
||||
an old coordinator talking to a new disk decodes absence as a plausible zero.
|
||||
No top-level proto field is required by the metadata-map UUID. If an ordered epoch or explicit expected-generation is ever added to proto, it uses **proto3 `optional`** (explicit presence). A non-optional scalar is forbidden: an old coordinator talking to a new disk decodes absence as a plausible zero.
|
||||
|
||||
### Mixed-version gate — one direction
|
||||
## Mixed-Version Gate: One Direction
|
||||
|
||||
When generation enforcement is not explicitly requested, or fleet confirmation
|
||||
is absent, behavior falls back to current semantics. Fail-closed is reserved for
|
||||
an explicit administrator-confirmed strict rollout.
|
||||
When generation enforcement is not explicitly requested, or fleet confirmation is absent, behavior falls back to current semantics. Fail-closed is reserved for an explicit administrator-confirmed strict rollout.
|
||||
|
||||
Current object transaction fencing follows that direction:
|
||||
| Flag (`crates/config/src/constants/object.rs`) | Default | Effect |
|
||||
|---|---|---|
|
||||
| `RUSTFS_OBJECT_TRANSACTION_FENCING_WRITE` | false | With either flag absent, PUT/MPU neither persists nor consumes the transaction UUID. |
|
||||
| `RUSTFS_OBJECT_TRANSACTION_FENCING_FLEET_CONFIRMED` | false | With both enabled, failure to obtain or retain the live fleet proof rejects the commit before rename. |
|
||||
|
||||
- `RUSTFS_OBJECT_TRANSACTION_FENCING_WRITE` and
|
||||
`RUSTFS_OBJECT_TRANSACTION_FENCING_FLEET_CONFIRMED` both default false.
|
||||
- With either flag absent, PUT/MPU does not persist or consume the transaction
|
||||
UUID.
|
||||
- With both flags enabled, failure to obtain or retain the live fleet proof
|
||||
rejects the commit before rename.
|
||||
The fleet proof is currently borrowed from the remote-version-state writer rollout. It proves membership/process-epoch convergence for that feature only; it does not prove an epoch type, per-disk CAS support, or RPC strict-mode convergence, and must not be treated as the final generation handshake.
|
||||
|
||||
This is an opt-in strict gate, not a negotiated generation capability. The
|
||||
proof is currently borrowed from the remote-version-state writer rollout. It
|
||||
proves current membership/process-epoch convergence for that feature, but does
|
||||
not prove an epoch type, per-disk generation CAS support, or RPC strict-mode
|
||||
convergence. Treating it as the final handshake is forbidden without an
|
||||
explicit proof mapping for those properties.
|
||||
### Capability negotiation (target)
|
||||
|
||||
## Capability negotiation
|
||||
Generation enforcement requires one **live fleet proof** containing at least: the selected authority version and comparison mode; the current membership/topology fingerprint and process epochs; support for every required disk mutation point; RPC signature/body/replay strict convergence; and the on-disk encoding version (the metadata-map UUID is version 1). Membership change or an old-node rejoin revokes the proof; revocation before commit fails an explicitly strict request and never rewrites or lowers a persisted generation. The proof may extend the authenticated fleet-proof machinery in `notification_sys` or the runtime capability contract; this document requires one shared token, not a mechanism.
|
||||
|
||||
Generation enforcement requires one **live fleet proof**, not independent
|
||||
boolean guesses in each consumer. The proof contract contains at least:
|
||||
## Open Decisions
|
||||
|
||||
1. the selected authority version and comparison mode (ordered or exact-CAS),
|
||||
2. the current membership/topology fingerprint and process epochs,
|
||||
3. support for every required disk mutation point,
|
||||
4. RPC signature/body/replay strict convergence, and
|
||||
5. the on-disk encoding version (the current metadata-map UUID is version 1).
|
||||
Blockers for calling the contract implemented:
|
||||
|
||||
The authoritative writer enables enforcement only while every target disk in
|
||||
the set is covered by a current proof. Membership change or an old node rejoin
|
||||
revokes that proof. Revocation before commit fails an explicitly strict request;
|
||||
when strict generation was never requested, the request remains on the legacy
|
||||
path. Revocation never rewrites or lowers an already-persisted generation.
|
||||
|
||||
The existing fleet-proof machinery in `notification_sys` may be reused if its
|
||||
authenticated statements are extended to cover the properties above. The
|
||||
runtime capability contract may instead expose the proof. This document does
|
||||
not choose the storage mechanism; it requires one token whose acquisition and
|
||||
revalidation semantics are shared by all consumers.
|
||||
|
||||
## Implementation order
|
||||
|
||||
Some original prerequisites have landed, but not in the originally proposed
|
||||
form. Remaining work follows this order:
|
||||
|
||||
1. **Resolve the authority mode in #1326.** Select total order or opaque
|
||||
exact-CAS, specify its atomic commit point, and audit PR #6077 against it.
|
||||
Do not retrofit ordering semantics onto the existing random UUID.
|
||||
2. **Define the generation fleet proof.** Map generation enablement to the RPC
|
||||
signature/body/replay strict proofs delivered by #1327/#1541/#1542 and to
|
||||
the selected per-disk comparison capability. Keep all strict defaults off
|
||||
until fallback metrics converge.
|
||||
3. **Implement #1313 generation-bound read leases.** The lease registry,
|
||||
cross-node visibility, TTL, and crash recovery must exist before old-dir GC
|
||||
can claim the full snapshot-lifetime guarantee. #1325 supplies the required
|
||||
multi-node failure tests.
|
||||
4. **Bind #1314 prepared reads.** A bundle binds the exact selected generation
|
||||
within its source pool. Cross-pool ordering is forbidden until a common
|
||||
authority is demonstrated. Validate rebalance and mixed-version fallback in
|
||||
the #1325 multi-pool harness.
|
||||
5. **Reconcile #1318 quota fencing.** Either bind reserve/settle/reconcile to
|
||||
the selected object generation or document and prove that its independent
|
||||
snapshot-lease fence is a separate arbitration domain that cannot settle a
|
||||
different generation.
|
||||
6. **Re-audit #1323 cleanup.** The existing UUID receipt remains valid crash
|
||||
cleanup, but full closure against active readers requires the #1313 lease
|
||||
check and the selected generation semantics.
|
||||
|
||||
## Open design decisions (pin before contract closure)
|
||||
|
||||
The following decisions remain blockers for calling the contract implemented:
|
||||
|
||||
- **Authority mode.** Choose total order or opaque exact-CAS. If total order is
|
||||
selected, define the type, per-object scope, persistence, overflow, and
|
||||
never-regress restart/minority-recovery tests. If opaque identity is selected,
|
||||
define the atomic expected-generation CAS and remove all ordered wording.
|
||||
- **Complete xl.meta-writer coverage.** Enumerate commit rename, rollback
|
||||
restore/delete, cleanup, heal, transition, restore, replication, and data
|
||||
movement. Each path must compare/carry the selected generation or be proved
|
||||
incapable of replacing the authoritative object identity.
|
||||
- **Rollback is an expected-generation CAS (#1312 B2).** The quorum-failure
|
||||
rollback in `rename_data` can restore backup metadata, not just remove a
|
||||
writer-private temporary file. It must execute only when the stored generation
|
||||
still matches the failed writer's expected generation. Panic, cancel, and
|
||||
timeout outcomes must be reaped into coordinator convergence rather than skip
|
||||
rollback through an early return.
|
||||
- **Sidecar is excluded unless proven atomic.** An epoch sidecar outside `xl.meta` is only admissible if it commits at the same atomic/CAS point as `xl.meta` with a defined recovery; otherwise it opens a crash gap and must be rejected in favor of the version-internal metadata map. The earlier "metadata map or sidecar" phrasing does not treat the two as equally safe.
|
||||
- **Generation capability proof.** Decide whether to extend the current
|
||||
authenticated fleet proof or the runtime capability contract. It must prove
|
||||
authority version, mutation coverage, topology/process epoch, and RPC strict
|
||||
convergence in one revalidatable token.
|
||||
- **Read-lease and GC crash recovery.** Select the cross-node registry, TTL
|
||||
reclamation, lease-holder crash behavior, and GC-executor recovery. The
|
||||
current cleanup receipt equality check does not answer these questions.
|
||||
- **Quota reserve → commit → settle binding.** The durable ledger's idempotency
|
||||
exists, but its independent mutation tokens must be related to the selected
|
||||
object generation with a concrete late-settle rejection test.
|
||||
- **PreparedPoolRead is pool-local only.** A #1314 bundle's generation validates freshness only within the pool that produced it. It cannot order commits across different pools unless a cross-pool common authority exists; absent that, the multi-pool wait cannot be short-circuited.
|
||||
- **Hot-path cost is a blocking metric.** Measure any additional consensus
|
||||
write, fsync, fleet-proof lookup, lease operation, or centralized serialization
|
||||
under 4 KiB and high-concurrency hot-key/hot-bucket A/B.
|
||||
- **Test infrastructure.** #1325 still lacks the complete 4-node × 4-drive,
|
||||
2-pool, directed network-fault, and large-object budget needed for restart,
|
||||
mixed-version, and cross-node lease acceptance.
|
||||
|
||||
## Acceptance for this contract
|
||||
|
||||
- [x] Architecture document exists and is linked from the architecture index.
|
||||
- [x] Transport signature, encoding, proto presence, mixed-version direction,
|
||||
and capability-proof requirements are defined once.
|
||||
- [x] Current implementations are separated from target guarantees; a closed
|
||||
child issue is not treated as proof of unified generation binding.
|
||||
- [ ] Authority mode and atomic comparison semantics are selected and tested.
|
||||
- [ ] #1312 / #1313 / #1314 / #1318 / #1323 bodies reference this document and
|
||||
use the selected authority terminology.
|
||||
- [ ] #1313 and #1314 bind the selected generation and pass #1325 multi-node /
|
||||
multi-pool failure tests.
|
||||
- [ ] #1318 either binds reserve/settle to the selected generation or provides
|
||||
an accepted proof that its separate fence cannot cross-settle generations.
|
||||
- [ ] #1323 reconciliation checks both committed generation and active
|
||||
generation-bound leases.
|
||||
- [ ] Generation strict enablement is backed by one live proof that includes RPC
|
||||
signature/body/replay strict convergence and per-disk comparison support.
|
||||
1. **Authority mode.** Total order or opaque exact-CAS. Do not retrofit ordering semantics onto the existing random UUID.
|
||||
2. **Complete `xl.meta`-writer coverage.** Enumerate commit rename, rollback restore/delete, cleanup, heal, transition, restore, replication, and data movement; each path compares/carries the selected generation or is proved incapable of replacing the authoritative identity.
|
||||
3. **Rollback as expected-generation CAS.** The quorum-failure rollback in `rename_data` restores backup metadata, not just a private temp file; it must run only when the stored generation still matches the failed writer's expectation.
|
||||
4. **Generation capability proof.** Extend the fleet proof or the runtime capability contract; one revalidatable token.
|
||||
5. **Read-lease and GC crash recovery.** Cross-node registry, TTL reclamation, lease-holder crash behavior, GC-executor recovery.
|
||||
6. **Quota reserve → commit → settle binding.** Relate the ledger's independent mutation tokens to the selected generation, with a concrete late-settle rejection test, or prove the fence is a separate arbitration domain that cannot cross-settle.
|
||||
7. **Prepared reads stay pool-local.** `PreparedPoolReadFallbackBarrier` validates freshness only within the pool that produced it; cross-pool ordering requires a common authority, and the multi-pool wait cannot be short-circuited without one.
|
||||
8. **Hot-path cost is a blocking metric.** Measure any added consensus write, fsync, fleet-proof lookup, lease operation, or centralized serialization under 4 KiB and hot-key/hot-bucket A/B.
|
||||
9. **Test infrastructure.** Multi-node, multi-pool, directed network-fault, and large-object budget for restart, mixed-version, and cross-node lease acceptance.
|
||||
|
||||
@@ -1,130 +1,40 @@
|
||||
# Workload Admission Contracts
|
||||
|
||||
This document records the `rustfs/backlog#660` PR-05 and PR-07 scheduler
|
||||
preservation and runtime workload-class contract slice.
|
||||
**Use this when:** adding a workload class or snapshot provider, consuming admission state from a background job, or deciding whether a job can "join" admission (it cannot; see Observation Surface Only).
|
||||
**Source of truth:** `WorkloadClass`, `AdmissionState`, `WorkloadAdmissionSnapshot`, `WorkloadAdmissionRegistrySnapshot`, `WorkloadAdmissionSnapshotProvider`, and `foreground_pressure` in `crates/concurrency/src/workload.rs`; the provider and consumer files named below.
|
||||
|
||||
## Preservation Coverage
|
||||
## Contract Shapes
|
||||
|
||||
The `rustfs-concurrency` tests pin the current reusable admission-facing
|
||||
behavior before later snapshot extraction:
|
||||
`rustfs-concurrency` owns the read-only shapes. `WorkloadClass` enumerates the admission categories (the variants are the source of truth); `AdmissionState`, `WorkloadAdmissionSnapshot`, and `WorkloadAdmissionRegistrySnapshot` are status shapes for runtime owners to fill. They do not replace the scheduler, request guards, scanner, heal, replication, or ECStore placement behavior. `GetObjectQueueSnapshot` permit semantics (saturated, over-available, zero-total) and worker-slot over-release clamping are pinned by `rustfs-concurrency` tests; scheduler buffer/priority behavior is pinned by `rustfs-io-core` and `rustfs/src/storage/concurrency/` tests.
|
||||
|
||||
- Worker slot over-release remains clamped by the configured worker limit.
|
||||
- `GetObjectQueueSnapshot` preserves saturated, over-available, and zero-total
|
||||
permit semantics.
|
||||
## Class To Provider Table
|
||||
|
||||
The former reusable scheduler and backpressure-pipe facades (and their
|
||||
preservation tests) were removed as zero-caller dead code in backlog#1025;
|
||||
scheduler buffer/priority behavior is now pinned by `rustfs-io-core` and
|
||||
`rustfs/src/storage/concurrency` tests.
|
||||
| Class | Provider (`impl WorkloadAdmissionSnapshotProvider`) | `active` / `queued` / `limit` source | Reports `Unknown` when |
|
||||
|---|---|---|---|
|
||||
| `ForegroundRead` | `ConcurrencyManager` in `rustfs/src/storage/concurrency/manager.rs` (source of truth); re-exposed unchanged by the RustFS runtime provider | disk-read permits in use / `None` (the semaphore exposes no waiter count) / configured max concurrent disk reads | the storage registry has no entry |
|
||||
| `ForegroundWrite` | none | none | always: no write-specific admission owner exposes a read-only surface yet |
|
||||
| `Metadata` | `RustFsWorkloadAdmissionSnapshotProvider` in `rustfs/src/workload_admission.rs` | `Open` once the bucket metadata runtime handle exists; no counts | bucket metadata runtime not initialized |
|
||||
| `Scanner` | same | scanner active work-unit counter / none / none | the counter is zero (idle and uninitialized are indistinguishable) |
|
||||
| `Repair` | same | heal active tasks / heal queue length / `None` (limits live behind the async heal manager state) | heal manager not initialized |
|
||||
| `Replication` | same | active regular + large-object + MRF workers / site replication queue count / `None` (limits owned by the async pool and resize policy) | replication runtime not initialized, or queue stats currently locked |
|
||||
|
||||
## Workload Class Contract
|
||||
## Observation Surface Only
|
||||
|
||||
`WorkloadClass` defines the required future admission categories:
|
||||
This is an observation surface only. Permit acquisition, priority assignment, buffer sizing, storage media detection, request guards, queue capacity, heal admission and priority merge/drop policy, replication worker resize and MRF handling, scanner cycle scheduling, bucket metadata loading and locks, and object write paths are unchanged by any provider. There is no runtime admission API for a background job to join; a job that needs bounded contention must bound it itself (see [kms-bulk-rekey-contract.md](kms-bulk-rekey-contract.md)).
|
||||
|
||||
- Foreground read.
|
||||
- Foreground write.
|
||||
- Metadata.
|
||||
- Scanner.
|
||||
- Repair.
|
||||
- Replication.
|
||||
Consumers that read the snapshot to self-throttle exist, and they do not change the owners' decisions:
|
||||
|
||||
`AdmissionState`, `WorkloadAdmissionSnapshot`, and
|
||||
`WorkloadAdmissionRegistrySnapshot` define read-only status shapes for later
|
||||
runtime owners. They do not replace the current scheduler, request guard,
|
||||
scanner, heal, replication, or ECStore placement behavior.
|
||||
| Consumer | File | Behavior |
|
||||
|---|---|---|
|
||||
| Data-movement backpressure (decommission, rebalance) | `crates/ecstore/src/data_movement/backpressure.rs` (`wait_for_data_movement_admission`, `foreground_pressure`) | Delays the next data-movement step while `ForegroundRead` or `ForegroundWrite` usage exceeds the configured high-water percent. ECStore receives the provider through `set_workload_admission_snapshot_provider` (`crates/ecstore/src/lib.rs`), published from `rustfs/src/startup_background.rs`; with no provider the step is admitted immediately. |
|
||||
| Heal manager mainline throttle | `crates/heal/src/heal/manager.rs` (`new_with_workload_provider`) | When `mainline_throttle_enable` is set, defers heal work while `ForegroundRead` or `ForegroundWrite` utilization exceeds the configured high-water percents; with no provider or the throttle disabled, heal pacing is unchanged. |
|
||||
|
||||
## Boundary Rules
|
||||
|
||||
- `rustfs-concurrency` owns this reusable contract surface.
|
||||
- The contract does not depend on `rustfs-ecstore` or RustFS binary runtime
|
||||
state.
|
||||
- No scheduler decision logic, queue capacity, Tokio runtime default, scanner
|
||||
admission, heal admission, replication admission, placement, membership, or
|
||||
NUMA behavior changes are part of this slice.
|
||||
- `rustfs-concurrency` owns the contract surface and does not depend on `rustfs-ecstore` or RustFS binary runtime state.
|
||||
- Adding a class or provider changes no scheduler decision logic, queue capacity, Tokio runtime default, scanner/heal/replication admission, placement, membership, or NUMA behavior.
|
||||
- Providers report `Unknown` rather than blocking or guessing when their owner is uninitialized or its stats are not immediately observable.
|
||||
|
||||
## Set-Local Snapshot Extraction
|
||||
## Provider Composition
|
||||
|
||||
The RustFS storage `ConcurrencyManager` now implements
|
||||
`WorkloadAdmissionSnapshotProvider` for local foreground-read admission:
|
||||
|
||||
- `ForegroundRead` reports local disk-read permit usage through
|
||||
`GetObjectQueueSnapshot`.
|
||||
- `active` is the number of disk-read permits currently in use.
|
||||
- `limit` is the configured maximum concurrent disk reads.
|
||||
- `queued` remains `None` because the current semaphore does not expose waiter
|
||||
counts.
|
||||
- Scanner, repair, replication, foreground write, and metadata entries remain
|
||||
`Unknown` until their owning runtime components expose read-only status.
|
||||
|
||||
This is an observation surface only. Permit acquisition, priority assignment,
|
||||
buffer sizing, storage media detection, request guards, and queue behavior are
|
||||
unchanged.
|
||||
|
||||
## Heal Repair Snapshot Extraction
|
||||
|
||||
The RustFS integration layer now exposes a read-only repair admission snapshot
|
||||
from the heal runtime counters:
|
||||
|
||||
- `Repair` reports the current heal active task count.
|
||||
- `queued` reports the current heal queue length.
|
||||
- `limit` remains `None` because the configured heal queue and concurrency
|
||||
limits live behind the async heal manager state.
|
||||
- Other workload classes remain `Unknown` in this provider until their owning
|
||||
runtime components expose read-only status.
|
||||
|
||||
This is an observation surface only. Heal request admission, queue capacity,
|
||||
priority merge/drop policy, task scheduling, retry handling, and repair
|
||||
behavior are unchanged.
|
||||
|
||||
## Replication Snapshot Extraction
|
||||
|
||||
The RustFS integration layer now exposes a read-only replication admission
|
||||
snapshot from the existing replication pool and queue statistics:
|
||||
|
||||
- `Replication` reports active regular, large-object, and MRF worker counts.
|
||||
- `queued` reports the current site replication queue count when queue stats
|
||||
are immediately observable.
|
||||
- `limit` remains `None` because replication worker limits remain owned by the
|
||||
async replication pool and resize policy.
|
||||
- If the replication runtime has not initialized, or queue stats are currently
|
||||
locked, the snapshot reports `Unknown` instead of blocking or guessing.
|
||||
|
||||
This is an observation surface only. Replication admission, queue channel
|
||||
capacity, worker resize behavior, MRF handling, target dispatch, and resync
|
||||
behavior are unchanged.
|
||||
|
||||
## RustFS Runtime Owner Snapshot Extraction
|
||||
|
||||
The RustFS integration layer now extends the workload admission registry with
|
||||
additional read-only owner mappings:
|
||||
|
||||
- `ForegroundRead` reuses the storage `ConcurrencyManager` disk-read permit
|
||||
snapshot so the RustFS-level provider exposes the same active and limit
|
||||
counts as the storage-local provider.
|
||||
- `Scanner` reports the existing scanner active work-unit counter. When the
|
||||
counter is zero, the snapshot remains `Unknown` because the current counter
|
||||
cannot distinguish an idle scanner from a scanner that has not initialized.
|
||||
- `Metadata` reports `Open` once the bucket metadata runtime handle is
|
||||
available, and `Unknown` before initialization.
|
||||
- `ForegroundWrite` remains `Unknown` until a write-specific admission owner
|
||||
exposes a read-only surface.
|
||||
|
||||
This is an observation surface only. Disk-read permit acquisition, scanner
|
||||
cycle scheduling, bucket metadata loading, metadata locks, object write paths,
|
||||
and queue behavior are unchanged.
|
||||
|
||||
## Provider Composition Boundary
|
||||
|
||||
`WorkloadAdmissionRegistrySnapshot::overlay` composes provider-owned registry
|
||||
snapshots without mutating runtime owners:
|
||||
|
||||
- The storage concurrency provider remains the source of truth for
|
||||
`ForegroundRead`.
|
||||
- The RustFS runtime owner provider overlays metadata, scanner, repair,
|
||||
replication, and foreground-write status on top of the storage registry.
|
||||
- Matching workload classes are replaced by the later provider snapshot; new
|
||||
classes are appended without reordering existing unrelated entries.
|
||||
|
||||
This keeps the later controller/status layer consuming a single read-only
|
||||
registry while preserving the existing storage, scanner, heal, replication, and
|
||||
metadata ownership boundaries.
|
||||
`WorkloadAdmissionRegistrySnapshot::overlay` composes provider-owned registries without mutating runtime owners: the storage concurrency provider is the source of truth for `ForegroundRead`; the RustFS runtime owner provider overlays metadata, scanner, repair, replication, and foreground-write status on top; matching classes are replaced by the later snapshot and new classes are appended without reordering. `workload_admission_registry_snapshot` in `rustfs/src/workload_admission.rs` is the single composed registry consumers read.
|
||||
|
||||
Reference in New Issue
Block a user