docs(knowledge-base): prune stale content and add agent-facing index (#7035)

This commit is contained in:
Zhengchao An
2026-09-02 08:26:59 +08:00
committed by GitHub
parent ceeff52229
commit 0a975f2fe2
99 changed files with 3312 additions and 10590 deletions
+3
View File
@@ -1,5 +1,8 @@
# Atomic object undo precondition
**Use this when:** you need the `x-rustfs-expected-current-version-id` precondition for undo-style CopyObject restores or delete-marker removal, or you are changing how it is parsed or enforced.
**Source of truth:** `rustfs/src/app/object/shared.rs` (`expected_current_version_id` header parser), `rustfs/src/app/object/copy.rs` (CopyObject enforcement), `crates/ecstore/src/set_disk/ops/object.rs` (`expected_current_version_id` checks under the namespace write lock).
RustFS supports a destination-side version precondition for the two S3
operations used to undo changes in a versioned bucket:
-240
View File
@@ -1,240 +0,0 @@
# Authing OIDC Integration Runbook
This runbook helps operators connect the RustFS Console to Authing through standard OpenID Connect. The examples use the default RustFS provider id, `default`.
## 1. Integration Model
RustFS expects a standards-compliant OpenID Connect provider, not an Authing-specific plugin. The Authing application must provide:
- issuer metadata through `.well-known/openid-configuration`
- authorization endpoint
- token endpoint
- JWKS or another verifiable ID token signature path
- authorization-code flow that returns an `id_token`
The RustFS browser login flow is:
1. The user opens the RustFS OIDC authorize endpoint.
2. RustFS creates `state`, `nonce`, and a PKCE S256 challenge.
3. The browser is redirected to Authing.
4. Authing redirects back to RustFS with `code` and `state`.
5. RustFS exchanges the code with `client_id`, `client_secret`, and the PKCE verifier.
6. RustFS validates the ID token signature, issuer, audience, expiry, and nonce.
7. RustFS reads identity and authorization claims from the ID token.
8. RustFS maps claim values to RustFS policy names and issues one-hour STS credentials for the Console.
## 2. Required Values
Collect these values before deployment:
| Value | Example | Notes |
| --- | --- | --- |
| Public RustFS browser origin | `https://rustfs.example.com` | The scheme and authority users open in the browser. |
| Provider id | `default` | This runbook uses the default provider. |
| RustFS callback URL | `https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default` | Register this exact URL in Authing. |
| Authing application domain | `https://example.authing.cn` | Use the value shown in the Authing application. |
| Authing issuer | `https://example.authing.cn/oidc` | Copy the issuer from Authing; do not guess the path. |
| Authing App ID | `<AUTHING_APP_ID>` | RustFS `client_id`. |
| Authing App Secret | `<AUTHING_APP_SECRET>` | RustFS `client_secret`. |
| RustFS scopes | `openid,profile,email,roles` | `openid` is required; include `roles` when Authing emits role claims. |
Authing deployments can use different issuer paths, such as `/oidc` or `/oauth/oidc`. Always copy the issuer from the Authing console and verify that discovery returns the same `issuer` value.
## 3. Authing Configuration
### 3.1 Create the Application
1. Open the Authing console.
2. Create a self-hosted application named `RustFS Console`.
3. Record the App ID, App Secret, application domain, issuer, and discovery URL.
### 3.2 Configure OIDC
Use these protocol settings:
| Setting | Value |
| --- | --- |
| Protocol | OpenID Connect |
| Grant type | Authorization Code |
| Response type | `code` |
| Token endpoint authentication | `client_secret_post` |
| PKCE | Allow or require `S256` |
| ID token signing algorithm | `RS256` recommended |
RustFS sends the client secret in the request body. Do not configure Authing to reject `client_secret_post`.
### 3.3 Register the Redirect URL
Add this exact callback URL in Authing:
```text
https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default
```
The scheme, host, port, path, and provider id must match the RustFS configuration.
### 3.4 Map Roles to RustFS Policies
RustFS does not call Authing authorization APIs. It reads `roles` or `groups` from the ID token and maps each value to a RustFS policy name.
Recommended policy names:
| Authing claim value | RustFS policy | Purpose |
| --- | --- | --- |
| `consoleAdmin` | `consoleAdmin` | Full Console, admin, KMS, and S3 access. |
| `readwrite` | `readwrite` | S3 read/write access. |
| `readonly` | `readonly` | S3 read-only access. |
| `writeonly` | `writeonly` | S3 write-only access. |
| `diagnostics` | `diagnostics` | Diagnostic admin access. |
For initial validation, assign a test user the `consoleAdmin` role and confirm that the ID token contains:
```json
{
"roles": ["consoleAdmin"]
}
```
`claim_prefix` only prepends a fixed string. It does not perform arbitrary role mapping. Keep Authing role values equal to RustFS policy names unless you already created policies with a fixed prefix.
## 4. RustFS Configuration
### 4.1 Environment Variables
Set the OIDC provider and the public browser origin:
```bash
export RUSTFS_BROWSER_REDIRECT_URL="https://rustfs.example.com"
export RUSTFS_IDENTITY_OPENID_ENABLE=on
export RUSTFS_IDENTITY_OPENID_CONFIG_URL="<AUTHING_ISSUER>"
export RUSTFS_IDENTITY_OPENID_CLIENT_ID="<AUTHING_APP_ID>"
export RUSTFS_IDENTITY_OPENID_CLIENT_SECRET="<AUTHING_APP_SECRET>"
export RUSTFS_IDENTITY_OPENID_SCOPES="openid,profile,email,roles"
export RUSTFS_IDENTITY_OPENID_REDIRECT_URI="https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default"
export RUSTFS_IDENTITY_OPENID_REDIRECT_URI_DYNAMIC=off
export RUSTFS_IDENTITY_OPENID_DISPLAY_NAME="Authing"
export RUSTFS_IDENTITY_OPENID_EMAIL_CLAIM="email"
export RUSTFS_IDENTITY_OPENID_USERNAME_CLAIM="preferred_username"
export RUSTFS_IDENTITY_OPENID_ROLES_CLAIM="roles"
```
For short-lived connectivity testing only, you may temporarily add:
```bash
export RUSTFS_IDENTITY_OPENID_ROLE_POLICY="consoleAdmin"
```
Do not keep `role_policy=consoleAdmin` in production unless every Authing user for this client should receive full Console access.
Restart RustFS after changing OIDC settings.
### 4.2 Admin Config
If the deployment manages OIDC through compatible admin configuration commands, set the provider like this:
```bash
mc admin config set rustfs identity_openid \
enable=on \
config_url="<AUTHING_ISSUER>" \
client_id="<AUTHING_APP_ID>" \
client_secret="<AUTHING_APP_SECRET>" \
scopes="openid,profile,email,roles" \
redirect_uri="https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default" \
redirect_uri_dynamic=off \
display_name="Authing" \
email_claim="email" \
username_claim="preferred_username" \
roles_claim="roles"
mc admin service restart rustfs
```
`RUSTFS_BROWSER_REDIRECT_URL` is a process environment variable, not an `identity_openid` provider key. Configure it in the RustFS service environment even when the provider itself is stored through admin config.
### 4.3 Redirect URL Priority
RustFS builds browser-facing URLs with this priority:
1. Provider `redirect_uri`, when configured, is used for the OIDC callback URL sent to Authing.
2. `RUSTFS_BROWSER_REDIRECT_URL`, when configured, is used as the public origin for OIDC callback generation when no provider `redirect_uri` exists, and for Console success redirects and logout fallback redirects.
3. Request headers are used only when provider dynamic redirects are enabled and no browser redirect URL is configured.
For reverse-proxy or load-balancer deployments, set `RUSTFS_BROWSER_REDIRECT_URL` to avoid depending on `Host` and `X-Forwarded-Proto` for Console redirects. OIDC authorize and callback requests must still reach the same RustFS node because in-flight OIDC `state` is local to the node.
## 5. Validation
### 5.1 Validate Authing Discovery
```bash
AUTHING_ISSUER="<AUTHING_ISSUER>"
curl -fsS "$AUTHING_ISSUER/.well-known/openid-configuration" | jq '{
issuer,
authorization_endpoint,
token_endpoint,
jwks_uri,
id_token_signing_alg_values_supported,
code_challenge_methods_supported,
token_endpoint_auth_methods_supported,
scopes_supported
}'
```
Check that:
- `issuer` exactly matches `RUSTFS_IDENTITY_OPENID_CONFIG_URL`
- `authorization_endpoint`, `token_endpoint`, and `jwks_uri` are present
- `code_challenge_methods_supported` includes `S256`
- `token_endpoint_auth_methods_supported` includes `client_secret_post`
- `scopes_supported` includes `openid`, `profile`, `email`, and any role scope you need
### 5.2 Validate RustFS Provider Visibility
```bash
curl -fsS "https://rustfs.example.com/rustfs/admin/v3/oidc/providers" | jq
```
The response should include the Authing provider unless `hide_from_ui` is enabled.
### 5.3 Test Browser Login
Open:
```text
https://rustfs.example.com/rustfs/admin/v3/oidc/authorize/default
```
Expected flow:
1. Browser redirects to Authing.
2. The user signs in.
3. Authing redirects to `/rustfs/admin/v3/oidc/callback/default?code=...&state=...`.
4. RustFS validates the ID token and issues STS credentials.
5. The browser lands on the RustFS Console and can use the expected permissions.
## 6. Troubleshooting
| Symptom | Common cause | Fix |
| --- | --- | --- |
| `/oidc/providers` does not show Authing | OIDC provider did not load, or RustFS was not restarted | Check environment variables and restart RustFS. |
| Authing reports redirect mismatch | Callback URL differs between Authing and RustFS | Use the exact `/rustfs/admin/v3/oidc/callback/default` URL. |
| RustFS reports missing `code` or `state` | Proxy dropped the query string | Preserve the full callback URL and query string. |
| Token exchange fails | Wrong client secret or unsupported token auth method | Confirm `client_secret_post` is allowed. |
| RustFS reports no `id_token` | Missing `openid` scope or non-OIDC OAuth flow | Include `openid` and use OIDC authorization code flow. |
| ID token verification fails | Issuer, audience, signing algorithm, or JWKS mismatch | Compare discovery metadata with RustFS config; prefer `RS256`. |
| Login succeeds but access is denied | No matching RustFS policy claim | Ensure `roles` or `groups` is in the ID token and equals a RustFS policy name. |
| Console redirects to an internal host | Missing `RUSTFS_BROWSER_REDIRECT_URL` or incorrect proxy headers | Set `RUSTFS_BROWSER_REDIRECT_URL` to the public browser origin. |
| Invalid or expired OIDC state | Callback reached a different RustFS node | Configure load-balancer session affinity for authorize and callback requests. |
## 7. Production Checklist
- [ ] RustFS and Authing use HTTPS.
- [ ] Authing redirect URL is exact, not a broad wildcard.
- [ ] `RUSTFS_BROWSER_REDIRECT_URL` is set to the public RustFS browser origin.
- [ ] `RUSTFS_IDENTITY_OPENID_REDIRECT_URI` matches the registered Authing callback URL.
- [ ] Authing emits role or group claims in the ID token.
- [ ] Claim values match RustFS policy names.
- [ ] `role_policy=consoleAdmin` is not used as a permanent production shortcut.
- [ ] The load balancer preserves query strings.
- [ ] OIDC authorize and callback requests have session affinity to the same RustFS node.
+48 -181
View File
@@ -1,203 +1,70 @@
# Container Resource Detection
# Container resource detection
RustFS automatically detects container resource limits (CPU and memory) from cgroup v1/v2. This ensures correct resource allocation and accurate metrics in containerized environments (Kubernetes, Docker, etc.).
**Use this when:** RustFS runs under a cgroup CPU or memory limit (Kubernetes, Docker) and you need to know which limit it detected, how to override it, or which log line and metrics expose it.
**Source of truth:** `rustfs/src/cgroup_resources.rs` (detection, `ContainerResources`, `container_resources()`), `rustfs/src/memory_observability.rs` (metrics), `rustfs/src/server/runtime.rs` (Tokio thread sizing consumer), `rustfs/src/startup_entrypoint.rs` (startup log call).
## Problem
RustFS resolves the CPU core count and the memory limit once at startup (cached in a `OnceLock`) and uses them for Tokio worker/blocking-thread sizing, the memory budget, and memory metrics. Precedence is override env var, then cgroup limit, then host value from `sysinfo`.
When RustFS runs in a container, the underlying system libraries report the **host's** total CPU cores and memory, not the container's limits. This leads to:
## Detection rules
1. **Over-provisioned Tokio threads**: Too many worker and blocking threads
2. **Incorrect memory metrics**: `rustfs_memory_usage_percent` shows host-based percentage
3. **Memory budget errors**: Object data cache sized to host RAM instead of container limit
4. **OOMKills**: Container exceeds its memory limit and gets killed
| Resource | Order | Source | Rule |
| --- | --- | --- | --- |
| CPU | 1 | cgroup v2 `/sys/fs/cgroup/cpu.max` | `"<quota> <period>"` (or bare `"<quota>"` with period 100000) gives `ceil(quota / period)`; `"max"` or a zero quota means no limit. |
| CPU | 2 | cgroup v1 `/sys/fs/cgroup/cpu/cpu.cfs_quota_us` with `cpu.cfs_period_us` | `ceil(quota / period)`; a zero, unparsable, or `u64::MAX` quota means no limit. |
| CPU | 3 | host | `sysinfo` CPU count, minimum 1. |
| Memory | 1 | cgroup v2 `/sys/fs/cgroup/memory.max` | bytes; `"max"` means no limit. |
| Memory | 2 | cgroup v1 `/sys/fs/cgroup/memory/memory.limit_in_bytes` | bytes; values `>= 1 << 62` mean no limit. |
| Memory | 3 | host | `sysinfo` total memory. |
## Solution
cgroup reads are compiled only for Linux; other platforms always take the host branch. `cgroup_detected` is true when at least one of the two cgroup reads returned a limit.
RustFS now detects cgroup limits directly from the filesystem:
## Environment variables
- **CPU**: `/sys/fs/cgroup/cpu.max` (v2) or `/sys/fs/cgroup/cpu/cpu.cfs_quota_us` (v1)
- **Memory**: `/sys/fs/cgroup/memory.max` (v2) or `/sys/fs/cgroup/memory/memory.limit_in_bytes` (v1)
Names are the constants `ENV_DISABLE_CGROUP_DETECTION`, `ENV_OVERRIDE_CPU_CORES`, and `ENV_OVERRIDE_MEMORY_BYTES` in `rustfs/src/cgroup_resources.rs`.
The effective resource limits are the **minimum** of host and cgroup values.
| Variable | Accepted values | Effect |
| --- | --- | --- |
| `RUSTFS_DISABLE_CGROUP_DETECTION` | `1` or `true` (case-insensitive) | Skip cgroup reads; host values apply unless overridden. |
| `RUSTFS_OVERRIDE_CPU_CORES` | integer `> 0` | Replaces the CPU core count regardless of cgroup or host. |
| `RUSTFS_OVERRIDE_MEMORY_BYTES` | integer `> 0`, bytes | Replaces the memory limit regardless of cgroup or host. |
## Detection Logic
Non-positive or unparsable override values are ignored. Changes take effect on process restart.
### CPU Detection
## Startup log
1. Read cgroup v2 `/sys/fs/cgroup/cpu.max`
- Format: `"$QUOTA $PERIOD"` or `"max"` (unlimited)
- Calculate: `cores = ceil(quota / period)`
2. Fallback to cgroup v1 `/sys/fs/cgroup/cpu/cpu.cfs_quota_us`
- Calculate: `cores = ceil(quota / period)`
3. Fallback to host CPU count from `sysinfo`
`log_container_resources` emits exactly one of these lines with `cpu_cores` and `memory_bytes` fields (the INFO variants also carry `memory_mib`):
### Memory Detection
1. Read cgroup v2 `/sys/fs/cgroup/memory.max`
- Value in bytes or `"max"` (unlimited)
2. Fallback to cgroup v1 `/sys/fs/cgroup/memory/memory.limit_in_bytes`
- Very large values (≥2^62) indicate unlimited
3. Fallback to host memory from `sysinfo`
## Environment Variables
### Disable Cgroup Detection
```bash
RUSTFS_DISABLE_CGROUP_DETECTION=1
```
Disables cgroup detection entirely. Useful for testing or when cgroup filesystem is not accessible.
### Override CPU Cores
```bash
RUSTFS_OVERRIDE_CPU_CORES=4
```
Overrides detected CPU cores. Takes precedence over cgroup detection.
### Override Memory Limit
```bash
RUSTFS_OVERRIDE_MEMORY_BYTES=2147483648
```
Overrides detected memory limit in bytes. Takes precedence over cgroup detection.
| Message | Level | Condition |
| --- | --- | --- |
| `container resources (overridden by environment variables)` | INFO | an override env var was applied; also carries `cgroup_detected` |
| `container resources (detected from cgroup)` | INFO | no override, at least one cgroup limit read |
| `container resources (using host values)` | DEBUG | neither override nor cgroup limit |
## Metrics
### New Metrics
Gauges emitted from `rustfs/src/memory_observability.rs`.
| Metric | Description |
|--------|-------------|
| `rustfs_memory_effective_total_bytes` | Effective memory total (host or cgroup) |
| `rustfs_cgroup_detected` | Whether cgroup limits were detected (1=yes, 0=no) |
| `rustfs_cgroup_cpu_cores_limit` | Detected CPU cores limit |
| `rustfs_cgroup_memory_limit_bytes` | Detected memory limit |
### Updated Metrics
| Metric | Change |
|--------|--------|
| `rustfs_memory_total_bytes` | Now uses effective memory (cgroup-aware) |
| `rustfs_memory_usage_percent` | Now calculated against effective memory |
## Startup Logging
RustFS logs detected container resources at startup:
```
INFO container resources (detected from cgroup) cpu_cores=2 memory_bytes=1073741824 memory_mib=1024
```
or
```
INFO container resources (overridden by environment variables) cpu_cores=4 memory_bytes=2147483648 memory_mib=2048
```
## Examples
### Kubernetes with Resource Limits
```yaml
resources:
limits:
cpu: "2"
memory: "1Gi"
requests:
cpu: "500m"
memory: "512Mi"
```
RustFS will detect:
- CPU cores: 2
- Memory: 1 GiB (1073741824 bytes)
### Docker with CPU and Memory Limits
```bash
docker run --cpus=2 --memory=1g rustfs/rustfs:latest
```
RustFS will detect:
- CPU cores: 2
- Memory: 1 GiB
### Manual Override
```bash
export RUSTFS_OVERRIDE_CPU_CORES=4
export RUSTFS_OVERRIDE_MEMORY_BYTES=2147483648
```
RustFS will use:
- CPU cores: 4
- Memory: 2 GiB
| Metric | Meaning |
| --- | --- |
| `rustfs_memory_effective_total_bytes{basis}` | Effective memory total. `basis` is `cgroup` when a cgroup limit was detected, else `host`; an override does not change the basis label. |
| `rustfs_container_cpu_cores` | Effective CPU cores. |
| `rustfs_container_memory_bytes` | Effective memory limit in bytes. |
| `rustfs_container_cgroup_detected` | `1` when a cgroup limit was read, else `0`. |
| `rustfs_container_overridden` | `1` when an override env var was applied, else `0`. |
| `rustfs_memory_total_bytes`, `rustfs_memory_usage_percent` | Computed against the effective total (`record_memory_usage` in `crates/io-metrics/src/lib.rs`). |
## Troubleshooting
### Cgroup Detection Not Working
When effective values look like the host rather than the container:
1. Check if cgroup filesystem is mounted:
```bash
ls -la /sys/fs/cgroup/
```
1. Confirm detection is not disabled: `env | grep RUSTFS_DISABLE_CGROUP_DETECTION`.
2. Find the startup line: `grep "container resources" <rustfs log>`. The `(using host values)` variant is DEBUG, so raise the log level if no variant appears.
3. Inspect the cgroup filesystem inside the container:
2. Check cgroup version:
```bash
stat -fc %T /sys/fs/cgroup/
```
- `cgroup2fs` = cgroup v2
- `tmpfs` = cgroup v1
```bash
stat -fc %T /sys/fs/cgroup/ # cgroup2fs = v2, tmpfs = v1
cat /sys/fs/cgroup/cpu.max /sys/fs/cgroup/memory.max # v2
cat /sys/fs/cgroup/cpu/cpu.cfs_quota_us /sys/fs/cgroup/memory/memory.limit_in_bytes # v1
```
3. Check if limits are set:
```bash
# cgroup v2
cat /sys/fs/cgroup/cpu.max
cat /sys/fs/cgroup/memory.max
# cgroup v1
cat /sys/fs/cgroup/cpu/cpu.cfs_quota_us
cat /sys/fs/cgroup/memory/memory.limit_in_bytes
```
### Metrics Show Host Values
If `rustfs_memory_effective_total_bytes` shows host memory instead of cgroup limit:
1. Verify cgroup detection is not disabled:
```bash
echo $RUSTFS_DISABLE_CGROUP_DETECTION
```
2. Check startup logs for cgroup detection:
```bash
grep "container resources" /logs/rustfs.log
```
3. Use environment variable override as workaround:
```bash
export RUSTFS_OVERRIDE_MEMORY_BYTES=1073741824
```
## Implementation Details
### Files Modified
- `rustfs/src/cgroup_resources.rs` - Core cgroup detection logic
- `rustfs/src/container_config.rs` - Container configuration with overrides
- `rustfs/src/memory_observability.rs` - Updated memory metrics
- `rustfs/src/server/runtime.rs` - Updated Tokio runtime configuration
- `rustfs/src/startup_entrypoint.rs` - Startup logging
### Performance Impact
- **Startup**: One-time detection adds ~1ms overhead
- **Runtime**: Cached values, no repeated filesystem reads
- **Memory**: Negligible (<1KB for cached values)
### Thread Safety
All detection functions are thread-safe and use `OnceLock` for caching.
4. If the runtime does not expose limits to the container, pin them with `RUSTFS_OVERRIDE_CPU_CORES` / `RUSTFS_OVERRIDE_MEMORY_BYTES` and confirm `rustfs_container_overridden` reads `1`.
+5 -3
View File
@@ -1,5 +1,8 @@
# dial9 Tokio Runtime Profiling
**Use this when:** you need Tokio runtime-level evidence (which task held a worker, long polls, park/unpark behaviour) that Prometheus metrics and `tracing` spans cannot provide, or you are building or running the opt-in `dial9` profiling binary.
**Source of truth:** `crates/obs/src/telemetry/dial9/mod.rs` (session setup), `crates/obs/src/metrics/collectors/dial9.rs` (metrics), `crates/config/src/constants/runtime.rs` (`RUSTFS_RUNTIME_DIAL9_*` and defaults), `.config/make/build.mak` (`build-profiling`), `crates/obs/build.rs` (feature/cfg pairing check).
`dial9-tokio-telemetry` records Tokio runtime-level events — poll start/end,
worker park/unpark, task spawn/terminate, and optionally async backtraces of
stalled tasks — into binary trace segments.
@@ -60,9 +63,8 @@ That is `cargo build --release --bin rustfs --features dial9` with
dumps and S3 upload are both unavailable, for the reasons given above and below.
`crates/obs/build.rs` fails the build if the `dial9` feature is enabled without
`--cfg tokio_unstable`. This is deliberate: an environment `RUSTFLAGS` *replaces*
the value from `.cargo/config.toml` rather than appending to it, so the flag used
to disappear silently whenever anything else set `RUSTFLAGS`.
`--cfg tokio_unstable`, so a mismatched `RUSTFLAGS` cannot produce a binary that
silently records nothing.
For CPU profiling with usable stacks, add `-C force-frame-pointers=yes`.
+11 -24
View File
@@ -1,5 +1,8 @@
# Drive Timeout Tuning
**Use this when:** `ListObjects`/`ListObjectsV2` on a large prefix fails with `Io error: timeout`, or RustFS runs on HDD-class, network, or throttled storage and you need to widen per-operation drive liveness budgets.
**Source of truth:** `crates/config/src/constants/drive.rs` (`DEFAULT_DRIVE_*_TIMEOUT_SECS`, `DRIVE_TIMEOUT_PROFILE_HIGH_LATENCY_SECS`), `crates/config/src/constants/object.rs` (`DEFAULT_OBJECT_DISK_READ_TIMEOUT`), `crates/config/src/constants/capacity.rs` (`DEFAULT_CAPACITY_MAX_TIMEOUT_SECS`), `crates/ecstore/src/cache_value/metacache_set.rs` (walk stall handling and `rustfs_list_path_raw_stall_total`).
This document describes the per-operation drive timeout knobs and the
drive-timeout profile. It is written for operators running RustFS on slow or
high-latency storage (HDD-class disks, network block devices, throttled
@@ -79,10 +82,11 @@ for the full list and defaults.
`ListObjects`/`ListObjectsV2` on a large prefix either:
- returns `500 InternalError` with `Io error: timeout`; or
- (on older builds) returns HTTP 200 with `IsTruncated=false` after fewer keys
than the bucket actually holds — a **silent** truncation that S3 clients
(`mc`, minio-go, SDK pagination loops) cannot detect, because
`IsTruncated=false` is the protocol's only end-of-listing signal.
- on builds without the failure contract below, returns HTTP 200 with
`IsTruncated=false` after fewer keys than the bucket actually holds — a
**silent** truncation that S3 clients (`mc`, minio-go, SDK pagination loops)
cannot detect, because `IsTruncated=false` is the protocol's only
end-of-listing signal.
Every "missing" object remains readable by exact key via `GetObject` /
`StatObject`; only the listing is affected.
@@ -97,21 +101,10 @@ single `readdir` exceed the budget on a perfectly healthy disk, especially on
HDD-class or throttled storage. That trips a drive timeout, which the listing
path escalates and surfaces to the client.
### The silent variant is fixed; the loud 500 is tuned away
### Failure contract
As of the walk-stall rework (merged to `main`, first released in **1.0.0-beta.9**):
- **The silent variant is eliminated.** A walk that dies mid-stream can no
longer be consumed as a clean end-of-listing. Once a walk has streamed any
entries and then stalls, the failure is recorded as a hard drive timeout and
escalated on that erasure set, so the client always sees an error — never a
well-formed short page. This is locked by the
`list_path_raw_returns_timeout_when_producer_fails_after_partial_entry`
regression test in `crates/ecstore/src/cache_value/metacache_set.rs`.
- **The remaining 500 is an operator-tunable, not a data-integrity bug.** A
genuinely wide flat directory can still exhaust the default 5s stall budget on
slow storage and fail the listing loudly. The supported mitigation is to widen
the budget.
- A walk that stalls after streaming any entries fails as a hard drive timeout escalated on that erasure set; the client always sees an error, never a well-formed short page. Locked by `list_path_raw_returns_timeout_when_producer_fails_after_partial_entry` in `crates/ecstore/src/cache_value/metacache_set.rs`.
- The remaining `500` on a genuinely wide flat directory is an operator tunable, not a data-integrity bug: widen the stall budget as below.
### Mitigation
@@ -129,12 +122,6 @@ Raise the walk stall budget, or select the high-latency profile:
-e RUSTFS_DRIVE_TIMEOUT_PROFILE=high_latency
```
> Note: on releases at or before `1.0.0-beta.8`, the foreground listing path was
> bounded by the *total* wall-clock knob `RUSTFS_DRIVE_WALKDIR_TIMEOUT_SECS`
> instead of the stall budget. If you cannot upgrade, raise that knob — but
> upgrading to `1.0.0-beta.9` or later is strongly preferred, because only the
> newer builds convert the *silent* truncation into a detectable error.
The most durable fix for pathologically wide directories is to shard keys under
additional prefix levels so no single directory holds an enormous flat child
set; the stall budget then never has to bound one giant `readdir`.
+14 -21
View File
@@ -1,5 +1,8 @@
# Durability modes (drive sync tiers)
**Use this when:** choosing or debugging the fsync tier (`strict|relaxed|none|legacy-off`) for a deployment or a single bucket, or changing any write-path sync behaviour.
**Source of truth:** `crates/ecstore/src/disk/local.rs` (`ENV_RUSTFS_DURABILITY_MODE`, `ENV_RUSTFS_DRIVE_SYNC_ENABLE`, mode resolution and per-write-point sync decisions), `crates/ecstore/src/bucket/durability.rs` (`ENV_NEW_BUCKET_DURABILITY_MODE`, per-bucket override), `rustfs/src/admin/handlers/durability.rs` (admin API), `crates/ecstore/src/bucket/metadata_sys.rs` (`BUCKET_METADATA_REFRESH_INTERVAL`).
RustFS lets operators choose how much fsync work runs on the object write
path. The default (`strict`) preserves the fully synced behavior RustFS has
always shipped; the relaxed tiers are **opt-in** trades of power-loss
@@ -81,13 +84,13 @@ failure:
MinIO's default posture (no per-object fsync) but means small objects have
the widest loss window.
- Durability of acknowledged writes therefore rests on **erasure-coded
redundancy across other nodes** plus the unclean-shutdown heal introduced
in PR #4221 converging the affected drive afterwards.
redundancy across other nodes** plus the unclean-shutdown heal converging
the affected drive afterwards.
Deployment rule for `relaxed`: only multi-node clusters whose nodes sit in
**independent power domains** (separate feeds/UPS). If all nodes can lose
power simultaneously — the exact incident class that motivated PR #4221
`relaxed` can lose recently acknowledged objects cluster-wide. Single-node
power simultaneously, `relaxed` can lose recently acknowledged objects
cluster-wide. Single-node
deployments must stay on `strict`.
**`none`.** No fsync on the object data path at all; acknowledged objects can
@@ -119,7 +122,7 @@ staged in tmp still commits with full `strict` durability.
The durability mode is server-side configuration only; it cannot be raised or
lowered by any request header.
## Per-bucket durability (phase 2)
## Per-bucket durability
A bucket can override the process-wide mode with its own tier. The override
is stored in the bucket's metadata (a `durability.json` entry in
@@ -146,8 +149,8 @@ configuration plane.
### New-bucket default
A bucket created after this feature ships gets a `relaxed` override **seeded
into its own metadata** at creation time (rustfs/backlog#1811), so it opts
A newly created bucket gets a `relaxed` override **seeded into its own
metadata** at creation time, so it opts
into MinIO's default posture (object data still fdatasynced; xl.meta and
directory-entry fsyncs left to the page cache) without touching the
process-wide default. This is a gradual migration:
@@ -226,20 +229,10 @@ bucket on power failure.
## Performance expectations
The often-quoted 26x PUT throughput delta was measured on macOS with the old
binary switch fully **off** (equivalent to `none`/`legacy-off`), where
`F_FULLFSYNC` heavily amplifies sync cost. `relaxed` keeps the per-shard
fdatasync, so its gain is necessarily smaller and must be measured on the
target platform (Linux ext4/xfs) before being relied on. Do not use `none`
numbers to size `relaxed`.
## Scope
Phase 1 (rustfs/backlog#926) shipped the global, per-process tier configured
by environment variable. Phase 2 (rustfs/backlog#938) adds the per-bucket
override described above, configured through the admin API and stored in
bucket metadata. `mc admin` integration for the per-bucket tier is a
follow-up.
Measure `relaxed` on the target platform (Linux ext4/xfs) before relying on a
number. Throughput deltas measured with sync fully **off** (`none` or
`legacy-off`) do not transfer: `relaxed` keeps the per-shard fdatasync, so its
gain is necessarily smaller. Do not use `none` numbers to size `relaxed`.
## Related Recovery Guides
@@ -1,35 +0,0 @@
# GET Path Experimental Performance Switches
This document records two experimental environment switches on the object GET
path. Both default to **off**, are read once at startup, and exist to support
staged performance work — they are not general tuning knobs. Until this
document existed they were referenced only by performance harness scripts,
which made them look like orphans during dead-code sweeps; they are kept
deliberately (rustfs/backlog#1832).
## RUSTFS_GET_SEEK_BUFFER_ENABLE
- Type: boolean (`true`/`false`), default `false`.
- Read once at startup in `rustfs/src/app/object_usecase.rs`.
- When enabled, small GET responses may be served through an in-memory seek
buffer, providing seek support without re-reading the object. The seek-buffer
code path is unit-test gated; whether the path stays or graduates to default
is a post-1.0 maintainer decision — do not remove either the switch or the
gated path as dead code.
## RUSTFS_GET_OUTPUT_HANDOFF_ATTRIBUTION_ENABLE
- Type: boolean (`true`/`false`), default `false`.
- Read once at startup in `rustfs/src/app/object_usecase.rs`.
- When enabled, GET responses attribute output-handoff stage timing in the GET
stage metrics, at a small per-request bookkeeping cost. Used by the A/B
performance runbooks (`scripts/run_get_codec_streaming_smoke.sh`,
`scripts/test_get_1mib_abba_stage_metrics.sh`) to compare handoff cost
between configurations.
## Operational guidance
Leave both switches unset in production. Enable them only when following a
performance runbook that asks for them, and unset them afterwards — both are
startup-latched, so changing a value requires a process restart to take
effect.
@@ -1,115 +0,0 @@
# Heal 并发安全说明(对象级 healing 标记对标审计结论)
对应 backlog rustfs/backlog#1874(父 #1862,HS-12)。本文回答一个问题:MinIO 在 heal
期间对对象打 `x-minio-healing:true` 元数据标记以防"heal 提交与并发删除/版本清理互毁"
cmd/xl-storage.go RenameData 的 healing 分支),RustFS 是否需要同款防御。
**结论:不需要。** RustFS 不存在 MinIO 用 healing 标记防御的那类竞争:所有会触达同一
`(bucket, object)` 提交面的路径都在同一把对象级 namespace 写锁上互斥,且 heal 的锁
guard 覆盖 rename 提交全程;MinIO 需要标记的根因(RenameData 提交内部与版本清理逻辑
交错)在 RustFS 的提交模型中不存在。RustFS 已有一个瞬态 healing 旗标用于另一目的
(见下文 §2),并有并发不变量回归测试锁定本结论(§5)。
## 1. 两个防御模型的对照
MinIOheal 时对对象写 `x-minio-healing:true`(持久元数据标记),后续任何 RenameData
提交看到该标记就跳过版本清理/legacy purge 逻辑——防御发生在锁外,靠元数据让路。
RustFS:三层防御,全部不依赖持久对象标记:
1. **锁内互斥**:heal 与一切前台/后台写路径的提交点在同一把 `(bucket, object)` ns 写锁
上串行(分布式部署为 quorum 锁 RPC,单机为进程内锁管理器;锁粒度是对象级,version
恒为 None)。
2. **提交模型隔离**rename_data 提交内没有会与 heal 交错的版本清理逻辑;被替换旧版本
的 data_dir 物理删除被移出提交临界区(commit tail),且只删已被新提交替换的 unshared
目录。
3. **瞬态 healing 旗标**`FileInfo::set_healing`crates/filemeta/src/fileinfo.rs)在
heal 提交的内存 FileInfo 上打 `"healing"` 内部键,rename_data 据此允许先清空 stale
目标 data_dir 再 rename——解决 heal 复用 data_dir 做 in-place 修复时 rename(2) 无法
替换非空目录的文件系统语义冲突(EEXIST/ENOTEMPTY)。该键是瞬态的,不落盘
`is_skip_meta_key`),与 MinIO 的持久标记目的不同。非 heal 提交撞上非空目标
data_dir 会显式失败,有测试锁定两个方向的行为。
## 2. 交点矩阵
中心路径:`heal_object_with_explicit_version_regen`crates/ecstore/src/set_disk/ops/heal.rs
下称 heal.rs)在入口取 `(bucket, object)` ns 写锁,guard 绑定到函数作用域末尾,覆盖
quorum 元数据读取 → EC 重建 → 逐盘 rename 提交 → tmp 清理 → HEAL_RENAME_INCOMPLETE
部分提交返回 → 孤儿 data_dir 回收的全过程。并发侧逐交点判定:
| # | 并发路径 | 并发侧锁 | 判定 | 关键证据 |
|---|---|---|---|---|
| 1 | PUT 对象提交 | `put_object_commit` 对象写锁,rename_data 在锁内 | 同锁串行 | ops/object.rs 提交锁段 + rename 调用点 |
| 2 | PUT 旧 data_dir tail 清理 | drop 对象锁后的 `commit_rename_data_dir`,无锁 | 无锁并发,语义安全(见 §3.1 | object.rs drop 后 tail 段;io_primitives.rs |
| 3 | DELETE 单对象/版本 | `delete_object` 对象写锁,delete_version 在锁内 | 同锁串行 | object.rs delete_object 锁段 |
| 4 | DELETE 批量 | 批量逐对象写锁(dist 走批量锁 RPC | 同锁串行 | object.rs delete_objects 锁段 |
| 5 | CompleteMultipart | 对象写锁 + upload 路径锁双锁,rename 在锁内 | 同锁串行 | ops/multipart.rs 提交锁段 |
| 6 | CompleteMultipart tail 清理 | drop 对象锁后的旧 data_dir 删除 | 无锁并发,语义安全(见 §3.1 | multipart.rs drop 后 tail 段 |
| 7 | AbortMultipart | 仅 multipart bucket 的 upload 路径锁 | 锁 key 不相交,但资源不相交(abort 不触对象 data_dir/xl.meta)→ 无实际交点 | multipart.rs abort 锁段 |
| 8 | ILM expiry(含 DeleteAllVersions | DeleteAllVersions 走 `delete_prefix_object=true` → 仍取对象锁;FreeVersionTask 显式取锁;noncurrent 批量走批量锁 | 同锁串行 | bucket_lifecycle_ops.rs 消费端链路 |
| 9 | 纯 prefix 删除(绕锁能力面) | `delete_prefix`-only 不取子对象锁 | 无锁并发,但生产调用方为零(见 §3.2 | object.rs delete_object 锁条件 |
| 10 | 孤儿 data_dir 回收 reclaim_orphan_data_dirs | 函数本体无锁;唯一生产调用方在 heal 锁内 | heal 流程内=锁内串行 | heal.rs 收尾调用;io_primitives.rs |
| 11 | 旧清理 receipt 对账 reconcile_old_data_cleanup_receipts | 函数本体无锁;调用点在 heal 锁内 + epoch fence 防误删 | 锁内串行 | object.rs 对账函数 |
| 12 | replication | 数据面为远端 HTTP 写(不落本地盘);本地元数据回写走对象锁 | 同锁串行 / 无交点 | replication_resyncer.rs 链路 |
| 13 | data_movement / rebalance / decommission 源清理 | 显式取对象锁 + 版本未变复核 + guard 复用(no_lock 只是复用已持锁) | 同锁串行 | data_movement/mod.rs 源清理 |
| 14 | copy_object | 目标对象锁 / 走 put 链锁 | 同锁串行 | object.rs copy_object 锁段 |
| 15 | 另一 heal 任务(跨 HealType/force_start | dedup key 跨类型不相交 + force_start 跳过去重 → 任务级可并发 | 最终在 ns 写锁上串行 | heal/manager.rs dedup key 构成 |
| 16 | admin `no_lock=true` heal | 客户端可控绕锁 | 无锁并发,明示运维选项(见 §3.3 | admin/handlers/heal.rs 透传 |
| 17 | stale multipart 清理 | multipart bucket 的 upload 路径锁 | 资源不相交 → 无交点 | bucket_lifecycle_ops.rs 清理链路 |
## 3. 残留窗口定性
### 3.1 PUT/CompleteMultipart commit tail(交点 2/6
写路径提交成功、释放对象锁之后,才 best-effort 删除被替换的旧 data_dir(注释明示有意
不阻塞下一操作)。该删除与并发 heal 对同一旧 data_dir 的读取/重建存在竞态窗口,但语义
安全:
- 删除目标是已被新提交替换的 unshared data_dirheal 的 canonical 元数据来自 quorum
仲裁(ETag/mod_time),此时 quorum 已指向新版本,heal 不会把已替换版本当作 canonical
复活;
- 竞态最坏后果 = heal 当轮对旧版本的一次 transient 失败/空转,重试轮自然收敛;清理
residue 会上报并重新入队 heal`report_old_data_dir_cleanup`);
- 换盘重建等长 heal 走 per-version 显式版本请求,quorum 元数据在锁内读取,不受 tail
影响。
### 3.2 纯 prefix 删除(交点 9
`delete_prefix && !delete_prefix_object` 的路径不取子对象锁(对象名空间锁无法保护前缀
递归删除),与并发 heal 存在理论复活窗口(heal 在 prefix 删除进行中依据旧 quorum 元
数据重建某版本)。全仓库核对结论:该路径的**生产调用方为零**——所有生产 `delete_prefix:
true` 调用点均同时设置 `delete_prefix_object: true`(从而取对象锁)或在测试模块内。这
是 API 能力面的暴露而非行为风险。若未来有调用方需要纯 prefix 删除,须在调用点证明与
heal/scanner 的隔离(例如 bucket 级停扫围栏)。
### 3.3 admin `no_lock=true`(交点 16
admin heal 请求可透传客户端 `nolock` 参数绕过 ns 锁(与 MinIO madmin 的同名选项对齐)。
这是运维明示选项:使用即自负与并发写的竞争责任。文档化即可,不建议收紧。
## 4. heal 侧自身的不变量保障
- dedup key 跨 HealType 不相交(object/metadata/mrf/ecdecode/prefix 各自键面)+ admin
`force_start` 可跳过去重 → 同对象可能同时存在多个 heal 任务,但它们的执行体全部在
`heal_object` 入口的 ns 写锁上串行(生产入口均 `no_lock=false`);
- read-repair 的本地 TTL 预留只去重自身来源,不拦截其他来源的 heal——同样由 ns 锁兜底;
- healing 旗标不落盘,故不存在"标记残留导致后续提交错误让路"的反向风险。
## 5. 回归测试
以下两个并发不变量测试随本审计加入 `crates/ecstore/src/set_disk/ops/heal.rs` 测试模块:
- `heal_racing_version_delete_never_resurrects_the_deleted_version`:注入 doomed 版本
shard 损坏后,版本化 DELETE 与 Deep heal 真并发(同一把锁争用),断言已删除版本不被
复活、存活版本完好;
- `heal_racing_unversioned_overwrites_preserves_the_last_commit`:非版本化覆盖提交(激活
commit tail 旧 data_dir 删除)与 Deep heal 循环竞态,断言最终 current 恰为最后一次
提交(etag 级一致)。
## 6. 结论
MinIO 的 `x-minio-healing` 是锁外元数据防御,前提是其 RenameData 提交内部存在与 heal
交错的版本清理逻辑;RustFS 的提交模型把这类交错从根上消除(提交面锁内互斥 + 清理外
移到 tail + tail 只删 unshared 旧目录),因此引入持久对象级 healing 标记没有对应的竞争
可防,反而会引入 FileInfo 落盘格式变更与标记残留清理两类新成本。维持现状,本对标疑点
关闭。
@@ -1,109 +0,0 @@
# Heal/Scanner 配置与语义对照(MinIO parity 决策记录)
对应 backlog rustfs/backlog#1878(父 #1862,批 HS-14/HS-16/HS-18)。本页沉淀三项"决策 + 文档化"结论:scanner idle 节流语义对照与迁移警告(HS-14)、单机默认扫描周期决策(HS-16)、stale multipart 与 tmp/.trash 清理三段核对(HS-18),并顺带收录 bitrot_cycle 与 alert_excess_folders 两项已确认的默认值差异。所有 MinIO 侧结论均于 2026-08 按 minio/minio master 逐源码核对(引用文件为上游路径),不转述二手资料。
运行时旋钮的完整清单、状态端点与调参流程见 [Scanner Runtime Controls](scanner-runtime-controls.md)excess 告警阈值差异见 [Scanner Excess Alerts](scanner-excess-alerts_zh.md)heal 并发模型对照见 [Heal 并发安全说明](heal-concurrency-safety-notes-zh.md)。
## 1. HS-14scanner idle 节流语义对照
### RustFS 当前语义(三因子)
RustFS 的 scanner 步进节流由三个因子共同决定(crates/scanner/src/sleeper.rs):
1. **总闸 `scanner.idle_mode` / `RUSTFS_SCANNER_IDLE_MODE`(默认 `true`**`false` 时所有节流 sleep 全部跳过,scanner 全速推进;`true` 时按下面两因子计算 sleep。
2. **速度档**`scanner.speed` / `RUSTFS_SCANNER_SPEED`,默认 `default`):档位表与 MinIO 完全一致(见下表)。目录级 sleep = `1ms × factor`(上限 `max_wait`);对象级 sleep = `本对象处理耗时 × factor`,下限 1ms、上限 `max_wait`
3. **前台读退避下限**`current_foreground_read_activity()` 取并发 GetObject 请求数(rustfs/src/storage/concurrency/request_guard.rs 的 `GetObjectGuard`)与流式读计数(`ForegroundReadGuard`)的较大值,换算为 `10ms × 活跃读数`、封顶 250ms 的下限;该下限对目录级与对象级 sleep 都生效(`.max(foreground_sleep)`),且**可以超过速度档的 `max_wait`**(自身封顶 250ms)。速度档为 `fastest`factor=0)时预设 sleep 为 0,但只要 `idle_mode=true`,前台读下限仍然生效。
| 速度档 | sleep factor | 单次 sleep 上限 | 周期间隔 |
|---|---:|---:|---:|
| `fastest` | 0 | 0 | 1s |
| `fast` | 1× | 100ms | 1m |
| `default` | 2× | 1s | 1m |
| `slow` | 10× | 15s | 1m |
| `slowest` | 100× | 15s | 30m |
实际行为矩阵(RustFS):
| `idle_mode` | 速度档 | 前台并发读 = 0 | 前台并发读 > 0 |
|---|---|---|---|
| `false` | 任意 | 完全不休眠,全速 | 完全不休眠,全速(前台退避也被总闸关闭) |
| `true` | `fastest` | 预设 sleep = 0,等效全速 | 每步 sleep = 前台读下限(10ms×读数,封顶 250ms) |
| `true` | 其余档 | 每步 sleep = 预设值(1ms~15s 封顶) | 每步 sleep = max(预设值, 前台读下限) |
周期间隔的解析优先级为 env `RUSTFS_SCANNER_CYCLE` > 持久化 `scanner.cycle` > `scanner.start_delay` > 启动期默认覆盖(当前恒无)> 速度档派生(crates/scanner/src/runtime_config.rs)。另有 `scanner.yield_every_n_objects`(默认 128)的协作式让出,与节流 sleep 相互独立。
### MinIO 当前语义(master 逐源码核对)
MinIO 的对应开关是 `scanner:idle_speed` / `MINIO_SCANNER_IDLE_SPEED`internal/config/scanner/scanner.go):取值为空串或 `on`(默认)时 `IdleMode=0`,取值 `off``IdleMode=1`。启动/配置加载时一次性写入 `scannerIdleMode`cmd/config-current.go),扫描侧闭包 `weSleep = scannerIdleMode.Load() == 0`cmd/xl-storage-disk-id-check.go):**`on`(默认)= 目录级与对象级节流 sleep 始终插入(按速度档 factorminSleep 100µs);`off` = 两条节流路径完全不 sleep,全速扫描**。当前上游没有任何按 S3 请求/磁盘活动动态调整节流的逻辑——这是静态开关。
命名具有误导性,是历史残留:2024-01 之前 `weSleep` 由磁盘活动驱动("Entire queue is full, so we sleep",即有并发 S3/heal 活动才 sleep),minio/minio#18734commit 7705605b)把该活动门替换为上述静态配置(初版取值 `throttled`/`full`,后改为 `on`/`off`),上游残留注释 "default is throttled when idle"、"Sleep always or based on incoming S3 requests" 均是替换前的语义描述,与现行代码不符。
### 对照与迁移警告
| 维度 | RustFS | MinIOmaster |
|---|---|---|
| 开关名 | `scanner.idle_mode` / `RUSTFS_SCANNER_IDLE_MODE` | `scanner:idle_speed` / `MINIO_SCANNER_IDLE_SPEED` |
| 取值 | 布尔 `true`/`false` | `on`/`off` |
| 默认 | `true`(节流开启) | `on`(节流开启) |
| 开 = | 节流总闸开:速度档 sleep + 前台读下限 | 节流总闸开:速度档 sleep |
| 关 = | 完全不休眠(含前台读下限一并失效) | 完全不休眠 |
| 活动耦合 | 有:前台并发读抬高 sleep 下限(10ms×读数,封顶 250ms) | 无(2024-01 起为静态开关) |
| 速度档表 | 两边完全一致(上表) | 同左 |
迁移警告:
- **环境变量名不可照搬**:RustFS 只读取 `RUSTFS_*` 前缀,不解析 `MINIO_SCANNER_*` 任何别名(crates/scanner、crates/utils 的 env 读取无别名链,测试还专门断言 `MINIO_SCANNER_SPEED`/`MINIO_SCANNER_CYCLE` 不泄漏生效)。照搬 `MINIO_SCANNER_IDLE_SPEED=off` 到 RustFS 会静默无效,必须改写成 `RUSTFS_SCANNER_IDLE_MODE=false`
- **取值词表不同**`on/off` vs `true/false`,不能原样复制。
- **方向澄清(修正父 issue 的预设)**:按当前上游源码,MinIO `idle_speed` 与 RustFS `idle_mode` 在"开=节流、关=全速"方向上是一致的,并非反向;父 issue 中"MinIO on=集群空闲才节流、off=始终按 delay 节流"的矩阵描述的是 2024-01 之前的活动耦合行为与反向解读,与 master 不符。真正需要写进迁移手册的差异是:MinIO 的 `idle_speed` 名称暗示"空闲时才慢"但实际是静态总闸;RustFS 的 `idle_mode=true` 在总闸之上还叠加了 MinIO 没有的前台读保护下限。
- **`false` 是大锤**RustFS `idle_mode=false` 会连前台读退避一起关闭,scanner 与前台读完全抢盘;仅在 benchmark 或可独占 IO 的窗口使用。
**决策(HS-14):保持现状。** RustFS 语义更直观(`idle_mode` = 节流总闸,`true` 即自适应限速),且比 MinIO 多一层前台读保护;不新增 `RUSTFS_SCANNER_IDLE_SPEED` 兼容别名(无社区强诉求不做,避免双入口漂移)。本节即对照表与迁移警告的正式落点。
## 2. bitrot_cycle 默认差异
| 项 | RustFS | MinIO |
|---|---|---|
| 键 | `heal.bitrot_cycle` / `RUSTFS_SCANNER_BITROT_CYCLE_SECS`scanner.bitrot_cycle 为兼容旧键) | `heal:bitrotscan` / `MINIO_HEAL_BITROTSCAN` |
| 默认 | 30 天(crates/config/src/constants/heal.rs 的 `DEFAULT_HEAL_BITROT_CYCLE_SECS`):按墙钟周期把扫描切深扫(deep bitrot | `off`internal/config/heal/heal.go 默认 `EnableOff`):不做周期性深扫,仅普通扫描 + 管理端手动深扫 |
| 对齐方式 | 迁移 MinIO 行为:`heal.bitrot_cycle=off``RUSTFS_SCANNER_BITROT_CYCLE_SECS=disabled` | 反向:`heal:bitrotscan=<秒>` |
RustFS 的 30 天默认是刻意的耐用性默认(周期性全量 bitrot 校验),代价是每 30 天一轮深扫 IO;单机场景另有清洁空闲退避封顶约 42 分钟的墙钟保护(见 scanner-runtime-controls.md)。这是行为差异而非缺陷,文档化即可。
## 3. alert_excess_folders 默认差异
RustFS 默认 65538(容纳 Proxmox Backup Server 每目录 65536 chunk 的布局),MinIO 默认 50000。差异原因、另两个 excess 阈值(versions=100 相同、version_size TiB vs TB)、事件名映射与冷却语义已完整记录在 [Scanner Excess Alerts](scanner-excess-alerts_zh.md),此处不重复。
## 4. HS-18stale multipart 与 tmp/.trash 清理三段核对
MinIO 把"清理已删除数据"拆成三段:stale upload 先 rename 进 `.minio.sys/tmp/.trash/<uuid>` 隔离(rename 快、原子);trash 由独立例程排空;tmp 下非 trash 的旧目录单独回收。逐段核对 RustFS:
| 段 | MinIO | RustFS | 判定 |
|---|---|---|---|
| stale multipart → 隔离 | `cleanupStaleUploadsOnDisk`cmd/erasure-multipart.go)逐盘列出 multipart 目录,按 uploadID 目录名里的 UnixNano 判龄,超过 `stale_uploads_expiry`(默认 24h)即 `renameAll``.minio.sys/tmp/.trash/<uuid>`,空 sha 目录、tmp 旧目录同法 | `cleanup_stale_multipart_uploads_in_set`crates/ecstore/src/bucket/lifecycle/bucket_lifecycle_ops.rs)发现候选后取 ns 写锁 + 重查(`lock_stale_multipart_cleanup`),`delete_all_with_quorum` 扇出逐盘递归删除,而 LocalDisk 的递归删除内部就是 `move_to_trash`crates/ecstore/src/disk/local.rs)把目录 rename 进 `.rustfs.sys/tmp/.trash/<uuid>` | 行为等价(都是先隔离后清理);RustFS 额外有写锁 + quorum 重查 + 锁丢失 fencecrates/ecstore/src/set_disk/ops/multipart.rs 的 `StaleMultipartCleanupGuard`),防并发 CompleteMultipartUpload 竞争,安全性强于 MinIO 的无锁 rename |
| trash 排空 | 每 `delete_cleanup_interval`(默认 5minternal/config/api/api.go)逐盘删 `.trash` 内条目,逐条以 `deleteCleanupSleeper`factor 5 / 25mscmd/globals.go)节流 | 每盘独立 `cleanup_deleted_objects_loop``DELETED_OBJECTS_CLEANUP_INTERVAL` = 5mcrates/ecstore/src/disk/local.rs),先排空 `.trash` 再回收 tmp 旧目录;排空为顺序 `remove_dir_all`/`remove_file`**无逐条 sleep 节流** | 基本等价;唯一差异是 RustFS 排空不节流,trash 积压大时单轮 IO 更突发(5m 周期天然限频),文档化,如实测出现清理风暴再补节流 |
| tmp 非 trash 旧目录 | 并在 `cleanupStaleUploadsOnDisk` 内:非 `.trash` 的 tmp 目录超过 `stale_uploads_expiry`24hrename 进 trash(随 6h 任务) | `cleanup_stale_tmp_objects`crates/ecstore/src/disk/local.rs)随 5m 循环执行:非 `.trash` 目录超过 `STALE_TMP_OBJECT_EXPIRY` = 24h 即 rename 进 trash;另有启动时 tmp → tmp-old 整体换名 + 后台删除的崩溃安全路径 | 行为等价(阈值同为 24h);RustFS 检查频率 5m vs MinIO 6h,回收更及时 |
周期与环境变量默认值对照(两边一致):
| 项 | RustFS | MinIO |
|---|---|---|
| stale upload 过期阈值 | `RUSTFS_API_STALE_UPLOADS_EXPIRY`,默认 24h | `MINIO_API_STALE_UPLOADS_EXPIRY`,默认 24h |
| stale multipart 清理周期 | `RUSTFS_API_STALE_UPLOADS_CLEANUP_INTERVAL`,默认 6h | `MINIO_API_STALE_UPLOADS_CLEANUP_INTERVAL`,默认 6h |
| trash 排空周期 | 5m(常量,暂无开关) | `MINIO_API_DELETE_CLEANUP_INTERVAL`,默认 5m |
关于 rustfs/src/delete_tail_activity.rs:它**不覆盖三段中的任何一段**。该模块是 delete 尾部活动的进程内指标计数(inflight gauge + 耗时 histogram),供 allocator 回收压力判断(rustfs/src/allocator_reclaim.rs)使用;生产代码目前只在对象复用路径使用 `Replication`/`Notify` 两个 stage 计数,`Tail`/`Cleanup` 枚举值暂无调用点。
崩溃残留窗口结论:
- trash 内部残留(排空中途崩溃):`.trash/<uuid>` 是自包含目录,下一轮 5m tick 重扫 `.trash` 自然收敛,与 MinIO 相同。
- 跨盘扇出中途崩溃(部分盘已 rename 进 trash、其余未动):若剩余盘数仍满足写 quorum,下一轮 6h 任务重新发现候选并重删,自然收敛;若已清理盘数超过 parity(剩余低于写 quorum),`check_multipart_upload_path_exists``FileNotFound` 不在 `OBJECT_OP_IGNORED_ERRS`crates/ecstore/src/disk/error_reduce.rs)而判 quorum 失败,候选被跳过,残留 uploadID 目录不会被该任务收敛(不可见于 S3 API,仅占盘空间)。该窗口极窄(逐盘 rename 为毫秒级,需恰在扇出中途且已过 parity 盘时进程死亡)。MinIO 同场景会收敛(逐盘独立处理、无 quorum 闸门)。**分级:有崩溃残留窗口(极窄)→ 登记后续修复**;修复需为清理守卫提供把"已不存在"计为达成终态的专用 quorum 变体(不能改共享的 `check_multipart_upload_path_exists` 语义,它同时服务 CompleteMultipartUpload),超出本批"几行小修"边界,不在本 PR 扩 scope。
## 5. HS-16:单机(ErasureSD)默认扫描周期决策
启动期曾有预留钩子 `single_disk_default_cycle_secs`,可按维护特征(lifecycle/replication/巡检失败)为单机覆盖默认周期,但从未接线、恒返回 `None`,已删除(本批 PR)。决策:**单机默认周期保持速度档派生(`default` 档 = 60s),不做特殊覆盖**。理由:其一,无任何实测依据表明单机冷启动 ILM 延迟需要更短周期,凭空缩短只会放大空闲扫描频次;其二,单机已有清洁空闲退避(连续干净周期间隔翻倍,默认 bitrot 窗口下封顶约 42 分钟,见 scanner-runtime-controls.md),空闲时的周期压力已被消化;其三,若确有诉求,用户可用 `RUSTFS_SCANNER_CYCLE` / `scanner.cycle` 显式配置,无需内置特殊路径。需要更激进短周期的场景应先拿实测数据再议。
## 6. 决策摘要
- HS-14:保持 `RUSTFS_SCANNER_IDLE_MODE` 现语义(true=节流总闸+前台读下限,false=全速),文档化对照表与迁移警告,不做兼容别名。
- HS-16:删除恒 `None` 的单机默认周期钩子,单机周期保持速度档派生 + 清洁空闲退避。
- HS-18:三段清理行为等价(trash 排空无逐条节流、tmp 回收频率 5m vs 6h 两处小差异文档化);跨盘扇出的极窄崩溃残留窗口登记后续;周期默认值 24h/6h/5m 与 MinIO 对齐。
+142 -128
View File
@@ -1,164 +1,178 @@
# Hotpath warp A/B runbook
# Hotpath warp runbook (A/B gate and ABBA evidence)
Relative-budget A/B gate for the hotpath series (rustfs/backlog#935 HP-14). It
runs the same warp workloads against a **baseline** binary and a **candidate**
binary, across the drive-sync on/off matrix, then applies a relative budget:
a metric regressing past the fail budget fails the gate, past the warn budget
warns. This is how the macOS profiling conclusions of the HP series get
confirmed or corrected on Linux — structural wins (call counts, read
amplification) should hold; absolute numbers are whatever the rig measures.
**Use this when:** you need S3-face (PutObject / GetObject / mixed) performance evidence for a code change: the nightly relative-budget gate, a quick local A/B, or a formal ABBA run whose numbers will be quoted in a PR.
**Source of truth:** `scripts/run_hotpath_warp_ab.sh` (quick A/B rig), `scripts/run_hotpath_warp_abba.sh` (ABBA runner; also what CI executes), `scripts/hotpath_warp_ab_gate.sh` (relative-budget gate), `scripts/run_object_batch_bench_enhanced.sh` (warp driver, medians, `baseline_compare.csv`), `.github/workflows/performance-ab.yml` (CI gate).
Pieces:
This runbook compares two binaries. To sweep one `RUSTFS_*` runtime knob at a time against a fixed binary, use [object-io-tuning-ab-matrix.md](object-io-tuning-ab-matrix.md) instead.
- `scripts/run_hotpath_warp_ab.sh` — orchestrator (baseline vs candidate,
workload × drive-sync matrix).
- `scripts/hotpath_warp_ab_gate.sh` — the budget gate over the
`baseline_compare.csv` deltas the load driver emits.
- `scripts/run_object_batch_bench_enhanced.sh` — the warp driver + median +
`baseline_compare.csv` (reused, not reimplemented).
- `.github/workflows/performance-ab.yml` — nightly on `main` (post-merge
detection) plus opt-in pre-merge via the `perf-ab` label.
Two entry points share one workload matrix, one gate, and one deploy-hook shape:
Metric directions: `reqps` (put obj/s) and `throughput` (get MiB/s) are
higher-is-better; `latency` / p99 (mixed) is lower-is-better. warp is assumed
pre-installed, as elsewhere in `scripts/`.
| Entry | Script | Legs per cell | Use |
| --- | --- | --- | --- |
| Quick A/B | `scripts/run_hotpath_warp_ab.sh` | baseline → candidate | local smoke, fast triage of a suspected regression |
| ABBA | `scripts/run_hotpath_warp_abba.sh` | A1 baseline → B1 candidate → B2 candidate → A2 baseline | formal evidence; the CI gate |
This rig is the only entry point for S3-face performance coverage. There is deliberately no in-process criterion benchmark for those operations: a criterion harness that stands up an embedded server measures the harness, not the S3 path. Micro-benchmarks stay at function level (EC encode, `xl.meta` parse, `rename_data`). Knob-level sweeps against a fixed binary are a different question; see [object-io-tuning-ab-matrix.md](object-io-tuning-ab-matrix.md).
## Shared prerequisites
- Linux host. A laptop is acceptable for a quick A/B smoke only; formal evidence needs a dedicated runner or a cluster.
- `warp` on `PATH`, or `--warp-bin <path>`.
- Two Linux release binaries, baseline and candidate (build each with `cargo build --release -p rustfs --bins` at its commit; cross-compile with `cargo zigbuild --release --target x86_64-unknown-linux-gnu -p rustfs --bins` for a cluster).
- Disposable disks or data root in local mode; an isolated benchmark bucket and credentials in cluster mode. Never bench against production data.
- Readiness polling (`--health-timeout`, default 180 s in both scripts) must outlast the server's own startup budget, `DEFAULT_STARTUP_READINESS_MAX_WAIT_SECS` in `crates/config/src/constants/health.rs`; a shorter poll misreports a slow cold start as a failure.
- Metric directions: `reqps` (put obj/s) and `throughput` (get MiB/s) are higher-is-better; `latency`/p99 (mixed) is lower-is-better.
## Workload matrix
Six workloads × the drive-sync on/off matrix × baseline/candidate = 24 cells:
Every workload runs with `RUSTFS_DRIVE_SYNC_ENABLE=true` and `=false`. Quick A/B: 6 × 2 × 2 = 24 cells; ABBA: 6 × 2 × 4 = 48 cells.
| Workload | mode | size | why |
| Workload | mode | size | Why |
| --- | --- | --- | --- |
| `put-4kib` / `get-4kib` | put / get | 4KiB | the #4221 fsync regression size (~-10% @4KiB) — previously invisible |
| `put-4mib` / `get-4mib` | put / get | 4MiB | bulk obj/s and MiB/s |
| `get-10mib` | get | 10MiB | the historical large-GET EOF size |
| `mixed-256k` | mixed | 256KiB | p99 latency |
| `put-4kib` / `get-4kib` | put / get | 4 KiB | small-object fsync-sensitive path |
| `put-4mib` / `get-4mib` | put / get | 4 MiB | bulk obj/s and MiB/s |
| `get-10mib` | get | 10 MiB | large-GET streaming path |
| `mixed-256k` | mixed | 256 KiB | p99 latency |
Sizes are passed to the load driver via `--sizes` (one size per cell); the
driver's `DEFAULT_SIZES` covers 1KiB..10MiB, so any of those can be added by
editing `WORKLOADS` in `scripts/run_hotpath_warp_ab.sh`. A 1KiB cell is left
out for now to keep the nightly matrix comfortably under budget; re-enable it
(one line in `WORKLOADS`) once perf-6 recalibrates the warp params.
Sizes are passed to the driver as `--sizes` (one per cell); any size in the driver's `DEFAULT_SIZES` (`scripts/run_object_batch_bench_enhanced.sh`) can be added by editing the `WORKLOADS` array in the chosen script.
CI runs a **short** warp matrix (`--duration`/`--rounds`/`--cooldown` tuned in
`.github/workflows/performance-ab.yml`) so all 24 cells fit the budget without
dropping cells. These params are deliberately noisy-but-fast for the Phase-0
"keep the pipeline alive" goal; perf-6 recalibrates them.
## Gate
## Baseline binary cache (CI)
`scripts/hotpath_warp_ab_gate.sh` compares each metric against baseline: a regression beyond `--fail-pct` (default `FAIL_PCT=10`) fails, beyond `--warn-pct` (default `WARN_PCT=5`) warns. `--allow-regression --exemption-reason "<why>"` records a FAIL as an exempted WARN and exits 0; use it only for a deliberate correctness trade (for example paying write cost to restore power-loss durability).
The nightly no longer builds both binaries from source. Every push to `main`
runs a `build-baseline-cache` job that builds the release binary once and stores
it in the actions cache under `rustfs-baseline-<sha>` (perf-3). The A/B job
restores the binary for `origin/main` by that key and passes it as
`--baseline-bin`; on the nightly, where the candidate commit equals the baseline
commit, the same cached binary serves both phases (`--skip-build`) and the run
does zero source builds — the common path finishes well under 50 minutes. A
cache miss (binary evicted, or not built for that SHA yet) transparently falls
back to the source double-build via `--baseline-ref origin/main`.
## Deploy-hook contract (external / cluster mode)
Each `gate.md` ends with a **Provenance** section recording the baseline and
candidate commit SHAs and whether each binary came from the cache or a source
build, plus the runner, warp version, and matrix params. `perf-5`/`perf-12`
reuse this contract for their archived baselines.
In external mode (`--endpoint <host:port>`) the rig never starts or restarts RustFS. Before each phase or leg it runs `--deploy-hook <cmd>` with the context below in the environment, then waits for `http://<endpoint><health-path>` (default `/health`). A non-zero hook exit aborts the run.
## Local mode (quick / CI smoke)
| Variable | Quick A/B | ABBA | Value |
| --- | --- | --- | --- |
| `HOTPATH_AB_PHASE` / `HOTPATH_ABBA_PHASE` | yes | yes | `baseline` or `candidate` |
| `HOTPATH_AB_BINARY` / `HOTPATH_ABBA_BINARY` | yes | yes | selected binary path (A/B: may be empty if the hook builds its own) |
| `HOTPATH_AB_DRIVE_SYNC` / `HOTPATH_ABBA_DRIVE_SYNC` | yes | yes | `true` or `false` for this cell |
| `HOTPATH_ABBA_LEG` | — | yes | `A1`, `B1`, `B2`, `A2` |
| `HOTPATH_ABBA_WORKLOAD`, `HOTPATH_ABBA_MODE`, `HOTPATH_ABBA_SIZE`, `HOTPATH_ABBA_CELL_ID` | — | yes | cell identity |
| `HOTPATH_ABBA_DATASET_NAMESPACE`, `HOTPATH_ABBA_BUCKET` | — | yes | run namespace and per-leg benchmark bucket the hook must provision or reset |
| `HOTPATH_ABBA_DEPLOY_EVIDENCE_FILE` | — | yes | path the hook must write a non-empty evidence file to |
Builds both binaries and runs a throwaway single-node server on local disks.
The ABBA runner refuses external mode without a hook unless `--allow-unmanaged-external` is passed, and output from that mode is not formal evidence.
Ansible-shaped hook (replace `/path/to/ansible` and the inventory group; the `config` tag must thread `RUSTFS_DRIVE_SYNC_ENABLE`, or the finer `RUSTFS_DURABILITY_MODE`, into the deployed unit):
```bash
--deploy-hook '
set -euo pipefail
cd /path/to/ansible
cp "${HOTPATH_ABBA_BINARY:?}" roles/rustfs/files/rustfs
export RUSTFS_DRIVE_SYNC_ENABLE="${HOTPATH_ABBA_DRIVE_SYNC:?}"
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags stop
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags config
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags binary-copy
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags start
'
```
For the quick A/B rig use the `HOTPATH_AB_*` names.
## Quick A/B (`run_hotpath_warp_ab.sh`)
```bash
# build both binaries (baseline from --baseline-ref, default origin/main) and run a throwaway single-node server on local disks
scripts/run_hotpath_warp_ab.sh --baseline-ref origin/main
# or with prebuilt binaries:
scripts/run_hotpath_warp_ab.sh --skip-build \
--baseline-bin ./rustfs-main --candidate-bin ./target/release/rustfs
# prebuilt binaries
scripts/run_hotpath_warp_ab.sh --skip-build --baseline-bin ./rustfs-main --candidate-bin ./target/release/rustfs
# print the plan only
scripts/run_hotpath_warp_ab.sh --dry-run --skip-build --baseline-bin /tmp/base --candidate-bin /tmp/cand
# external cluster
scripts/run_hotpath_warp_ab.sh --endpoint "$CLUSTER_ENDPOINT" --deploy-hook '<see above>' \
--baseline-bin /path/to/rustfs-main --candidate-bin ./target/x86_64-unknown-linux-gnu/release/rustfs
```
Preview the full plan without running anything:
Outputs under `target/hotpath-ab/<ts>/`: `gate.md` (ends with a **Provenance** section: baseline/candidate SHAs, binary source, runner, warp version, matrix params; extend with `--provenance-note`) and `server-logs/<phase>-sync-<sync>.{log,env}`. On a health-check failure the rig prints the last 50 server log lines; in local mode it fails fast if the server process exits before becoming healthy.
## ABBA evidence run (`run_hotpath_warp_abba.sh`)
`B1` and `B2` are compared with `A1` for the candidate delta; `A2` is compared with `A1` for baseline drift. Required flags: `--baseline-bin`, `--candidate-bin`, `--baseline-revision`, `--candidate-revision`. The script enforces `--rounds >= 3`; prefer `--rounds 5` or more for formal evidence when budget allows.
```bash
scripts/run_hotpath_warp_ab.sh --dry-run --skip-build \
--baseline-bin /tmp/base --candidate-bin /tmp/cand
# local Linux runner (throwaway data root; a reused run namespace is rejected)
scripts/run_hotpath_warp_abba.sh \
--baseline-bin /tmp/rustfs-baseline --candidate-bin /tmp/rustfs-candidate \
--baseline-revision "$(git rev-parse origin/main)" --candidate-revision "$(git rev-parse HEAD)" \
--address 127.0.0.1:9000 --data-root /var/tmp/rustfs-hotpath-abba --disks 4 \
--duration 120s --rounds 3 --cooldown 30 --concurrency 16 \
--out-dir target/hotpath-abba/linux-local
# production-like cluster
scripts/run_hotpath_warp_abba.sh \
--baseline-bin /srv/rustfs-binaries/rustfs-baseline --candidate-bin /srv/rustfs-binaries/rustfs-candidate \
--baseline-revision <sha> --candidate-revision <sha> \
--endpoint rustfs-bench.example.internal:9000 --deploy-hook '<see above>' \
--duration 180s --rounds 5 --cooldown 45 --concurrency 32 \
--out-dir target/hotpath-abba/cluster-pr-XXXX
```
## External mode (real cluster, ansible-deployed)
Add `--dry-run` to print the schedule without starting servers or warp. Output layout:
For the production-representative run, warp targets an already-running cluster
and a `--deploy-hook` swaps in each phase's binary and durability config
between the baseline and candidate phases. The hook receives context via the
environment:
```text
<out-dir>/
manifest.env
abba_schedule.csv
candidate_gate.md
baseline_drift_gate.md
summary.md
<workload>/<sync>/<leg>/median_summary.csv
<workload>/<sync>/<leg>/baseline_compare.csv
```
- `HOTPATH_AB_PHASE``baseline` or `candidate`
- `HOTPATH_AB_BINARY` — binary path (or empty; the hook may build its own)
- `HOTPATH_AB_DRIVE_SYNC``true` or `false` for this matrix cell
Attach to the PR or issue: `summary.md`, `candidate_gate.md`, `baseline_drift_gate.md`, `abba_schedule.csv`, every `median_summary.csv` and `baseline_compare.csv` for a failed or borderline workload, and the host telemetry used to explain saturation. Preserve the output directory unmodified.
This maps directly onto the team's ansible harness. Build the candidate with
the cross toolchain, stage both binaries, then let the hook drive
`rustfs-manage.yml`:
### Interpretation
| Candidate gate | A2 drift gate | Interpretation |
| --- | --- | --- |
| PASS | PASS | Candidate acceptable for the measured matrix. |
| WARN | PASS | Small measurable signal; inspect telemetry and decide whether it is expected. |
| FAIL | PASS | Candidate likely regressed the workload; investigate before merge. |
| FAIL | FAIL on the same workload | Environment drift is high; rerun on a quieter runner or raise duration and rounds. |
| PASS | FAIL | Rig unstable; do not quote the numbers as proof of improvement. |
Rules that override the table: a candidate result is actionable only when the `A2` drift for the same workload passes or is materially smaller than the `B1`/`B2` delta; never report a win or loss for a workload whose drift gate failed without a rerun. When `B1` and `B2` disagree, the cell is inconclusive even if the gate passes. Report only measured facts: deltas, drift, saturation, failed workloads.
### CPU and memory evidence
Warp output says whether throughput or latency changed; host telemetry says why. Collect it for the whole run and stop the collectors after the script exits.
| Tool | Command | Answers |
| --- | --- | --- |
| `pidstat` | `pidstat -durh 5 > <out-dir>/telemetry/pidstat.txt &` | per-process CPU, memory, disk |
| `mpstat` | `mpstat 5 > <out-dir>/telemetry/mpstat.txt &` | CPU saturation and steal |
| `iostat` | `iostat -xz 5 > <out-dir>/telemetry/iostat.txt &` | device queue depth and latency |
| `perf` | `perf record -F 99 -g -- sleep 180` around one representative cell, then `perf report --stdio` | CPU attribution after the gate shows an effect |
samply against a running RustFS process goes through the bounded helper, one attach window per leg or focused cell:
```bash
# 1. Build the candidate (cross-compile for the cluster target).
cargo zigbuild --release --target x86_64-unknown-linux-gnu -p rustfs --bins
# 2. Run the A/B against the cluster; the hook deploys the phase's binary and
# applies the drive-sync config, then restarts, before each phase.
scripts/run_hotpath_warp_ab.sh \
--endpoint "$CLUSTER_ENDPOINT" \
--deploy-hook '
set -euo pipefail
cd /home/xiaomage/xiaomage/ansible
# Select the phase binary and the drive-sync value for this cell.
cp "${HOTPATH_AB_BINARY:?}" ./roles/rustfs/files/rustfs
export RUSTFS_DRIVE_SYNC_ENABLE="$HOTPATH_AB_DRIVE_SYNC"
ansible-playbook -f 4 -l testing rustfs-manage.yml --tags stop
ansible-playbook -f 4 -l testing rustfs-manage.yml --tags config
ansible-playbook -f 4 -l testing rustfs-manage.yml --tags binary-copy
ansible-playbook -f 4 -l testing rustfs-manage.yml --tags start
' \
--baseline-bin /path/to/rustfs-main \
--candidate-bin ./target/x86_64-unknown-linux-gnu/release/rustfs
scripts/run_samply_attach_window.sh --pid "$RUSTFS_PID" --duration-secs 180 \
--output <out-dir>/telemetry/samply-A1-get-4mib.json.gz
```
The `config` tag is responsible for threading `RUSTFS_DRIVE_SYNC_ENABLE` (or
the finer `RUSTFS_DURABILITY_MODE`) into the deployed unit — the hook exports
it so the config template can pick it up. The rig itself never restarts the
cluster; lifecycle stays with ansible.
After each window confirm the `.json.gz` profile and its `.syms.json` sidecar are non-empty, no `samply` process is still attached to the PID, and any temporary `perf_event_paranoid` change is restored; reject the cell otherwise.
## Budget and exemptions
Instrumented builds (features in `rustfs/Cargo.toml`): `--features hotpath-alloc` for allocation attribution, `--features hotpath-cpu` for CPU hotpath sections. Compare instrumented binaries only with other builds of the same mode; never use them for throughput acceptance, because the instrumentation changes what is measured.
Default budget: a metric regressing more than **10%** vs baseline fails,
more than **5%** warns. Tune with `--fail-pct` / `--warn-pct`.
## CI gate (`performance-ab.yml`)
Some regressions are the correct trade — #4221 deliberately paid a large write
cost to restore power-loss durability. For those, run with
`--allow-regression` (or add the `perf-deliberate-tradeoff` label in CI): the
FAIL is recorded and rendered as an exempted WARN, and the gate exits 0.
| Aspect | Value |
| --- | --- |
| Triggers | `schedule` (nightly cron `31 6 * * *` UTC against `main`) and `workflow_dispatch` (inputs `duration`, default `12s`; `allow_regression`, boolean). No `pull_request` trigger and no label gating. |
| Jobs | `warp-ab` (runner `sm-standard-2`, `timeout-minutes: 180`); `alert-on-failure` (opens the scheduled-failure issue; scheduled runs only). |
| Baseline commit | scheduled: head of the last successful scheduled run (falls back to the candidate itself when there is none); dispatch: `origin/main`. The baseline must be an ancestor of the candidate. |
| Binary cache | `actions/cache` inside `warp-ab`, key `rustfs-baseline-<baseline_sha>`. Miss: source build (same-commit runs build once and reuse the binary for both phases); a run whose gate passes saves the candidate binary under `rustfs-baseline-<candidate_sha>` for the next night. |
| Command | `scripts/run_hotpath_warp_abba.sh --duration <input> --rounds 3 --cooldown 5 --health-timeout 180 --baseline-revision <sha> --candidate-revision <sha> --baseline-bin ... --candidate-bin ...` |
| Exemption | `allow_regression=true` on dispatch adds `--allow-regression --exemption-reason "workflow dispatch override"`. |
| Artifacts | `hotpath-warp-ab-<run_number>` containing `target/hotpath-abba/` (14-day retention); the step summary renders `candidate_gate.md`, or the server-log tails when the rig failed before the gate. The `Enforce gate` step fails the job on a non-zero rig exit. |
## Diagnosing a failed run
Each phase's server log and its startup environment are written under the run's
output dir (`target/hotpath-ab/<ts>/server-logs/<phase>-sync-<sync>.{log,env}`)
and uploaded in the `hotpath-warp-ab-<run>` artifact, so a failure is
diagnosable after the fact. On a health-check failure the rig also dumps the
last 50 log lines into the job log and the CI job writes the failing phase (or
the gate table) into the GitHub step summary.
Readiness polling waits up to `--health-timeout` seconds (default **180**),
which must outlast the server's own startup-readiness budget
(`RUSTFS_STARTUP_READINESS_MAX_WAIT_SECS`, default 120s) — a shorter poll on a
slow shared runner misreports a slow cold start as a failure. In local mode the
rig also fails fast if the server process exits before becoming healthy instead
of polling out the full budget.
## Scope note
The gate logic is unit-validated across pass/warn/fail/exempt outcomes; the
orchestrator and workflow are shellcheck- and `--dry-run`-validated. The first
real warp measurement belongs on a Linux runner or the ansible cluster — there
is no warp/multi-disk rig in the repo's local checkout.
This warp A/B gate is the **only** entry point for S3-face (PutObject /
GetObject / ListObjects) performance coverage. There is deliberately no
in-process criterion benchmark for those operations: a criterion harness that
stands up an embedded server measures the harness, not the S3 path, so it would
report a number without guarding anything. Micro-benchmarks stay at the
function level (EC encode, `xl.meta` parse, `rename_data`; perf-8).
The nightly detects a regression within a day of landing; it does not block a merge. For pre-merge evidence run the ABBA procedure above and attach the outputs to the PR.
@@ -1,296 +0,0 @@
# Hotpath warp ABBA validation runbook
This runbook describes how to collect formal Linux or production-cluster
evidence for hotpath performance changes. Use it when a short local A/B smoke
run is too noisy to decide whether a regression is real.
The ABBA runner executes each workload and drive-sync cell as:
```text
A1 baseline -> B1 candidate -> B2 candidate -> A2 baseline
```
`B1` and `B2` are compared with `A1` to measure the candidate delta. `A2` is
also compared with `A1` to measure baseline drift. Treat a candidate regression
as actionable only when the `A2` drift is passing or materially smaller than
the `B1` and `B2` delta for the same workload.
## Scope
Use this runbook for hotpath profiling and performance validation of RustFS
object I/O changes, especially when CPU, memory allocation, lock/channel wait
time, request throughput, or tail latency is the review question.
The script validates the same workload matrix as the hotpath warp A/B gate:
| Workload | mode | size |
| --- | --- | --- |
| `put-4kib` | put | 4KiB |
| `put-4mib` | put | 4MiB |
| `get-4kib` | get | 4KiB |
| `get-4mib` | get | 4MiB |
| `get-10mib` | get | 10MiB |
| `mixed-256k` | mixed | 256KiB |
Each workload runs with `RUSTFS_DRIVE_SYNC_ENABLE=true` and
`RUSTFS_DRIVE_SYNC_ENABLE=false`, so a full ABBA pass produces 48 measurement
cells: 6 workloads x 2 drive-sync modes x 4 ABBA legs.
## Prerequisites
Run the formal pass on Linux, not on a laptop smoke environment.
Required tools on the bench host:
- `bash`, `curl`, `git`, and core GNU userland.
- `warp` on `PATH`, or pass `--warp-bin`.
- Two RustFS Linux binaries: one baseline and one candidate.
- Enough isolated disks or directories for the local runner, or an externally
managed RustFS cluster for production-like validation.
- Stable host telemetry collection such as `pidstat`, `mpstat`, `iostat`,
`sar`, `perf`, `heaptrack`, or the platform's equivalent observability stack.
Cluster-mode requirements:
- A deploy hook that can replace the RustFS binary on every node.
- The hook must apply `RUSTFS_DRIVE_SYNC_ENABLE` for the current ABBA leg.
- The hook must restart RustFS and return only after the rollout command has
been accepted. The ABBA script performs the HTTP readiness wait.
- The benchmark client should run outside the RustFS nodes when possible.
- Do not run against a production data set unless the workload bucket and test
credentials are isolated and approved for destructive benchmark traffic.
## Build the binaries
Build the baseline from the comparison commit, usually `origin/main` or the
previous accepted release:
```bash
git fetch origin main
git switch --detach origin/main
cargo build --release -p rustfs --bins
cp target/release/rustfs /tmp/rustfs-baseline
```
Build the candidate from the PR commit:
```bash
git switch <candidate-branch>
cargo build --release -p rustfs --bins
cp target/release/rustfs /tmp/rustfs-candidate
```
For cross-compiled cluster binaries, keep both outputs on the bench host and
make the deploy hook copy the selected binary to the cluster. The ABBA runner
passes the selected binary path through `HOTPATH_ABBA_BINARY`.
## Local Linux runner
Use local mode for a dedicated Linux runner with disposable data paths. This is
not a substitute for a production-like cluster, but it is useful before spending
cluster time.
```bash
scripts/run_hotpath_warp_abba.sh \
--baseline-bin /tmp/rustfs-baseline \
--candidate-bin /tmp/rustfs-candidate \
--address 127.0.0.1:9000 \
--data-root /var/tmp/rustfs-hotpath-abba \
--disks 4 \
--duration 120s \
--rounds 3 \
--cooldown 30 \
--concurrency 16 \
--out-dir target/hotpath-abba/linux-local
```
The script starts and stops RustFS for each ABBA leg. The data root is
throwaway and should not contain important data.
## Production-like cluster runner
Use external mode when RustFS lifecycle is managed by ansible, systemd, a
cluster scheduler, or a dedicated deployment harness. In this mode the ABBA
script does not start RustFS directly; it calls `--deploy-hook` before each leg
and then waits for `http://<endpoint><health-path>`.
The deploy hook receives:
| Environment variable | Value |
| --- | --- |
| `HOTPATH_ABBA_LEG` | `A1`, `B1`, `B2`, or `A2` |
| `HOTPATH_ABBA_PHASE` | `baseline` or `candidate` |
| `HOTPATH_ABBA_BINARY` | selected baseline or candidate binary path |
| `HOTPATH_ABBA_DRIVE_SYNC` | `true` or `false` |
Example ansible-shaped command:
```bash
scripts/run_hotpath_warp_abba.sh \
--baseline-bin /srv/rustfs-binaries/rustfs-baseline \
--candidate-bin /srv/rustfs-binaries/rustfs-candidate \
--endpoint rustfs-bench.example.internal:9000 \
--deploy-hook '
set -euo pipefail
cd /srv/rustfs-ansible
cp "${HOTPATH_ABBA_BINARY:?}" roles/rustfs/files/rustfs
export RUSTFS_DRIVE_SYNC_ENABLE="${HOTPATH_ABBA_DRIVE_SYNC:?}"
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags stop
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags config
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags binary-copy
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags start
' \
--duration 180s \
--rounds 5 \
--cooldown 45 \
--concurrency 32 \
--out-dir target/hotpath-abba/cluster-pr-XXXX
```
For formal evidence, prefer `--rounds 5` or higher when the cluster budget
allows it. The script enforces `--rounds >= 3`.
## CPU and memory evidence
ABBA warp output answers whether the candidate changed throughput or latency.
Collect host telemetry at the same time to explain why.
Recommended minimum:
```bash
mkdir -p target/hotpath-abba/cluster-pr-XXXX/telemetry
pidstat -durh 5 > target/hotpath-abba/cluster-pr-XXXX/telemetry/pidstat.txt &
PIDSTAT_PID=$!
mpstat 5 > target/hotpath-abba/cluster-pr-XXXX/telemetry/mpstat.txt &
MPSTAT_PID=$!
iostat -xz 5 > target/hotpath-abba/cluster-pr-XXXX/telemetry/iostat.txt &
IOSTAT_PID=$!
```
Stop the collectors after the ABBA script exits:
```bash
kill "$PIDSTAT_PID" "$MPSTAT_PID" "$IOSTAT_PID"
```
For deeper CPU attribution, run `perf record` around one representative
workload after the ABBA gate identifies a candidate regression or improvement:
```bash
perf record -F 99 -g -- sleep 180
perf report --stdio > target/hotpath-abba/cluster-pr-XXXX/telemetry/perf-report.txt
```
When using samply against an already-running RustFS service, attach through the
bounded helper instead of calling `samply record -p` directly:
```bash
scripts/run_samply_attach_window.sh \
--pid "$RUSTFS_PID" \
--duration-secs 180 \
--output target/hotpath-abba/cluster-pr-XXXX/telemetry/samply-A1-get-4mib.json.gz
```
Run one attach window per ABBA leg or focused verification cell. After every
window, confirm that the `.json.gz` profile and `.syms.json` sidecar are
non-empty, that no `samply` process is still attached to the RustFS PID, and
that any temporary `perf_event_paranoid` change has been restored before the
next cell starts.
For allocation profiling, build the candidate with:
```bash
cargo build --release -p rustfs --bins --features hotpath-alloc
```
Then run the same ABBA command with that binary. Compare allocation-heavy
function sections only within the same build mode. Do not compare
`hotpath-alloc` binaries directly with default release binaries for throughput
acceptance, because allocation instrumentation intentionally changes what is
measured.
For CPU hotpath sections emitted by hotpath, build with:
```bash
cargo build --release -p rustfs --bins --features hotpath-cpu
```
Use the CPU-enabled report to explain hotspots after the default or plain
`hotpath` ABBA gate shows a real effect.
## Output layout
The ABBA runner writes:
```text
<out-dir>/
manifest.env
abba_schedule.csv
candidate_gate.md
baseline_drift_gate.md
summary.md
<workload>/<sync>/<leg>/median_summary.csv
<workload>/<sync>/<leg>/baseline_compare.csv
```
Attach or link at least these files in the issue or PR:
- `summary.md`
- `candidate_gate.md`
- `baseline_drift_gate.md`
- `abba_schedule.csv`
- every `median_summary.csv` and `baseline_compare.csv` for a failed or
borderline workload
- host telemetry files used to explain CPU, memory, or disk saturation
## Interpretation
Use this decision table:
| Candidate gate | A2 drift gate | Interpretation |
| --- | --- | --- |
| PASS | PASS | Candidate is acceptable for the measured matrix. |
| WARN | PASS | Candidate has a small measurable signal; inspect telemetry and decide if it is expected. |
| FAIL | PASS | Candidate likely regressed the affected workload; investigate before merge. |
| FAIL | FAIL on the same workload | Environment drift is high; rerun on a quieter runner or increase duration and rounds. |
| PASS | FAIL | Candidate did not exceed the budget, but the rig was unstable; avoid using the numbers as proof of improvement. |
When `B1` and `B2` disagree, treat the result as inconclusive even if the gate
passes. Increase duration, rounds, cooldown, or runner isolation before drawing
a conclusion.
## AI execution checklist
When delegating the run to an AI agent or an automation runner, provide these
inputs explicitly:
- repository checkout and candidate branch or commit;
- baseline commit or binary path;
- candidate binary path;
- runner type: local Linux or external cluster;
- endpoint, access key, secret key source, and region;
- deploy hook path or exact command for cluster mode;
- output directory;
- required duration, rounds, cooldown, concurrency, and fail/warn budgets;
- where to upload artifacts after the run.
The AI agent should execute this sequence:
1. Confirm `uname -a`, RustFS commits, binary SHA256 sums, `warp --version`,
CPU model, memory size, disk layout, and whether the run is local or cluster.
2. Run `scripts/run_hotpath_warp_abba.sh --dry-run` with the final arguments.
3. Run the real ABBA command with `--rounds >= 3`.
4. For samply CPU attribution, use `scripts/run_samply_attach_window.sh` for
each bounded attach window and reject the cell if the profile is empty or a
stale `samply` process remains.
5. Preserve the full output directory without editing generated CSV files.
6. Read `summary.md`, `candidate_gate.md`, and `baseline_drift_gate.md`.
7. Summarize only measured facts: candidate deltas, baseline drift, CPU or
memory saturation, and any failed workloads.
8. Post the summary and artifact location to the tracking issue or PR.
Do not report a performance win or loss when the baseline drift gate failed on
the same workload and no rerun was collected.
@@ -1,146 +1,94 @@
# Internode gRPC Optimization — A/B Benchmark Runbook
# Internode gRPC A/B benchmark runbook
Reproducible procedure to collect **before/after** artifacts for each internode gRPC
optimization stage (grpc-optimization P0P3). Every stage is env-gated, so "before" and
"after" are the *same binary* with different env — no rebuild between runs.
**Use this when:** collecting before/after evidence for an env-gated internode gRPC transport stage (P0 transport tuning, P1 channel isolation, P2 msgpack-only codec, P3 prewarm/offline bypass) on a real multi-node cluster.
**Source of truth:** `scripts/run_internode_grpc_ab_bench.sh` (stage/phase driver), `crates/config/src/constants/internode.rs` (`ENV_INTERNODE_*` / `DEFAULT_INTERNODE_*`), `crates/io-metrics/src/internode_metrics.rs` (metric names), [internode-msgpack-json-convergence-runbook.md](internode-msgpack-json-convergence-runbook.md) (the P2 gate).
> Live runs need a multi-node cluster (Docker or ≥2 rustfs endpoints), a load tool
> (`warp` or `s3bench`), and a Prometheus scrape of `/metrics`. They are not runnable in a
> single-process sandbox. Capture artifacts on a real cluster.
Every stage is env-gated, so "before" and "after" run the *same binary* with different server env; no rebuild between runs. Live runs need a multi-node cluster (Docker compose or two or more endpoints), a load tool (`warp` or `s3bench`), and a metrics sink; they are not runnable in a single-process sandbox.
## One-click driver
## Prerequisites
`scripts/run_internode_grpc_ab_bench.sh --stage <p0|p1|p2|p3> --phase <before|after|request-only|canary|rollback> [-- <bench args>]`
wraps the env matrix below: it writes the stage/phase **server** env to
`<out-dir>/server-env.sh`, then runs the right underlying bench into
`target/bench/internode-transport/<stage>-<phase>/`.
| Requirement | Detail |
| --- | --- |
| RPC secret | Internode RPC fails closed: remote endpoints with default credentials and no `RUSTFS_RPC_SECRET` abort startup with `store init aborted: endpoints include remote nodes but ...` (`crates/ecstore/src/store/init.rs`). Set a non-default `RUSTFS_RPC_SECRET`, identical on every node. |
| systemd start timeout | `deploy/build/rustfs.service` is `Type=notify` and ships `TimeoutStartSec=120s`; READY fires only after quorum. If freshly purged disks need longer, raise it in a drop-in rather than lowering it. |
| Metrics export | RustFS has no Prometheus pull endpoint (`/admin/v3/metrics` is NDJSON, see `rustfs/src/admin/handlers/metrics.rs`); it pushes OTLP. Run an otel-collector (OTLP receiver → Prometheus exporter) and set `RUSTFS_OBS_ENDPOINT`, `RUSTFS_OBS_METRICS_EXPORT_ENABLED=true`, `RUSTFS_OBS_METER_INTERVAL=5`. For lock p99 also set `RUSTFS_OBJECT_LOCK_DIAG_ENABLE=true` (default off). |
| Server env | `RUSTFS_INTERNODE_*` are **server** env. For p0/p1/p2, source the emitted `server-env.sh` on every node and restart before the run; the driver cannot mutate a running server. |
## Driver
```bash
# P1 A/B (restart the cluster with each phase's server-env.sh between the two runs):
scripts/run_internode_grpc_ab_bench.sh --stage p1 --phase before -- --access-key AK --secret-key SK --metrics-url http://node1:9000/metrics
scripts/run_internode_grpc_ab_bench.sh --stage p1 --phase after -- --access-key AK --secret-key SK --metrics-url http://node1:9000/metrics
# P3 failover A/B (docker four-node):
scripts/run_internode_grpc_ab_bench.sh --stage p3 --phase after
# P2 rollout gates (env preview for request-only rehearsal, canary, and rollback):
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase request-only --dry-run
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase canary --dry-run
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase rollback --dry-run
scripts/run_internode_grpc_ab_bench.sh --stage <p0|p1|p2|p3> --phase <before|after|request-only|canary|rollback> [--dry-run] [-- <bench args>]
```
`RUSTFS_INTERNODE_*` are **server** env: for the load-driven stages (p0/p1/p2) source the emitted `server-env.sh` on every node and restart rustfs *before* the run — the driver cannot mutate an already-running server. The P2 `canary` phase is the exception: source it only on the selected canary node after the release-window counters and fleet-support checks pass, and keep the rest of the fleet on `before` or `request-only` while observing fallback/decode-error counters. Use `--dry-run` to preview the env and command.
The driver writes the stage/phase server env to `<out-dir>/server-env.sh` and runs the underlying bench into `target/bench/internode-transport/<stage>-<phase>/`. `p0/p1/p3` accept `before|after`; `p2` also accepts `request-only|canary|rollback`. `--dry-run` prints the env and command only.
## Harness
```bash
# P1 A/B: restart the cluster with each phase's server-env.sh between the two runs
scripts/run_internode_grpc_ab_bench.sh --stage p1 --phase before -- --access-key AK --secret-key SK --metrics-url http://node1:9000/metrics
scripts/run_internode_grpc_ab_bench.sh --stage p1 --phase after -- --access-key AK --secret-key SK --metrics-url http://node1:9000/metrics
# P3 failover A/B (docker four-node)
scripts/run_internode_grpc_ab_bench.sh --stage p3 --phase after
# P2 env previews
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase canary --dry-run
```
- Throughput / latency: `scripts/run_internode_transport_baseline.sh` (drives
`run_object_batch_bench.sh`; writes `target/bench/internode-transport-<ts>/`). Pass
`--metrics-url <prometheus>` to also capture internode metric deltas.
- Failover / offline: `scripts/run_four_node_cluster_failover_bench.sh` (spins up a 4-node
compose cluster, kills `FAILOVER_NODE`, benchmarks; writes
`target/bench/four-node-failover-<ts>/`).
The one-click driver writes each run to `target/bench/internode-transport/<stage>-<phase>/`
(e.g. `p0-before/`, `p0-after/`, `p1-before/`, `p1-after/`, `p3-before/`, `p3-after/`), each
containing the emitted `server-env.sh` plus the underlying bench artifacts. `target/` is
gitignored — attach the paired directories to the PR / issue.
## Metrics to capture (Prometheus)
| Metric | Stage signal |
|---|---|
| `rustfs_system_network_internode_operation_duration_ms{operation,backend}` | control-plane RTT (P0), lock/bulk latency |
| `rustfs_system_network_internode_operation_payload_bytes` | payload size distribution (P0/P1 sizing) |
| `rustfs_system_network_internode_operation_large_payloads_total` | large unary RPCs sharing the channel (P1 target) |
| `rustfs_system_network_internode_dial_avg_time_nanos`, `..._dial_errors_total` | connect cost (P3 prewarm) |
| `rustfs_system_network_internode_msgpack_json_decode_total{direction,message,codec}` | must be **>0** for each expected P2 series before zero fallback/error readings are meaningful |
| `rustfs_system_network_internode_msgpack_json_fallback_total{direction,message}` | must be **0** before enabling both msgpack-only gates (P2) |
| `rustfs_system_network_internode_msgpack_json_decode_error_total{direction,message,codec}` | must be **0** before enabling both msgpack-only gates (P2) |
| `rustfs_cluster_servers_offline_total` | offline detection correctness (P3 bypass) |
| lock p99 (lock metrics) | P1 head-of-line-blocking win |
Underlying benches: `scripts/run_internode_transport_baseline.sh` (throughput/latency via `scripts/run_object_batch_bench.sh`; `--metrics-url <prometheus>` also captures internode metric deltas) and `scripts/run_four_node_cluster_failover_bench.sh` (4-node compose cluster, kills `FAILOVER_NODE`). `target/` is gitignored; attach the paired directories to the PR or issue.
## Per-stage env matrix
Run **before** with the stage's env at its baseline column, **after** with the enabled
column, everything else at defaults. Roll a restart between runs.
Run **before** at the baseline column and **after** at the enabled column, everything else at defaults, with a restart between. Defaults are the `DEFAULT_INTERNODE_*` constants in `crates/config/src/constants/internode.rs`.
| Stage | Env | before (baseline) | after (enabled) |
|---|---|---|---|
| P0 nodelay | `RUSTFS_INTERNODE_RPC_TCP_NODELAY` | `false` | `true` (default) |
| P0 stream window | `RUSTFS_INTERNODE_RPC_HTTP2_STREAM_WINDOW_SIZE` | `0` | unset (1 MiB) |
| P0 conn window | `RUSTFS_INTERNODE_RPC_HTTP2_CONN_WINDOW_SIZE` | `0` | unset (2 MiB) |
| P0 msg limit | `RUSTFS_INTERNODE_RPC_MAX_MESSAGE_SIZE` | `4194304` | unset (100 MiB) |
| Stage | Env | before | after |
| --- | --- | --- | --- |
| P0 nodelay | `RUSTFS_INTERNODE_RPC_TCP_NODELAY` | `false` | unset (`DEFAULT_INTERNODE_RPC_TCP_NODELAY`) |
| P0 stream window | `RUSTFS_INTERNODE_RPC_HTTP2_STREAM_WINDOW_SIZE` | `0` | unset (`DEFAULT_INTERNODE_RPC_HTTP2_STREAM_WINDOW_SIZE`) |
| P0 conn window | `RUSTFS_INTERNODE_RPC_HTTP2_CONN_WINDOW_SIZE` | `0` | unset (`DEFAULT_INTERNODE_RPC_HTTP2_CONN_WINDOW_SIZE`) |
| P0 msg limit | `RUSTFS_INTERNODE_RPC_MAX_MESSAGE_SIZE` | `4194304` (tonic default) | unset (RustFS default, see `rustfs/src/server/http.rs`) |
| P1 isolation | `RUSTFS_INTERNODE_CHANNEL_ISOLATION` | `false` (default) | `true` |
| P1 bulk pool | `RUSTFS_INTERNODE_BULK_CHANNELS` | `1` | `2``4` |
| P2 msgpack-only request | `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY` | `false` (default) | `true` (only after fallback counter = 0 across a window) |
| P2 fleet confirmation | `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED` | `false` (default) | `true` (only after mixed-version, fallback-zero, soak, and rollback gates pass) |
| P1 bulk pool | `RUSTFS_INTERNODE_BULK_CHANNELS` | `1` | unset (`DEFAULT_INTERNODE_BULK_CHANNELS`) or higher |
| P2 msgpack-only | `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY` + `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED` | both `false` (default) | per the convergence runbook |
| P3 prewarm | `RUSTFS_INTERNODE_PREWARM` | `false` (default) | `true` |
| P3 offline bypass | `RUSTFS_INTERNODE_OFFLINE_BYPASS` | `false` (default) | `true` |
| P3 reprobe / threshold | `RUSTFS_INTERNODE_OFFLINE_REPROBE_SECS` / `RUSTFS_INTERNODE_OFFLINE_FAILURE_THRESHOLD` | defaults | `5` / `3` |
| P3 reprobe / threshold | `RUSTFS_INTERNODE_OFFLINE_REPROBE_SECS` / `RUSTFS_INTERNODE_OFFLINE_FAILURE_THRESHOLD` | defaults | defaults (`DEFAULT_INTERNODE_OFFLINE_REPROBE_SECS`, `DEFAULT_INTERNODE_OFFLINE_FAILURE_THRESHOLD`) |
## Procedure per stage
## Metrics to capture
1. **Baseline**: start the cluster with the stage's env at the *before* column. Run the
relevant bench; save to `.../baseline/` (or `.../after-P{n-1}/` when chaining stages).
2. **After**: restart with the *after* column; re-run the identical bench; save to
`.../after-P{n}/`.
3. Diff the object-bench summaries and the metric deltas.
All names are defined in `crates/io-metrics/src/internode_metrics.rs`.
- **P0** — `run_internode_transport_baseline.sh` with `--sizes 4KiB,1MiB,16MiB,128MiB` and
`--concurrencies 1,16,64`. Expect: small-RPC `duration_ms` (DiskInfo/Ping) down (nodelay),
large-metadata (ReadMultiple/BatchReadVersion) throughput up (windows). Functional: a
`>4 MiB` multi-version `xl.meta` no longer fails `out_of_range`.
- **P1** — mixed workload (large `ReadAll` + high-frequency `Refresh`). Acceptance gate from
the design doc: **lock p99 down ≥ 20%** with `RUSTFS_INTERNODE_CHANNEL_ISOLATION=true`.
- **P2** — observe `msgpack_json_decode_total`, `msgpack_json_fallback_total`, and `msgpack_json_decode_error_total` across a release window; every expected `codec="msgpack"` decode series must be **>0**, while fallback and decode-error series must stay **0** before flipping both `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true` and `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED=true` (see the msgpack convergence runbook). Codec allocation via a `dhat`/`heaptrack` micro-run.
- **P3** — cold-start: first cross-node op latency should drop ~one connect RTT with prewarm.
Failover: `run_four_node_cluster_failover_bench.sh`, kill a node with
`RUSTFS_INTERNODE_OFFLINE_BYPASS=true`; expect faster failover and a correct
`rustfs_cluster_servers_offline_total` (1 while the node is down, back to 0 after recovery).
| Metric | Stage signal |
| --- | --- |
| `rustfs_system_network_internode_operation_duration_ms{operation,backend}` | control-plane RTT (P0), lock/bulk latency (P1), first-op latency (P3) |
| `rustfs_system_network_internode_operation_payload_bytes` | payload size distribution (P0/P1 sizing) |
| `rustfs_system_network_internode_operation_large_payloads_total` | large unary RPCs sharing a channel (P1 target) |
| `rustfs_system_network_internode_dial_avg_time_nanos`, `rustfs_system_network_internode_dial_errors_total` | connect cost and failures (P3) |
| `rustfs_system_network_internode_msgpack_json_decode_total{direction,message,codec}`, `..._msgpack_json_fallback_total`, `..._msgpack_json_decode_error_total` | P2 gate inputs |
| `rustfs_cluster_servers_offline_total` | offline detection correctness (P3 bypass) |
| lock p99 (lock metrics, needs `RUSTFS_OBJECT_LOCK_DIAG_ENABLE=true`) | P1 head-of-line-blocking win |
## Acceptance gates & artifact layout
## Acceptance gates
Each stage's paired run must satisfy an explicit gate before its numbers are accepted. Record
the gate verdict (pass/fail + measured delta) in the paired directory's `summary` and attach it.
Record the verdict (pass/fail plus measured delta) in the paired directory's summary.
| Stage | Bench | Acceptance gate | Primary metric(s) |
|---|---|---|---|
| **P0** | `run_internode_transport_baseline.sh` | small-RPC `duration_ms` (DiskInfo/Ping) **down**; large-metadata (ReadMultiple/BatchReadVersion) throughput **up**; a `>4 MiB` multi-version `xl.meta` no longer fails `out_of_range` (functional). | `..._operation_duration_ms{operation}`, `..._operation_payload_bytes`, object-bench throughput |
| **P1** | `run_internode_transport_baseline.sh` (mixed: large `ReadAll` + high-frequency `Refresh`) | **lock p99 down ≥ 20%** with `RUSTFS_INTERNODE_CHANNEL_ISOLATION=true` vs baseline. | lock p99 (lock metrics), `..._operation_large_payloads_total` |
| **P3 cold-start** | `run_internode_transport_baseline.sh` (fresh cluster, first cross-node op) | first cross-node op latency **drops ~one connect RTT** with `RUSTFS_INTERNODE_PREWARM=true`. | `..._dial_avg_time_nanos`, first-op `..._operation_duration_ms` |
| **P3 offline** | dedicated *sustained-offline + survivor cross-node access* experiment (the standard four-node failover bench is **not** sensitive to the bypass — quorum holds, `recovery_seconds=0`) | with `RUSTFS_INTERNODE_OFFLINE_BYPASS=true`, survivor cross-node op latency to the downed peer **fast-fails** instead of hanging the dial timeout; `rustfs_cluster_servers_offline_total` = **1** while down, back to **0** after recovery. | `rustfs_cluster_servers_offline_total`, survivor cross-node `..._operation_duration_ms`, `..._dial_errors_total` |
| Stage | Bench | Gate | Primary metrics |
| --- | --- | --- | --- |
| P0 | `run_internode_transport_baseline.sh --sizes 4KiB,1MiB,16MiB,128MiB --concurrencies 1,16,64` | small-RPC `duration_ms` (DiskInfo/Ping) down; large-metadata (ReadMultiple/BatchReadVersion) throughput up; a `>4 MiB` multi-version `xl.meta` no longer fails `out_of_range` | `..._operation_duration_ms{operation}`, `..._operation_payload_bytes`, object-bench throughput |
| P1 | `run_internode_transport_baseline.sh` with a mixed workload (large `ReadAll` plus high-frequency `Refresh`) | lock p99 down by at least 20% with `RUSTFS_INTERNODE_CHANNEL_ISOLATION=true` | lock p99, `..._operation_large_payloads_total` |
| P2 | none (not a throughput gate) | operational gate defined once in [internode-msgpack-json-convergence-runbook.md](internode-msgpack-json-convergence-runbook.md): expected `codec="msgpack"` decode series non-zero, fallback and decode-error series zero across a full observation window before both flags are enabled. Optionally a `dhat`/`heaptrack` micro-run for codec allocation. | the three `msgpack_json_*` counters |
| P3 cold-start | `run_internode_transport_baseline.sh` on a fresh cluster, first cross-node op | first cross-node op latency drops by about one connect RTT with prewarm | `..._dial_avg_time_nanos`, first-op `..._operation_duration_ms` |
| P3 offline | sustained-offline plus survivor cross-node access (below) | with bypass on, survivor ops to the downed peer fast-fail instead of hanging for the dial timeout; `rustfs_cluster_servers_offline_total` is `1` while down and `0` after recovery | `rustfs_cluster_servers_offline_total`, survivor `..._operation_duration_ms`, `..._dial_errors_total` |
> **P2 is not a throughput gate.** Its acceptance is operational: every expected `msgpack_json_decode_total{codec="msgpack"}` series must have non-zero traffic, while `msgpack_json_fallback_total` and `msgpack_json_decode_error_total` must read **0** across a full release window before both msgpack-only env gates are enabled (see the msgpack convergence runbook). Do not benchmark P2 as before/after throughput.
P3 offline method: the standard four-node failover bench is not sensitive to the bypass (quorum holds, `recovery_seconds=0`). Instead: all nodes up → stop one node and keep it down → warm up until offline detection trips → drive warp against the survivors only (`--host` excludes the dead node) → compare survivor op p99 and the offline gauge with `RUSTFS_INTERNODE_OFFLINE_BYPASS` off and on.
Artifact layout per stage (attach both halves + the diff):
Artifact layout:
```
```text
target/bench/internode-transport/
p0-before/ p0-after/ # server-env.sh + object-bench summaries + metric deltas
p1-before/ p1-after/ # + lock p99 delta (the ≥20% gate)
p2-before/ p2-request-only/ p2-canary/ p2-after/ p2-rollback/ # env gate artifacts + fallback/decode-error observations
p3-before/ p3-after/ # cold-start + sustained-offline experiment + offline gauge trace
p0-before/ p0-after/ # server-env.sh + object-bench summaries + metric deltas
p1-before/ p1-after/ # + lock p99 delta
p2-before/ p2-request-only/ p2-canary/ p2-after/ p2-rollback/ # env previews + counter observations
p3-before/ p3-after/ # cold-start + sustained-offline + offline gauge trace
```
## Bench-host prerequisites (ansible bare-metal)
Captured while running the first real A/B on a 4-node ansible cluster; needed before any live run:
- **RPC secret is mandatory on current `main`.** Internode RPC fails closed: default creds
(`RUSTFS_SECRET_KEY=rustfsadmin`) with no `RUSTFS_RPC_SECRET` → the node aborts startup immediately
with `store init aborted: endpoints include remote nodes but RPC authentication secret is not
configured` (builds before the preflight instead showed `No valid auth token` and never reached
`storage_quorum`). Set a **non-default** `RUSTFS_RPC_SECRET`, identical on every node.
- **systemd start timeout.** The install unit is `Type=notify`; READY only fires after quorum, which on
freshly-purged disks exceeds a 30 s `TimeoutStartSec` → crash loop. Use a drop-in `TimeoutStartSec=infinity`.
- **Server-side metrics need OTLP.** RustFS has no Prometheus pull endpoint (`/admin/v3/metrics` is NDJSON,
not exposition); it only pushes via OTLP. To capture lock/offline/internode metrics, run an
otel-collector (OTLP receiver → Prometheus exporter) and set `RUSTFS_OBS_ENDPOINT`,
`RUSTFS_OBS_METRICS_EXPORT_ENABLED=true`, `RUSTFS_OBS_METER_INTERVAL=5`. For lock p99 also set
`RUSTFS_OBJECT_LOCK_DIAG_ENABLE=true` (default off).
- **P3 offline method.** Standard failover is quorum-insensitive to the bypass. Use a *sustained-offline +
survivor cross-node access* run instead: all nodes up → stop one node (sustained) → warm up to trip
offline detection → drive warp on the survivors only (`--host` excludes the dead node) → compare
survivor op p99 and `rustfs_cluster_servers_offline_total` for `RUSTFS_INTERNODE_OFFLINE_BYPASS` off/on.
## Rollback
Every stage rolls back by setting the env back to its baseline column and restarting. P2 rolls back by unsetting either msgpack-only gate, or by setting both `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=false` and `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED=false`; no wire format is broken because the JSON fields, `_bin` fields, and proto field numbers remain additive. Do not remove or reuse JSON proto fields as part of this benchmark stage.
Every stage rolls back by restoring the baseline column and restarting. P2 rollback and its wire-format guarantees are in the [convergence runbook's rollback matrix](internode-msgpack-json-convergence-runbook.md#rollback-matrix); do not remove or reuse JSON proto fields as part of any benchmark stage.
@@ -1,75 +1,37 @@
# Internode msgpack/JSON Convergence Runbook
# Internode msgpack/JSON convergence runbook
Operational runbook for retiring the redundant JSON compatibility fields on internode
gRPC metadata RPCs (grpc-optimization **P2-1**). This is a **cross-version** change: it
proceeds strictly by observation-gated stages, never in one step.
**Use this when:** operating or changing the staged retirement of the JSON compatibility fields on internode gRPC metadata RPCs, deciding whether a fleet may enable `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY`, or adding a new `*_bin` proto field.
**Source of truth:** `crates/config/src/constants/internode.rs` (`ENV_INTERNODE_RPC_MSGPACK_ONLY`, `ENV_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED` and their compile-time default-off asserts), `crates/protos/src/node.proto` (`*_bin` fields), `decode_msgpack_or_json` in `rustfs/src/storage/rpc/node_service/disk.rs` (server side) and `crates/ecstore/src/cluster/rpc/remote_disk.rs` (client side), `crates/io-metrics/src/internode_metrics.rs` (counters).
This is a cross-version change and proceeds by observation-gated stages, never in one step.
## Background
Internode RPCs dual-encode each metadata value as **both**:
Internode RPCs dual-encode each metadata value as a msgpack binary field (`*_bin`, for example `file_info_bin`) and a JSON compatibility string (for example `file_info`). Decoders prefer `_bin` and fall back to JSON only when `_bin` is empty. The dual-write costs bandwidth and CPU; before the JSON fields can be dropped, the fallback branch must be proven unused across the fleet, otherwise a mixed-version rolling upgrade could read an emptied field.
- a msgpack binary field (`*_bin`, e.g. `file_info_bin`), and
- a JSON compatibility string (e.g. `file_info`).
## Observation counters
Decoders prefer the `_bin` payload and fall back to the JSON string only when `_bin` is
empty (`decode_msgpack_or_json`). The dual-write costs bandwidth and CPU. Before the JSON
fields can be dropped, the fallback branch must be proven **unused** in production —
otherwise a rolling upgrade with mixed node versions could read an emptied field.
| Counter | Labels | Increments when |
| --- | --- | --- |
| `rustfs_system_network_internode_msgpack_json_decode_total` | `direction`, `message`, `codec` | a msgpack or JSON decode succeeded |
| `rustfs_system_network_internode_msgpack_json_fallback_total` | `direction`, `message` | a decode fell back to JSON because `_bin` was empty |
| `rustfs_system_network_internode_msgpack_json_decode_error_total` | `direction`, `message`, `codec` | either codec failed to decode (`codec` = the failed codec) |
## The observation metric (already shipped)
`direction="request"` is a server decoding a peer's request (`node_service/disk.rs`); `direction="response"` is a client decoding a peer's response (`cluster/rpc/remote_disk.rs`), including the list-level `ReadMultiple` / `BatchReadVersion` fallbacks. `message` is the value name (`FileInfo`, `RawFileInfo`, `ReadMultipleResp`, ...).
```
rustfs_system_network_internode_msgpack_json_fallback_total{direction, message}
rustfs_system_network_internode_msgpack_json_decode_total{direction, message, codec}
rustfs_system_network_internode_msgpack_json_decode_error_total{direction, message, codec}
```
The decode counter increments after a successful msgpack or JSON compatibility decode. The fallback counter increments whenever a decode falls back to the JSON field because the msgpack payload was absent. The decode-error counter increments whenever either codec fails to decode.
- `direction="request"` — a server decoding a peer's request (`node_service/disk.rs`).
- `direction="response"` — a client decoding a peer's response (`cluster/rpc/remote_disk.rs`),
including the list-level `ReadMultiple` / `BatchReadVersion` fallbacks.
- `message` — the value name, e.g. `FileInfo`, `RawFileInfo`, `ReadMultipleResp`.
- `codec` — the failed codec for decode errors: `msgpack` for corrupt non-empty `_bin`, or `json` for corrupt legacy fallback JSON.
## Stage 0 — Observe (current stage)
Ship the current release (which contains these counters) and let it run for **at least one
full release window** across the whole fleet. The fallback and decode-error counters must stay at **zero**.
First confirm each expected message/direction has real traffic in the observation window:
Gate query template, run over the whole observation window (adjust `[30d]`):
```promql
sum by (direction, message, codec) (
increase(rustfs_system_network_internode_msgpack_json_decode_total[30d])
)
sum by (direction, message, codec) (increase(<counter>[30d]))
```
For every convergence-ready message/direction below, the `codec="msgpack"` series must be non-zero before a zero fallback result is meaningful. Missing series, zero traffic, counter reset, or scrape gaps make the gate inconclusive rather than passed.
| Counter | Required reading | Meaning of a violation |
| --- | --- | --- |
| `..._decode_total` | every convergence-ready `{direction, message}` has a non-zero `codec="msgpack"` series | no traffic, a counter reset, or scrape gaps make the gate inconclusive, not passed |
| `..._fallback_total` | `0` for every series | some peer still sends an empty `_bin` (old node, or a sender that does not fill `_bin`); investigate the labels |
| `..._decode_error_total` | `0` for every series | `codec="msgpack"`: corrupt or incompatible `_bin` bytes; `codec="json"`: corrupt legacy fallback. Either blocks convergence and rollback confidence |
Confirm zero across the observation window (adjust `[30d]` to the window length):
```promql
sum by (direction, message) (
increase(rustfs_system_network_internode_msgpack_json_fallback_total[30d])
)
```
Every series must be `0`. A non-zero value means some peer is still emitting an empty
`_bin` (an old node, or a message whose sender does not fill `_bin`) — investigate the
`{direction, message}` label before proceeding.
Decode errors must also stay at zero across the observation window:
```promql
sum by (direction, message, codec) (
increase(rustfs_system_network_internode_msgpack_json_decode_error_total[30d])
)
```
A non-zero `codec="msgpack"` series means a peer sent corrupt or incompatible `_bin` bytes; it must fail closed and block convergence. A non-zero `codec="json"` series means the legacy fallback field was corrupt or semantically incompatible; it also blocks convergence and rollback confidence.
Standing alert (keep enabled through all stages):
Standing alerts (keep enabled through every stage):
```yaml
- alert: InternodeMsgpackJsonFallback
@@ -90,96 +52,70 @@ Standing alert (keep enabled through all stages):
## Field → peer-decoder audit
The send-side change (Stage 1) may only empty a JSON field whose **peer decodes `_bin`
first**. The following mapping is verified against the current code.
Stage 1 may only empty a JSON field whose peer decodes `_bin` first. Any `*_bin` field not listed here must be mapped to a confirmed `_bin`-first peer decoder before it joins the convergence set.
### Convergence-ready (peer decodes `_bin` first)
Convergence-ready:
| Direction | Message / field | Peer decoder |
|---|---|---|
| request | `WriteMetadata.file_info` | `FileInfo` (node_service/disk.rs) |
| --- | --- | --- |
| request | `WriteMetadata.file_info` | `FileInfo` (`node_service/disk.rs`) |
| request | `UpdateMetadata.file_info` | `FileInfo` |
| request | `UpdateMetadata.opts` | `UpdateMetadataOpts` |
| request | `RenameData.file_info` | `FileInfo` |
| request | `ReadMultiple.read_multiple_req` | `ReadMultipleReq` |
| request | `BatchReadVersion.batch_read_version_req` | `BatchReadVersionReq` |
| request | `Read*.opts` | `ReadOptions` |
| response | `ReadVersion.file_info` | `FileInfo` (cluster/rpc/remote_disk.rs) |
| response | `ReadVersion.file_info` | `FileInfo` (`cluster/rpc/remote_disk.rs`) |
| response | `ReadXL.raw_file_info` | `RawFileInfo` |
| response | `RenameData.rename_data_resp` | `RenameDataResp` |
| response | `ReadMultiple` resp list | per-item + list fallback |
| response | `BatchReadVersion` resp list | per-item + list fallback |
| response | `ReadMultiple` response list | per-item plus list fallback |
| response | `BatchReadVersion` response list | per-item plus list fallback |
### `_bin` support added THIS release — converge after their own window
The `DeleteVersion`/`DeleteVersions` protos had **no `_bin` fields**. They gained additive
`*_bin` fields plus bin-first server decoders in this release, and the client now dual-writes
them. They are **kept out** of the msgpack-only set (always dual-write) until their own
fallback counter has read zero across a window with the new decoders fully deployed.
Delete messages, kept on dual-write until their own window reads zero with `_bin`-first decoders deployed fleet-wide (the client always dual-writes these regardless of the flags):
| Direction | Message / field | Status |
|---|---|---|
| request | `DeleteVersion.file_info` (`FileInfo`) | `_bin` added; dual-write; converge after window |
| request | `DeleteVersion.opts` (`DeleteOptions`) | `_bin` added; dual-write; converge after window |
| request | `DeleteVersions.versions` (`FileInfoVersions`) | `_bin` added; dual-write; converge after window |
| request | `DeleteVersions.opts` (`DeleteOptions`) | `_bin` added; dual-write; converge after window |
| --- | --- | --- |
| request | `DeleteVersion.file_info` (`FileInfo`) | `_bin` present; dual-write; converge after its own window |
| request | `DeleteVersion.opts` (`DeleteOptions`) | same |
| request | `DeleteVersions.versions` (`FileInfoVersions`) | same |
| request | `DeleteVersions.opts` (`DeleteOptions`) | same |
> Any `*_bin` proto field not in the tables above must be mapped to a confirmed `_bin`-first
> peer decoder before it is added to the convergence set.
### Still JSON-only (no `_bin` field)
Still JSON-only:
| Direction | Message / field | Note |
|---|---|---|
| response | `DeleteVersion.raw_file_info` | proto has no `_bin`; needs an additive proto field before it can converge. |
| --- | --- | --- |
| response | `DeleteVersion.raw_file_info` | proto has no `_bin`; needs an additive proto field before it can converge |
## Stage 1Stop writing JSON (env-gated, after Stage 0 reads zero)
## Stage 0Observe
The send-side lever is **implemented** and requires two default-off env flags:
Run the fleet with both flags at their default (`false`) for at least one full observation window and evaluate the three gate rows above. All three must hold before Stage 1.
- `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true`
- `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED=true`
## Stage 1 — Stop writing JSON (env-gated)
The first flag only requests msgpack-only. The second flag is the explicit proof gate that the fleet has passed the release-window fallback, capability, and rollback checks. If only `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true` is set, RustFS keeps dual-writing JSON compatibility fields so old JSON-only peers remain compatible. When both flags are enabled, the convergence-ready fields above send only `_bin` and leave the JSON string empty; the `_bin` payload is always sent and decoders keep the JSON read fallback unchanged. The delete fields are excluded (dual-write) per the section above.
Flag semantics are defined on `ENV_INTERNODE_RPC_MSGPACK_ONLY` and `ENV_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED` in `crates/config/src/constants/internode.rs`; both default off and only their conjunction stops the JSON write for convergence-ready fields. Decoders keep the JSON fallback in every state.
Only enable it **after** Stage 0 has read zero for a full window across the fleet:
1. Rehearse `request-only`: `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true`, `..._FLEET_CONFIRMED=false`. Behaviour is unchanged (still dual-write) and safe with old peers; any fallback or decode-error increment still blocks.
2. Canary: set both flags `true` on one node, restart that node only, and watch the fallback and decode-error counters for a soak period with real internode traffic.
3. Fleet: if the counters stay zero, enable both flags fleet-wide with a rolling restart.
4. Rollback: set either flag to `false` (or unset it) and restart; no wire format changed, so rollback is immediate.
1. Ship with the flag **off** (no behavior change).
2. Rehearse the `request-only` gate with `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true` and `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED=false`. This must keep writing JSON compatibility fields and is safe with unsupported/old peers; any fallback or decode-error increment still blocks convergence.
3. Enable it on one canary node with both `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true` and `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED=true`, then restart that node only and watch the fallback and decode-error counters for a soak period while real internode traffic is present. If they stay zero, enable fleet-wide.
4. **Rollback:** set either `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=false` or `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED=false` (or unset either flag) and restart. No
wire-format was broken in this stage, so rollback is immediate and safe.
The benchmark driver pins these operator states as dry-run phases (`before`, `request-only`, `canary`, `after`, `rollback`); see [internode-grpc-benchmark-runbook.md](internode-grpc-benchmark-runbook.md).
The benchmark helper pins these operator states as dry-run phases:
## Stage 2 — Remove the proto JSON fields
```bash
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase before --dry-run
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase request-only --dry-run
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase canary --dry-run
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase after --dry-run
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase rollback --dry-run
```
Only after Stage 1 has been stable fleet-wide for a full window with the counters still zero.
## Stage 2 — Remove the proto JSON fields (next release, N+1)
Only after Stage 1 has been stable with the flag on for a full window and the counter is
still zero.
1. Mark the retired text fields `reserved` in `crates/protos/src/node.proto` (never reuse
the field numbers) and delete the JSON read-fallback branches; codec becomes msgpack-only.
2. This is a hard wire-format change — it requires the mixed-version upgrade rehearsal
(four-node scripts) to pass, and it cannot be rolled back by env alone.
1. Mark the retired text fields `reserved` in `crates/protos/src/node.proto` (never reuse field numbers) and delete the JSON read-fallback branches.
2. This is a hard wire-format change: it requires the mixed-version upgrade rehearsal (four-node scripts) to pass and cannot be rolled back by env alone.
## Rollback matrix
| Stage | Wire-format broken? | Rollback |
|---|---|---|
| 0 Observe | no | n/a (metric only) |
| 1 msgpack-only send | no | unset either msgpack-only env flag + restart |
| 2 remove fields | yes | redeploy prior release; field numbers stay `reserved` |
| Stage | Wire format broken? | Rollback |
| --- | --- | --- |
| 0 Observe | no | n/a (metrics only) |
| 1 msgpack-only send | no | unset either flag and restart |
| 2 remove fields | yes | redeploy the prior release; field numbers stay `reserved` |
## Related
- Codec + counter implementation: commit `feat(internode): P2 msgpack/JSON codec observability + encode buffer presizing`.
- Decoders: `decode_msgpack_or_json` in `crates/ecstore/src/cluster/rpc/remote_disk.rs` (client)
and `rustfs/src/storage/rpc/node_service/disk.rs` (server).
- Transport and codec observability landed in `feat(internode): optimize gRPC transport (#4337)`.
@@ -1,305 +0,0 @@
# Keycloak OIDC Integration Runbook
This runbook describes how to connect the RustFS Console to Keycloak by using OpenID Connect Authorization Code Flow. The examples use the default RustFS provider id, `default`.
## 1. Integration Model
RustFS supports standard OpenID Connect for Console login:
- RustFS sends an authorization-code request with PKCE S256.
- Keycloak redirects back with `code` and `state`.
- RustFS exchanges the code at the token endpoint.
- RustFS requires an `id_token` and verifies signature, issuer, audience, expiry, and nonce.
- RustFS maps ID token claim values to local RustFS IAM policies.
RustFS does not call Keycloak Authorization Services for object or admin authorization. Authorization is handled by RustFS policies after claims are mapped.
## 2. Example Values
Replace these values for your environment:
| Value | Example | Notes |
| --- | --- | --- |
| Keycloak base URL | `https://keycloak.example.com` | Public Keycloak URL. |
| Realm | `rustfs` | Keycloak realm name. |
| Keycloak issuer | `https://keycloak.example.com/realms/rustfs` | RustFS `config_url`. |
| Discovery URL | `https://keycloak.example.com/realms/rustfs/.well-known/openid-configuration` | Used to validate metadata. |
| Public RustFS browser origin | `https://rustfs.example.com` | The scheme and authority users open in the browser. |
| Provider id | `default` | This runbook uses the default provider. |
| RustFS callback URL | `https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default` | Register this exact URL in Keycloak. |
| Keycloak client id | `rustfs-console` | OIDC client used by RustFS. |
| Keycloak client secret | `<KEYCLOAK_CLIENT_SECRET>` | Confidential client secret. |
| RustFS scopes | `openid,profile,email` | Add custom scopes if they emit authorization claims. |
| RustFS groups claim | `groups` | Recommended flat array claim. |
| RustFS roles claim | `roles` | Optional flat array claim. |
## 3. Keycloak Configuration
### 3.1 Create or Select the Realm
1. Open the Keycloak Admin Console.
2. Create or select the `rustfs` realm.
3. Verify discovery:
```bash
curl -fsS "https://keycloak.example.com/realms/rustfs/.well-known/openid-configuration" \
| jq '.issuer,.authorization_endpoint,.token_endpoint,.jwks_uri'
```
The `issuer` should be:
```text
https://keycloak.example.com/realms/rustfs
```
### 3.2 Create the RustFS Client
In the Keycloak Admin Console:
1. Open `Clients` and create a client.
2. Set `Client type` to `OpenID Connect`.
3. Set `Client ID` to `rustfs-console`.
4. Enable `Client authentication`.
5. Enable `Standard flow`.
6. Disable unused flows such as `Implicit flow`, `Direct access grants`, and `Service accounts roles`.
7. Set `Valid redirect URIs` to:
```text
https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default
```
8. Set `Web origins` to:
```text
https://rustfs.example.com
```
9. Set `Proof Key for Code Exchange Code Challenge Method` to `S256`.
10. Save and copy the client secret from `Credentials`.
RustFS submits the client secret in the token request body. Do not use a client policy that disables `client_secret_post`.
### 3.3 Map Groups or Roles to RustFS Policies
RustFS policy names are the final authorization source. Common built-in policies are:
| Policy | Purpose |
| --- | --- |
| `consoleAdmin` | Full Console, admin, KMS, and S3 access. |
| `readwrite` | S3 read/write access. |
| `readonly` | S3 read-only access. |
| `writeonly` | S3 write-only access. |
| `diagnostics` | Diagnostic admin access. |
Recommended production setup:
1. Create Keycloak groups such as `consoleAdmin` and `readonly`.
2. Add users to the groups.
3. Add a group membership mapper for the `rustfs-console` client.
4. Emit a flat top-level ID token claim named `groups`.
5. Keep group values equal to RustFS policy names.
### 3.4 Group Claim Mapper
Create a `Group Membership` mapper in the dedicated client scope:
| Mapper field | Value |
| --- | --- |
| Name | `rustfs-groups` |
| Token Claim Name | `groups` |
| Full group path | `Off` |
| Add to ID token | `On` |
| Add to access token | `On` |
| Add to userinfo | `On` |
| Multivalued | `On` |
Keep `Full group path` disabled. RustFS policy names cannot contain `/`, so `/consoleAdmin` will not map to the `consoleAdmin` policy.
### 3.5 Optional Role Claim Mapper
If the deployment uses Keycloak roles:
1. Assign realm or client roles such as `consoleAdmin`.
2. Add a `User Realm Role` or `User Client Role` mapper.
3. Emit a flat top-level claim named `roles`.
4. Set `RUSTFS_IDENTITY_OPENID_ROLES_CLAIM=roles`.
RustFS does not parse Keycloak's default nested `realm_access.roles` claim. Emit a flat `roles` array when role mapping is required.
## 4. RustFS Configuration
### 4.1 Environment Variables
Configure the provider and the public browser origin:
```bash
export RUSTFS_BROWSER_REDIRECT_URL="https://rustfs.example.com"
export RUSTFS_IDENTITY_OPENID_ENABLE=on
export RUSTFS_IDENTITY_OPENID_CONFIG_URL="https://keycloak.example.com/realms/rustfs"
export RUSTFS_IDENTITY_OPENID_CLIENT_ID="rustfs-console"
export RUSTFS_IDENTITY_OPENID_CLIENT_SECRET="<KEYCLOAK_CLIENT_SECRET>"
export RUSTFS_IDENTITY_OPENID_SCOPES="openid,profile,email"
export RUSTFS_IDENTITY_OPENID_REDIRECT_URI="https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default"
export RUSTFS_IDENTITY_OPENID_REDIRECT_URI_DYNAMIC=off
export RUSTFS_IDENTITY_OPENID_DISPLAY_NAME="Keycloak"
export RUSTFS_IDENTITY_OPENID_GROUPS_CLAIM="groups"
export RUSTFS_IDENTITY_OPENID_ROLES_CLAIM="roles"
export RUSTFS_IDENTITY_OPENID_EMAIL_CLAIM="email"
export RUSTFS_IDENTITY_OPENID_USERNAME_CLAIM="preferred_username"
```
If RustFS reaches Keycloak through an internal URL while tokens use a public issuer, configure both values:
```bash
export RUSTFS_IDENTITY_OPENID_CONFIG_URL="http://keycloak.keycloak.svc.cluster.local:8080/realms/rustfs/.well-known/openid-configuration"
export RUSTFS_IDENTITY_OPENID_ISSUER="https://keycloak.example.com/realms/rustfs"
export RUSTFS_OUTBOUND_ALLOW_ORIGINS="http://keycloak.keycloak.svc.cluster.local:8080"
```
Discovery and issuer-relative JWKS requests use the internal `CONFIG_URL` base. ID token issuer validation still uses `ISSUER`.
The outbound allowlist entry is the exact internal origin only; do not include the realm or discovery path. RustFS reads this process setting at startup, so restart every RustFS node after changing it.
Use HTTPS with a trusted CA for the internal URL whenever possible. Discovery and JWKS define the token-signing trust root; use HTTP only on a network where DNS and traffic cannot be tampered with, because a compromised response can authorize forged tokens.
For short-lived connectivity testing only, you may temporarily add:
```bash
export RUSTFS_IDENTITY_OPENID_ROLE_POLICY="consoleAdmin"
```
Do not use the temporary `role_policy` shortcut as a permanent production authorization model.
Restart RustFS after changing OIDC settings.
### 4.2 Admin Config
If the deployment uses compatible admin configuration commands:
```bash
mc admin config set rustfs identity_openid \
enable=on \
config_url="https://keycloak.example.com/realms/rustfs" \
client_id="rustfs-console" \
client_secret="<KEYCLOAK_CLIENT_SECRET>" \
scopes="openid,profile,email" \
redirect_uri="https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default" \
redirect_uri_dynamic=off \
display_name="Keycloak" \
groups_claim="groups" \
roles_claim="roles" \
email_claim="email" \
username_claim="preferred_username"
mc admin service restart rustfs
```
`RUSTFS_BROWSER_REDIRECT_URL` is a process environment variable, not an `identity_openid` provider key. Configure it in the RustFS service environment even when the provider is stored through admin config.
### 4.3 Named Provider
To use a provider id such as `keycloak`, register this callback URL in Keycloak:
```text
https://rustfs.example.com/rustfs/admin/v3/oidc/callback/keycloak
```
Then suffix the provider-specific environment variables:
```bash
export RUSTFS_IDENTITY_OPENID_ENABLE_keycloak=on
export RUSTFS_IDENTITY_OPENID_CONFIG_URL_keycloak="https://keycloak.example.com/realms/rustfs"
export RUSTFS_IDENTITY_OPENID_CLIENT_ID_keycloak="rustfs-console"
export RUSTFS_IDENTITY_OPENID_CLIENT_SECRET_keycloak="<KEYCLOAK_CLIENT_SECRET>"
export RUSTFS_IDENTITY_OPENID_SCOPES_keycloak="openid,profile,email"
export RUSTFS_IDENTITY_OPENID_REDIRECT_URI_keycloak="https://rustfs.example.com/rustfs/admin/v3/oidc/callback/keycloak"
export RUSTFS_IDENTITY_OPENID_REDIRECT_URI_DYNAMIC_keycloak=off
export RUSTFS_IDENTITY_OPENID_DISPLAY_NAME_keycloak="Keycloak"
export RUSTFS_IDENTITY_OPENID_GROUPS_CLAIM_keycloak="groups"
```
`RUSTFS_BROWSER_REDIRECT_URL` remains global and is not suffixed per provider.
### 4.4 Redirect URL Priority
RustFS builds browser-facing URLs with this priority:
1. Provider `redirect_uri`, when configured, is used for the OIDC callback URL sent to Keycloak.
2. `RUSTFS_BROWSER_REDIRECT_URL`, when configured, is used as the public origin for OIDC callback generation when no provider `redirect_uri` exists, and for Console success redirects and logout fallback redirects.
3. Request headers are used only when provider dynamic redirects are enabled and no browser redirect URL is configured.
For reverse-proxy or load-balancer deployments, set `RUSTFS_BROWSER_REDIRECT_URL` to avoid depending on `Host` and `X-Forwarded-Proto` for Console redirects. OIDC authorize and callback requests must still reach the same RustFS node because in-flight OIDC `state` is local to the node.
## 5. Validation
### 5.1 Validate Discovery
```bash
curl -fsS "https://keycloak.example.com/realms/rustfs/.well-known/openid-configuration" | jq '{
issuer,
authorization_endpoint,
token_endpoint,
jwks_uri,
code_challenge_methods_supported,
token_endpoint_auth_methods_supported
}'
```
Check that:
- `issuer` equals `RUSTFS_IDENTITY_OPENID_ISSUER` when set; otherwise it matches the issuer derived from `RUSTFS_IDENTITY_OPENID_CONFIG_URL`
- `authorization_endpoint`, `token_endpoint`, and `jwks_uri` are present
- `code_challenge_methods_supported` includes `S256`
- the token endpoint accepts client secret authentication compatible with request-body submission
### 5.2 Validate ID Token Claims
After a test login, decode the ID token and confirm:
- `iss` matches the Keycloak issuer
- `aud` includes the RustFS client id
- `email` and `preferred_username` are present when configured
- `groups` or `roles` contains RustFS policy names if fine-grained authorization is enabled
### 5.3 Test Browser Login
Open:
```text
https://rustfs.example.com/rustfs/admin/v3/oidc/authorize/default
```
Expected flow:
1. Browser redirects to Keycloak.
2. The user signs in.
3. Keycloak redirects to `/rustfs/admin/v3/oidc/callback/default?code=...&state=...`.
4. RustFS validates the ID token and issues STS credentials.
5. The browser lands on the RustFS Console and can use the expected permissions.
## 6. Troubleshooting
| Symptom | Common cause | Fix |
| --- | --- | --- |
| Keycloak reports `invalid redirect_uri` | Valid Redirect URIs does not match the RustFS callback URL | Use the exact callback URL and provider id. |
| Callback reports missing `code` or `state` | Proxy dropped the query string | Preserve the full callback URL and query string. |
| Token exchange fails | Client secret or client authentication policy mismatch | Confirm the client is confidential and accepts request-body secret auth. |
| RustFS reports no `id_token` | Missing `openid` scope or disabled Standard Flow | Include `openid` and enable Standard Flow. |
| ID token verification fails | Issuer, client id, audience, or JWKS mismatch | Compare discovery metadata and client settings. |
| Login succeeds but access is denied | No RustFS policy claim was mapped | Emit `groups` or `roles` as a flat ID token claim matching RustFS policy names. |
| Groups appear as `/consoleAdmin` | Keycloak `Full group path` is enabled | Disable `Full group path`. |
| Console redirects to an internal host | Missing `RUSTFS_BROWSER_REDIRECT_URL` or incorrect proxy headers | Set `RUSTFS_BROWSER_REDIRECT_URL` to the public browser origin. |
| Invalid or expired OIDC state | Callback reached a different RustFS node | Configure load-balancer session affinity for authorize and callback requests. |
| OIDC provider or login button is missing after upgrading to beta.12+ | The internal Keycloak origin is blocked by the outbound policy | Add the exact `scheme://host:port` origin to `RUSTFS_OUTBOUND_ALLOW_ORIGINS` and restart every RustFS node. |
## 7. Production Checklist
- [ ] Keycloak and RustFS use HTTPS.
- [ ] Keycloak Valid Redirect URIs uses exact callback URLs.
- [ ] `RUSTFS_BROWSER_REDIRECT_URL` is set to the public RustFS browser origin.
- [ ] `RUSTFS_IDENTITY_OPENID_REDIRECT_URI` matches the registered Keycloak callback URL.
- [ ] PKCE S256 is enabled or required.
- [ ] Users receive `groups` or `roles` claims that match RustFS policy names.
- [ ] `role_policy=consoleAdmin` is not used as a permanent production shortcut.
- [ ] The load balancer preserves query strings.
- [ ] OIDC authorize and callback requests have session affinity to the same RustFS node.
- [ ] Internal Keycloak origins are listed exactly in `RUSTFS_OUTBOUND_ALLOW_ORIGINS` on every RustFS node.
+42 -54
View File
@@ -1,68 +1,56 @@
# KMS admin API contract and client handoff
# KMS admin API contract
This page is the server-side handoff for rustfs/backlog#1639. Response-shape snapshots for the key and metadata handlers live beside the producers under `rustfs/src/admin/handlers/snapshots/` (PR #5626). This matrix records the remaining route-level handoff contract without duplicating those snapshots. The management `KmsStatusResponse` shape remains an identified gap and will be pinned after the status-handler changes in #1636 land.
**Use this when:** you are wiring a client (CLI, console, automation) to the KMS admin endpoints and need the IAM action, risk class, per-key scope, and key-listing paging rules for each route.
**Source of truth:** `rustfs/src/admin/route_policy.rs` (action and risk per route, asserted by `rustfs/src/admin/route_registration_test.rs`); `crates/kms/src/backends/mod.rs` (`DEFAULT_LIST_KEYS_PAGE_SIZE`, `MAX_LIST_KEYS_PAGE_SIZE`, `list_keys_page_size`); response shapes pinned by snapshots under `crates/kms/src/snapshots/` and `rustfs/src/admin/handlers/snapshots/`.
The wire prefix is `/rustfs/admin/v3`. Request and response field names for the producer snapshots are pinned in #5626; fields in this matrix are the client handoff reference. `GET /kms/status` and `GET /kms/service-status` intentionally use different response types; `capabilities` on `/kms/status` is additive and optional.
The wire prefix is `/rustfs/admin/v3`. `GET /kms/status` and `GET /kms/service-status` return different response types; `capabilities` on `/kms/status` is additive and optional. The **Per-key** column says whether the route authorizes against the key it names (see [Per-key KMS authorization](kms-per-key-authorization.md)); `no` means the route matches any KMS resource in the caller's policy.
| Method and endpoint | Action / risk | Per-key | rc | console | Handoff |
| --- | --- | ---: | --- | --- | --- |
| `POST /kms/configure` | `kms:Configure` / high | no | supported | supported | none |
| `POST /kms/reconfigure` | `kms:Configure` / high | no | supported | supported | none |
| `POST /kms/start` | `kms:ServiceControl` / high | no | supported | supported | none |
| `POST /kms/stop` | `kms:ServiceControl` / high | no | supported | supported | none |
| `POST /kms/reload` | `kms:ServiceControl` / high | no | pending | pending | Re-reads the cluster-persisted configuration without resubmitting secrets; response reuses the configure shape. |
| `GET /kms/config` | `kms:Configure` / sensitive | no | no | supported | Redact operational paths before display. |
| `POST /kms/clear-cache` | `kms:ClearCache` / high | no | no | supported | Keep the current `{status,message}` response stable. |
| `POST /kms/keys` | `kms:Configure` / high | no | supported | supported | none |
| `GET /kms/keys` | `kms:ListKeys` / sensitive | no | supported | supported | none |
| `GET /kms/keys/{key_id}` | `kms:DescribeKey` / sensitive | yes | supported | supported | none |
| `DELETE /kms/keys/delete` | `kms:DeleteKey` / critical | yes | supported | supported | Preserve immediate-delete confirmations. |
| `POST /kms/keys/cancel-deletion` | `kms:DeleteKey` / high | yes | supported | supported | none |
| `POST /kms/create-key` | `kms:Configure` / high | no | no | no | Legacy `mc` alias; do not add a second client command. |
| `POST /kms/key/create` | `kms:Configure` / high | no | no | no | Legacy `mc` alias; do not add a second client command. |
| `GET /kms/describe-key` | `kms:DescribeKey` / sensitive | yes | no | no | Legacy `mc` alias. |
| `GET /kms/key/status` | `kms:DescribeKey` / sensitive | yes | supported | no | `rc key status` uses this legacy-compatible shape. |
| `GET /kms/list-keys` | `kms:ListKeys` / sensitive | no | supported | no | `rc key list` uses this legacy-compatible shape. |
| `POST /kms/generate-data-key` | `kms:GenerateDataKey` / high | yes | do not expose | do not expose | Programmatic primitive; never print plaintext key material. |
| `GET /kms/status` | `kms:ServiceControl` / sensitive | no | supported | supported | Keep `capabilities` optional for old servers. |
| `POST /kms/status` | `kms:ServiceControl` / high | no | no | no | Internal compatibility route; not a client command. |
| `GET /kms/service-status` | `kms:ServiceControl` / sensitive | no | no | supported | Do not conflate this type with `/kms/status`. |
| `POST /kms/keys/enable` | `kms:EnableKey` / high | yes | pending | pending | Add an explicit lifecycle command/UI action. |
| `POST /kms/keys/disable` | `kms:DisableKey` / high | yes | pending | pending | Add an explicit lifecycle command/UI action. |
| `POST /kms/keys/rotate` | `kms:RotateKey` / high | yes | pending | pending | Add an explicit lifecycle command/UI action. |
| `POST /kms/keys/update-description` | `kms:UpdateKeyDescription` / high | yes | pending | pending | Add a metadata mutation command/UI action. |
| `POST /kms/keys/tag` | `kms:TagResource` / high | yes | pending | pending | Add a metadata mutation command/UI action. |
| `POST /kms/keys/untag` | `kms:UntagResource` / high | yes | pending | pending | Add a metadata mutation command/UI action. |
| `GET /kms/backup` | `kms:Backup` / sensitive | no | pending | pending | Status/readiness only; never expose KEK material. |
| `POST /kms/backup` | `kms:Backup` / high | no | pending | pending | Preserve `backup_id` and metadata-only response. |
| `POST /kms/restore/dry-run` | `kms:Restore` / sensitive | no | pending | pending | Dry-run must be the default and show differences. |
| `POST /kms/restore` | `kms:Restore` / high | no | pending | pending | Require `confirm_backup_id` and `confirm_conflict_policy`; no blanket `--yes`. |
| `POST /kms/restore/abort` | `kms:Restore` / high | no | pending | pending | Require `confirm_target_key_dir`. |
## Endpoint matrix
| Method and endpoint | IAM action | Risk | Per-key | Notes |
| --- | --- | --- | --- | --- |
| `POST /kms/configure` | `kms:Configure` | high | no | Persists to cluster storage, switches the local node, broadcasts a best-effort peer reload |
| `POST /kms/reconfigure` | `kms:Configure` | high | no | Same contract as configure |
| `POST /kms/start` | `kms:ServiceControl` | high | no | |
| `POST /kms/stop` | `kms:ServiceControl` | high | no | |
| `POST /kms/reload` | `kms:ServiceControl` | high | no | Re-reads the persisted configuration without resubmitting secrets; reuses the configure response shape |
| `GET /kms/status` | `kms:ServiceControl` | sensitive | no | Backend type plus capability matrix |
| `POST /kms/status` | `kms:ServiceControl` | high | no | Compatibility route; not a client command |
| `GET /kms/service-status` | `kms:ServiceControl` | sensitive | no | Carries `cluster_config` fingerprints and the `consistent` flag |
| `GET /kms/config` | `kms:Configure` | sensitive | no | Contains operational paths; redact before display |
| `POST /kms/clear-cache` | `kms:ClearCache` | high | no | `KmsClearCacheResponse` (`{status,message}`) |
| `POST /kms/keys` | `kms:Configure` | high | no | Key creation shares the configure action |
| `GET /kms/keys` | `kms:ListKeys` | sensitive | no | See the key listing contract below |
| `GET /kms/keys/{key_id}` | `kms:DescribeKey` | sensitive | yes | `?impact=true` opts into the configuration-reference report |
| `DELETE /kms/keys/delete` | `kms:DeleteKey` | critical | yes | JSON body; `force_immediate` also requires `confirm_key_id` and the server-side `RUSTFS_KMS_ALLOW_IMMEDIATE_DELETION` gate |
| `POST /kms/keys/cancel-deletion` | `kms:DeleteKey` | high | yes | |
| `POST /kms/keys/enable` | `kms:EnableKey` | high | yes | |
| `POST /kms/keys/disable` | `kms:DisableKey` | high | yes | |
| `POST /kms/keys/rotate` | `kms:RotateKey` | high | yes | Subject to the rotation constraints in [KMS backend security properties](kms-backend-security.md#master-key-rotation-retention-destruction-and-upgrade-ordering) |
| `POST /kms/keys/rekey` | `kms:Rekey` | high | no | Bulk DEK rekey sweep; cluster-scoped, see [`kms-bulk-rekey-contract.md`](../architecture/kms-bulk-rekey-contract.md) |
| `GET /kms/keys/rekey/status` | `kms:Rekey` | sensitive | no | |
| `POST /kms/keys/rekey/cancel` | `kms:Rekey` | high | no | |
| `POST /kms/keys/update-description` | `kms:UpdateKeyDescription` | high | yes | |
| `POST /kms/keys/tag` | `kms:TagResource` | high | yes | |
| `POST /kms/keys/untag` | `kms:UntagResource` | high | yes | |
| `POST /kms/generate-data-key` | `kms:GenerateDataKey` | high | yes | Response carries a base64 plaintext data key; never surface it in a UI or CLI |
| `GET /kms/backup` | `kms:Backup` | sensitive | no | Status and readiness only; no KEK material |
| `POST /kms/backup` | `kms:Backup` | high | no | Returns `backup_id` and metadata only |
| `POST /kms/restore/dry-run` | `kms:Restore` | sensitive | no | Preflight; writes nothing |
| `POST /kms/restore` | `kms:Restore` | high | no | Requires `confirm_backup_id` and `confirm_conflict_policy` |
| `POST /kms/restore/abort` | `kms:Restore` | high | no | Requires `confirm_target_key_dir` |
| `POST /kms/create-key`, `POST /kms/key/create` | `kms:Configure` | high | no | Legacy `mc` aliases of `POST /kms/keys` |
| `GET /kms/describe-key`, `GET /kms/key/status` | `kms:DescribeKey` | sensitive | yes | Legacy aliases of `GET /kms/keys/{key_id}` |
| `GET /kms/list-keys` | `kms:ListKeys` | sensitive | no | Legacy alias of `GET /kms/keys`; same listing contract |
## Key listing contract
Both listing routes (`GET /kms/keys` and the legacy `GET /kms/list-keys`) share one contract.
`limit` is optional. When it is absent the server applies its own default page size of 100. When it is present it must parse as a non-negative integer: `limit=abc`, `limit=-1` and a value-less `limit` are refused with `400`, not silently read as "use the default". `limit=0` is a well-formed request for an empty page. Any page size above 1000 is served as 1000 — the response is `truncated` with a usable `next_marker`, so a client that pages until `truncated` is false still reaches every key. Clients must not assume a page is the size they asked for.
`limit` is optional. When it is absent the server applies `DEFAULT_LIST_KEYS_PAGE_SIZE` (100). When it is present it must parse as a non-negative integer: `limit=abc`, `limit=-1` and a value-less `limit` are refused with `400`, not silently read as "use the default". `limit=0` is a well-formed request for an empty page. Any page size above `MAX_LIST_KEYS_PAGE_SIZE` (1000) is served as 1000 — the response is `truncated` with a usable `next_marker`, so a client that pages until `truncated` is false still reaches every key. Clients must not assume a page is the size they asked for.
`marker` is opaque to the client: treat it as a cursor to hand back unchanged, never as a value to construct. On the Local, Vault KV2, Vault Transit and Static backends it happens to be an exclusive lower bound on the key identifier, which is what makes paging survive keys being created or destroyed mid-listing; on the AWS backend it is AWS's own pagination token, and sending a key id there is rejected. An empty `marker` means the same thing as no marker at all. Filters are applied after the page is cut, so a filtered page can be short — even empty — while more keys remain. Page until `truncated` is false, never until a page comes back short.
`unreadable_key_ids` is present only when the server listed a key whose record it could not describe — a record written by a newer build, or damaged material. The identifiers are reported rather than omitted, so a listing never quietly understates the key set; a client displaying an inventory should surface them as damaged rather than dropping them, and paging always advances past a damaged key. A failure that says nothing about a specific key (timeout, `5xx`, permission denied) still fails the whole listing instead of appearing here.
One case is deliberately an error rather than a report: a listing that covered the entire key set — no `marker`, and not `truncated` — in which nothing was readable. An empty `keys` array there would be indistinguishable, to any client written before this field existed, from a deployment that has no keys, and the usual response to that is to provision a new one. Such a listing returns `500` instead, naming the first failure; the individual identifiers are in the server log. A truncated page, or one resumed from a marker, always reports rather than failing, so a damaged key can never strand the keys behind it.
## Server-side snapshot coverage
The merged #5626 producer snapshots cover the nine modern/legacy key response types and the metadata response type served by `kms_keys.rs` and `kms_key_metadata.rs`: create, describe, list, generate-data-key, delete, cancel-deletion, update-description, tag, and untag. The four dynamic responses served verbatim by `kms_dynamic.rs` are covered in `crates/kms/src/snapshots/`: configure, start, stop, and the `service-status` response. `POST /kms/reload` serves the same `ConfigureKmsResponse` type the configure snapshot pins; it adds no new wire shape.
`POST /kms/clear-cache` now has a named `KmsClearCacheResponse` and a producer snapshot beside the others; its serialized bytes are unchanged from the inline JSON it replaced.
The remaining wire-shape gaps are intentionally documented rather than duplicated here: the management `KmsStatusResponse` (`GET|POST /kms/status`, pending #1636), `KmsConfigResponse`, all three lifecycle responses, and the backup/restore response family. Adding producer snapshots for those gaps is a separate server test task; it must not be inferred from the client matrix.
## Client handoff gaps
The `rc` client currently has status, key list/status/create/delete/cancel-deletion, configure/reconfigure/start/restart/stop, and diagnostic/roundtrip entry points. It has no lifecycle enable/disable/rotate, key metadata, backup/restore, or reload commands; `POST /kms/reload` is the recovery path when a restarted server reports not-configured while a persisted configuration exists, so it is a client delivery item alongside the lifecycle gaps. The console currently calls service-status, configure/reconfigure/start/stop/config, clear-cache, status, and the modern key CRUD routes. It has no lifecycle, metadata, or backup/restore UI. These pending cells are delivery items for `rustfs/cli` and `rustfs/console`; they are not implemented in this repository. A read-only issue search on 2026-08-02 found no matching KMS issue in either client repository, so the client handoff still needs issue creation there.
`POST /kms/generate-data-key` is deliberately marked “do not expose” for both clients: its response contains a base64 plaintext data key. `GET /kms/config` and backup status/restore responses contain operational paths and identifiers, not key material, but still require UI/CLI redaction and confirmation handling.
The producer response snapshots in #5626 and this matrix do not imply that rustfs/backlog#1639 is complete. The pending client cells must be closed in their respective repositories before the parent delivery item can be marked complete.
+108 -170
View File
@@ -1,56 +1,38 @@
# KMS backend security properties
RustFS ships several KMS backends. They differ not only in deployment effort but in **where master key material lives and who can read it**. Pick a backend based on the confidentiality boundary you need, not on the name alone.
**Use this when:** choosing a KMS backend, scheduling or debugging master key rotation, planning a rolling upgrade of a cluster with KMS enabled, or auditing where master key material lives and who can read it.
**Source of truth:** `crates/kms/src/config.rs` (`KmsBackend`, `ENV_KMS_*` constants), `crates/kms/src/backends/{local,vault,vault_transit,aws}.rs`, `crates/kms/src/encryption/dek.rs` (`DataKeyEnvelope`), `rustfs/src/admin/route_policy.rs` (KMS route actions).
For how the Vault backends authenticate (static token, AppRole, Kubernetes, Vault Agent token file) and how credential refresh and the fail-closed window behave, see the [Vault KMS authentication runbook](vault-kms-authentication.md). For what may be claimed about the cryptographic implementations themselves, see [Cryptographic compliance positioning](kms-cryptographic-compliance.md). For which RustFS identities may manage or use a given key, see [Per-key KMS authorization](kms-per-key-authorization.md). If you are migrating from MinIO, read [Migrating from MinIO: encrypted objects do not carry over](#migrating-from-minio-encrypted-objects-do-not-carry-over) first.
RustFS ships several KMS backends. They differ not only in deployment effort but in **where master key material lives and who can read it**. Pick a backend based on the confidentiality boundary you need, not on the name alone. Related: [Vault KMS authentication runbook](vault-kms-authentication.md) (credential sources, refresh, fail-closed window), [Cryptographic compliance positioning](kms-cryptographic-compliance.md), [Per-key KMS authorization](kms-per-key-authorization.md), [KMS admin API contract](kms-admin-contract.md), [KMS observability runbook](kms-observability-runbook.md).
## Backend comparison
| Backend | Config tag | Master key material location | At-rest protection of key material | Durability | Rotation | Intended use |
| --- | --- | --- | --- | --- | --- | --- |
| Local | `Local` | Files under `key_dir`, encrypted with the configured local master key | Local master key (AES-GCM) + file permissions | Crash-durable commits on local filesystems only; see [Local backend durability and deployment support matrix](#local-backend-durability-and-deployment-support-matrix) | Rejected by design (single material, development backend) | Development, testing and demos only; not supported for production |
| Local | `Local` | Files under `key_dir`, encrypted with the configured local master key | Local master key (AES-GCM) + file permissions | Crash-durable commits on local filesystems only; see [Local backend durability and deployment support matrix](#local-backend-durability-and-deployment-support-matrix) | Rejected by design (single material) | Development, testing and demos only; not supported for production |
| Static | `Static` | Provided out-of-band via environment/file; never persisted by RustFS | Operator-managed secret distribution | No state persisted by RustFS | Rejected (read-only backend) | Development and testing with an externally supplied key; not supported for production |
| Vault KV2 | `VaultKV2` (legacy alias `Vault`) | Stored **directly** in Vault KV v2 (Base64-encoded plaintext) | Vault ACLs + KV v2 at-rest encryption + TLS only | Delegated to Vault storage | Versioned retention (immutable per-version records + current pointer) | Deployments that accept Vault KV ACLs as the sole confidentiality boundary |
| Vault Transit | `VaultTransit` | Key-encryption keys never leave Vault; only Transit ciphertext is visible outside | Vault Transit engine (cryptographic isolation) | Delegated to Vault storage | Via Vault Transit key versioning | Deployments that need key material to be unreadable through storage APIs |
| AWS KMS | `AWS` (alias `AwsKms`) | Key material never leaves AWS KMS; RustFS mirrors no key state | AWS KMS (cryptographic isolation) + IAM | Delegated to AWS | On-demand `RotateKeyOnDemand`; prior backing keys stay usable for decryption | Deployments already rooted in AWS IAM that want AWS as the cryptographic root — read [AWS KMS: deviations from the shared backend contract](#aws-kms-deviations-from-the-shared-backend-contract) first |
| AWS KMS | `AWS` (alias `AwsKms`) | Key material never leaves AWS KMS; RustFS mirrors no key state | AWS KMS (cryptographic isolation) + IAM | Delegated to AWS | On-demand `RotateKeyOnDemand`; prior backing keys stay usable for decryption | Deployments rooted in AWS IAM — read [AWS KMS: deviations from the shared backend contract](#aws-kms-deviations-from-the-shared-backend-contract) first |
## Migrating from MinIO: encrypted objects do not carry over
> **Warning: RustFS does not currently support reading objects that MinIO encrypted.**
> This applies to SSE-S3, SSE-KMS, and SSE-C, in every released binary and container image, and it holds regardless of which KMS backend you configure. Configuring the `Static` backend with the same key material MinIO used does **not** make those objects readable — MinIO wraps data keys in a different envelope format that no RustFS backend produces or accepts (`crates/kms/src/config.rs:304-308`). Plan for this **before** moving data. Tracked in rustfs/backlog#1638.
> **Warning: default RustFS builds fail closed on objects that MinIO encrypted.** This applies to SSE-S3, SSE-KMS, and SSE-C, whichever KMS backend you configure; configuring `Static` with MinIO's key material does not make them readable. Such objects list and HEAD normally (their `xl.meta` parses), and only the payload read fails — with S3 `InvalidObjectState`, never plaintext. Read a sample of encrypted objects, not just their listings, before decommissioning the MinIO deployment.
The read does fail closed — ciphertext is never served as plaintext. MinIO's internal encryption headers mark the object as encrypted (`crates/utils/src/http/header_compat.rs:50-67`), so the read path demands encryption material and refuses when none resolves (`crates/ecstore/src/object_api/readers.rs:559-568`). Two properties still make the problem easy to discover late:
- **The error does not say what happened.** It surfaces as a 500 `InternalError`, which reads as a RustFS fault rather than "another implementation encrypted this object".
- **Surrounding metadata migrates fine.** The object's `xl.meta` parses, so encrypted objects list and HEAD normally and report plausible sizes. The failure appears only when something reads the payload.
Read a sample of encrypted objects, not just their listings, before decommissioning the MinIO deployment.
Current options for a migration whose source contains encrypted objects:
- Decrypt on the MinIO side first, migrate plaintext, then let RustFS re-encrypt with its own KMS.
- Copy through the S3 API rather than moving drives — MinIO decrypts on read, and RustFS encrypts on write. This re-encrypts rather than preserving ciphertext and costs a full data transfer.
- Leave encrypted objects on MinIO and migrate only unencrypted data.
Inventory the source before choosing: bucket default-encryption settings mean objects can be encrypted without any client having sent SSE headers.
The same limitation applies in reverse — objects RustFS encrypts are not readable by MinIO. For the code-level breakdown of which seams block each SSE mode, see [MinIO file-format interoperability, Part C](../architecture/minio-file-format-compat.md#part-c--server-side-encryption-sse).
The migration warning is not a wire-protocol promise: the **AWS KMS wire protocol** and **MinIO KES wire protocol** are explicit non-targets for this document. The AWS backend uses the AWS SDK client path (`crates/kms/src/backends/aws.rs:830`), and KES remains outside the MinIO on-disk interop scope. Track those ecosystem evaluations and the MinIO/RustFS SSE compatibility matrix in the [#1562 Production Ready exit gate](https://github.com/rustfs/backlog/issues/1562); #1638 alone does not satisfy that gate.
The read path exists behind the `rio-v2` feature as a migration-only build, the reverse direction (RustFS-written SSE objects read by MinIO) is unsupported, and the migration options are enumerated in [MinIO file-format interoperability, Part C](../architecture/minio-file-format-compat.md#part-c--server-side-encryption-sse). The AWS KMS and MinIO KES wire protocols are non-targets of that document and of this one.
## Vault KV2: what the backend does and does not do
The Vault KV2 backend uses Vault purely as a **secure storage** service:
- Master key material is generated by RustFS and written to KV v2 as a Base64-encoded value (`encrypted_key_material` is an encoding, not a ciphertext).
- The backend never calls the Vault Transit engine. The `mount_path` configuration field and the `RUSTFS_KMS_VAULT_MOUNT_PATH` environment variable are deprecated leftovers: they are accepted for compatibility and ignored.
- The backend never calls the Vault Transit engine. The `mount_path` configuration field and the `RUSTFS_KMS_VAULT_MOUNT_PATH` environment variable are deprecated leftovers: accepted for compatibility and ignored.
- Data-encryption keys (DEKs) handed to the object-encryption path are still wrapped with AES-256-GCM under the master key; the statement above concerns the master key's storage in Vault, not the DEK envelope.
- No runtime interface reports this boundary. `GET /rustfs/admin/v3/kms/status` names the active backend (`backend_type: vault-kv2`) and returns a `capabilities` matrix, but that matrix enumerates only the operations the backend supports — nothing in it describes where master key material lives or who can read it. Determining which confidentiality boundary is in force means reading `backend_type` and applying the comparison table above; this document is the only statement of the boundary an operator can consult.
- The `at_rest_protection: storage-only` field carried by a KMS backup manifest is a different thing: it declares the protection state of key material inside a backup bundle, not a property the running backend reports about itself.
- Key rotation retains every historical master key version as an immutable record under `{prefix}/{key_id}/versions/{N}` and only then moves the current-version pointer; see [Master key rotation](#master-key-rotation-retention-destruction-and-upgrade-ordering) for the retention preconditions and the cluster-upgrade ordering constraint.
- No runtime interface reports this boundary. `GET /rustfs/admin/v3/kms/status` names the active backend (`backend_type: vault-kv2`) and returns a `capabilities` matrix, but that matrix enumerates only supported operations. Determining the confidentiality boundary in force means reading `backend_type` and applying the comparison table above.
- The `at_rest_protection: storage-only` field carried by a KMS backup manifest declares the protection state of key material inside a backup bundle, not a property the running backend reports about itself.
- Key rotation retains every historical master key version as an immutable record under `{prefix}/{key_id}/versions/{N}` and only then moves the current-version pointer; see [Master key rotation](#master-key-rotation-retention-destruction-and-upgrade-ordering).
> **Warning: KV read access is equivalent to holding the master keys.**
> Any Vault identity (token, AppRole, or policy) that can `read` the RustFS key path in KV v2 can recover the plaintext master key material and decrypt every object protected by those keys. Treat KV read grants on that path with the same care as handing out the keys themselves. If this is not acceptable, use the Vault Transit backend instead.
> **Warning: KV read access is equivalent to holding the master keys.** Any Vault identity (token, AppRole, or policy) that can `read` the RustFS key path in KV v2 can recover the plaintext master key material and decrypt every object protected by those keys. If this is not acceptable, use the Vault Transit backend instead.
## Minimal Vault policy for the KV2 backend
@@ -67,63 +49,49 @@ path "secret/metadata/rustfs/kms/keys/*" {
}
```
Notes:
- The trailing wildcards also cover the per-version material records that rotation creates under `.../keys/{key_id}/versions/{N}`; no extra policy paths are needed.
- `delete` on the metadata path is required for permanent key deletion (`force_immediate`); drop it if you never hard-delete keys. RustFS refuses `force_immediate` unless the server sets `RUSTFS_KMS_ALLOW_IMMEDIATE_DELETION=true`, so leaving that gate off keeps the capability unreachable no matter what the Vault policy allows.
- Do not attach `sudo`, wildcard mounts, or Transit paths to this policy; the KV2 backend does not use them.
- Auditing KV reads on the key prefix is strongly recommended: every read event is a potential master-key disclosure.
- The trailing wildcards also cover the per-version material records under `.../keys/{key_id}/versions/{N}`; no extra policy paths are needed.
- `delete` on the metadata path is required only for permanent key deletion (`force_immediate`); drop it if you never hard-delete keys. RustFS refuses `force_immediate` unless the server sets `RUSTFS_KMS_ALLOW_IMMEDIATE_DELETION=true`, so leaving that gate off keeps the capability unreachable whatever the Vault policy allows.
- Do not attach `sudo`, wildcard mounts, or Transit paths; the KV2 backend does not use them.
- Audit KV reads on the key prefix: every read event is a potential master-key disclosure.
## Master key rotation: retention, destruction, and upgrade ordering
Rotation support differs per backend. Local and Static advertise no `rotate` capability `capabilities.rotate` is false in the `kms/status` response and reject rotation with `UnsupportedCapability`; their single key material is never overwritten. Vault Transit delegates rotation to the Transit engine's own key versioning (ciphertext is version-prefixed, e.g. `vault:v1:...`). Vault KV2 rotates by retaining every historical version, as described below. Rotation is reachable through the admin API as `POST /rustfs/admin/v3/kms/keys/rotate`, which the route policy classifies as high risk and gates behind `kms:RotateKey`; it is not exposed through the S3 surface. The upgrade ordering constraint below therefore applies to an operator action, not only to a call from inside the process.
Rotation is reachable through `POST /rustfs/admin/v3/kms/keys/rotate` (`kms:RotateKey`, high risk); it is not exposed on the S3 surface. Local and Static advertise no `rotate` capability (`capabilities.rotate` is false in the `kms/status` response) and reject rotation with `UnsupportedCapability`. Vault Transit delegates rotation to the Transit engine's own versioning (ciphertext is version-prefixed, e.g. `vault:v1:...`). Vault KV2 rotates by retaining every historical version, as described below.
### Rotation drivers and scheduling, per backend
The rotate endpoint is one API over three very different mechanisms, and which component actually performs the rotation decides how periodic rotation must be scheduled — on two backends it cannot be scheduled at all.
The rotate endpoint is one API over three different mechanisms, and which component performs the rotation decides how periodic rotation must be scheduled.
| Backend | Can rotate | Who performs the rotation | How to schedule periodic rotation |
| --- | --- | --- | --- |
| Local | No | Nobody — the backend advertises no `rotate` capability and the rotate endpoint is refused with `UnsupportedCapability` | Cannot be scheduled. Migrating to a rotating backend is the only path to rotation |
| Static | No | Nobody — same refusal as Local; the material is supplied out-of-band and read-only | Cannot be scheduled. Migrate to a rotating backend |
| Vault KV2 | Yes | **RustFS** owns the whole rotation protocol: freeze the outgoing material as an immutable version record, persist the new version's material, then move the current pointer with a check-and-set write | An **external scheduler** (cron, Kubernetes CronJob, your automation platform) calling `POST /rustfs/admin/v3/kms/keys/rotate`. RustFS deliberately ships no built-in rotation timer — see below |
| Vault Transit | Yes | **Vault's Transit engine** RustFS only forwards the call to Transit's rotate endpoint and records the version bump in its own metadata | Vault's native `auto_rotate_period` on the Transit key. Do **not** additionally point an external scheduler at the RustFS rotate endpoint — see below |
| AWS KMS | Yes | **AWS** — the RustFS rotate endpoint maps to `RotateKeyOnDemand` | AWS's native automatic rotation, configured on the AWS side. Do **not** drive periodic rotation through the RustFS endpoint — see below |
| Backend | Can rotate | Who performs the rotation | How to schedule periodic rotation | Wrap ceiling |
| --- | --- | --- | --- | --- |
| Local | No | Nobody — the rotate endpoint is refused with `UnsupportedCapability` | Cannot be scheduled; migrating to a rotating backend is the only path | Unmitigable |
| Static | No | Nobody — same refusal; the material is supplied out-of-band and read-only | Cannot be scheduled; migrate | Unmitigable |
| Vault KV2 | Yes | **RustFS** owns the protocol: freeze the outgoing material as an immutable version record, persist the new material, move the current pointer with a check-and-set write | Exactly **one external scheduler** (cron, Kubernetes CronJob) calling the rotate endpoint with credentials scoped to `kms:RotateKey` | Reset by each rotation |
| Vault Transit | Yes | **Vault's Transit engine**; RustFS forwards the call and records the version bump | Vault's native `auto_rotate_period` on the Transit key. Do **not** also drive the RustFS endpoint: two owners of the version cadence means neither configured period holds. The reported key version advances only through RustFS, so on an auto-rotating key treat it as a floor | Not applicable (wraps inside Vault) |
| AWS KMS | Yes | **AWS** — the endpoint maps to `RotateKeyOnDemand` | AWS's native automatic rotation. Do **not** drive periodic rotation through the RustFS endpoint: AWS caps lifetime on-demand rotations, so a scheduler exhausts the quota and then fails forever. Keep the endpoint for incidents. RustFS neither enables nor observes AWS automatic rotation and records no rotation timestamp, so its rotation-age signals measure key age on this backend | Not applicable (wraps inside AWS) |
**Local and Static: the wrap ceiling is unmitigable.** These backends wrap every DEK with AES-256-GCM under their single master key using a random 96-bit nonce, and NIST SP 800-38D caps AES-GCM at 2^32 invocations per key when nonces are chosen at random. Each encrypted object write wraps a DEK, so the invocation count tracks the number of encrypted-object writes over the deployment's lifetime. On a rotating backend that count restarts whenever new master key material takes over; on Local and Static it can never restart, because there is no rotation to restart it. The only mitigation is migrating to a backend that rotates. The same 2^32 bound applies to the KV2 backend's wrapping — RustFS wraps DEKs locally there too — but there each rotation mints fresh master key material and resets the count, which is one more reason to actually schedule KV2 rotation rather than merely support it.
**Wrap ceiling.** Where RustFS wraps DEKs locally (Local, Static, Vault KV2) every DEK is wrapped with AES-256-GCM under the master key using a random 96-bit nonce, and NIST SP 800-38D caps AES-GCM at 2^32 invocations per key under random nonces. Each encrypted-object write is one wrap, so the count tracks lifetime encrypted writes. Rotation installs fresh material and restarts the count; on Local and Static there is no rotation, so the ceiling can only be escaped by migrating.
**Vault KV2: bring your own scheduler, deliberately.** RustFS performs the rotation but does not decide when: there is no built-in rotation worker, by design rather than omission. A timer inside the server cannot verify the [cluster-upgrade precondition](#upgrade-before-first-rotation-hard-constraint) before firing, and rotation is not idempotent without leader election, N nodes running the same schedule would perform N rotations per period, advancing the key version N times. Run exactly one external scheduler, point it at the admin rotate endpoint with credentials scoped to `kms:RotateKey`, and use the [rotation readiness fields](#rotation-readiness-reported-never-acted-on) plus the `KmsKeyRotationOverdue` alert in the [KMS observability runbook](kms-observability-runbook.md#kmskeyrotationoverdue) to verify the schedule is actually keeping up.
**Vault Transit: exactly one owner of the version cadence.** Configure `auto_rotate_period` on the Transit key and let Vault own the schedule. Layering an external scheduler that calls the RustFS rotate endpoint on top of `auto_rotate_period` creates two competing owners of the key's version cadence, and the effective rotation period stops being the one either owner was configured with. The data path is indifferent to who rotates — Transit ciphertext self-describes the version that wrapped it, so envelopes never pin a version RustFS tracked — but the key version RustFS reports only advances when rotation goes through RustFS, so on an auto-rotating key treat the reported version as a floor, not the truth.
**AWS KMS: native automatic rotation for cadence, `RotateKeyOnDemand` for incidents.** The RustFS rotate endpoint maps to AWS `RotateKeyOnDemand`, and AWS enforces a lifetime limit on the number of on-demand rotations a key may receive (see the AWS KMS documentation) — a periodic scheduler driving the RustFS endpoint will exhaust that quota and then fail forever. Configure AWS's automatic rotation for periodic cadence and keep the RustFS endpoint for what on-demand rotation is for: incident response and one-off rotations. Note that RustFS neither enables nor observes AWS automatic rotation, and it records no rotation timestamp for AWS keys, so the readiness fields and the rotation-age gauge measure key age on this backend — verify the actual cadence in AWS, not through RustFS.
**Why there is no built-in rotation timer (KV2).** A timer inside the server cannot verify the [upgrade-before-first-rotation constraint](#upgrade-before-first-rotation-hard-constraint), and rotation is not idempotent: without leader election, N nodes on the same schedule would advance the key version N times per period. Verify that your scheduler keeps up with the [rotation readiness fields](#rotation-readiness-reported-never-acted-on) and the `KmsKeyRotationOverdue` alert in the [KMS observability runbook](kms-observability-runbook.md#kmskeyrotationoverdue).
**Pre-rotation checklist** (before the first rotation of any key, and before enabling any schedule):
1. Every node in the cluster runs a build that understands the `master_key_version` envelope field — the [hard upgrade-ordering constraint](#upgrade-before-first-rotation-hard-constraint) below. A timer cannot check this; you must.
1. Every node runs a build that understands the `master_key_version` envelope field — the [hard constraint](#upgrade-before-first-rotation-hard-constraint) below. A timer cannot check this; you must.
2. No rolling upgrade is in progress — see [Do not do these during a mixed-version window](#do-not-do-these-during-a-mixed-version-window).
3. The [retention and destruction preconditions](#retention-and-destruction-preconditions) are understood: every version record a stored DEK envelope references must remain readable forever, and no retention tooling prunes the version subtree.
4. For KV2, exactly one scheduler exists, so no two callers race the same rotation period.
3. The [retention and destruction preconditions](#retention-and-destruction-preconditions) are understood: every version record a stored DEK envelope references must remain readable, and no retention tooling prunes the version subtree.
4. For KV2, exactly one scheduler exists.
5. `RUSTFS_KMS_ROTATION_MAX_AGE_SECS` is set to the rotation period your policy requires, so the per-key `rotation_due` verdict and the rotation-age alert verify the schedule instead of assuming it.
### Rotation readiness: reported, never acted on
RustFS does not rotate keys on a schedule. There is no built-in rotation worker, deliberately: rotation is a policy decision with a per-backend cost and a hard upgrade-ordering constraint (see below), and a server that rotated on its own would make that decision on an operator's behalf at a moment it did not choose. What the server does instead is tell you which keys have outlived a period you configure.
RustFS reports which keys have outlived a period you configure; nothing consults the verdict before encrypting or decrypting, and it has no effect on readiness or liveness.
Set `RUSTFS_KMS_ROTATION_MAX_AGE_SECS` to that period in whole seconds. Unset — the default — leaves the verdict unreported rather than assuming a policy: how often keys must be rotated is a compliance decision, and a built-in default would report keys as overdue against a rule nobody wrote. An unparsable value is treated the same way, with a warning, instead of silently falling back to a number the operator did not choose. Values below one hour are raised to one hour, because a threshold of seconds reports every key as overdue moments after it was rotated and teaches operators to ignore the signal.
| Setting | Meaning | Unset or unparsable | Floor |
| --- | --- | --- | --- |
| `RUSTFS_KMS_ROTATION_MAX_AGE_SECS` | Rotation period in whole seconds; keys older than this report `rotation_due` with reason `age` or `never_rotated` | No age verdict is reported (a warning is logged for an unparsable value) — how often keys must rotate is a compliance decision, not a built-in default | 1 hour |
| `RUSTFS_KMS_ROTATION_MAX_WRAPS` | Data keys one key's material may wrap before `rotation_due` with reason `wraps` | No wrap verdict is reported | 1,000,000 (wraps are accounted in reserved blocks of that size) |
A second, independent threshold covers the cryptographic bound rather than the policy one. `RUSTFS_KMS_ROTATION_MAX_WRAPS` is the number of data keys one key's material may wrap before the verdict reports `rotation_due` with reason `wraps`. It follows the same discipline — unset or unparsable leaves the verdict unreported, and values below one million are raised to one million because wraps are accounted in reserved blocks of that size, so a smaller threshold would trip on the first reservation. Only backends where RustFS wraps locally and can rotate report a count (Vault KV2 today); Transit and AWS wrap externally and report none, so the wrap half stays silent there rather than guessing. When both thresholds are crossed the reported reason is `wraps`: the AES-GCM random-nonce ceiling is not negotiable, while the age period is a policy an operator chose.
`GET /rustfs/admin/v3/kms/keys` then carries two additional fields per key:
- `rotation_due` — whether the key has outlived the configured period.
- `rotation_due_reason``age` when the key was rotated but longer ago than the period, `never_rotated` when it has never been rotated and has been in use longer than the period, and `unsupported` when the backend cannot rotate at all. Absent when there is no verdict.
The verdict is advisory in the strongest sense: nothing consults it before encrypting or decrypting, a key reported as due keeps serving traffic unchanged, and it has no effect on readiness or liveness. It is computed in one place, from the backend's declared rotation capability plus the key's own timestamps, so no two backends can disagree about what "overdue" means — and a backend that cannot rotate is reported as `unsupported` rather than being told to do something it cannot.
`GET /rustfs/admin/v3/kms/keys/{key_id}` does **not** carry these fields. Its response type records a creation date but no rotation timestamp, so a verdict computed there could not tell a key rotated last week from one never rotated at all, and reporting `never_rotated` for a key that was in fact rotated would be worse than reporting nothing. Read the verdict from the listing.
Driving the rotation itself remains external: call `POST /rustfs/admin/v3/kms/keys/rotate` from your own scheduler, having first satisfied the upgrade-ordering constraint below — and only on the backend where that is the right scheduling model; see [Rotation drivers and scheduling, per backend](#rotation-drivers-and-scheduling-per-backend).
`GET /rustfs/admin/v3/kms/keys` carries `rotation_due` and `rotation_due_reason` per key: `age` (rotated, but longer ago than the period), `never_rotated` (in use longer than the period, never rotated), `wraps` (wrap budget exceeded; wins over `age` when both hold, because the AES-GCM ceiling is not negotiable), or `unsupported` (the backend cannot rotate). Only backends where RustFS wraps locally and can rotate count wraps (Vault KV2); Transit and AWS report no count. `GET /rustfs/admin/v3/kms/keys/{key_id}` does **not** carry these fields — its response records a creation date but no rotation timestamp, so read the verdict from the listing.
### Vault KV2 versioned retention model
@@ -133,59 +101,53 @@ Decryption loads exactly the version recorded in the envelope and fails closed w
### Retention and destruction preconditions
- Every version record that any stored DEK envelope references must remain readable. The bulk rekey sweep (`POST /rustfs/admin/v3/kms/keys/rekey`, gated on `kms:Rekey`) rewraps stored envelopes onto the current version; until a sweep has completed with zero failures after the last rotation, assume **every** version of a rotated key is referenced: destroying a version record permanently orphans all objects whose DEKs it wrapped. A completed sweep is evidence, not authority — the deletion gate stays the decision point. Replication strips encryption metadata in transit, so a sweep never propagates to a replica site: each site runs its own.
- Version records are ordinary KV v2 secrets under the key subtree. Never run `kv metadata delete` or `kv destroy` against `{prefix}/{key_id}/versions/*`, and do not apply `delete-version-after` or retention tooling to that subtree. RustFS-managed retention does not rely on KV2's own secret versioning (each version record has a single KV revision), so KV `max-versions` settings do not protect or endanger history — but metadata deletion always removes a record entirely.
- Permanent key deletion through RustFS (`force_immediate` after `PendingDeletion`) purges the key's version records together with the key record; that is the only supported way to remove them. It is refused by default: the server must set `RUSTFS_KMS_ALLOW_IMMEDIATE_DELETION=true`, and the request must be a `DELETE` with a JSON body that sets `force_immediate` and echoes the key id back as `confirm_key_id` the query-parameter form (`?force_immediate=true`) is refused outright, whatever the gate is set to. Leave the gate off unless you are actively destroying keys, and turn it off again afterwards — the pending-deletion window plus `CancelKeyDeletion` is the only recovery path for objects encrypted under the key.
- Every version record that any stored DEK envelope references must remain readable. The bulk rekey sweep (`POST /rustfs/admin/v3/kms/keys/rekey`, `kms:Rekey`) rewraps stored envelopes onto the current version; until a sweep has completed with zero failures after the last rotation, assume **every** version of a rotated key is referenced. A completed sweep is evidence, not authority — the deletion gate stays the decision point. Replication strips encryption metadata in transit, so each replica site runs its own sweep.
- Version records are ordinary KV v2 secrets under the key subtree. Never run `kv metadata delete` or `kv destroy` against `{prefix}/{key_id}/versions/*`, and do not apply `delete-version-after` or retention tooling to that subtree. Each version record has a single KV revision, so KV `max-versions` settings neither protect nor endanger history — but metadata deletion removes a record entirely.
- Permanent key deletion through RustFS (`force_immediate` after `PendingDeletion`) purges the key's version records together with the key record; that is the only supported way to remove them. It requires `RUSTFS_KMS_ALLOW_IMMEDIATE_DELETION=true` on the server and a `DELETE` with a JSON body that sets `force_immediate` and echoes the key id as `confirm_key_id`; the query-parameter form is refused outright. Leave the gate off except while actively destroying keys — the pending-deletion window plus `CancelKeyDeletion` is the only recovery path for objects encrypted under the key.
- `force_immediate` is refused with `409 Conflict` while any bucket's default encryption configuration names the key or the key is the KMS service default key. A scheduled deletion is not refused for that reason: it destroys nothing and stays cancellable, and the background sweep re-checks the same references before destroying material.
- For Vault Transit, retention is governed by the Transit key's `min_decryption_version`: never raise it above the oldest version that may still protect live ciphertext.
- `force_immediate` is additionally refused, with a `409 Conflict`, while any bucket's default encryption configuration still names the key, or while the key is the KMS service default key. A scheduled deletion is not refused for that reason: it destroys nothing and stays cancellable, and the background sweep re-checks the same references before it destroys the material.
### Reading the `impact` section
`DeleteKey` responses always carry an `impact` section listing the configuration that currently points at the key — the buckets whose default encryption names it, and whether it is the service default key — so the references that will refuse the destruction are visible when the deletion is scheduled rather than only in a server-side log once the window has run out.
`DeleteKey` responses always carry an `impact` section listing the configuration that points at the key (buckets whose default encryption names it, and whether it is the service default key). `DescribeKey` (`GET /rustfs/admin/v3/kms/keys/{key_id}`) returns the same section only when asked with `impact=true`, because collecting it lists every bucket; a value other than `true`/`false` is rejected with `400`. **An absent section means "not collected", never "nothing references this key".**
`DescribeKey` (`GET /rustfs/admin/v3/kms/keys/{key_id}`) can return the same section, but only when the request asks for it with `impact=true`. It is opt-in there because collecting it lists every bucket and `DescribeKey` is polled; without the parameter the endpoint does exactly the work it did before and returns no `impact` field at all. A value other than `true` or `false` is rejected with `400` rather than treated as `false`, so a typo can never answer a request for the section with a response that merely lacks one. **An absent section means "not collected", never "nothing references this key".**
Read it for what it says and nothing more. `coverage.scanned` names the sources that were read; `coverage.not_scanned` names the ones that were not, which currently includes every object encrypted under the key. `completeness` is `exact` only over the scanned sources, and `unavailable` when a source could not be read at all — an unavailable report is not an empty one, and both an unreadable source and an outstanding reference will stop the sweep from destroying the material.
**An empty `references` list does not mean the key is unused.** No object metadata is consulted, so a key with no configuration references can still protect an arbitrary amount of live data, which stays readable only until the material is gone. There is no field in the response that asserts otherwise, and none should be inferred from one.
`coverage.scanned` names the sources that were read and `coverage.not_scanned` the ones that were not — which currently includes every object encrypted under the key. `completeness` is `exact` only over the scanned sources and `unavailable` when a source could not be read; both an unreadable source and an outstanding reference stop the sweep from destroying material. **An empty `references` list does not mean the key is unused:** no object metadata is consulted, so a key with no configuration references can still protect live data.
### Upgrade before first rotation (hard constraint)
Do not rotate any key until **every** RustFS node in the cluster runs a build that understands the `master_key_version` envelope field. Older binaries ignore the field and always decrypt with the current material: harmless while nothing has been rotated, but after a rotation they will fail to decrypt every object wrapped by an earlier key version. Complete the rolling upgrade of the entire cluster first, then rotate.
This is the sharpest instance of a broader class of constraints; the rest are collected in [Mixed-version clusters during a rolling upgrade](#mixed-version-clusters-during-a-rolling-upgrade).
Do not rotate any key until **every** RustFS node runs a build that understands the `master_key_version` envelope field. Older binaries ignore the field and always decrypt with the current material: harmless while nothing has been rotated, but after a rotation they fail to decrypt every object wrapped by an earlier key version. Complete the rolling upgrade of the entire cluster first, then rotate. The rest of this constraint class is collected in [Mixed-version clusters during a rolling upgrade](#mixed-version-clusters-during-a-rolling-upgrade).
## Mixed-version clusters during a rolling upgrade
During a rolling upgrade the cluster runs two RustFS builds at once. That window matters more for KMS than for most subsystems, because KMS state is shared three ways: **Vault** holds the key records and Transit metadata, **cluster storage** holds the persisted KMS configuration, and **each node's process memory** holds caches and the live backend instance. Nodes on different builds agree on the first, may disagree on the third, and — for configuration — can disagree for as long as the operator leaves them running, because the reload broadcast that converges configuration is one of the things an older build rejects.
This section states only what is true of the current implementation. It is written for the KV2 and Transit backends; the Local backend is unsupported for multi-node deployments regardless of version (see the [deployment support matrix](#deployment-support-matrix)).
During a rolling upgrade KMS state is shared three ways: **Vault** holds key records and Transit metadata, **cluster storage** holds the persisted KMS configuration, and **each node's process memory** holds caches and the live backend instance. Nodes on different builds agree on the first, may disagree on the third, and can disagree on configuration for as long as the operator leaves them running, because the reload broadcast that converges configuration is one of the things an older build rejects. This section is written for the KV2 and Transit backends; the Local backend is unsupported for multi-node deployments regardless of version (see the [deployment support matrix](#deployment-support-matrix)).
### Persisted formats are backward compatible in both directions
Nothing in this list requires a coordinated format cutover. The compatibility is deliberate and is covered by decode tests.
No coordinated format cutover is required; the compatibility is deliberate and covered by decode tests.
- **DEK envelopes.** `DataKeyEnvelope::master_key_version` is optional and omitted when absent, so envelopes written by non-rotating backends stay byte-identical to the historical seven-field JSON shape. An upgraded node reading a pre-versioning envelope resolves `None` to the key's recorded baseline version, or — for a key that was never rotated, and so has no baseline — to the current version, which is exactly the pre-versioning behavior. Unknown values are skipped while parsing; a bounded field-name sample is emitted at a progressively rate-limited `warn` level, and `rustfs_kms_persisted_unknown_fields_total{record_kind="data-key-envelope"}` counts every observed field.
- **Local key records.** Each `<key_id>.key` record carries `format_version: 1`; records written before that field existed default to version 1 when read, and the pre-version reader ignores the added v1 marker. A reader accepts a record whose version is at most the version it understands, and rejects a newer version with `UnsupportedFormatVersion` before it attempts to decrypt key material. Unknown fields remain accepted for rollback compatibility. Their values are ignored while parsing; a bounded field-name sample is emitted at a progressively rate-limited `warn` level, and `rustfs_kms_persisted_unknown_fields_total{record_kind="local-key-record"}` counts every observed field. Once a future version greater than 1 has written a key record, do not roll back to a build that predates this marker: such a build cannot reject that future version before interpreting the rest of the record.
- **KV2 key records.** `baseline_version` is read with a serde default, so records written by older builds deserialize unchanged, and `None` correctly means "never rotated".
- **Transit metadata records.** Metadata persisted in KV v2 by either build decodes on the other.
The one-way hazard is the rotation constraint above: an older binary reading a *new* envelope silently ignores the version field and decrypts with the current material.
| Record | Compatibility mechanism | Caveat |
| --- | --- | --- |
| DEK envelopes | `DataKeyEnvelope::master_key_version` is optional and omitted when absent, so envelopes from non-rotating backends stay byte-identical to the historical seven-field JSON. An upgraded node resolves a pre-versioning envelope to the key's `baseline_version`, or to the current version for a never-rotated key. Unknown fields are skipped; a bounded field-name sample is logged at a rate-limited `warn` and counted by `rustfs_kms_persisted_unknown_fields_total{record_kind="data-key-envelope"}` | An older binary reading a *new* envelope ignores the version field and decrypts with the current material — the rotation constraint above |
| Local key records | Each `<key_id>.key` carries `format_version: 1`; records without it default to 1. A reader accepts a version at most the one it understands and rejects newer with `UnsupportedFormatVersion` before decrypting. Unknown fields are accepted and counted by `rustfs_kms_persisted_unknown_fields_total{record_kind="local-key-record"}` | Once a future version greater than 1 has written a record, do not roll back to a build predating the marker |
| KV2 key records | `baseline_version` is read with a serde default; `None` means "never rotated" | An old build drops the field on write-back (see below) |
| Transit metadata records | Decode on either build | — |
### DEK envelope context binding (`RUSTFS_KMS_ENVELOPE_AAD`)
Historically the KV2 and Local backends sealed only the DEK plaintext; the `encryption_context` rode in the envelope unauthenticated and was checked by field comparison alone, so a party able to rewrite the stored envelope could rewrite the context to match whatever it presented. With `RUSTFS_KMS_ENVELOPE_AAD=true`, newly wrapped envelopes bind the canonical context bytes as AES-GCM additional data and carry `context_binding: 1`; rewriting the stored context, or stripping the flag, then fails authentication. (Static, Vault Transit and AWS already bound the context through their own mechanisms and are unaffected.)
Historically the KV2 and Local backends sealed only the DEK plaintext; the `encryption_context` rode in the envelope unauthenticated and was checked by field comparison alone, so a party able to rewrite the stored envelope could rewrite the context. With `RUSTFS_KMS_ENVELOPE_AAD=true`, newly wrapped envelopes bind the canonical context bytes as AES-GCM additional data and carry `context_binding: 1`; rewriting the stored context, or stripping the flag, then fails authentication. Static, Vault Transit and AWS already bound the context through their own mechanisms and are unaffected.
Rollout constraint: **reading bound envelopes needs no switch, but a node that predates the field cannot open them** — its unwrap runs without the additional data and fails authentication. Enable the switch only after every node in the cluster runs a release that understands `context_binding`; the default stays off for one release for exactly this reason, mirroring the `RUSTFS_ENCRYPTION_FRAME_V2` rollout. Rewrap migrates existing envelopes: with the switch on, a rewrap sweep upgrades unbound envelopes to the bound format (converging to zero writes on re-run), and a bound envelope never regresses to the unbound shape whatever the switch says. An envelope carrying an unrecognized `context_binding` value is refused rather than decrypted without its binding.
Rollout constraint: reading bound envelopes needs no switch, but **a node that predates the field cannot open them** — its unwrap runs without the additional data and fails authentication. The switch defaults off (`ENV_KMS_ENVELOPE_AAD` in `crates/kms/src/config.rs`); enable it only after every node runs a release that understands `context_binding`, mirroring the `RUSTFS_ENCRYPTION_FRAME_V2` rollout. With the switch on, a rewrap sweep upgrades unbound envelopes to the bound format (converging to zero writes on re-run); a bound envelope never regresses to the unbound shape, and an envelope carrying an unrecognized `context_binding` value is refused rather than decrypted without its binding.
### Guarantees that hold only once every node is upgraded
These are properties of the upgraded code, so a single node left behind removes them for the whole cluster.
These are properties of builds from `1.0.0-rc.1` onward; a single older node removes them for the whole cluster.
- **Check-and-set lifecycle writes.** Upgraded builds write every KV2 lifecycle mutation — create, enable, disable, tag metadata, schedule deletion, cancel deletion — as a versioned read followed by a check-and-set write, retrying on conflict by re-reading and re-validating the state gate (rustfs/rustfs#5518). Transit metadata writes got the same treatment (rustfs/rustfs#5520). Builds older than those write blind. A blind write from an old node can overwrite a check-and-set commit from an upgraded node without any conflict being reported, which is precisely the lost update the change was made to eliminate.
- **`baseline_version` survives a write-back.** The KV2 key record does not deny unknown fields, so an old build reads a new record without error — and drops `baseline_version` when it writes that record back for any reason. A key that loses its baseline resolves pre-versioning envelopes to the current version again, which after a rotation means the wrong master key material. Any lifecycle operation issued to an old node is enough to trigger this.
- **`wrap_budget_reserved` keeps overestimating.** The KV2 key record's approximate wrap counter (`wrap_budget_reserved`, behind the `rustfs_kms_max_key_wrap_operations` gauge) is dropped the same way when an old build rewrites the record, regressing the count toward zero — the one way this deliberately overestimate-only counter can understate the wraps actually performed. Nothing breaks: the counter is advisory, and the next block reservation from an upgraded node re-establishes a floor. Just do not trust a *low* gauge reading taken during or shortly after a mixed-version window.
- **Version-record awareness.** Rotation stores each historical version under `{prefix}/{key_id}/versions/{N}` as a create-only record (check-and-set of 0), so two nodes racing the same version number produce exactly one creator; the loser adopts the persisted, never-current material or fails without touching the current pointer. Old builds have no concept of that sub-path: they never read or write it, and their key listing reports the KV2 directory entry (`my-key/`) as though it were a key, because the directory filter only exists in upgraded builds.
| Guarantee | Upgraded behaviour | What an old node does |
| --- | --- | --- |
| Check-and-set lifecycle writes | Every KV2 lifecycle mutation (create, enable, disable, tag, schedule/cancel deletion) and every Transit metadata write is a versioned read followed by a check-and-set write, retried on conflict | Writes blind, so it can overwrite a check-and-set commit without any conflict being reported — the lost update the change eliminated |
| `baseline_version` survives write-back | Preserved | Reads the record without error and drops `baseline_version` on any write-back; the key then resolves pre-versioning envelopes to the current version, which after a rotation is the wrong material |
| `wrap_budget_reserved` only overestimates | The KV2 record's wrap counter (behind `rustfs_kms_max_key_wrap_operations`) is reserved in blocks and never understates | Drops the field on write-back, regressing the count toward zero. Nothing breaks — the counter is advisory and the next reservation re-establishes a floor — but do not trust a *low* reading taken during or shortly after a mixed-version window |
| Version-record awareness | Version records under `{prefix}/{key_id}/versions/{N}` are create-only (check-and-set of 0), so two nodes racing a version number produce exactly one creator | Never reads or writes the sub-path, and its key listing reports the KV2 directory entry (`my-key/`) as though it were a key |
### Windows in which nodes can legitimately disagree
@@ -193,79 +155,65 @@ Even with every node on the same build, some state is process-local. These windo
| What can diverge | Bound | Mechanism |
| --- | --- | --- |
| Transit key lifecycle state used by the `encrypt` and `generate_data_key` gates | ≤ 300 s (`METADATA_CACHE_TTL`) | Each node caches Transit metadata in process, TTL- and capacity-bounded, with targeted invalidation when a data-path call reports the key is gone server-side. A disable or schedule-deletion performed on one node is enforced on the others within one TTL at the latest, sooner if they hit that signal. |
| `describe_key` output | One metadata cache TTL: 300 s by default, otherwise whatever `cache_ttl_seconds` was configured with, clamped to 24 h | The manager-level key metadata cache, built from the configured cache settings. This is a reporting cache; the KV2 state gates do not read it. |
| KV2 key lifecycle state | None | The KV2 backend re-reads the key record from Vault for every lifecycle and data-key operation, so a committed disable is effective on every upgraded node immediately. |
| Active KMS configuration | One best-effort reload broadcast; unbounded for any peer that did not apply it | See below. |
Builds older than rustfs/rustfs#5520 held the Transit metadata cache with no TTL and no capacity bound. On such a node the divergence window is not 300 seconds but "until the process restarts": it can keep encrypting under a key that another node disabled, indefinitely.
The `describe_key` bound is the only one on that list an operator sets, so compute it rather than assuming the default: the window is the `cache_ttl_seconds` the KMS configure request was given, 300 s when it was omitted, clamped down to 24 h at use if it is larger (clamped rather than rejected, so an oversized setting still starts). Zero is refused outright while caching is enabled. `kms service-status` and the KMS configuration endpoint report the effective, post-clamp value, so the number the admin API shows is the number the cache honours. Note that this is the Transit row's neighbour and not its equal: `METADATA_CACHE_TTL` above is a separate, deliberately non-tunable 300 s, because that cache does gate cryptographic operations.
One upgrade caveat: builds older than rustfs/rustfs#5569 ignored `cache_ttl_seconds` and ran a hardcoded 300 s, while their configure converters persisted 3600 s as the default value. A cluster configured through the admin API before that fix therefore widens its `describe_key` staleness window from an effective 300 s to the 3600 s already stored in `config/kms_config.json`, with no configuration change of its own. Read the reported value back after upgrading instead of assuming it stayed at 300 s. No cryptographic or authorization path widens with it — encrypt, decrypt and data-key generation go straight to the backend and never read this cache.
| Transit key lifecycle state used by the `encrypt` and `generate_data_key` gates | ≤ `METADATA_CACHE_TTL` (300 s, not tunable — this cache gates cryptographic operations) | Per-node in-process Transit metadata cache, TTL- and capacity-bounded, with targeted invalidation when a data-path call reports the key gone server-side. Builds older than `1.0.0-rc.1` held this cache with no TTL and no capacity bound: on such a node the window is "until the process restarts" |
| `describe_key` output | One metadata cache TTL: `cache_ttl_seconds` from the KMS configure request, 300 s when omitted, clamped down to 24 h at use (clamped rather than rejected; zero is refused while caching is enabled) | Manager-level key metadata cache; a reporting cache the KV2 state gates never read. `kms service-status` and the configuration endpoint report the effective post-clamp value. Builds older than `1.0.0-rc.1` ignored `cache_ttl_seconds` and ran 300 s while persisting 3600 s as the default, so a cluster configured before then widens to the stored 3600 s on upgrade — read the reported value back |
| KV2 key lifecycle state | None | The KV2 backend re-reads the key record from Vault for every lifecycle and data-key operation |
| Active KMS configuration | One best-effort reload broadcast; unbounded for any peer that did not apply it | See below |
### Configuration changes converge through a best-effort peer reload
`POST /rustfs/admin/v3/kms/configure` and `POST /rustfs/admin/v3/kms/reconfigure` persist the new configuration to cluster storage at `config/kms_config.json`, switch the KMS service **on the node that handled the request**, and then broadcast a reload signal to every peer. A peer that accepts the signal re-reads the persisted configuration and reconfigures itself, so a runtime change normally reaches the whole cluster without any restart. A peer already running that exact configuration treats the signal as a no-op.
`POST /rustfs/admin/v3/kms/configure` and `/kms/reconfigure` persist the new configuration to cluster storage at `config/kms_config.json`, switch the KMS service **on the node that handled the request**, and broadcast a reload signal once to every peer. A peer that accepts the signal re-reads the persisted configuration and reconfigures itself; a peer already running that configuration treats it as a no-op. Convergence is best effort and the request never fails on account of a peer:
Convergence is best effort by contract, and the request never fails on account of a peer: the local node has already switched, and KMS configuration has no quorum or authoritative holder to roll back to. What that leaves:
- There is no background retry. A peer that is unreachable, whose build predates the KMS subsystem, or whose reload fails keeps its previous configuration until a later `reconfigure` reaches it or it restarts. For those peers the split is unbounded.
- The admin response reports success either way, but its message names every peer that did not converge, and the server logs one `kms_peer_config_reload_failed` warning per peer.
- While a split lasts, both configurations are live: if the change switched backends, or changed the Vault mount or key prefix, nodes write new key material to different places and a key created through one node is invisible to the others.
- The broadcast is sent **once**, with no background retry. A peer that is unreachable, that rejects the signal because its build predates the KMS subsystem, or whose reload itself fails keeps serving its previous configuration until a later `reconfigure` reaches it, or until it restarts and loads the persisted configuration during startup. For those peers the split window is still unbounded.
- The admin response reports success either way, but its message names every peer that did not converge, and the server logs one `kms_peer_config_reload_failed` warning per peer. Read the message: an operation that reports success can still have left the cluster split.
- For as long as a split lasts, both configurations are live. If the change switched backends, or changed the Vault mount or key prefix, different nodes write new key material to different places, and a key created through one node is invisible to the others.
`GET /rustfs/admin/v3/kms/service-status` makes the split observable from a single request: it returns a `cluster_config` object holding one redacted configuration fingerprint per node plus a `consistent` flag. `consistent` is true only when every node answered with the same fingerprint — an unreachable peer, a peer whose build reports no fingerprint, and a node with no configuration at all each read as divergent rather than as agreement. Secrets are substituted out before a configuration is fingerprinted, so two nodes on the same backend holding different credentials still fingerprint alike; the field detects a configuration split, not a credential split.
Treat a `configure` or `reconfigure` whose response names unconverged peers as an unfinished cluster-wide operation: re-issue it once those peers are reachable, or restart them.
`GET /rustfs/admin/v3/kms/service-status` returns a `cluster_config` object holding one redacted configuration fingerprint per node plus a `consistent` flag, true only when every node answered with the same fingerprint (an unreachable peer, a build reporting no fingerprint, and an unconfigured node each read as divergent). Secrets are substituted out before fingerprinting, so the field detects a configuration split, not a credential split. Treat a `configure` or `reconfigure` whose response names unconverged peers as unfinished: re-issue it once those peers are reachable, or restart them.
### Recommended rolling upgrade order
Follow the node-at-a-time procedure in the [multi-node restart runbook](rolling-restart.md); this adds the KMS-specific sequencing around it.
Follow the node-at-a-time procedure in the [multi-node restart runbook](rolling-restart.md); this adds the KMS-specific sequencing.
1. **Freeze KMS administrative traffic** for the duration: no key creation, enable, disable, tagging, schedule-deletion, cancel-deletion, rotation, or reconfiguration. Object read and write traffic continues normally.
2. **Upgrade one node at a time**, waiting for each to report ready before starting the next.
3. **Verify no node is left behind** before unfreezing. A single old node is enough to reintroduce blind writes and to strip `baseline_version` on its next lifecycle write.
3. **Verify no node is left behind** before unfreezing. A single old node reintroduces blind writes and strips `baseline_version` on its next lifecycle write.
4. **Resume administrative traffic.**
5. **Only then perform the first rotation of any key.** Once the whole cluster understands `master_key_version`, rotation is safe; before that it is not.
6. **If the KMS configuration was changed at any point**, confirm `cluster_config.consistent` is true in the `service-status` response, and re-issue the change — or restart the node — for every peer still reporting a different fingerprint. A peer whose build predates the reload signal never converges on its own.
5. **Only then perform the first rotation of any key.**
6. **If the KMS configuration was changed at any point**, confirm `cluster_config.consistent` is true in the `service-status` response, and re-issue the change — or restart the node — for every peer still reporting a different fingerprint.
### Do not do these during a mixed-version window
- **Rotate any key.** This is the hard constraint stated above; a rotation is unrecoverable for objects an old node must read.
- **Issue any KV2 lifecycle write to an old node.** Its blind write can clobber a concurrent check-and-set commit and will drop `baseline_version` from the record.
- **Create the same key ID from two nodes.** The create path is create-only on upgraded builds, but an old node's blind write does not honor that: the later writer's material wins and every DEK already wrapped with the earlier material becomes permanently unwrappable.
- **Assume a disable or schedule-deletion took effect cluster-wide.** Old Transit nodes cache lifecycle state without expiry; confirm per node, or restart the old nodes, before treating a key as no longer in use.
- **Reconfigure the KMS backend and consider it done.** The reload broadcast is exactly what an old build rejects, so during a mixed-version window the change reaches only the node that served it and the already-upgraded peers. Check the response message and `cluster_config.consistent` before assuming otherwise.
- **Rotate any key.** Unrecoverable for objects an old node must read.
- **Issue any KV2 lifecycle write to an old node.** Its blind write can clobber a concurrent check-and-set commit and drops `baseline_version`.
- **Create the same key ID from two nodes.** The create path is create-only on upgraded builds, but an old node's blind write does not honor that: the later writer's material wins and every DEK wrapped with the earlier material becomes permanently unwrappable.
- **Assume a disable or schedule-deletion took effect cluster-wide.** Old Transit nodes cache lifecycle state without expiry; confirm per node, or restart the old nodes.
- **Reconfigure the KMS backend and consider it done.** The reload broadcast is exactly what an old build rejects; check the response message and `cluster_config.consistent`.
- **Delete or prune version records** under `{prefix}/{key_id}/versions/*` for any reason. This is never safe, mixed-version or not; see [Retention and destruction preconditions](#retention-and-destruction-preconditions).
## Choosing between Vault KV2 and Vault Transit
Use **Vault Transit** (`VaultTransit`) when key material must be cryptographically isolated from anyone holding storage-level read access: Transit keeps key-encryption keys inside Vault and only ever returns ciphertext, and supports server-side key versioning/rotation.
Use **Vault KV2** only when you accept that the Vault ACL on the key path *is* the confidentiality boundary and you want the operational simplicity of a single KV mount.
Use **Vault Transit** (`VaultTransit`) when key material must be cryptographically isolated from anyone holding storage-level read access: Transit keeps key-encryption keys inside Vault, only ever returns ciphertext, and supports server-side key versioning and rotation. Use **Vault KV2** only when you accept that the Vault ACL on the key path *is* the confidentiality boundary and want the operational simplicity of a single KV mount.
## AWS KMS: deviations from the shared backend contract
Select it with `RUSTFS_KMS_BACKEND=aws`. Credentials and region resolution are delegated entirely to the standard `aws-config` provider chain (environment, shared profile, container/IMDS role), so RustFS never stores, persists, or redacts AWS credential material of its own. Only two non-credential settings are read: `RUSTFS_KMS_AWS_REGION` and `RUSTFS_KMS_AWS_ENDPOINT_URL`. A plaintext (`http://`) endpoint override would expose every KMS request including plaintext data keys, so it is refused unless the development opt-in is set.
Select it with `RUSTFS_KMS_BACKEND=aws`. Credentials and region resolution are delegated entirely to the standard `aws-config` provider chain (environment, shared profile, container/IMDS role), so RustFS never stores, persists, or redacts AWS credential material. Only two non-credential settings are read: `RUSTFS_KMS_AWS_REGION` and `RUSTFS_KMS_AWS_ENDPOINT_URL`. A plaintext (`http://`) endpoint override would expose every KMS request including plaintext data keys, so it is refused unless the development opt-in is set.
AWS owns key state, backing-key rotation, and the deletion window, and this backend mirrors none of it locally. That makes four behaviours differ from every RustFS-managed backend. Verify each against your operational assumptions before switching:
AWS owns key state, backing-key rotation, and the deletion window, and this backend mirrors none of it locally. Four behaviours therefore differ from every RustFS-managed backend:
| Behaviour | RustFS-managed backends | AWS KMS backend |
| --- | --- | --- |
| Decryption with a `Disabled` or `PendingDeletion` key | Kept working, so disabling a key never breaks reads of objects already encrypted under it | **Refused by AWS.** Objects encrypted under a key that is later disabled become unreadable until it is re-enabled |
| Key deletion | Physical deletion available | **No physical delete.** `ScheduleKeyDeletion` is the only removal path; AWS destroys the material when the 7-30 day window elapses. RustFS never destroys AWS-held material, and `force_immediate` is refused |
| Cancelling a scheduled deletion | Key returns to `Enabled` | Key is left **`Disabled`**; enable it explicitly to make it usable again |
| Creating a key under a caller-chosen name | The requested name becomes the key id | **Refused.** AWS assigns identifiers and this backend does not manage aliases, so a named create would produce a key unreachable by that name |
| Key deletion | Physical deletion available | **No physical delete.** `ScheduleKeyDeletion` is the only removal path; AWS destroys the material when the 7-30 day window elapses. `force_immediate` is refused |
| Cancelling a scheduled deletion | Key returns to `Enabled` | Key is left **`Disabled`**; enable it explicitly |
| Creating a key under a caller-chosen name | The requested name becomes the key id | **Refused.** AWS assigns identifiers and this backend does not manage aliases |
Two consequences follow from that last row: **SSE-S3 key auto-creation and the synthetic KMS probe are unavailable on this backend**, because both address a key by a name they choose. Pre-create keys in AWS and reference them by AWS key id or ARN.
Consequences of the last row: **SSE-S3 key auto-creation and the synthetic KMS probe are unavailable on this backend**, because both address a key by a name they choose. Pre-create keys in AWS and reference them by AWS key id or ARN.
The AWS backend is intentionally exempt from `backends::contract_tests::assert_state_machine_contract`. That shared driver assumes that disabled and pending-deletion keys still decrypt, that cancelling deletion returns a key to `Enabled`, and that creation accepts a caller-assigned key name. AWS rejects decryption for the first case, leaves a cancelled key `Disabled`, and assigns key identifiers itself, so running the driver would encode the wrong behavior. The exemption is pinned by the offline `aws_backend_shared_contract_exemption_is_pinned` test in `crates/kms/src/backends/aws.rs`; if AWS changes any of these semantics, integrate the backend into the shared driver and remove this exemption rather than weakening the shared assertions.
The AWS backend is exempt from `backends::contract_tests::assert_state_machine_contract`, whose assumptions (disabled keys still decrypt, cancel returns to `Enabled`, caller-assigned names) AWS violates. The exemption is pinned by the offline `aws_backend_shared_contract_exemption_is_pinned` test in `crates/kms/src/backends/aws.rs`; if AWS changes any of these semantics, integrate the backend into the shared driver and remove the exemption rather than weakening the shared assertions.
Key versions are opaque. AWS addresses backing keys internally and picks the right one to decrypt with, so RustFS reports `key_version` as 1 and cannot enumerate versions. Rotation uses `RotateKeyOnDemand`, which retains prior backing keys for decryption; AWS's separate automatic yearly rotation is neither enabled nor reported on by RustFS.
Key versions are opaque: RustFS reports `key_version` as 1 and cannot enumerate versions. Rotation uses `RotateKeyOnDemand`, which retains prior backing keys for decryption; AWS's automatic yearly rotation is neither enabled nor reported on by RustFS.
The KMS admin API accepts the AWS backend as `"backend_type": "AWS"` (aliases `aws`, `aws-kms`, `aws_kms`, `AwsKms`) on `/v3/kms/configure` and `/v3/kms/reconfigure`. The body carries `region` (**required**), and optionally `endpoint_url`, `default_key_id`, and the shared timeout/retry/cache settings. It accepts no credential fields at all — unknown fields are rejected because every node resolves credentials through its own provider chain.
`region` is mandatory on this path even though `RUSTFS_KMS_AWS_REGION` is optional at startup: the admin configuration is persisted once and replayed on every node, so a request that left the region to each node's ambient chain would let nodes address different regions, and therefore different keys, while reporting an identical configuration. `default_key_id` must be an AWS key id or ARN that already exists — this backend never creates keys by name.
The KMS admin API accepts the backend as `"backend_type": "AWS"` (aliases `aws`, `aws-kms`, `aws_kms`, `AwsKms`) on `/v3/kms/configure` and `/v3/kms/reconfigure`. The body carries `region` (**required**), and optionally `endpoint_url`, `default_key_id`, and the shared timeout/retry/cache settings; credential fields are rejected as unknown, because every node resolves credentials through its own provider chain. `region` is mandatory even though `RUSTFS_KMS_AWS_REGION` is optional at startup: the admin configuration is replayed on every node, and a request that left the region to each node's ambient chain would let nodes address different regions — and therefore different keys — under an identical configuration. `default_key_id` must be an AWS key id or ARN that already exists.
## Vault TLS: custom CA and mutual TLS
@@ -273,40 +221,34 @@ Both Vault backends (KV2 and Transit) support a private certificate authority an
| Setting | Environment variable | Admin configure field | Meaning |
| --- | --- | --- | --- |
| CA bundle | `RUSTFS_KMS_VAULT_CA_CERT` | `ca_cert_path` | Path to a PEM CA bundle trusted for the Vault connection, in addition to nothing else: when set, only this bundle is trusted |
| CA bundle | `RUSTFS_KMS_VAULT_CA_CERT` | `ca_cert_path` | Path to a PEM CA bundle; when set, only this bundle is trusted |
| Client certificate | `RUSTFS_KMS_VAULT_CLIENT_CERT` | `client_cert_path` | Path to a PEM client certificate presented to Vault; requires the client key |
| Client key | `RUSTFS_KMS_VAULT_CLIENT_KEY` | `client_key_path` | Path to the PEM private key matching the client certificate |
| Skip verification | `RUSTFS_KMS_VAULT_SKIP_TLS_VERIFY` | `skip_tls_verify` | Disables server certificate verification; gated on the insecure development defaults opt-in |
Paths are read on the node applying the configuration, so the files must exist at the same path on every node. The certificate and key must be configured together; configuration validation rejects one without the other, and the files are read and parsed when the backend starts, so a bad path or malformed PEM fails the configuration instead of a later request. The `kms/status` backend summary reports `has_custom_ca` and `has_client_identity` booleans (never the file contents).
The Vault client library would otherwise fall back to the `VAULT_CACERT`, `VAULT_CAPATH`, `VAULT_CLIENT_CERT` and `VAULT_CLIENT_KEY` process environment variables; RustFS always sets the trust roots and identity explicitly - to the configured values or to empty - so stray Vault environment variables cannot splice TLS material into the connection behind the KMS configuration.
Paths are read on the node applying the configuration, so the files must exist at the same path on every node. Certificate and key must be configured together; the files are read and parsed when the backend starts, so a bad path or malformed PEM fails the configuration rather than a later request. The `kms/status` backend summary reports `has_custom_ca` and `has_client_identity` booleans, never file contents. RustFS always sets the trust roots and identity explicitly — to the configured values or to empty — so the `VAULT_CACERT`, `VAULT_CAPATH`, `VAULT_CLIENT_CERT` and `VAULT_CLIENT_KEY` process environment variables cannot splice TLS material into the connection behind the KMS configuration.
## Local backend durability and deployment support matrix
The Local backend stores one JSON record per key (`<key_id>.key`) plus an Argon2id salt file (`.master-key.salt`) inside the configured `key_dir`. This section documents which deployments that layout supports and how the backend recovers from a crash or power loss. For where the key material lives and who can read it, see the [backend comparison](#backend-comparison) above.
The Local backend stores one JSON record per key (`<key_id>.key`) plus an Argon2id salt file (`.master-key.salt`) inside the configured `key_dir`. For where the key material lives and who can read it, see the [backend comparison](#backend-comparison).
### Positioning
The facts today:
- `Local` is the current default backend (`kms_backend` defaults to `local`).
- The RustFS Kubernetes operator places the key directory on a PersistentVolumeClaim, so the keys survive pod rescheduling.
- The in-code documentation labels the backend "for development and testing only", and configuration validation enforces stricter rules outside explicit development mode: a master key is required and `key_dir` must not live under the process temp directory.
- `Local` is the default backend (`kms_backend` defaults to `local`) and is a development, testing and demo backend; it is not supported for production. Activating a backend whose capabilities report `production_supported: false` logs a `kms_backend_positioning` warning on every start, restart and reconfigure, and the `kms/status` capability matrix carries the same flag. The positioning is a warning, not a gate.
- Configuration validation enforces stricter rules outside explicit development mode: a master key is required and `key_dir` must not live under the process temp directory.
- The RustFS Kubernetes operator places the key directory on a PersistentVolumeClaim, so keys survive pod rescheduling.
- Production multi-node deployments should use the Vault Transit backend.
The backend's positioning is settled (owner decision, 2026-08): `Local` is a development, testing and demo backend and is not supported for production. The runtime now states this itself — activating a backend whose capabilities report `production_supported: false` logs a `kms_backend_positioning` warning on every start, restart and reconfigure, and the `kms/status` capability matrix carries the same flag for consoles and tooling. Existing deployments are not blocked: the positioning is a warning, not a gate. This section describes what the implementation guarantees for those who accept that positioning.
### Deployment support matrix
| Deployment | Supported | Notes |
| --- | --- | --- |
| Local filesystem (ext4, XFS, APFS, ...) | Yes | The commit protocol relies on POSIX `rename`/`hard_link` atomicity and `fsync` durability, which local filesystems provide |
| Local filesystem (ext4, XFS, APFS, ...) | Yes | The commit protocol relies on POSIX `rename`/`hard_link` atomicity and `fsync` durability |
| Kubernetes PVC | Yes | Only when the PersistentVolume is backed by a local or block filesystem; this is how the RustFS operator provisions the key directory |
| NFS or other shared/network filesystems | No | Network filesystems do not reliably provide the atomicity and fsync semantics the commit protocol depends on; an NFS-backed PersistentVolume is this case, not the PVC case above |
| Multiple RustFS processes sharing one `key_dir` | No | Concurrent key **creation** is linearized (`hard_link` refuses to clobber an existing key), but every other write — status updates, deletion, cancellation — is a read-modify-write with no cross-process lock, so concurrent writers can silently lose updates |
| NFS or other shared/network filesystems | No | Network filesystems do not reliably provide the atomicity and fsync semantics the protocol depends on; an NFS-backed PersistentVolume is this case |
| Multiple RustFS processes sharing one `key_dir` | No | Concurrent key **creation** is linearized (`hard_link` refuses to clobber), but every other write is a read-modify-write with no cross-process lock, so concurrent writers can silently lose updates |
Within a single process, per-key write locks serialize read-modify-write updates, so concurrent API calls against one RustFS instance are safe.
Within a single process, per-key write locks serialize read-modify-write updates.
### Crash recovery behavior
@@ -317,27 +259,23 @@ Every mutation of the key directory uses a durable commit protocol:
3. The file is published atomically: `rename` to replace an existing file, `hard_link` to create a new one without clobbering.
4. The parent directory is fsynced so the new directory entry is durable.
Deletion mirrors the tail of the protocol (`remove_file` followed by a parent directory fsync), so a deleted key cannot resurface after power loss. A crash at any step leaves either the complete old state or the complete new state, plus at most an unpublished temp file.
Deletion mirrors the tail of the protocol (`remove_file` followed by a parent directory fsync). A crash at any step leaves either the complete old state or the complete new state, plus at most an unpublished temp file. On startup the backend:
On startup the backend then:
- **Removes orphaned commit temp files.** The matcher is strict (`<prefix>.tmp-<uuid>`, never anything ending in `.key`), so published key files — including a key the user named to look like a temp file — are never touched. Publishing is atomic, so a matching leftover can only be an unpublished remnant of an interrupted commit.
- **Removes orphaned commit temp files.** The matcher is strict (`<prefix>.tmp-<uuid>`, never anything ending in `.key`), so published key files are never touched.
- **Validates every published `.key` file.** A record that fails to decode fails startup rather than being silently skipped.
- **Guards the salt file.** If `.master-key.salt` is missing but the directory contains keys marked `encrypted-master-key`, initialization fails closed with a configuration error naming the salt path. A regenerated salt derives a different master key and can never decrypt those keys, so the correct recovery is to **restore the salt file (or the whole directory) from backup**, never to let a fresh salt be generated. The guard is equally strict about a record it cannot read or cannot interpret for example one written by a newer RustFS that names an at-rest protection this build does not implement: such a directory's protection state is unknown, so no replacement salt is generated for it either. Recovery is to restore the salt file, run a build that understands the record, or move the unrecognized file out of `key_dir` after confirming it is not needed. An empty directory, or a legacy directory predating the salt file, still initializes normally.
- **Guards the salt file.** If `.master-key.salt` is missing but the directory contains keys marked `encrypted-master-key`, initialization fails closed naming the salt path. A regenerated salt derives a different master key and can never decrypt those keys, so the recovery is to **restore the salt file (or the whole directory) from backup**, never to let a fresh salt be generated. The guard is equally strict about a record it cannot read or interpret (for example one written by a newer RustFS naming an at-rest protection this build does not implement): no replacement salt is generated for such a directory either. An empty directory, or a legacy directory predating the salt file, initializes normally.
### Filesystem permissions and the boundaries the protocol assumes
The key directory is held at `0o700` and every file published into it — key records, the salt, and the files a restore stages and cuts over — is written owner-only. The requested mode is applied and re-read on the open file *before* the content becomes durable, so the process umask cannot widen it, and an unspecified `file_permissions` resolves to owner-only inside the commit protocol rather than at each call site, so no write path can leave it to the umask.
- The key directory is held at `0o700` and every file published into it — key records, the salt, restore staging — is written owner-only. The requested mode is applied and re-read on the open file before the content becomes durable, so the process umask cannot widen it.
- A directory wider than `0o700` is **narrowed on every start** (and re-read to confirm), not refused: kubelet creates `emptyDir` at `0o777`, several PVC provisioners `mkdir -m 0777`, and a `--tmpfs` mount lands at `1777`. Only a directory this process cannot secure is fatal. Narrowing is logged with the previous mode whenever it was reachable beyond the owner.
- Publishing never writes through a symlink: `hard_link` refuses any existing destination (including a dangling symlink) and `rename` replaces the link itself. Startup removes anything wearing a commit-temp name that is not a directory, symlinks included; the protocol only ever creates temps with `create_new`, so such an entry is either its own leftover or something planted.
A directory wider than `0o700` is **narrowed on every start**, and the result is re-read to confirm it took effect. It is not refused: the mode is far more often the platform's than the operator's — kubelet creates an `emptyDir` `0o777`, several PVC provisioners `mkdir -m 0777`, a `--tmpfs` mount lands at `1777` — and refusing would turn each of those into a server that will not start while leaving the exposure in place on the way out. Narrowing removes it. Only a directory this process cannot secure is fatal, because at that point the mode is both dangerous and outside our control. Narrowing is logged with the previous mode whenever it was reachable beyond the owner.
Two boundaries are **not** verified, and deployments should not assume them:
Publishing never writes through a symlink. `hard_link` refuses any destination that already exists — including a dangling symlink — so a create cannot adopt an inode it did not write, and `rename` replaces the link itself rather than the file it points at. Startup removes anything wearing a commit-temp name that is not a directory, symlinks included; the protocol only ever creates temps with `create_new`, so such an entry is either its own leftover or something planted.
Two boundaries in this area are **not** verified, and deployments should not assume them:
- **Cross-device operations.** The temp file is always created in the destination's own directory, so `rename` and `hard_link` never cross a filesystem and `EXDEV` is unreachable by construction. That invariant is tested; a real cross-device attempt is not, because it needs a second filesystem. The restore staging directory is always `.restore-staging` inside `key_dir`, so this holds by construction unless that subdirectory is separately bind-mounted onto another filesystem — do not do that.
- **The key directory being replaced mid-commit.** Every path is re-resolved from the directory name rather than held as a directory file descriptor. If `key_dir` is swapped between the `rename` and the parent `fsync`, the fsync lands on the replacement and the new directory entry is never made durable, while the call still reports success. Reaching this requires write access to the key directory's **parent**, which nothing here checks — the mode enforcement above covers `key_dir` itself and says nothing about what encloses it. Keep the parent owner-writable too. Closing this properly means moving the protocol to `renameat`/`linkat` against a held directory descriptor; it is a real gap, recorded as one.
- **Cross-device operations.** The temp file is always created in the destination's own directory, so `rename` and `hard_link` never cross a filesystem; that invariant is tested, a real cross-device attempt is not. The restore staging directory is always `.restore-staging` inside `key_dir` — do not bind-mount it onto another filesystem.
- **The key directory being replaced mid-commit.** Paths are re-resolved from the directory name rather than held as a directory descriptor. If `key_dir` is swapped between the `rename` and the parent `fsync`, the fsync lands on the replacement and the call still reports success. Reaching this requires write access to the key directory's **parent**, which nothing checks — keep the parent owner-writable too. Closing it properly means moving to `renameat`/`linkat` against a held descriptor; it is a recorded gap.
### Backing up the key directory
Back up `key_dir` as a whole, including the hidden `.master-key.salt` file. A key file on its own is not restorable: decrypting it requires the master key derived from the configured `master_key` **and** the persisted salt. Restoring a partial directory — key files without the salt, or the salt without the key files — leaves the backend unable to decrypt, and the salt guard above will (correctly) refuse to start with encrypted keys and no salt. Losing the salt file with no backup means every key encrypted under it is unrecoverable.
Back up `key_dir` as a whole, including the hidden `.master-key.salt` file. A key file on its own is not restorable: decrypting it requires the master key derived from the configured `master_key` **and** the persisted salt. Restoring a partial directory leaves the backend unable to decrypt, and the salt guard will (correctly) refuse to start. Losing the salt file with no backup means every key encrypted under it is unrecoverable. The rehearsal procedure is in the [KMS disaster-recovery drill](kms-disaster-recovery-drill.md).
+30 -76
View File
@@ -1,36 +1,33 @@
# Cryptographic compliance positioning
This document records where RustFS stands on cryptographic module validation, what may and may not be said about it in external material, and what each possible route to a stronger position would actually cost. It exists so that the question is answered once, from the code, instead of being re-litigated from assumptions about crate names and feature flags.
**Use this when:** writing README, CHANGELOG, release notes, marketing, RFP, or security-questionnaire text that touches FIPS, or reasoning about the `rustfs-crypto` `fips` feature and algorithm deprecation.
**Source of truth:** `scripts/check_fips_wording.sh` (the enforced guard); `crates/crypto/Cargo.toml` (`fips` feature); `crates/crypto/src/encdec/id.rs` (`ID` algorithm bytes); `rustfs/src/startup_runtime_hooks.rs` (`install_default_crypto_provider`).
For where master key material lives per backend and how rotation retention works, see [KMS backend security properties](kms-backend-security.md).
This document records where RustFS stands on cryptographic module validation and what may and may not be said about it, so the question is answered once, from the code. For where master key material lives per backend, see [KMS backend security properties](kms-backend-security.md).
## Status: not FIPS 140-3 validated
**RustFS is not FIPS 140-3 (or 140-2) validated, and no component it links is running as a validated cryptographic module.** There is no CMVP certificate covering RustFS or the libraries it uses in the shipped configuration.
This is a deliberate position, not an oversight. It is also not a statement about algorithm strength: the algorithms in use are standard, well-reviewed AEADs. Validation is a property of a specific module build, its documented boundary, and a certificate — none of which RustFS has or currently pursues.
**RustFS is not FIPS 140-3 (or 140-2) validated, and no component it links runs as a validated cryptographic module.** There is no CMVP certificate covering RustFS or the libraries it uses in the shipped configuration. This is a deliberate position and not a statement about algorithm strength: the algorithms in use are standard, well-reviewed AEADs. Validation is a property of a specific module build, its documented boundary, and a certificate — none of which RustFS has or pursues.
### What the process actually links
The table below is the audited inventory as of this document's writing. "Validated module" asks only whether the code performing the operation is a FIPS-validated cryptographic module; the answer is uniformly no.
| Layer | Where | Implementation | Primitives | Validated module |
| --- | --- | --- | --- | --- |
| TLS (S3 server, internode, outbound clients) | Process-wide default provider installed by `install_default_crypto_provider` in `rustfs/src/startup_runtime_hooks.rs` | `rustls` with the `aws-lc-rs` provider | TLS 1.2/1.3 suites, `prefer-post-quantum` hybrid key exchange | No — this is the ordinary `aws-lc-rs` build, not the `aws-lc-fips-sys`-backed FIPS variant |
| TLS (S3 server, internode, outbound clients) | Process-wide default provider installed by `install_default_crypto_provider` in `rustfs/src/startup_runtime_hooks.rs` | `rustls` with the `aws-lc-rs` provider | TLS 1.2/1.3 suites, `prefer-post-quantum` hybrid key exchange | No — the ordinary `aws-lc-rs` build, not the `aws-lc-fips-sys`-backed FIPS variant |
| Object data path AEAD (SSE) | `crates/kms/src/encryption/ciphers.rs`, `crates/rio/src/encrypt_reader.rs`, `crates/rio-v2/src/encrypt_reader.rs` | RustCrypto `aes-gcm`, `chacha20poly1305` | AES-256-GCM, ChaCha20-Poly1305 | No |
| DEK wrapping | `crates/kms/src/encryption/dek.rs` | RustCrypto `aes-gcm` | AES-256-GCM | No |
| Local KMS backend master key | `crates/kms/src/backends/local.rs` | RustCrypto `argon2`, `aes-gcm` | Argon2id KDF, AES-256-GCM | No |
| Config and IAM blobs at rest | `crates/crypto/src/encdec/` (`rustfs-crypto`) | RustCrypto `pbkdf2`/`argon2`, `aes-gcm`, `chacha20poly1305`, `sha2` | see [the `fips` feature](#the-rustfs-crypto-fips-feature-what-it-actually-does) | No |
| JWT signing and verification | `jsonwebtoken` with the `aws_lc_rs` feature (`crates/crypto`, `crates/iam`, `crates/policy`) | AWS-LC through `aws-lc-rs` | Non-FIPS build | No |
Two consequences follow directly from the table and are worth stating explicitly, because both are commonly assumed the other way:
Two consequences are commonly assumed the other way:
- **AWS-LC being present does not imply FIPS.** `aws-lc-rs` has a FIPS variant; the workspace does not enable it. Every `aws-lc-rs` dependency in the workspace is the default, non-FIPS build.
- **The data path never touches AWS-LC.** Every byte of object plaintext is encrypted by RustCrypto software implementations. Swapping the TLS provider would not change that; see [route 1](#route-1-adopt-the-aws-lc-rs-fips-variant) for what would.
- **AWS-LC being present does not imply FIPS.** `aws-lc-rs` has a FIPS variant; the workspace does not enable it anywhere.
- **The data path never touches AWS-LC.** Every byte of object plaintext is encrypted by RustCrypto software implementations. Swapping the TLS provider would not change that; see route 1 below.
## Terminology red lines for external material
These rules apply to the README, CHANGELOG, release notes, marketing pages, sales decks, RFP responses, and security questionnaires. Claiming validation RustFS does not have is a false statement of fact with regulatory and contractual consequences, not a marketing overreach.
These rules apply to the README, CHANGELOG, release notes, marketing pages, sales decks, RFP responses, and security questionnaires. Claiming validation RustFS does not have is a false statement of fact with regulatory and contractual consequences.
### Never use
@@ -38,22 +35,20 @@ These rules apply to the README, CHANGELOG, release notes, marketing pages, sale
- "FIPS mode", "runs in FIPS mode", "FIPS-enabled"
- "NIST certified", "NIST approved", "CMVP certificate", any certificate number
- "meets FIPS requirements", "satisfies FIPS", or any phrasing a reader would reasonably read as validation
- The internal Cargo feature name `fips` as a product capability. It is a build-time algorithm selector (see below), and surfacing it as a feature name invites exactly the misreading this section exists to prevent.
- The internal Cargo feature name `fips` as a product capability. It is a build-time algorithm selector (see below).
### Permitted, with the qualifier attached
- **"FIPS-preferred algorithms"** — permitted only when accompanied, in the same paragraph or table cell, by an explicit non-validation statement. The defined meaning is: *the default algorithm selection is restricted to algorithms on the FIPS 140-3 approved list, implemented by software that has not been validated as a cryptographic module.*
- Naming specific primitives factually ("AES-256-GCM", "ChaCha20-Poly1305", "PBKDF2-HMAC-SHA256") is always fine. Algorithm names carry no validation claim.
- **"FIPS-preferred algorithms"** — only when accompanied, in the same paragraph or table cell, by an explicit non-validation statement. Defined meaning: *the default algorithm selection is restricted to algorithms on the FIPS 140-3 approved list, implemented by software that has not been validated as a cryptographic module.*
- Naming specific primitives factually ("AES-256-GCM", "ChaCha20-Poly1305", "PBKDF2-HMAC-SHA256") is always fine; algorithm names carry no validation claim.
Suggested boilerplate when the topic cannot be avoided:
Boilerplate when the topic cannot be avoided:
> RustFS encrypts object data with AES-256-GCM and supports ChaCha20-Poly1305. These are FIPS-approved algorithms, but the implementations are not FIPS 140-3 validated cryptographic modules and RustFS makes no FIPS validation claim.
### Guard
`README.md` and `CHANGELOG.md` currently contain no FIPS-related wording; `scripts/check_fips_wording.sh` is the grep guard for that public baseline. Any future occurrence of the banned strings in either file should be treated as a defect and either removed or brought under the qualifier rule above. This document intentionally contains the terminology needed to define the policy and is not part of that narrow outward-material scan.
The same script carries a second block for the adjacent over-claim: no file under `crates/kms` may describe the Vault KV2 backend as wrapping key material through Vault's Transit engine. `KmsBackend::VaultKv2` stores RustFS-wrapped key material in Vault's KV v2 engine and never calls Transit, so that wording would tell an operator their key material is cryptographically isolated inside Vault when it is not. Use the `VaultTransit` backend when that isolation is the requirement.
`scripts/check_fips_wording.sh` greps `README.md` and `CHANGELOG.md` for the banned phrases above, and separately rejects any wording under `crates/kms` that describes the Vault KV2 backend as wrapping key material through Vault's Transit engine (`KmsBackend::VaultKv2` stores RustFS-wrapped material in KV v2 and never calls Transit; use `VaultTransit` when cryptographic isolation is the requirement). This document is intentionally outside the scan: it needs the terminology to define the policy.
## The `rustfs-crypto` `fips` feature: what it actually does
@@ -64,87 +59,46 @@ The same script carries a second block for the adjacent over-claim: no file unde
| enabled (default) | `ID::Pbkdf2AESGCM` (`0x02`) | PBKDF2-HMAC-SHA256, 8192 iterations | AES-256-GCM |
| disabled | `ID::Argon2idAESGCM` (`0x00`) or `ID::Argon2idChaCHa20Poly1305` (`0x01`), chosen at runtime by CPU AES support | Argon2id (64 MiB, t=1, p=4) | AES-256-GCM or ChaCha20-Poly1305 |
The selection sites are `crates/crypto/src/encdec/encrypt.rs` and `crates/crypto/src/encdec/stream_io.rs`; the algorithm identifiers and their KDF parameters live in `crates/crypto/src/encdec/id.rs`.
Selection sites are `crates/crypto/src/encdec/encrypt.rs` and `crates/crypto/src/encdec/stream_io.rs`; identifiers and KDF parameters live in `crates/crypto/src/encdec/id.rs`.
Three properties matter for anyone reasoning about this feature:
- **It affects writes only.** The decrypt path accepts all three identifiers unconditionally, and every ciphertext carries its identifier byte, so toggling the feature never orphans existing data.
- **It does not select a different implementation.** Both branches call RustCrypto; the feature cannot move RustFS toward or away from validation.
- **It is a trade-off, not an upgrade.** PBKDF2-HMAC-SHA256 at 8192 iterations is a work factor well below current password-hashing guidance, whereas the non-FIPS branch uses memory-hard Argon2id. Against an offline attack on the passphrase of a stolen config or IAM blob, the default branch is the weaker of the two.
- **It affects writes only.** The decrypt path in `crates/crypto/src/encdec/id.rs` accepts all three identifiers unconditionally, and every ciphertext carries its identifier byte. Toggling the feature therefore never orphans existing data in either direction.
- **It does not select a different implementation.** Both branches call RustCrypto. There is no validated module on either side of the switch, so the feature cannot move RustFS toward or away from validation.
- **It is a trade-off, not an upgrade.** The FIPS-preferred branch uses PBKDF2-HMAC-SHA256 at 8192 iterations, a work factor well below current password-hashing guidance, whereas the non-FIPS branch uses memory-hard Argon2id. Against an attacker who has obtained an encrypted config or IAM blob and is attacking the passphrase offline, the default branch is the weaker of the two. Enabling the feature buys approved-algorithm alignment, not more resistance.
### Rename recommendation
The name `fips` states a compliance property the feature does not provide, and `rustfs-crypto` is published, so the name is visible to downstream consumers. Recommended direction:
1. Introduce `fips-preferred-algs` as the real feature name, carrying the current behavior.
2. Redefine `fips = ["fips-preferred-algs"]` so existing consumers keep building, and mark it deprecated in the crate documentation with a pointer to this document.
3. Drop the `fips` alias after one release cycle.
4. While renaming, raise the PBKDF2 iteration count or document the trade-off above at the feature definition, so the choice is explicit rather than inherited.
This is a naming and documentation change only; no ciphertext format changes, because the identifier bytes stay as they are.
**Known naming debt.** The feature name `fips` states a compliance property the feature does not provide, and `rustfs-crypto` is published. The intended fix is to introduce `fips-preferred-algs` as the real name, keep `fips` as a deprecated alias for one release cycle, and revisit the PBKDF2 iteration count at the same time; none of this has been done, and no ciphertext format changes when it is.
## Routes to a stronger position, and what each costs
### Route 1: adopt the `aws-lc-rs` FIPS variant
Switch the whole process to `aws-lc-rs`'s FIPS build (backed by `aws-lc-fips-sys`) so cryptographic operations run inside a validated module boundary.
**Scope.** The TLS provider swap is the small part — one feature flag plus the provider install sites. The substantial work is the data path: every AEAD call in `crates/kms/src/encryption/ciphers.rs`, `crates/kms/src/encryption/dek.rs`, `crates/rio/src/encrypt_reader.rs`, `crates/rio-v2/src/encrypt_reader.rs`, `crates/kms/src/backends/local.rs`, and `crates/crypto/src/encdec/` would have to be re-implemented against `aws-lc-rs` primitives. Anything the validated module does not expose has to be dropped or moved out of the boundary: Argon2id has no FIPS status, so the Local backend's KDF and the non-FIPS branch of `rustfs-crypto` would need a compatibility story (read-only support for existing records, PBKDF2 for new ones), and ChaCha20-Poly1305 would become non-approved for new writes.
**Build and platform cost.** `aws-lc-fips-sys` builds a pinned, validated source release and needs CMake, a C toolchain, and Go at build time; it supports a narrower target set than the ordinary crate. The platform matrix cost of plain AWS-LC is already documented and non-hypothetical: rustfs/backlog#883 records that the static musl release build compiles AWS-LC's `getentropy` entropy backend, which aborts on Linux kernels older than 3.17 (the Synology class of device), and that upstream considers this by design with no plan to fix it. The FIPS variant constrains the buildable matrix strictly harder than that, and pins upgrades to whatever the certified source revision allows.
**What it would and would not buy.** Linking the validated module makes the accurate claim "cryptographic operations are performed by a FIPS 140-3 validated module", not "RustFS is FIPS validated". A product-level claim additionally requires a documented module boundary, approved-mode enforcement, power-on self-tests, key zeroization, and entropy-source documentation, plus the operational procedures to keep them true across releases.
**Verdict.** Heavy, and it re-opens a platform-support question that is already an open problem. Justified only by a concrete customer or regulatory commitment that names FIPS as a requirement.
### Route 2: let an externally validated KMS carry key operations
Keep RustFS as-is and place key management inside someone else's validated boundary: the Vault Transit backend against a Vault deployment whose seal/HSM is validated, or an equivalent managed KMS.
**Scope.** Mostly already built. The Transit backend (`VaultTransit`) never lets key-encryption key material leave Vault; RustFS only ever holds Transit ciphertext. What remains is configuration guidance, a supported-deployment statement, and the operational documentation that says which parts of the system are covered.
**What it buys.** Master key generation, wrapping, unwrapping, and rotation happen inside the external module. That is a real, defensible partial answer to "where do keys live and who validated that": it covers the key operations, which is often the part an auditor actually asks about.
**What it does not buy.** The object data path is untouched. DEKs are used for bulk AEAD by RustCrypto inside the RustFS process, and TLS still runs the non-FIPS AWS-LC build. The honest formulation is "key management operations are performed by an externally validated module; the object data path is not validated".
**Verdict.** The nearest partial step, with no code rewrite and no platform-matrix risk. This is the route to point customers at when the requirement is about key custody rather than about a certificate covering the storage layer.
### Route 3: make no validation claim (current default)
Document the position, hold the terminology line, and revisit only when a requirement with a name attached shows up.
**Cost.** This document plus the grep guard. Nothing else.
**Verdict.** The current decision. FIPS 140-3 validation is explicitly not a roadmap target, and adjacent items (PKCS#11, KMIP, BYOK, signing keys) are deferred for lack of demand and because HSM-dependent paths cannot be exercised in CI.
| Route | Scope | What it buys | Verdict |
| --- | --- | --- | --- |
| 1. Adopt the `aws-lc-rs` FIPS variant (`aws-lc-fips-sys`) | The TLS provider swap is the small part. Every AEAD call in the data path (`crates/kms/src/encryption/`, `crates/rio*/src/encrypt_reader.rs`, `crates/kms/src/backends/local.rs`, `crates/crypto/src/encdec/`) would be re-implemented against `aws-lc-rs` primitives; Argon2id (no FIPS status) and ChaCha20-Poly1305 would need read-only compatibility stories. Build needs CMake, a C toolchain, and Go, on a narrower target set — the static musl release already hits AWS-LC's `getentropy` abort on kernels older than 3.17, and the FIPS variant constrains the matrix strictly harder | The accurate claim becomes "cryptographic operations are performed by a FIPS 140-3 validated module", not "RustFS is FIPS validated"; a product-level claim additionally needs a documented boundary, approved-mode enforcement, self-tests, zeroization, and entropy documentation | Heavy; re-opens an open platform-support problem. Justified only by a named customer or regulatory commitment |
| 2. Let an externally validated KMS carry key operations | Mostly built: the `VaultTransit` backend never lets key-encryption key material leave Vault, so master key generation, wrapping, unwrapping, and rotation happen inside whatever module Vault's seal/HSM is validated against. Remaining work is configuration guidance and a supported-deployment statement | A defensible partial answer to "where do keys live and who validated that". The object data path stays RustCrypto and TLS stays non-FIPS AWS-LC: "key management operations are performed by an externally validated module; the object data path is not validated" | Nearest partial step, no code rewrite. Point customers here when the requirement is key custody rather than a certificate covering the storage layer |
| 3. Make no validation claim (current default) | This document plus the grep guard | Nothing further | The current decision. FIPS 140-3 validation is not a roadmap target; adjacent items (PKCS#11, KMIP, BYOK, signing keys) are deferred for lack of demand and because HSM-dependent paths cannot be exercised in CI |
## Algorithm disablement and migration policy
Retiring an algorithm from a storage system is not a code change; it is a data migration with a code change at each end. This section fixes the sequence so that no future deprecation removes a decrypt path while data still depends on it.
Retiring an algorithm from a storage system is a data migration with a code change at each end. This section fixes the sequence so that no deprecation removes a decrypt path while data still depends on it.
### Every persisted artifact is self-describing
The precondition for safe migration already holds: nothing relies on a global "current algorithm" setting to be decodable.
- `rustfs-crypto` blobs carry the `ID` byte (`crates/crypto/src/encdec/id.rs`) immediately after the salt.
- KMS ciphers are selected from the recorded `EncryptionAlgorithm` (`crates/kms/src/types.rs`).
- DEK envelopes record which master key version wrapped them in `DataKeyEnvelope::master_key_version` (`crates/kms/src/encryption/dek.rs`).
So for any stored object it is decidable, from the object alone, which algorithm and which key version it needs.
For any stored object it is therefore decidable, from the object alone, which algorithm and key version it needs.
### Deprecation classes
Retirement moves an algorithm through these states, never skipping one:
1. **Write-disabled, read-supported.** New writes select a replacement; existing data decrypts unchanged. This is the only step that is cheap and reversible.
1. **Write-disabled, read-supported.** New writes select a replacement; existing data decrypts unchanged. The only cheap, reversible step.
2. **Read-deprecated.** Reads still work but are counted and warned on, so the remaining population is measurable.
3. **Read-removed.** The decrypt path is deleted. Permitted only once the remaining population is provably zero.
### Sequencing rules
- Never advance to read-removed on the strength of an argument that data "should have been" migrated. Removal requires evidence that nothing references the algorithm, not an elapsed-time policy.
- A change to default algorithm selection is a compatibility event: it changes what new nodes write, which matters in a mixed-version cluster. Record it in the release notes and in the relevant crate's feature documentation, and check it against the [mixed-version constraints](kms-backend-security.md#mixed-version-clusters-during-a-rolling-upgrade).
- Roll out write-disablement before the corresponding read change, and let the cluster fully converge in between. A build that cannot read what a peer is still writing is the failure mode to avoid.
- A change to default algorithm selection is a compatibility event: it changes what new nodes write, which matters in a mixed-version cluster. Record it in the release notes and the crate's feature documentation, and check it against the [mixed-version constraints](kms-backend-security.md#mixed-version-clusters-during-a-rolling-upgrade).
- Roll out write-disablement before the corresponding read change, and let the cluster fully converge in between.
### Known gap
Step 3 is partially reachable for object data: the bulk rekey sweep (`POST /rustfs/admin/v3/kms/keys/rekey`) migrates stored DEK envelopes off superseded **master key versions** without touching object bodies. It does not re-encrypt object data, so migrating off a data-encryption **algorithm** still has no supported path — treat every algorithm that has ever been written as permanently read-required, and confine algorithm deprecation to step 1.
The bulk rekey sweep (`POST /rustfs/admin/v3/kms/keys/rekey`, see [`kms-bulk-rekey-contract.md`](../architecture/kms-bulk-rekey-contract.md)) migrates stored DEK envelopes off superseded **master key versions** without touching object bodies. It does not re-encrypt object data, so migrating off a data-encryption **algorithm** has no supported path — treat every algorithm that has ever been written as permanently read-required, and confine algorithm deprecation to step 1.
@@ -1,12 +1,13 @@
# KMS disaster-recovery drill
A KMS backup that has never been restored is a hypothesis. This runbook turns it into evidence: it rehearses the complete loop — back up, lose the persistence layer, preflight, restore, and read historical objects again — and files a machine-readable evidence bundle for each run. For what each backend's backup actually covers, see [KMS backend security properties](kms-backend-security.md); for the metrics and alerts around KMS operations, see the [KMS observability runbook](kms-observability-runbook.md).
**Use this when:** rehearsing a KMS backup-and-restore against a lost key directory (Local backend), producing an evidence bundle for an audit, or restoring a Vault-backed KMS after Vault's own snapshot restore.
**Source of truth:** `crates/kms/examples/kms_dr_drill.rs` (operator entry point); `crates/kms/src/backup/{capability,drill,local_export,local_restore}.rs` (`DrillEvidence`, disaster matrix, restore commit marker).
The acceptance criterion of a drill is not that files came back. It is that objects encrypted before the disaster decrypt after the restore. The harness keeps the ciphertext and encryption metadata of every object it sealed before the disaster and, once the restore is complete, decrypts each one through a freshly opened backend and compares against the pre-disaster digest. Anything less proves only that a bundle is well formed.
The drill rehearses the complete loop — back up, lose the persistence layer, preflight, restore, read historical objects again — and files a machine-readable evidence bundle per run. Its acceptance criterion is not that files came back but that objects encrypted before the disaster decrypt after the restore: the harness keeps the ciphertext and encryption metadata of every object it sealed, and after the restore decrypts each through a freshly opened backend and compares against the pre-disaster digest. For what each backend's backup covers, see [KMS backend security properties](kms-backend-security.md); for KMS metrics and alerts, see the [KMS observability runbook](kms-observability-runbook.md).
## Scope
The drill covers the **Local** backend, which is the only backend RustFS produces a full-material bundle for. The responsibility split is deliberate and is described in `crates/kms/src/backup/capability.rs`:
The drill covers the **Local** backend, the only backend RustFS produces a full-material bundle for. The responsibility split is described in `crates/kms/src/backup/capability.rs`:
| Backend | What a RustFS bundle carries | What restores it |
| --- | --- | --- |
@@ -15,7 +16,7 @@ The drill covers the **Local** backend, which is the only backend RustFS produce
| Vault KV2 + Transit | KV metadata and Transit ciphertext references | Vault's native snapshot restore, then the RustFS orchestration |
| Vault Transit | Metadata, configuration references, verification data | Vault's native snapshot restore, then the RustFS orchestration |
For the Vault backends there is no RustFS-side export, so there is no loop for a drill to close end to end: the cryptographic root is non-exportable and comes back through Vault's own disaster-recovery flow. What RustFS owns there is the refusal to proceed before that has happened, plus the ordering of everything after it. Rehearse it with the Vault section below.
For the Vault backends there is no RustFS-side export: the cryptographic root is non-exportable and comes back through Vault's own disaster-recovery flow. RustFS owns the refusal to proceed before that has happened and the ordering of everything after it — see the Vault section below.
## What the drill measures
@@ -45,7 +46,7 @@ Optional variables: `RUSTFS_KMS_DRILL_DISASTER` (see below), `RUSTFS_KMS_DRILL_I
## Disaster matrix
Run all three; they exercise different failure surfaces and converge on the same procedure, which is the point — an operator does not have to diagnose the failure mode before acting.
Run all three; they exercise different failure surfaces and converge on the same procedure, so an operator does not have to diagnose the failure mode before acting.
| `RUSTFS_KMS_DRILL_DISASTER` | Simulates |
| --- | --- |
+88 -87
View File
@@ -1,6 +1,9 @@
# KMS observability runbook
This runbook covers the KMS metrics, the Grafana dashboard that visualizes them, and the response procedure for each Prometheus alert shipped in `.docker/observability/prometheus-rules/rustfs-kms-alerts.yml`. It is the `runbook_url` target for those alerts. For what each KMS backend protects and how Vault authentication behaves, see the [KMS backend security properties](kms-backend-security.md) and the [Vault KMS authentication runbook](vault-kms-authentication.md).
**Use this when:** a `Kms*` Prometheus alert fires (this file is their `runbook_url` target), you are building dashboards or alerts on KMS metrics, or KMS reports not-configured after a restart.
**Source of truth:** `.docker/observability/prometheus-rules/rustfs-kms-alerts.yml` (alert names, thresholds); `crates/kms/src/policy.rs` (backend operation metrics), `crates/kms/src/cache.rs`, `crates/kms/src/deletion_worker.rs`, `crates/kms/src/backends/vault_credentials.rs`, `crates/kms/src/probe.rs`; dashboard `deploy/observability/grafana/rustfs-kms-observability.json`.
For what each backend protects and how Vault authentication behaves, see [KMS backend security properties](kms-backend-security.md) and the [Vault KMS authentication runbook](vault-kms-authentication.md).
## Metric reference
@@ -19,21 +22,21 @@ All six are emitted at the single operation-policy choke point (`crates/kms/src/
| `rustfs_kms_backend_in_flight` | gauge | `backend`, `scope` | External backend attempts currently in flight after admission |
| `rustfs_kms_backend_circuit_open` | gauge | `backend`, `scope` | Open or half-open circuits; `0` means closed |
Label values:
| Label | Values |
| --- | --- |
| `backend` | `vault-kv2`, `vault-transit`, `aws`, and `vault-restore` (calls a restore makes against a Vault bundle's trust root). Operation names are shared across backends (each has a `decrypt`), so this label separates a Transit latency regression from an AWS one. Vault credential logins and renewals report their backend's name and are told apart by `operation`; the `scope` label appears only on the two gauges. Local and Static serve from process memory, never enter the operation policy, and emit no `backend` series |
| `outcome` | `success`; `fatal` (non-retryable failure on first observation); `budget_exhausted` (attempt budget ran out on retryable failures); `deadline_exceeded` (operation deadline ran out before another attempt could complete); `backpressure_timeout` (deadline elapsed before capacity admission); `backpressure_rejected` (active capacity and the bounded queue were full or unavailable); `circuit_open` (a retryable failure opened the breaker, or an open breaker rejected the operation); `cancelled` (shutdown or caller cancellation) |
| `op_class` | `read_idempotent` (safe to retry); `mutating_non_idempotent` (never replayed — a retryable failure terminates after one attempt because the server may have processed the request); `auth` (login and token renewal) |
| `error_class` | `retryable_conn` (dial, TLS, broken connection); `retryable_status` (retryable backend status, e.g. Vault 5xx or a sealed Vault's 503); `attempt_timeout` (per-attempt timeout; retried like a connection failure); `fatal` (authentication, permissions, malformed request, missing key or version) |
| `operation` | Static per-call-site names, e.g. `vault_kv2_read_key_version`, `vault_kv2_cas_write_key`, `vault_transit_encrypt`, `vault_transit_decrypt`, `vault_login`, `vault_token_renew` |
- `backend`: the backend that served the call — `vault-kv2`, `vault-transit`, `aws`, and `vault-restore` for the calls a restore makes against a Vault bundle's trust root. Operation names are shared across backends (every one of them has a `decrypt`), so without this label a Vault Transit latency regression and an AWS one land in the same series. Vault credential logins and renewals report their backend's own name and are told apart by the operation, not by a separate `backend` value; the `scope` label that distinguishes them appears only on the two gauges. Local and Static serve from process memory and never enter the operation policy, so they emit no `backend` series at all.
- `outcome`: `success`, `fatal` (a non-retryable failure ended the operation on first observation), `budget_exhausted` (the attempt budget ran out on retryable failures), `deadline_exceeded` (the operation deadline ran out before another attempt could complete), `backpressure_timeout` (the deadline elapsed before capacity admission completed), `backpressure_rejected` (active capacity and the bounded queue were full or unavailable), `circuit_open` (a retryable failure opened the breaker or an open breaker rejected the operation), `cancelled` (shutdown or caller cancellation).
- `op_class`: `read_idempotent` (safe to retry), `mutating_non_idempotent` (never replayed — a retryable failure terminates after a single attempt because the server may have processed the request), `auth` (login and token renewal).
- `error_class`: `retryable_conn` (connection-level failure: dial, TLS, broken connection), `retryable_status` (retryable backend status, e.g. Vault 5xx or a sealed Vault's 503), `attempt_timeout` (the per-attempt timeout cut the attempt off; retried like a connection failure because the server may still have processed the request), `fatal` (non-retryable: authentication, permissions, malformed request, missing key or version).
- `operation`: static per-call-site names, e.g. `vault_kv2_read_key_version`, `vault_kv2_cas_write_key`, `vault_transit_encrypt`, `vault_transit_decrypt`, `vault_login`, `vault_token_renew`.
Admission sharing follows two boundaries. Total active backend capacity is shared by backend identity and capped at `DEFAULT_MAX_CONCURRENT_OPERATIONS`; ordinary operations may use that minus `RESERVED_CREDENTIAL_OPERATIONS`, so login and renewal always retain a reserved slot. Each backend configuration generation owns fresh bounded queues and circuit breakers for its policy scopes, so a failed reconfiguration candidate cannot inherit or mutate the running generation's admission state.
Admission sharing follows two different boundaries. Total active backend capacity is shared by backend identity and capped at 64; ordinary operations are limited to 63 so login and renewal always retain one reserved slot without exceeding the total cap. Each backend configuration generation owns fresh bounded queues and circuit breakers for its policy scopes, so a failed reconfiguration candidate cannot inherit or mutate the running generation's admission state.
Instrumentation boundary: the Local and Static backends do not flow through the choke point and emit no operation metrics; bringing them under the same instrumentation is tracked separately (rustfs/backlog#1569). Absence of these six series on a cluster using those backends is expected, not an outage. The families below sit above the backend layer and are emitted regardless.
Instrumentation boundary: the Local and Static backends do not flow through the choke point and emit no operation metrics. Absence of these six series on a cluster using those backends is expected, not an outage. The families below sit above the backend layer and are emitted regardless.
### Key metadata cache metrics
Emitted by the manager-level key metadata cache (`crates/kms/src/cache.rs`), which every backend shares. Publication is gated by the cache's `enable_metrics` setting, which defaults to on and which no configure-request field sets today, so in practice these are always published. The counters behind the admin status API are maintained either way, so the switch could never blind `kms service-status`.
Emitted by the manager-level key metadata cache (`crates/kms/src/cache.rs`), which every backend shares. Publication is gated by the cache's `enable_metrics` setting, which defaults to on and which no configure-request field sets, so in practice these are always published. The counters behind the admin status API are maintained either way.
| Metric | Type | Labels | Meaning |
| --- | --- | --- | --- |
@@ -41,48 +44,55 @@ Emitted by the manager-level key metadata cache (`crates/kms/src/cache.rs`), whi
| `rustfs_kms_metadata_cache_evictions_total` | counter | `cause` | Entries dropped from the cache, by removal cause |
| `rustfs_kms_metadata_cache_entries` | gauge | — | Entries the cache currently holds |
`cause` is `expired` (TTL), `size` (capacity), `explicit` (invalidated by a key lifecycle operation), or `replaced` (overwritten by a newer value). Only `expired` and `size` are true evictions — a sustained `explicit`/`replaced` rate is lifecycle traffic, not cache pressure.
| `cause` | Meaning |
| --- | --- |
| `expired` | TTL — a true eviction |
| `size` | Capacity — a true eviction |
| `explicit` | Invalidated by a key lifecycle operation; a sustained rate is lifecycle traffic, not cache pressure |
| `replaced` | Overwritten by a newer value; same reading as `explicit` |
The entry gauge is republished from every write path and from lookups that miss, because TTL expiry drops entries without any write taking place; a cache that goes completely idle can therefore hold a stale value until the next lookup. Note also that this cache only serves key metadata reads such as `describe_key` — encrypt, decrypt and data key generation never consult it, so a low hit ratio is not a data-path problem.
The entry gauge is republished from every write path and from lookups that miss, because TTL expiry drops entries without any write taking place; a cache that goes completely idle can hold a stale value until the next lookup. This cache only serves key metadata reads such as `describe_key` — encrypt, decrypt and data key generation never consult it, so a low hit ratio is not a data-path problem.
### Key lifecycle metrics
Published by the background deletion worker (`crates/kms/src/deletion_worker.rs`) at the end of each sweep, derived from the pages the sweep already walks, so observing the lifecycle costs no extra backend call. The worker only runs on backends whose capabilities include `schedule_deletion`, so a deployment on a backend without it emits none of these.
Published by the background deletion worker (`crates/kms/src/deletion_worker.rs`) at the end of each sweep, derived from the pages the sweep already walks. The worker only runs on backends whose capabilities include `schedule_deletion`, so a deployment on a backend without it emits none of these.
| Metric | Type | Labels | Meaning |
| --- | --- | --- | --- |
| `rustfs_kms_pending_deletion_keys` | gauge | — | Keys scheduled for deletion whose deadline has not passed |
| `rustfs_kms_deletion_tombstone_keys` | gauge | — | Keys left tombstoned by an interrupted removal, still awaiting the sweep |
| `rustfs_kms_oldest_key_rotation_age_seconds` | gauge | — | Seconds since the least recently rotated usable key was rotated, counting from creation for keys with no recorded rotation; `0` when there are none |
| `rustfs_kms_max_key_wrap_operations` | gauge | — | Largest reserved wrap-operation count across usable keys; published only by backends that count wraps (Vault KV2 today) |
| `rustfs_kms_deletion_sweep_keys_total` | counter | `outcome` | Keys the sweep acted on, by outcome: `removed`, `blocked`, `skipped`, `failed`, `unreadable` |
| `rustfs_kms_max_key_wrap_operations` | gauge | — | Largest reserved wrap-operation count across usable keys; published only by backends that count wraps (Vault KV2) |
| `rustfs_kms_deletion_sweep_keys_total` | counter | `outcome` | Keys the sweep acted on, by outcome |
`outcome` is `removed`, `blocked` (live configuration — the default key, or a reference reported by the injected checker — still points at the key, so the sweep refuses to remove it), `skipped` (pending but not yet due, or the state changed between inspection and removal), `failed` (the removal attempt failed and is retried next sweep), or `unreadable` (the backend listed a key record this build cannot describe — a record written by a newer build, or damaged material). Every series is emitted at zero from the first sweep on, so a `rate()` over it is defined immediately.
| `outcome` | Meaning |
| --- | --- |
| `removed` | Material destroyed |
| `blocked` | Live configuration (the default key, or a reference reported by the injected checker) still points at the key; the sweep refuses to remove it |
| `skipped` | Pending but not yet due, or the state changed between inspection and removal |
| `failed` | The removal attempt failed; retried next sweep. Also reported, with no key ids, when the listing itself failed |
| `unreadable` | The backend listed a key record this build cannot describe — written by a newer build, or damaged material |
A non-zero `unreadable` rate does not stop the sweep — the expired keys it *can* read are still destroyed — but it does suppress the lifecycle gauges for that round, because a census taken over a partially readable key set would quietly undercount. Sustained `unreadable` therefore shows up as gauges that stop advancing; investigate the named key ids from the sweep's log line before trusting a rotation-age or pending-deletion reading again.
Every series is emitted at zero from the first sweep on, so a `rate()` over it is defined immediately. A non-zero `unreadable` rate does not stop the sweep, but it suppresses the lifecycle gauges for that round, because a census over a partially readable key set would undercount; sustained `unreadable` therefore shows up as gauges that stop advancing investigate the key ids named in the sweep's log line before trusting a rotation-age or pending-deletion reading again. When *no* key in a complete listing is readable, the backend fails the listing outright (see the [key listing contract](kms-admin-contract.md#key-listing-contract)), so the sweep reports `outcome="failed"` with the listing error and names no key ids: `failed` climbing while `unreadable` stays at zero and the gauges freeze means the whole key set is unreadable on this node — a mixed-version node, or a credential that cannot open any record. Gauges are republished only by a sweep that saw the whole key set; keys already on their way out are excluded from the rotation-age and wrap gauges.
Total damage looks different, and it is worth knowing which you are seeing. When *no* key in a complete listing is readable, the backend fails the listing outright rather than returning an empty page (see the key listing contract in the admin contract page), so the sweep never gets a page to count: it reports `outcome="failed"` with the listing error in its `warn!` line and names no key ids. So `failed` climbing while `unreadable` stays at zero and the gauges freeze means the whole key set is unreadable on this node — a mixed-version node, or a credential that cannot open any record — not that individual removals are failing.
`rustfs_kms_max_key_wrap_operations` tracks the AES-GCM wrap ceiling described under [Rotation drivers and scheduling, per backend](kms-backend-security.md#rotation-drivers-and-scheduling-per-backend). The value is a reservation-based approximation that by design *overestimates*: nodes reserve wrap budget from the key record in blocks of one million and count individual wraps in memory only, so a crash discards unused budget, never a counted wrap. Alert on it approaching 2^32 and rotate the key. It can understate in two bounded, logged cases: a node whose reservation writes keep failing continues wrapping under the warn `Vault KMS wrap budget reservation failed`, and an old build rewriting the key record during a mixed-version window drops the field (see the [mixed-version notes](kms-backend-security.md#mixed-version-clusters-during-a-rolling-upgrade)). Transit and AWS wrap inside the KMS and Local/Static cannot rotate, so none of them publish this series.
The gauges are republished only by a sweep that saw the whole key set; a sweep that could not finish listing leaves the previous, complete values standing rather than understating them. Keys already on their way out are excluded from the rotation-age and wrap gauges, so neither stays pinned high by a key that will never be rotated — or wrap — again.
`rustfs_kms_max_key_wrap_operations` exists because AES-256-GCM caps one key at 2^32 encryptions under random nonces (NIST SP 800-38D), and the KV2 backend wraps every DEK locally with the key's current material — so wraps track encrypted-object writes and the bound is real. The value is a reservation-based approximation that by design *overestimates*: nodes reserve wrap budget from the key record in blocks of one million and count individual wraps in memory only, so a crash discards unused budget, never a counted wrap. Alert on it approaching 2^32 and rotate the key — rotation installs fresh material and resets the counter. Two ways it can understate, both bounded and logged: a node whose reservation writes keep failing continues wrapping under a warn (`Vault KMS wrap budget reservation failed`), and an old build rewriting the key record during a mixed-version window drops the field (see the [mixed-version notes](kms-backend-security.md#mixed-version-clusters-during-a-rolling-upgrade)). Backends that do not wrap locally with rotatable material publish nothing here: Transit and AWS wrap inside the KMS, and Local/Static cannot rotate, so a counter would be an alarm with no remediation.
The rotation age comes from whatever the backend reports as the last rotation, and backends only report a rotation they recorded themselves. Today only the Vault KV2 backend persists that timestamp — it is stamped in the same check-and-set write that commits the rotation (`crates/kms/src/backends/vault.rs`), so it exists if and only if the rotation did. Vault Transit and AWS KMS record no rotation timestamp at all: their key listings always report the rotation time as absent, so on those backends every key ages from creation permanently, the gauge measures key age rather than rotation age, and rotating does not reset it. A KV2 key rotated before the timestamp existed likewise ages from creation until its next rotation stamps the record. In every case the gauge overstates rather than invents — it can report an already-rotated key as overdue, never a stale key as fresh — so an alert on it fires early rather than late. Backends that cannot rotate at all (Local, Static) age every key from creation by construction.
The rotation age comes from whatever the backend reports as the last rotation, and only Vault KV2 persists that timestamp — stamped in the same check-and-set write that commits the rotation (`crates/kms/src/backends/vault.rs`). Vault Transit and AWS KMS record none, so on those backends every key ages from creation permanently, the gauge measures key age rather than rotation age, and rotating does not reset it. A KV2 key rotated before the timestamp existed likewise ages from creation until its next rotation. In every case the gauge overstates rather than invents, so an alert on it fires early rather than late.
### Vault credential metrics
Published by the Vault credential provider (`crates/kms/src/backends/vault_credentials.rs`), so they exist only on Vault-backed backends. Both are label-less: there is exactly one credential generation to describe, and the Vault address, mount, auth path and token are all off limits as label values.
Published by the Vault credential provider (`crates/kms/src/backends/vault_credentials.rs`), so they exist only on Vault-backed backends. Both are label-less: there is exactly one credential generation to describe, and the Vault address, mount, auth path and token are off limits as label values.
| Metric | Type | Labels | Meaning |
| --- | --- | --- | --- |
| `rustfs_kms_vault_token_ttl_seconds` | gauge | — | Seconds left before the active Vault token expires; `0` once it has |
| `rustfs_kms_vault_credentials_fail_closed` | gauge | — | `1` while the provider refuses to hand out its token because it is inside the fail-closed safety window, `0` otherwise |
The renewal loop republishes both on a 10-second cadence while it waits, generating no extra Vault traffic, so a scrape landing between refresh cycles never reads a TTL frozen at the last refresh. `rustfs_kms_vault_credentials_fail_closed` at `1` is the metric form of the fail-closed window described in the [Vault KMS authentication runbook](vault-kms-authentication.md): while it is set, Vault-backed operations fail rather than run on a credential that may already be invalid.
The renewal loop republishes both on a 10-second cadence while it waits, generating no extra Vault traffic. `rustfs_kms_vault_credentials_fail_closed` at `1` is the metric form of the fail-closed window described in the [Vault KMS authentication runbook](vault-kms-authentication.md): while it is set, Vault-backed operations fail rather than run on a credential that may already be invalid.
### Synthetic probe metrics
Published by the background probe worker (`crates/kms/src/probe.rs`), which generates a data key under a reserved probe key, decrypts it, and compares the material. It runs every `RUSTFS_KMS_PROBE_INTERVAL_SECS` seconds (default 60, raised to a floor of 5, `0` disables the probe entirely), and the status it publishes is what KMS readiness reads.
Published by the background probe worker (`crates/kms/src/probe.rs`), which generates a data key under a reserved probe key, decrypts it, and compares the material. It runs every `RUSTFS_KMS_PROBE_INTERVAL_SECS` seconds (default `DEFAULT_PROBE_INTERVAL`, raised to a floor of `MIN_PROBE_INTERVAL`, `0` disables the probe), and the status it publishes is what KMS readiness reads.
| Metric | Type | Labels | Meaning |
| --- | --- | --- | --- |
@@ -92,17 +102,22 @@ Published by the background probe worker (`crates/kms/src/probe.rs`), which gene
| `rustfs_kms_probe_last_success_timestamp_seconds` | gauge | — | Unix timestamp of the most recent successful round |
| `rustfs_kms_probe_consecutive_failures` | gauge | — | Rounds that have failed since the last success |
`failure_kind` is `key_provisioning` (the probe key could not be described or created), `generate`, `decrypt`, or `mismatch` — the last means both calls answered but the material did not survive the round trip, which is as serious as an outage and is reported as loudly.
| `failure_kind` | Meaning |
| --- | --- |
| `key_provisioning` | The probe key could not be described or created |
| `generate` | Data key generation failed |
| `decrypt` | Decryption failed |
| `mismatch` | Both calls answered but the material did not survive the round trip — as serious as an outage |
`unsupported` means the backend cannot host the probe key. It is deliberately counted as its own result and never as a failure, and the worker stops after recording it, so failure-counter alerts stay silent on such deployments; the AWS KMS backend is the case in practice, because it refuses a caller-named create. Note also that `rustfs_kms_probe_last_success_timestamp_seconds` only ever moves forward on a success, so while the probe fails its age keeps growing — alert on that age, not on the presence of a failure counter.
`unsupported` means the backend cannot host the probe key (the AWS KMS backend, which refuses a caller-named create). It is counted as its own result, never as a failure, and the worker stops after recording it, so failure-counter alerts stay silent on such deployments. `rustfs_kms_probe_last_success_timestamp_seconds` only moves forward on a success, so while the probe fails its age keeps growing — alert on that age, not on the presence of a failure counter.
Export path: the `metrics` facade feeds the OTel recorder in `crates/obs`, which exports over OTLP to the collector scraped by Prometheus. Histograms therefore appear in Prometheus as `_bucket`/`_sum`/`_count` series. None of these metrics carry the RustFS `server` label used by the node observability dashboard — distinguish nodes through your scrape topology (`job`/`instance` or promoted OTel resource attributes such as `service_instance_id`).
## Dashboard
Import `deploy/observability/grafana/rustfs-kms-observability.json` into Grafana and select a Prometheus data source that scrapes RustFS metrics. The dashboard has two variables: `datasource` (Prometheus data source) and `operation` (multi-select over the `operation` label). In the docker-compose observability stack (`.docker/observability/`), dashboards are provisioned from a directory (`grafana/provisioning/dashboards/dashboard.yml` points at `/etc/grafana/dashboards`), so no per-file registration is needed there.
Import `deploy/observability/grafana/rustfs-kms-observability.json` into Grafana and select a Prometheus data source that scrapes RustFS metrics. The dashboard has two variables: `datasource` and `operation` (multi-select over the `operation` label). In the docker-compose observability stack (`.docker/observability/`), dashboards are provisioned from a directory (`grafana/provisioning/dashboards/dashboard.yml` points at `/etc/grafana/dashboards`), so no per-file registration is needed.
The shipped dashboard covers the backend operation metrics only. Its "Planned Panels (TODO)" text panel still describes the cache, lifecycle, Vault credential and probe families as not landed — that panel is stale: the emitting code is merged and the metric names, types and label values are in [Metric reference](#metric-reference) above. Until real panels replace it, query those families ad hoc; nothing in the shipped dashboard or alert rules reads them. See [Coverage gaps](#coverage-gaps).
The shipped dashboard covers the backend operation metrics only. Its "Planned Panels (TODO)" text panel is stale: the cache, lifecycle, Vault credential and probe families are emitted and documented in [Metric reference](#metric-reference) above. Until panels replace it, query those families ad hoc. See [Coverage gaps](#coverage-gaps).
## Alert rules
@@ -114,14 +129,12 @@ Every threshold in that file is a conservative default chosen without a producti
### KmsBackendFatalErrors
Meaning: attempts are failing with `error_class="fatal"` — failures the policy never retries. Each one is a KMS backend call that failed permanently (authentication, permissions, malformed request, or a missing key/version), so callers are seeing errors right now. This is the highest-signal KMS alert: fatal failures do not appear as background noise in a healthy system.
Investigation:
Meaning: attempts are failing with `error_class="fatal"` — failures the policy never retries (authentication, permissions, malformed request, or a missing key/version), so callers are seeing errors right now. This is the highest-signal KMS alert: fatal failures do not appear as background noise in a healthy system.
1. Break the rate down by operation: `sum by (operation) (rate(rustfs_kms_backend_attempt_failures_total{error_class="fatal"}[5m]))`.
2. If the failing operations are `vault_login` or `vault_token_renew` (`op_class="auth"`), the Vault credentials are invalid or expired. Follow the [Vault KMS authentication runbook](vault-kms-authentication.md) — note that credential refresh is fail-closed, so a broken credential eventually takes down all Vault-backed operations, not just auth. Look for the `Vault token renewal failed; falling back to a fresh login` and `Vault credential refresh failed; retrying until the credentials recover` warnings in the RustFS logs; `rustfs_kms_vault_credentials_fail_closed` at `1`, or `rustfs_kms_vault_token_ttl_seconds` at or near `0`, confirms that state without reading logs.
3. If the failing operations are `vault_kv2_*` or `vault_transit_*`, check for Vault permission denials: compare the token's policy against the minimal policy in [KMS backend security properties](kms-backend-security.md) (a policy that drifted or was re-scoped produces 403s that classify as fatal), and check the Vault audit log for the corresponding denied requests.
4. A fatal `KeyVersionNotFound` on decrypt-path operations means a DEK envelope references a key version whose record is missing. Decryption deliberately fails closed with no fallback — see the rotation retention preconditions in [KMS backend security properties](kms-backend-security.md) and verify nobody destroyed version records under the key subtree.
2. If the failing operations are `vault_login` or `vault_token_renew` (`op_class="auth"`), the Vault credentials are invalid or expired. Follow the [Vault KMS authentication runbook](vault-kms-authentication.md) — credential refresh is fail-closed, so a broken credential eventually takes down all Vault-backed operations. Look for the `Vault token renewal failed; falling back to a fresh login` and `Vault credential refresh failed; retrying until the credentials recover` warnings; `rustfs_kms_vault_credentials_fail_closed` at `1`, or `rustfs_kms_vault_token_ttl_seconds` at or near `0`, confirms that state without reading logs.
3. If the failing operations are `vault_kv2_*` or `vault_transit_*`, check for Vault permission denials: compare the token's policy against the [minimal policy](kms-backend-security.md#minimal-vault-policy-for-the-kv2-backend) (a re-scoped policy produces 403s that classify as fatal), and check the Vault audit log for the denied requests.
4. A fatal `KeyVersionNotFound` on decrypt-path operations means a DEK envelope references a key version whose record is missing. Decryption deliberately fails closed with no fallback — see the [retention and destruction preconditions](kms-backend-security.md#retention-and-destruction-preconditions) and verify nobody destroyed version records under the key subtree.
5. Confirm blast radius with the outcome view: `sum by (operation) (rate(rustfs_kms_backend_operations_total{outcome="fatal"}[5m]))`.
Related signals: the "Attempt Failure Rate by Error Class" and "Backend Operation Rate by Outcome" dashboard panels; Vault server audit and server logs; S3-level 5xx on encrypted buckets.
@@ -130,41 +143,35 @@ Related signals: the "Attempt Failure Rate by Error Class" and "Backend Operatio
Meaning: more than 5% of KMS operations are terminating without success (`fatal`, `budget_exhausted`, `deadline_exceeded`, `backpressure_timeout`, `backpressure_rejected`, or `circuit_open`; `cancelled` is excluded because shutdown windows legitimately produce it). A traffic guard suppresses the alert below ~0.02 ops/s so a single failure on a near-idle cluster does not page.
Investigation:
1. Break the failures down by outcome: `sum by (outcome) (rate(rustfs_kms_backend_operations_total{outcome!~"success|cancelled"}[5m]))`.
2. If `fatal` dominates, follow [KmsBackendFatalErrors](#kmsbackendfatalerrors).
3. If `budget_exhausted` or `deadline_exceeded` dominates, follow [KmsBackendRetryBudgetExhausted](#kmsbackendretrybudgetexhausted) — the backend is unavailable or too slow for longer than the retry policy can bridge.
4. If `backpressure_timeout` or `backpressure_rejected` dominates, compare `rustfs_kms_backend_in_flight` by `backend` and `scope`; total active capacity is shared by backend identity, one slot is reserved for credential refresh, and each configuration generation has fresh scope-local bounded queues.
3. If `budget_exhausted` or `deadline_exceeded` dominates, follow [KmsBackendRetryBudgetExhausted](#kmsbackendretrybudgetexhausted).
4. If `backpressure_timeout` or `backpressure_rejected` dominates, compare `rustfs_kms_backend_in_flight` by `backend` and `scope`; total active capacity is shared by backend identity with one slot reserved for credential refresh, and each configuration generation has fresh scope-local bounded queues.
5. If `circuit_open` dominates, follow [KmsBackendCircuitOpen](#kmsbackendcircuitopen).
6. Correlate with client impact: encrypted-object PUT/GET failures and S3 error rates on buckets with encryption configured.
Related signals: the "Non-Success Outcome Ratio" dashboard panel; the KMS-related warnings listed under the other alerts in this runbook.
Related signals: the "Non-Success Outcome Ratio" dashboard panel; the KMS-related warnings listed under the other alerts.
### KmsBackendP99LatencyHigh
Meaning: the p99 wall-clock duration of KMS operations is sustained above 2s. The histogram includes retries and backoff sleeps, so a high p99 with a healthy p50 usually means a slow retry tail (a subset of calls failing and being retried), not a uniform slowdown.
Meaning: the p99 wall-clock duration of KMS operations is sustained above 2s. The histogram includes retries and backoff sleeps, so a high p99 with a healthy p50 usually means a slow retry tail, not a uniform slowdown.
Investigation:
1. Compare p50 and p99 on the "Operation Duration p50 / p99" panel. Flat p50 with elevated p99 points at retries; both elevated points at the backend or the network path being uniformly slow.
2. Split by backend and operation with `histogram_quantile(0.99, sum by (le, backend, operation) (rate(rustfs_kms_backend_operation_duration_seconds_bucket[5m])))` to see whether one backend call or all of them regressed.
3. Check the attempts histogram: an average meaningfully above 1 confirms the latency is retry-driven; follow [KmsBackendAttemptFailureSpike](#kmsbackendattemptfailurespike) for the failure classes.
4. If latency is not retry-driven, check the network path to Vault (TLS handshakes, DNS, proxies) and Vault's own telemetry (storage backend latency, load).
5. Remember that this latency sits inside S3 request latency for encrypted objects: sustained p99 near the operation deadline will start converting into `deadline_exceeded` outcomes.
1. Compare p50 and p99 on the "Operation Duration p50 / p99" panel. Flat p50 with elevated p99 points at retries; both elevated points at the backend or network path being uniformly slow.
2. Split by backend and operation: `histogram_quantile(0.99, sum by (le, backend, operation) (rate(rustfs_kms_backend_operation_duration_seconds_bucket[5m])))`.
3. Check the attempts histogram: an average meaningfully above 1 confirms retry-driven latency; follow [KmsBackendAttemptFailureSpike](#kmsbackendattemptfailurespike) for the failure classes.
4. If not retry-driven, check the network path to Vault (TLS handshakes, DNS, proxies) and Vault's own telemetry (storage backend latency, load).
5. This latency sits inside S3 request latency for encrypted objects: sustained p99 near the operation deadline starts converting into `deadline_exceeded` outcomes.
Related signals: the "Operation Duration p99 by Operation" and "Operation Attempts Distribution" panels; `KMS backend attempt failed with a retryable error; backing off before retry` warnings (fields: `operation`, `attempt`, `error_class`, `backoff`).
### KmsBackendAttemptFailureSpike
Meaning: individual attempts are failing at a sustained rate across all error classes. The retry policy may still be absorbing these — operations can keep succeeding while this alert fires — but the system is burning retry budget and running degraded, and a small further degradation will surface to callers.
Investigation:
Meaning: individual attempts are failing at a sustained rate across all error classes. The retry policy may still be absorbing them — operations can keep succeeding while this fires — but the system is burning retry budget and a small further degradation will surface to callers.
1. Break the rate down by class: `sum by (error_class) (rate(rustfs_kms_backend_attempt_failures_total[5m]))`.
2. `retryable_conn`: network-level failures — check connectivity, TLS, DNS, and whether Vault is down or restarting.
3. `retryable_status`: the backend answered with a retryable error — check Vault health and seal status (a sealed Vault returns 503, which lands here), and Vault-side rate limiting.
4. `attempt_timeout`: attempts are being cut off by the per-attempt timeout — either the backend is slow (correlate with [KmsBackendP99LatencyHigh](#kmsbackendp99latencyhigh)) or the configured attempt timeout is too tight for the deployment's network path.
3. `retryable_status`: the backend answered with a retryable error — check Vault health and seal status (a sealed Vault returns 503), and Vault-side rate limiting.
4. `attempt_timeout`: attempts are cut off by the per-attempt timeout — either the backend is slow (correlate with [KmsBackendP99LatencyHigh](#kmsbackendp99latencyhigh)) or the configured attempt timeout is too tight for the network path.
5. `fatal`: follow [KmsBackendFatalErrors](#kmsbackendfatalerrors).
6. Grep RustFS logs for `KMS backend attempt failed with a retryable error; backing off before retry` — the structured fields (`operation`, `attempt`, `error_class`, `backoff`) identify which call sites are cycling.
@@ -174,71 +181,65 @@ Related signals: the "Attempt Failure Rate by Error Class" panel; the attempts h
Meaning: operations are terminating as `budget_exhausted` or `deadline_exceeded` — every individual failure was retryable, but the backend stayed unhealthy for longer than the retry policy could bridge, so callers received hard failures.
Investigation:
1. Identify the failing operations: `sum by (operation) (rate(rustfs_kms_backend_operations_total{outcome=~"budget_exhausted|deadline_exceeded"}[5m]))`.
2. Establish how long the underlying failure has persisted from the attempt-failure rate history; follow [KmsBackendAttemptFailureSpike](#kmsbackendattemptfailurespike) for the class-specific diagnosis.
3. Note the by-design case: `mutating_non_idempotent` operations (e.g. `vault_kv2_cas_write_key`, `vault_transit_create_key`) are never replayed, so a single retryable failure terminates them as `budget_exhausted` after one attempt. A spike confined to mutating operations means write-path failures, not an exhausted retry loop.
4. `deadline_exceeded` clustering with duration p99 near the operation deadline means the budget is being spent on slow attempts rather than fast failures — treat as a latency problem first.
5. Confirm client impact and, if the backend outage is confirmed external (Vault down), coordinate recovery there; RustFS will resume without intervention once the backend recovers.
3. By-design case: `mutating_non_idempotent` operations (e.g. `vault_kv2_cas_write_key`, `vault_transit_create_key`) are never replayed, so a single retryable failure terminates them as `budget_exhausted` after one attempt. A spike confined to mutating operations means write-path failures, not an exhausted retry loop.
4. `deadline_exceeded` clustering with duration p99 near the operation deadline means the budget is spent on slow attempts rather than fast failures — treat as a latency problem first.
5. Confirm client impact; if the backend outage is external (Vault down), coordinate recovery there RustFS resumes without intervention once the backend recovers.
Related signals: the "Backend Operation Rate by Outcome" panel; retry-backoff warnings in RustFS logs; Vault availability monitoring.
Related signals: the "Backend Operation Rate by Outcome" panel; retry-backoff warnings; Vault availability monitoring.
### KmsBackendCircuitOpen
Meaning: `rustfs_kms_backend_circuit_open` has remained above `0` for a `backend` and `scope` for one minute. This direct gauge alert does not depend on operation traffic: it remains visible when the circuit is open and rejecting calls, and while the single half-open recovery probe is running.
Investigation:
Meaning: `rustfs_kms_backend_circuit_open` has remained above `0` for a `backend` and `scope` for one minute. This direct gauge alert does not depend on operation traffic: it stays visible while the circuit rejects calls and while the single half-open recovery probe runs.
1. Identify the affected scope with `rustfs_kms_backend_circuit_open > 0`.
2. Break recent rejections down by operation: `sum by (operation) (rate(rustfs_kms_backend_operations_total{outcome="circuit_open"}[5m]))`.
3. Check `sum by (error_class) (rate(rustfs_kms_backend_attempt_failures_total{error_class=~"retryable_conn|retryable_status|attempt_timeout"}[5m]))` to distinguish transport failures, retryable backend responses such as a sealed Vault, and attempt timeouts. An attempt timeout counts toward the breaker as a retryable connection failure.
4. After the open interval, the next eligible operation is the only half-open probe. A success or non-retryable failure closes the circuit; a retryable failure reopens it. A non-retryable probe still fails as `fatal`, so follow [KmsBackendFatalErrors](#kmsbackendfatalerrors) even after the circuit gauge clears. Do not restart RustFS just to clear the state.
5. Remember the sharing boundary: each configuration generation has fresh scope-local breaker and queue state, while total active capacity is shared by backend identity with one slot reserved for credential refresh. Check other scopes for capacity pressure even when their circuits remain closed.
4. After the open interval, the next eligible operation is the only half-open probe. A success or non-retryable failure closes the circuit; a retryable failure reopens it. A non-retryable probe still fails as `fatal`, so follow [KmsBackendFatalErrors](#kmsbackendfatalerrors) even after the gauge clears. Do not restart RustFS just to clear the state.
5. Each configuration generation has fresh scope-local breaker and queue state, while total active capacity is shared by backend identity with one slot reserved for credential refresh. Check other scopes for capacity pressure even when their circuits remain closed.
Related signals: `circuit_open`, `backpressure_timeout`, and `backpressure_rejected` on the "Backend Operation Rate by Outcome" panel; `rustfs_kms_backend_in_flight`; Vault availability and seal status.
### KmsKeyRotationOverdue
Meaning: `rustfs_kms_oldest_key_rotation_age_seconds` — seconds since the least recently rotated usable key was rotated, counting from creation for keys with no recorded rotation — has been above 400 days for an hour. This is a compliance and hygiene signal, not an outage: encryption and decryption continue unchanged, and nothing in RustFS acts on the verdict. But the longer master key material stays in service the larger the blast radius of its compromise, and on backends where RustFS wraps DEKs locally (Local, Static, Vault KV2) the AES-GCM random-nonce invocation ceiling (NIST SP 800-38D: at most 2^32 wraps under one key) is consumed by every encrypted object write and only ever resets through rotation.
Meaning: `rustfs_kms_oldest_key_rotation_age_seconds` — seconds since the least recently rotated usable key was rotated, counting from creation for keys with no recorded rotation — has been above 400 days for an hour. This is a compliance and hygiene signal, not an outage: encryption and decryption continue unchanged, and nothing in RustFS acts on the verdict. The reasons rotation matters (blast radius, and the AES-GCM wrap ceiling on backends where RustFS wraps DEKs locally) are stated once in [Rotation drivers and scheduling, per backend](kms-backend-security.md#rotation-drivers-and-scheduling-per-backend).
Investigation:
1. Find which keys are due. The gauge deliberately names no key — a per-key label would carry key identifiers into the metric stream — so read the per-key verdict from the listing: `GET /rustfs/admin/v3/kms/keys` carries `rotation_due` and `rotation_due_reason` (`age`, `never_rotated`, `wraps`, or `unsupported`) per key, computed against `RUSTFS_KMS_ROTATION_MAX_AGE_SECS` and `RUSTFS_KMS_ROTATION_MAX_WRAPS`. A `wraps` reason means the key's material has wrapped more data keys than the configured budget — the AES-GCM random-nonce ceiling rather than an age policy, so it is not satisfied by relaxing the age threshold. The verdict appears only on the listing, not on single-key describe. If `RUSTFS_KMS_ROTATION_MAX_AGE_SECS` is unset, set it to your policy's rotation period so the per-key verdict and this alert agree on what "overdue" means.
2. If the reason is `unsupported`, the backend cannot rotate at all (Local, Static). There is no key-level response; the decision is a backend migration, and the wrap ceiling above is the reason it cannot be deferred forever. See the [rotation drivers and scheduling matrix](kms-backend-security.md#rotation-drivers-and-scheduling-per-backend).
3. On a backend that can rotate, act per the driver matrix: on **Vault KV2**, check why your external rotation scheduler did not run (or set one up — RustFS deliberately ships none) and satisfy the [pre-rotation checklist](kms-backend-security.md#rotation-drivers-and-scheduling-per-backend) before rotating, above all the [upgrade-ordering hard constraint](kms-backend-security.md#upgrade-before-first-rotation-hard-constraint) — never respond to this alert by rotating in the middle of a rolling upgrade. On **Vault Transit**, check `auto_rotate_period` on the key in Vault. On **AWS KMS**, check the key's automatic rotation status in AWS — and do not schedule rotation through the RustFS endpoint, which maps to quota-limited `RotateKeyOnDemand`.
4. Know the gauge's blind spot on Transit and AWS before chasing a rotation that already happened: only KV2 persists a rotation timestamp, so Transit and AWS keys age from creation permanently and this alert will not clear after a rotation there. Confirm the real cadence at the owning system — the Transit key's version history in Vault, or the key's rotation status in AWS — and treat a confirmed-healthy cadence as a known overstatement of this gauge rather than an overdue key.
1. Find which keys are due. The gauge deliberately names no key, so read the per-key verdict from the listing: `GET /rustfs/admin/v3/kms/keys` carries `rotation_due` and `rotation_due_reason` (`age`, `never_rotated`, `wraps`, or `unsupported`) per key, computed against `RUSTFS_KMS_ROTATION_MAX_AGE_SECS` and `RUSTFS_KMS_ROTATION_MAX_WRAPS` — see [Rotation readiness](kms-backend-security.md#rotation-readiness-reported-never-acted-on). A `wraps` reason is not satisfied by relaxing the age threshold. If `RUSTFS_KMS_ROTATION_MAX_AGE_SECS` is unset, set it to your policy's rotation period so the per-key verdict and this alert agree.
2. If the reason is `unsupported`, the backend cannot rotate at all (Local, Static); the only response is a backend migration.
3. On a backend that can rotate, act per the driver matrix in [Rotation drivers and scheduling, per backend](kms-backend-security.md#rotation-drivers-and-scheduling-per-backend) — external scheduler for Vault KV2, `auto_rotate_period` for Vault Transit, AWS automatic rotation for AWS KMS — and satisfy its pre-rotation checklist first, above all the [upgrade-ordering hard constraint](kms-backend-security.md#upgrade-before-first-rotation-hard-constraint). Never respond to this alert by rotating in the middle of a rolling upgrade.
4. Know the gauge's blind spot on Transit and AWS: only KV2 persists a rotation timestamp, so Transit and AWS keys age from creation permanently and this alert will not clear after a rotation there. Confirm the real cadence at the owning system (the Transit key's version history, or the key's rotation status in AWS) and treat a confirmed-healthy cadence as a known overstatement of this gauge.
5. If a KV2 key was genuinely rotated and the gauge stays high, remember the gauge is republished only by a sweep that saw the whole key set: check `rustfs_kms_deletion_sweep_keys_total` for `unreadable` or `failed` outcomes freezing the lifecycle gauges (see [Key lifecycle metrics](#key-lifecycle-metrics)), and that the deletion worker is running at all — it only runs on backends with the `schedule_deletion` capability, which is also why the Static backend never emits this series.
Related signals: `rotation_due` / `rotation_due_reason` on the key listing; `rustfs_kms_deletion_sweep_keys_total{outcome=~"unreadable|failed"}` (a frozen gauge is stale, not healthy); the [rotation drivers and scheduling matrix](kms-backend-security.md#rotation-drivers-and-scheduling-per-backend) and pre-rotation checklist in the backend security properties document.
Related signals: `rotation_due` / `rotation_due_reason` on the key listing; `rustfs_kms_deletion_sweep_keys_total{outcome=~"unreadable|failed"}` (a frozen gauge is stale, not healthy).
## Startup persisted-configuration load
KMS configured through the admin API is persisted to cluster storage and restored on every startup. The load result is visible in two places; check both before concluding that KMS "was never configured":
- **Startup log**, `event="kms_persisted_config_lookup"` (`target: rustfs::init`): `state="found"` means the persisted configuration was loaded and applied; `state="not_found"` means no persisted configuration exists on disk; `state="load_failed"` means one exists but reading, unsealing, or decoding it failed.
- **`GET /rustfs/admin/v3/kms/service-status`**: `"NotConfigured"` matches `not_found` (nothing persisted — configuring from scratch is the correct response), while a status of `Error("Failed to load persisted KMS configuration: ...")` or `Error("Failed to apply persisted KMS configuration: ...")` matches `load_failed`. The two states call for different operator actions; do not resubmit a full configuration to recover from `load_failed`.
| Startup log `event="kms_persisted_config_lookup"` (`target: rustfs::init`) | `GET /rustfs/admin/v3/kms/service-status` | Meaning and action |
| --- | --- | --- |
| `state="found"` | configured | Persisted configuration loaded and applied |
| `state="not_found"` | `"NotConfigured"` | Nothing persisted; configuring from scratch is the correct response |
| `state="load_failed"` | `Error("Failed to load persisted KMS configuration: ...")` or `Error("Failed to apply persisted KMS configuration: ...")` | A configuration exists but reading, unsealing, or decoding it failed. Do not resubmit a full configuration; use reload |
To recover from `load_failed` — or from any state where the server runs but its in-memory KMS lags the persisted configuration — call `POST /rustfs/admin/v3/kms/reload` (requires `kms:ServiceControl`). It re-reads the persisted configuration from cluster storage and reconfigures the service without resubmitting secrets, then broadcasts the reload to peer nodes. If reload keeps failing, check cluster storage health first (the read needs quorum), then `RUSTFS_KMS_CONFIG_SECRET`: an unseal error means the secret is missing or differs from the one that sealed the persisted copy — it must be identical on every node.
To recover from `load_failed` — or from any state where the server runs but its in-memory KMS lags the persisted configuration — call `POST /rustfs/admin/v3/kms/reload` (`kms:ServiceControl`). It re-reads the persisted configuration from cluster storage and reconfigures the service without resubmitting secrets, then broadcasts the reload to peer nodes. If reload keeps failing, check cluster storage health first (the read needs quorum), then `RUSTFS_KMS_CONFIG_SECRET`: an unseal error means the secret is missing or differs from the one that sealed the persisted copy — it must be identical on every node.
A separate event, `kms_config_load_skipped` with `reason="storage_uninitialized"`, comes from the ambient loader used by the peer-reload RPC path; during normal startup the loader receives the store explicitly, so seeing this event outside a peer reload indicates a request arrived before storage initialization finished.
A separate event, `kms_config_load_skipped` with `reason="storage_uninitialized"`, comes from the ambient loader used by the peer-reload RPC path; seeing it outside a peer reload indicates a request arrived before storage initialization finished.
## Threshold calibration
Every numeric traffic or latency threshold in `rustfs-kms-alerts.yml` (5% error ratio, 2s p99, 0.5/s attempt failures, 0.05/s budget exhaustion) is a conservative default chosen without a production baseline, biased toward not paging on healthy-but-busy systems. Before relying on these alerts for paging: run the workload in staging for at least a week, record the steady-state values of the expressions above, then tighten thresholds to sit clearly above observed peaks. `KmsBackendCircuitOpen` is different: its gauge is direct state, and the one-minute hold only suppresses a circuit that recovers immediately. `KmsKeyRotationOverdue` is different in the other direction: its 400-day threshold is a policy default (sitting above a common one-year rotation period), not a traffic default — calibrate it against the rotation period your compliance policy requires and against `RUSTFS_KMS_ROTATION_MAX_AGE_SECS`, not against a staging baseline. Once a stable baseline exists, consider converting `KmsBackendAttemptFailureSpike` to a baseline-relative form (`offset 1d` ratio, see `.docker/observability/prometheus-rules/rustfs-get-optimization-alerts.yaml` for the pattern). Formal SLO targets for KMS operations are deliberately out of scope until that baseline exists (rustfs/backlog#1584).
Every numeric traffic or latency threshold in `rustfs-kms-alerts.yml` (5% error ratio, 2s p99, 0.5/s attempt failures, 0.05/s budget exhaustion) is a conservative default chosen without a production baseline, biased toward not paging on healthy-but-busy systems. Before relying on these alerts for paging: run the workload in staging for at least a week, record the steady-state values of the expressions above, then tighten thresholds to sit clearly above observed peaks. `KmsBackendCircuitOpen` is different: its gauge is direct state, and the one-minute hold only suppresses a circuit that recovers immediately. `KmsKeyRotationOverdue` is different in the other direction: its 400-day threshold is a policy default sitting above a common one-year rotation period — calibrate it against the rotation period your compliance policy requires and against `RUSTFS_KMS_ROTATION_MAX_AGE_SECS`, not against a staging baseline. Once a stable baseline exists, consider converting `KmsBackendAttemptFailureSpike` to a baseline-relative form (`offset 1d` ratio; see `.docker/observability/prometheus-rules/rustfs-get-optimization-alerts.yaml` for the pattern). Formal SLO targets for KMS operations are deliberately out of scope until that baseline exists.
## Coverage gaps
The four metric families designed under rustfs/backlog#1584 — key-cache effectiveness, key lifecycle, Vault credentials, synthetic probe — have all landed and are documented in [Metric reference](#metric-reference). What is still missing:
- **No dashboard panels for those four families, and an alert rule for only one of them.** The key lifecycle family has one rule — [`KmsKeyRotationOverdue`](#kmskeyrotationoverdue) on the rotation-age gauge — while the cache, Vault credential, and probe families are emitted but neither visualized nor alerted on, so they surface only in ad-hoc queries. Building against them is safe now: the names and label values above are what the code emits.
- **The Local and Static backends emit no operation metrics**, because they do not flow through the operation-policy choke point; bringing them under the same instrumentation is tracked separately (rustfs/backlog#1569). Their cache metrics are emitted normally.
- **No formal SLO targets**, deliberately, until a production baseline exists — see [Threshold calibration](#threshold-calibration).
When a panel or alert rule for one of the landed families is added, replace the corresponding TODO bullet in the dashboard's "Planned Panels" text panel and update this section.
- The cache, Vault credential, and probe metric families have no dashboard panels and no alert rules; the lifecycle family has only [`KmsKeyRotationOverdue`](#kmskeyrotationoverdue). Building against them is safe: the names and label values above are what the code emits.
- The Local and Static backends emit no operation metrics, because they do not flow through the operation-policy choke point. Their cache metrics are emitted normally.
- No formal SLO targets until a production baseline exists — see [Threshold calibration](#threshold-calibration).
## Related documents
- [KMS backend security properties](kms-backend-security.md) — backend trust boundaries, minimal Vault policies, rotation retention preconditions.
- [KMS backend security properties](kms-backend-security.md) — backend trust boundaries, minimal Vault policies, rotation drivers and retention preconditions.
- [Vault KMS authentication runbook](vault-kms-authentication.md) — credential sources, refresh behavior, and the fail-closed window.
- [KMS admin API contract](kms-admin-contract.md) — endpoint actions and the key listing contract.
- `deploy/observability/README.md` — dashboard import notes for all RustFS dashboards.
+4 -3
View File
@@ -1,8 +1,9 @@
# Per-key KMS authorization
RustFS authorizes KMS access with identity policies. A statement may name the keys it applies to, so a grant such as `kms:DisableKey` no longer implies every key in the cluster, and an SSE-KMS request is checked against the key it actually resolves to.
**Use this when:** writing IAM policies that grant KMS actions on specific keys, choosing a canned KMS role, or enabling SSE-KMS data-path enforcement (`RUSTFS_KMS_ENFORCE_SSE_KEY_POLICY`).
**Source of truth:** `crates/policy/src/policy/resource.rs` (KMS ARN grammar, `KMS_ALIAS_SEGMENT`); `crates/policy/src/policy/action.rs` (`KmsAction`); the canned `KMSKeyAdministrator` / `KMSKeyUser` / `KMSAuditor` policies in `crates/policy`; `rustfs/src/admin/route_policy.rs` (which admin routes are per-key, see [KMS admin API contract](kms-admin-contract.md)).
This page covers the resource grammar, the built-in role templates, the two enforcement planes, and the migration path. For where master key material lives per backend, see [KMS backend security properties](kms-backend-security.md).
RustFS authorizes KMS access with identity policies. A statement may name the keys it applies to, so a grant such as `kms:DisableKey` no longer implies every key in the cluster, and an SSE-KMS request is checked against the key it actually resolves to. This page covers the resource grammar, the built-in role templates, the two enforcement planes, and the migration path. For where master key material lives per backend, see [KMS backend security properties](kms-backend-security.md).
## Resource grammar
@@ -81,7 +82,7 @@ Scope and exemptions:
- **SSE-KMS only.** SSE-S3 wraps its data key with a server-owned key the caller never names, and SSE-C never reaches KMS; both are exempt, matching AWS.
- **The resolved key**, not the header. A bucket default encryption rule naming a KMS key is authorized the same way an explicit `x-amz-server-side-encryption-aws-kms-key-id` header is.
- **Anonymous requests are denied.** An anonymous caller has no identity policy and therefore holds no `kms` grants, so under enforcement every anonymous read or write of an SSE-KMS object fails with `AccessDenied` even when a bucket policy makes the bucket public. This matches AWS, where anonymous requests cannot use SSE-KMS objects at all, and it keeps the per-key gate meaningful: were anonymous requests exempt, any denied identity could bypass the gate on a public bucket by simply dropping its credentials. **A public bucket serving SSE-KMS objects is incompatible with enforcement** — serve public content unencrypted or under SSE-S3 instead. With enforcement off (the default), anonymous access to SSE-KMS objects remains governed by bucket policy alone. The server warns once per process when it first denies an anonymous request; per-request denials appear on audit entries (`kmsOutcome=failure`, `kmsErrorClass=access_denied`, empty requester identity) and at debug level.
- **Anonymous requests are denied.** An anonymous caller has no identity policy and therefore no `kms` grants, so under enforcement every anonymous read or write of an SSE-KMS object fails with `AccessDenied`, even on a bucket a bucket policy makes public (matching AWS, and closing the bypass of dropping credentials on a public bucket). **A public bucket serving SSE-KMS objects is incompatible with enforcement** — serve public content unencrypted or under SSE-S3. With enforcement off (the default), anonymous access to SSE-KMS objects is governed by bucket policy alone. The server warns once per process on the first anonymous denial; per-request denials appear on audit entries (`kmsOutcome=failure`, `kmsErrorClass=access_denied`, empty requester identity) and at debug level.
- **Internal work is exempt.** Replication, lifecycle transitions, healing and the scanner run as the system principal.
- **Authorization runs before key state is checked**, so a denial cannot be used to probe whether a key exists, is disabled, or is pending deletion. The response is always `AccessDenied`.
- **Multipart uploads are authorized at create time**, where the session data key is generated. Part uploads and completion reuse that envelope and are not re-authorized against the destination key.
+50 -52
View File
@@ -1,66 +1,66 @@
# 日志故障诊断(`rustfs diagnose`)
# Log diagnosis (`rustfs diagnose`)
对客户/现场提供的日志文件做离线故障归因:解析 RustFS 的 JSON 日志
(含 `kubectl logs` / docker compose / journald 采集前缀与 stderr panic 块),
匹配内置故障规则库,输出按严重度排序的诊断报告。设计与规则清单见
rustfs/backlog#1281(总纲)。
**Use this when:** you have RustFS log files from a customer or a cluster (plain, rotated, archived, or `kubectl logs` output) and need an offline root-cause report, or you need to add or override a diagnosis rule without waiting for a release.
**Source of truth:** `rustfs/src/config/cli.rs` (`DiagnoseOpts`, `DiagnoseFormat`); `rustfs/src/diagnose.rs` (exit codes, time parsing); `crates/log-analyzer/src/rules/model.rs` (`Severity`, `Matcher`, `Rule`); `crates/log-analyzer/src/rules/external.rs` (`EXTERNAL_FILE` mirrors the example below).
不启动存储、不联网、跑完即退,任何装有 `rustfs` 二进制的机器都可用。
`rustfs diagnose` parses RustFS JSON logs — including `kubectl logs`, docker compose, and journald collection prefixes, and stderr panic blocks — matches them against the built-in rule library, and prints findings sorted by severity. It starts no storage, opens no network connection, and exits once the report is written; any host with the `rustfs` binary can run it.
## 用法
## Usage
```bash
# 分析一个目录(自动处理 .zst/.gz 轮转归档)
# Analyze a directory (rotated .zst/.gz archives are handled automatically)
rustfs diagnose /var/log/rustfs/
# 客户打包的多节点日志(zip/tar.gz 自动递归展开,第一层目录名作为节点标签)
# Multi-node bundle from a customer (zip/tar.gz are expanded recursively;
# the first-level directory name becomes the node label)
rustfs diagnose customer-logs.zip
# 从 stdin 读(容器场景)
# Read from stdin (container workflows)
kubectl logs rustfs-0 | rustfs diagnose -
# 只看最近 24 小时,输出 Markdown 直接贴工单
# Last 24 hours only, Markdown output ready to paste into a ticket
rustfs diagnose logs.tar.gz --since 24h --format md > report.md
```
常用参数:`--format text|json|md`(默认 text;JSON 带稳定 `schema_version`)、
`--since/--until`(RFC-3339 或相对时间 `30m`/`24h`/`7d`)、`--min-level`
`--top N`(未识别模式条数)、`--samples N`(每个 finding 的样本行数)。
| Flag | Meaning | Default |
| --- | --- | --- |
| `<paths>...` | Files, directories, archives (`.zip`, `.tar`, `.tar.gz`, `.zst`, `.gz`), or `-` for stdin | required |
| `--format text\|json\|md` | Output format; JSON carries a stable `schema_version` | `text` |
| `--since`, `--until` | Time window bounds: RFC-3339, or relative (`30m`, `24h`, `7d`) counted back from now | unbounded |
| `--min-level` | Minimum level to analyze (`trace\|debug\|info\|warn\|error`) | all levels |
| `--top N` | Number of unrecognized error patterns to list | 20 |
| `--samples N` | Sample lines per finding | 3 |
| `--redact` | Hash customer identifiers in the report | off |
| `--rules <file.json>` | Extra rules file; same-id rules override built-ins | none |
## 报告解读
Exit codes: `0` — diagnosis completed, with or without findings; `2` — a rejected argument value (`--since`, `--until`, `--min-level`, `--rules`) or no readable input; `1` — a clap usage error (missing path, unknown flag). Findings never fail the process: this is a diagnosis tool, not a CI gate.
- **发现(按严重度)**:P0 数据风险 → P1 服务不可用 → P2 降级 → P3 客户端侧
→ P4 提示。每条带诊断结论、建议动作、证据字段与样本行。
- **因果折叠**:当症状类发现(如 quorum 刷屏)在时间上跟随其已知根因
(如盘 faulty)出现时,报告把症状折叠进根因块的"级联症状"行,根因块
提升到两者中更高的严重度位置——报告首块直接回答"最可能的原因"。
JSON 输出保留全部 findings(`collapsed_into` / `caused` 字段标注关系)。
- **时间线异常(提示)**:三个确定性启发,均为提示不定罪——混合 UTC 偏移
(伴签名错误时点名时钟偏移)、节点时间范围完全不重叠、时间线断档
(断档后紧跟 startup 类发现时升级为"重启证据")。
- **低频提示**:命中数低于规则阈值的匹配(如零星的签名错误),仅供参考。
- **未识别的高频错误**:规则库未覆盖的 WARN/ERROR 消息模板聚类。这一节是
规则库迭代的输入——反复出现的新模板请提给维护者补规则
(`crates/log-analyzer/src/rules/seed/`)。
- **跳过的输入与时区提示**:所有被跳过的文件(二进制/超限)逐条披露;
混合 UTC 偏移会显式提醒(时钟偏移本身就是 `SignatureDoesNotMatch`
的常见根因)。
## Reading the report
## `--redact`(报告需要转发时)
| Section | Content |
| --- | --- |
| Findings, by severity | `P0 data risk``P1 unavailable``P2 degraded``P3 client side``P4 info`. Each finding carries a diagnosis, a suggested action, evidence fields, and sample lines |
| Causal folding | When a symptom finding (for example a burst of quorum errors) follows its known root cause (for example a faulty disk) in time, the symptom is folded into the root-cause block as a cascaded symptom and that block is promoted to the higher severity of the two, so the first block answers "most likely cause". JSON output keeps every finding and marks the relation with `collapsed_into` / `caused` |
| Timeline anomalies (hints) | Three deterministic heuristics, advisory only: mixed UTC offsets (naming clock skew when signature errors coincide), node time ranges that do not overlap, and log gaps (upgraded to restart evidence when a startup finding follows the gap) |
| Low-confidence hits | Matches below a rule's `min_count` threshold (for example sporadic signature errors), for reference only |
| Unrecognized high-frequency errors | Clustered WARN/ERROR message templates the rule library does not cover. Recurring new templates are the input for new rules under `crates/log-analyzer/src/rules/seed/` |
| Skipped inputs and timezone hints | Every skipped file (binary, or over the size cap) is listed; mixed UTC offsets are called out explicitly because clock skew is a common root cause of `SignatureDoesNotMatch` |
把 bucket/object/AK、IPv4/IPv6、peer 与磁盘路径、节点标签、来源文件路径等客户标识替换为稳定哈希(同值同哈希,保持可关联);规则 id、诊断文本、模块 target 与 panic 源码位置(RustFS 自身代码,非客户数据)保留。样本会连同其完整 `fields` 一起脱敏。
## `--redact`
覆盖是尽力而为,不是绝对保证:结构化字段与 `key=value`/IP 形态的文本都会被处理,但散落在自由文本里、既非字段形态也非 IP 的标识(例如句子里顺带提到的一个 bucket 名)可能仍有残留。转发前建议再抽查一遍。
Replaces customer identifiers — bucket, object, and access-key names, IPv4/IPv6 addresses, peer and disk paths, node labels, source file paths — with stable hashes: equal values hash equally, so correlation survives. Rule ids, diagnosis text, module targets, and panic source locations (RustFS code, not customer data) are kept. Samples are redacted together with their full `fields`.
## 自定义规则(`--rules <file.json>`)
Coverage is best-effort, not a guarantee: structured fields and `key=value` / IP-shaped text are handled, but an identifier embedded in free text that is neither field-shaped nor an IP (for example a bucket name mentioned mid-sentence) can survive. Spot-check the report before forwarding it.
支持团队可以在不等发版的情况下补规则,或热修一条误报的内置规则:
## Custom rules (`--rules <file.json>`)
Support teams can add rules, or hot-fix a false positive in a built-in rule, without waiting for a release:
```bash
rustfs diagnose customer.zip --rules extra-rules.json
```
文件格式(`Rule` 的 JSON 表示与内置规则完全一致):
The file is the JSON form of `Rule`, identical to the built-in rules. Rule text fields (`title`, `diagnosis`, `suggestion`) are operator-facing Chinese by design; this example is the one mirrored by `EXTERNAL_FILE` in `crates/log-analyzer/src/rules/external.rs`:
```json
{
@@ -79,20 +79,18 @@ rustfs diagnose customer.zip --rules extra-rules.json
}
```
- `severity`:`p0_data_risk | p1_unavailable | p2_degraded | p3_client_side | p4_info`;
- `matcher`:`message_prefix` / `message_contains` / `message_regex` /
`field_equals {name, value}` / `target_prefix` / `is_panic` / `min_level` /
`all [..]` / `any [..]`,与内置规则同一套类型;
- 可选字段:`evidence_fields``min_count`(默认 1)`implies_root_cause`
(参与因果折叠)、`anchors`;
- **同 id 覆盖内置规则**(用于热修误报);合并后的规则集整体校验,任何错误
(坏 regex、重复 id、空 matcher 组)会逐条打印并以退出码 2 失败——不会
带着半坏的规则集分析;
- 外部规则的 `anchors` 不受 CI 锚点守卫约束,质量由文件作者自担。
| Field | Values |
| --- | --- |
| `severity` | `p0_data_risk`, `p1_unavailable`, `p2_degraded`, `p3_client_side`, `p4_info` |
| `matcher` | `message_prefix`, `message_contains`, `message_regex`, `field_equals {name, value}`, `target_prefix`, `is_panic`, `min_level`, `all [..]`, `any [..]` — the same types the built-in rules use |
| Optional fields | `evidence_fields`, `min_count` (default 1), `implies_root_cause` (participates in causal folding), `anchors` |
## 已知边界
- A rule with the same `id` as a built-in rule replaces it (this is the hot-fix path for a false positive).
- The merged rule set is validated as a whole; every error (bad regex, duplicate id, empty matcher group) is printed and the command exits with code 2 rather than analyzing with a half-broken set.
- `anchors` in external rules are not checked by the CI anchor guard (`scripts/check_log_analyzer_rules.sh`); their quality is the author's responsibility.
- panic 只出现在 stderr:若客户只采集了 stdout,panic 不会出现在日志里,
`rwlock ... poisoned` 类发现会提示"曾发生 panic"。
- 审计日志(camelCase 外发流)不在本工具范围内。
- 规则锚点由 CI 守卫(rustfs/backlog#1289)保证与源码日志文案同步。
## Known boundaries
- Panics appear only on stderr. If a customer collected stdout only, the panic itself is absent, but a `rwlock ... poisoned` finding hints that a panic occurred.
- Audit logs (the camelCase outbound stream) are out of scope.
- Built-in rule anchors are kept in sync with the source log messages by the CI guard `scripts/check_log_analyzer_rules.sh`.
+79 -314
View File
@@ -1,39 +1,25 @@
# NATS JetStream Operations Guide
The guide covers enabling and operating the JetStream publish path for the NATS
notify and audit targets in RustFS. It is written for operators who need
at-least-once delivery of events to NATS, must pre-provision and size the stream
the targets publish to, and need to diagnose delivery failures.
**Use this when:** you need at-least-once delivery of notification or audit events to NATS, must pre-provision and size the stream RustFS publishes to, or are diagnosing JetStream delivery failures.
**Source of truth:** `crates/config/src/constants/targets.rs` (`NATS_JETSTREAM_*` keys and the ack-timeout bounds), `crates/config/src/notify/nats.rs` and `crates/config/src/audit/nats.rs` (env names), `crates/targets/src/target/nats/{validation,jetstream}.rs` (stream validation, `retry_lifetime`, outcome classification), `crates/targets/src/runtime/mod.rs` (`REPLAY_MAX_RETRIES`, `REPLAY_BASE_RETRY_DELAY`), `crates/targets/src/store.rs` (`FAILED_STORE_MAX_ENTRIES`, `FAILED_STORE_TTL`).
## What the JetStream Path Does
## What the JetStream path does
By default the NATS notify and audit targets publish with NATS Core, which
returns once the message is written to the socket. A process or server failure
in the gap between the socket write and the server persisting the message loses
the event. The JetStream path closes that gap. With it enabled, each event is
published to a JetStream stream and the local store-and-forward queue entry
clears only after the server returns a publish acknowledgement, which means the
stream leader has accepted and sequenced the message, or after a terminal
rejection has been recorded in the failed-events store. A server restart
mid-flight loses nothing, because an unacknowledged event stays in the local
queue and replays after the server returns.
By default the NATS notify and audit targets publish with NATS Core, which returns once the message is written to the socket; a failure between that write and the server persisting the message loses the event. With JetStream enabled, each event is first written to the local store-and-forward queue and then published to a JetStream stream with a stable `Nats-Msg-Id` header; the queue entry clears only after the server returns a publish acknowledgement (the stream leader has accepted and sequenced the message) or after a terminal rejection has been recorded in the failed-events store. An unacknowledged event survives a process restart on disk and replays.
The path is opt-in and off by default. With it off, behaviour is the NATS Core
path. It applies to both the notify NATS target and the audit NATS target.
The path is opt-in and off by default, and applies to both the notify NATS target and the audit NATS target. RustFS never creates the stream: the operator owns the stream and its retention, storage, and replication policy. RustFS validates that the stream exists and is writable, reports a validation failure otherwise, and a store-backed target keeps queueing events through a failed validation until the stream is repaired.
RustFS gets each event to the server without losing it before the
acknowledgement. The operator owns the stream and its policy. RustFS validates
that the stream exists and is writable and reports a validation failure
otherwise. A store-backed target keeps running through a failed validation,
holding queued events until the stream is repaired. RustFS never creates the
stream. Stream retention, storage, and replication policy stay under operator
control.
## Configuration
## Enabling JetStream: Recommended Configuration
Each key has a configuration-key form and an environment-variable form per target; the audit target uses the `RUSTFS_AUDIT_NATS_` prefix in place of `RUSTFS_NOTIFY_NATS_`.
Three keys turn the path on and tune it. Each is available on both the notify
NATS target and the audit NATS target, and each has a configuration-key form and
an environment-variable form per target.
| Key | Env (notify) | Required when enabled | Default / range | Meaning |
| --- | --- | --- | --- | --- |
| `jetstream_enable` | `RUSTFS_NOTIFY_NATS_JETSTREAM_ENABLE` | — | off | Turns the JetStream path on for the target |
| `jetstream_stream_name` | `RUSTFS_NOTIFY_NATS_JETSTREAM_STREAM_NAME` | yes | none | Pre-provisioned stream to publish to |
| `jetstream_ack_timeout_secs` | `RUSTFS_NOTIFY_NATS_JETSTREAM_ACK_TIMEOUT_SECS` | no | `NATS_JETSTREAM_ACK_TIMEOUT_DEFAULT_SECS` (30); accepted range `NATS_JETSTREAM_ACK_TIMEOUT_MIN_SECS`..=`NATS_JETSTREAM_ACK_TIMEOUT_MAX_SECS` (10..=120) | How long a publish waits for an acknowledgement before it is treated as timed out and retried. Each attempt, including connection establishment, is bounded by this deadline |
| `queue_dir` | `RUSTFS_NOTIFY_NATS_QUEUE_DIR` | yes | none | Local store-and-forward queue directory; durability needs a local store to replay from |
| `queue_limit` | `RUSTFS_NOTIFY_NATS_QUEUE_LIMIT` | no | target default | Maximum live-queue entries; see sizing below |
```bash
RUSTFS_NOTIFY_NATS_JETSTREAM_ENABLE=true
@@ -42,315 +28,94 @@ RUSTFS_NOTIFY_NATS_JETSTREAM_ACK_TIMEOUT_SECS=30
RUSTFS_NOTIFY_NATS_QUEUE_DIR=/var/lib/rustfs/notify-nats
```
The audit target takes the same keys under the RUSTFS_AUDIT_NATS_ prefix.
Enabling the path without a stream name or without a queue directory, or with an out-of-range acknowledgement timeout, is rejected at startup and in the admin validation path (`validate_jetstream_settings`). The default 30 s timeout suits production: a server replicating to several replicas acknowledges only after replication and the deferred fsync, which legitimately takes well over 100 ms at the tail.
- jetstream_enable turns the JetStream path on for the target. Off by default.
Environment forms RUSTFS_NOTIFY_NATS_JETSTREAM_ENABLE and
RUSTFS_AUDIT_NATS_JETSTREAM_ENABLE.
- jetstream_stream_name is the name of the pre-provisioned stream to publish to.
Required when enable is on, with no default. Environment forms
RUSTFS_NOTIFY_NATS_JETSTREAM_STREAM_NAME and
RUSTFS_AUDIT_NATS_JETSTREAM_STREAM_NAME.
- jetstream_ack_timeout_secs is how long a publish waits for an acknowledgement
before it is treated as timed out and retried. Range 10 to 120, default 30.
Environment forms RUSTFS_NOTIFY_NATS_JETSTREAM_ACK_TIMEOUT_SECS and
RUSTFS_AUDIT_NATS_JETSTREAM_ACK_TIMEOUT_SECS.
- queue_dir is the local store-and-forward queue directory and is required when
enable is on. Durability is not achievable without a local store to replay
from. Environment forms RUSTFS_NOTIFY_NATS_QUEUE_DIR and
RUSTFS_AUDIT_NATS_QUEUE_DIR.
A mistyped key name is not a value error and escapes that check: the key reads as absent and the target silently stays on the NATS Core path. After enabling, confirm the target reports its JetStream fields on startup and logs the stream-validation success line naming the configured stream; absence of that line for a target meant to be enabled indicates an unrecognised key name.
A configuration that enables the path without a stream name or without a queue
directory is rejected at startup and in the admin validation path. An
acknowledgement timeout outside the 10 to 120 second range is rejected the same
way.
## Stream requirements
The 30 second acknowledgement timeout suits production. A server replicating to
multiple replicas acknowledges only after the replication and the deferred fsync,
which legitimately takes well over 100 milliseconds at the tail, so a short
timeout would spuriously fail a publish mid-replication.
RustFS reads the stream once at target init and in the admin validation path, and reports a validation failure rather than publishing into a stream where writes would silently fail. Revalidation runs while the verdict is unset and stops once one validation passes, resuming only after the verdict is reset (reconnect, TLS rotation, an acknowledgement naming an unexpected stream, or a stream-not-found publish outcome). The stream must:
After enabling, confirm the target is on the JetStream path before relying on it.
Value errors, an out-of-range acknowledgement timeout or a missing stream name
while enable is on, are already rejected loudly at startup and in the admin
validation path. A mistyped configuration key name is not a value error, so it
escapes that check: the key reads as absent, and the target stays on the NATS
Core path with no durability. Verify the enable took effect by confirming the
target reports its JetStream fields on startup and logs the stream-validation
success line naming the configured stream. Absence of that line for a target that
was meant to be enabled indicates a key name the server did not recognise.
1. Exist and be writable. A missing or unreachable stream fails validation; a stream provisioned after the path is enabled picks up the queued events on the next replay, because stream-not-found is retryable.
2. Capture the configured publish subject in its subject filter, by literal match or NATS wildcard.
3. Acknowledge writes (`no_ack` false); otherwise the queue would never clear.
4. Not be sealed.
5. Set a duplicate window of at least the worst-case retry span (next section).
## How Delivery Works
Choose retention, storage type, and replica count for the durability the deployment needs. Use file storage if a server restart must preserve already-persisted events. For durability across a node failure, provision at least 3 replicas: an acknowledgement returns after the stream leader commits, and a leader that acknowledges and then fails before replicating to a quorum can lose that message on failover, so a single-replica stream is durable only against a clean restart.
An event is queued to the local store-and-forward queue first, then published
through JetStream. The publish carries a stable Nats-Msg-Id header and the path
awaits the real publish acknowledgement.
## Duplicate window and the retry span
- The queue entry clears only after the durable acknowledgement returns, or
after a terminal rejection has been recorded in the failed-events store. A
timeout, or any retryable error, retains the entry for a retry. The entry is
never cleared without one of those two outcomes and never silently dropped. A
process interruption between the failed-record write and the live-entry removal
can leave a record whose event a later replay still delivers, a diagnostic
residue and not a lost event.
- The Nats-Msg-Id is minted once when the event is queued and stored with it. It
is identical across every retry and replay of that entry, and unique per
entry, so the server collapses retries and replays of the same event within
the stream duplicate window. A retry after a slow acknowledgement reuses the
same identifier and is not delivered twice.
- A crash before the acknowledgement leaves the entry in the queue. The entry
replays after restart. Closing the target releases the cached connection and
context but never the queued entries, so an entry queued before close survives
on disk and replays on the next start.
- A broker outage detected at connection establishment keeps queued events on
disk. Replay retries with backoff while the broker is unreachable, and the
queued events deliver when the connection recovers, keeping entries queued like
the NATS Core path rather than replicating its mechanics byte for byte. Every
retryable publish failure on an established connection is treated the same way,
kept on the live queue and retried until it delivers, which includes a
connection dropping during the stream-validation lookup. Only a non-retryable
rejection is moved to the failed-events store.
The `Nats-Msg-Id` is minted once when the event is queued and reused on every retry and replay of that entry, so the server collapses duplicates within the stream duplicate window. The window must therefore cover the worst-case span from the first publish attempt of a stored entry to its last within one replay cycle, or a late retry of an already-persisted event is delivered twice.
For durability across a node failure, provision the stream with a replica count
of at least 3. An acknowledgement returns after the stream leader commits. A
leader that acknowledges and then fails before the message replicates to a quorum
can lose that message on failover, so a single-replica stream is durable only
against a clean restart, not against the loss of the node holding the data.
`retry_lifetime` (`crates/targets/src/target/nats/jetstream.rs`) computes that span, and validation rejects a stream whose duplicate window is below it, naming the configured and required window in the log line:
## Pre-provisioning the Stream
```text
duplicate_window >= REPLAY_MAX_RETRIES * ack_timeout_secs
+ inter_attempt_backoff_sum(REPLAY_MAX_RETRIES)
+ replay_backoff_term(REPLAY_MAX_RETRIES)
```
The stream is provisioned by the operator before the path is enabled. RustFS
reads the stream once at target init and in the admin validation path, asserts
the requirements below, and reports a validation failure rather than publishing
into a stream where writes would silently fail. A failed validation does not
stop a store-backed target. Events keep queueing to the local store. Revalidation
runs while the validation verdict is unset and stops once one validation passes,
resuming only after the verdict is reset, and the queued events deliver once the
stream is repaired. The stream is not auto-created. A missing stream is an
error.
- `REPLAY_MAX_RETRIES` is the number of publish attempts per replay cycle.
- The sleep before the retry at shift `n` is `replay_backoff_term(n) = REPLAY_BASE_RETRY_DELAY * 2^n`; `inter_attempt_backoff_sum` adds the terms for shifts `1..REPLAY_MAX_RETRIES` (no sleep follows the last attempt).
- The final `replay_backoff_term(REPLAY_MAX_RETRIES)` is a deliberate headroom term above the realized span.
The stream must:
With the constants as shipped (5 attempts, 2 s base) the backoff sum is 60 s and the headroom 64 s, so the default 30 s acknowledgement timeout requires a window of at least 274 s; each extra second of acknowledgement timeout adds `REPLAY_MAX_RETRIES` seconds to the required window. Raise the stream duplicate window in step whenever you raise `jetstream_ack_timeout_secs`.
- Exist and be writable. A missing or unreachable stream fails validation.
- Capture the configured publish subject in its subject filter, by literal match
or by a NATS wildcard. A stream that does not capture the subject fails
validation.
- Acknowledge writes (no_ack false). A stream with no_ack set never returns an
acknowledgement, so the queue would never clear. It fails validation.
- Not be sealed. A sealed stream rejects writes and fails validation.
- Set a duplicate window of at least the retry span (see the next section). A
window below that span fails validation.
The validated window covers one retry cycle. An entry that exhausts a cycle without delivering stays on the live queue and is retried on a later cycle, so an entry surviving across cycles can be delivered again — consistent with at-least-once delivery. Throttling `RUSTFS_NOTIFY_TARGET_STREAM_CONCURRENCY` below the number of concurrently backlogged targets inserts untimed waits between the retries of one entry and can push a late retry past the window; keep the default or add window margin when lowering it.
Choose retention, storage type, and replica count to match the durability the
deployment needs. Use file storage if a server restart must preserve
already-persisted events. These are operator decisions and RustFS does not set
them. If the stream is provisioned shortly after the path is enabled, the queued
events deliver on the next replay rather than being lost, because a
stream-not-found error is retryable and keeps the events on the live queue until
the stream exists.
## Retryable versus terminal outcomes
## Setting the Duplicate Window
The acknowledgement timeout and the stream duplicate window are linked. The
duplicate window must cover the worst-case retry span, or a late retry of an
event the server already persisted is delivered a second time instead of being
recognised as a duplicate.
The worst-case retry span is the retry count times the acknowledgement timeout,
plus the realized backoff sleeps across the retries and a final headroom term:
duplicate_window >= (retry_count * ack_timeout_secs) + realized_backoff + headroom
There are 5 retry attempts. The realized backoff sleeps sum to 60 seconds
(4 + 8 + 16 + 32), and no sleep follows the last attempt, so a 64 second headroom
term is added above the realized span. At the default 30 second acknowledgement
timeout the required window is:
(5 * 30) + 60 + 64 = 274 seconds
Set the stream duplicate window to at least 274 seconds when the acknowledgement
timeout is at the 30 second default. A configuration whose duplicate window is
below this span is rejected at validation, with the configured and required
window named in the log line.
The span scales with the acknowledgement timeout. Raising
jetstream_ack_timeout_secs requires raising the stream duplicate window in step
using the formula above. The 60 second realized backoff and the 64 second headroom
stay fixed, so each extra second of acknowledgement timeout adds 5 seconds to the
required window. At the 120 second maximum timeout the required window is
(5 * 120) + 60 + 64 = 724 seconds. Each publish attempt, including connection
establishment, is bounded by a single deadline equal to the acknowledgement
timeout. The formula is a conservative upper bound on the retry span, and the
final headroom term holds the required window above the realized span.
The validated window covers one retry cycle: the worst-case span from the first
publish attempt of a stored entry to its last within a single replay cycle. An
entry that exhausts a cycle without delivering stays on the live queue and is
retried on a later cycle, so an entry surviving across cycles can spend longer in
the queue than the duplicate window and be delivered again, consistent with
at-least-once delivery.
## The Failed-events Store
An event that cannot be delivered is recorded, not silently dropped. Only one
case produces a failed event:
- A terminal rejection. A message that exceeds the maximum payload size, a wrong
expected last message identifier or last sequence, or a sealed stream. These do
not improve with a retry, so they fail fast.
Every other rejection is retryable and never produces a failed event. A timeout,
backpressure, no responders, a missing stream, a subject the stream does not
capture, or an authorization failure keeps the event on the live queue and
retries it until it delivers or an operator intervenes. A wrong subject and wrong
credentials surface through the validation and connection path as retryable, so
they hold the events on the live queue rather than moving them to the failed
store.
A failed event is moved to an on-disk failed store that sits in a failed
child directory inside the queue directory, one per target. The directory is created
lazily on the first failed write, so a target that never fails terminally leaves
no failed directory. The entry preserves the event body, its routing metadata, the
deduplication identifier, an error class tag of terminal, the failure time, and
the retry count. An error-level log line is written for each failed event, naming
the bucket, object, event name, and the error, so a broken integration is visible
rather than hidden. A record in the failed store is diagnostic only and is never
republished, so a condition repaired later, for example a raised broker payload
limit, delivers only events still on the live queue. The NATS Core path without
JetStream instead retries such an event until the limit allows it.
The failed store is bounded by count (10000 entries per target) and by age (a 72
hour retention), and is kept separate from the live queue limit so an
accumulation of failures cannot crowd out new events. The count is maintained as a
cached value, seeded at startup and reconciled to the directory on each
maintenance interval, so it stays accurate without a directory scan on the hot
path. Failed-store writes and the maintenance scan run under one exclusive guard,
so the at-bound check and the write stay atomic and the bound holds against
concurrent writers. A change made to the directory outside the store can drift the
cached count until the next maintenance interval reconciles it. When the count
bound is reached the oldest failed entry is dropped, with a warning naming the
trimmed entry, so a newer failure is never lost in favour of an older one. Entries
past the retention bound are removed as expired on the replay maintenance tick.
Because a retryable failure keeps events on the live queue, a long outage grows
the queue toward its configured queue_limit bound. Once the queue reaches that
bound, new events are rejected at ingest with a logged error rather than
overwriting queued events, so the backlog is bounded and visible.
Size queue_limit for the longest outage the deployment must survive without
rejecting new events. Multiply the peak event rate in events per second by the
outage window in seconds. A target that averages 50 events per second and must
ride out a one hour broker outage needs a queue_limit of at least 50 * 3600 =
180000 entries. Add headroom above the calculated figure, and provision the
queue directory storage for the resulting entry count.
## Publish Outcome Handling
Every publish outcome falls into one of two families. Transient conditions are
retried until they deliver and the entry stays on the live queue, so a transient
condition never reaches the failed store. Permanent conditions move to the failed
store immediately. Transient conditions classify toward retry because a terminal
misclassification risks losing an entry once the failed store retention lapses.
Permanent conditions move immediately so a poison message cannot block the queue.
Both families cover publish outcomes on an established connection. A failure to
establish the connection at all, a refused connection or unreadable TLS
material, is handled before either family applies: the entry retries with
backoff and stays queued until the connection recovers. A publish-level failure
on an established connection, including a connection that drops during the
stream-validation lookup, is retried on the live queue when it is transient and
moved to the failed store only when it is a permanent rejection.
Every publish outcome falls into one of two families, classified once in `crates/targets/src/target/nats/jetstream.rs`. Retryable conditions keep the entry on the live queue and retry until it delivers or an operator intervenes; they never reach the failed store, because a misclassification there would lose an entry once the failed-store retention lapses. Terminal conditions move to the failed store immediately so a poison message cannot block the queue. A failure to establish the connection at all (refused connection, unreadable TLS material) is handled before either family: the entry retries with backoff and stays queued until the connection recovers.
| Outcome family | Examples | Handling |
| --- | --- | --- |
| Connectivity | connection lost mid-publish, broken pipe | Retried until delivered, live queue |
| Timeouts | no acknowledgement within the timeout, attempt deadline reached | Retried until delivered, live queue |
| Cluster in transition | no leader elected, peer membership changing | Retried until delivered, live queue |
| Stream offline | stream or JetStream subsystem temporarily offline, stream not found | Retried until delivered, live queue |
| Resource and quota exhaustion | insufficient server resources, storage, memory, or account quota reached | Retried until delivered, live queue |
| Server errors | any rejection reporting a 5xx status, with or without a specific error code | Retried until delivered, live queue |
| Permanent rejections | payload too large, wrong expected sequence or message identifier, sealed stream | Failed store immediately |
| Connectivity | Connection lost mid-publish, broken pipe, connection dropped during the stream-validation lookup | Retryable |
| Timeouts | No acknowledgement within the timeout, attempt deadline reached | Retryable |
| Cluster in transition | No leader elected, peer membership changing | Retryable |
| Stream offline | Stream or JetStream subsystem temporarily offline, stream not found | Retryable |
| Resource and quota exhaustion | Insufficient server resources, storage, memory, or account quota | Retryable |
| Server errors | Any rejection reporting a 5xx status | Retryable |
| Wrong subject, wrong credentials, no responders, authorization failure | Surface through the validation and connection path | Retryable — they hold events on the live queue until the configuration is fixed |
| Permanent rejections | Payload exceeds the maximum size, wrong expected last message id or last sequence, sealed stream | Terminal — failed store immediately |
A sealed stream is classified at three points, with three outcomes. Startup validation and the admin
validation path reject a sealed stream before the target serves traffic, so init fails and no event
is queued against it. A stream sealed while the target runs is caught by the next validation on the
publish path, which classifies it retryable and keeps the entry on the live queue while validation
keeps failing. A publish that reaches an already-cached validation pass and is then rejected with the
sealed-stream code is terminal and moves to the failed store immediately. The permanent-rejections
row lists that last outcome.
A sealed stream is classified at three points: startup and admin validation reject it before the target serves traffic; a stream sealed while the target runs is caught by the next validation on the publish path and treated as retryable; a publish that reaches an already-cached validation pass and is then rejected with the sealed-stream code is terminal.
A process interruption between the failed-record write and the live-entry removal can leave a failed record whose event a later replay still delivers — a diagnostic residue, not a lost event.
## The failed-events store
A terminal rejection is recorded, not silently dropped. The failed store is an on-disk `failed` child directory inside the queue directory, one per target, created lazily on the first terminal write. Each entry preserves the event body, routing metadata, the deduplication identifier, an error class tag of `terminal`, the failure time, and the retry count, and an error-level log line names the bucket, object, event name, and error. Records are diagnostic only and are never republished: a condition repaired later (for example a raised broker payload limit) delivers only events still on the live queue, whereas the NATS Core path would have kept retrying such an event.
The store is bounded by `FAILED_STORE_MAX_ENTRIES` per target and by `FAILED_STORE_TTL` (`crates/targets/src/store.rs`), separately from the live queue limit so failures cannot crowd out new events. When the count bound is reached the oldest failed entry is dropped with a warning naming it; entries past the TTL are removed on the replay maintenance tick. The count is a cached value seeded at startup and reconciled to the directory on each maintenance interval; writes and the maintenance scan share one exclusive guard so the bound holds against concurrent writers, and a change made to the directory outside the store drifts the cached count until the next reconciliation.
## Sizing `queue_limit`
Because retryable failures keep events on the live queue, a long outage grows the queue toward `queue_limit`. At the bound, new events are rejected at ingest with a logged error rather than overwriting queued events. Size it for the longest outage the deployment must survive without rejecting: peak events per second times the outage window in seconds — a target averaging 50 events/s that must ride out a one-hour broker outage needs at least 50 * 3600 = 180000 entries — plus headroom, and provision the queue directory storage accordingly.
## Observability
The failed_store_length gauge reports the number of entries in the failed-events
store per target, next to the existing failed_messages and queue_length gauges,
on both the notify and audit metric paths. A rising failed_store_length points to
terminal rejections accumulating for a target. An exhaustion warning marks each
cycle where an entry spends its full retry budget without delivering, after which
the entry stays queued and retries on the next scan. The warning repeats once per
retry cycle while the entry stays queued, so a persistent delivery problem is
visible at warn level without per-attempt noise. The failed_messages count
advances only on terminal and dropped events, never on a retry or an exhaustion,
so it counts entries that left the queue for good rather than entries still
retrying.
| Gauge | Meaning |
| --- | --- |
| `queue_length` | Entries on the live queue (existing gauge) |
| `failed_store_length` | Entries in the failed-events store per target; a rising value means terminal rejections are accumulating |
| `failed_messages` | Advances only on terminal and dropped events, never on a retry or an exhaustion — entries that left the queue for good |
All three are emitted on both the notify and audit metric paths (`crates/targets/src/runtime/mod.rs`). An exhaustion warning marks each cycle in which an entry spends its full retry budget without delivering; the entry stays queued and the warning repeats once per cycle, so a persistent delivery problem is visible at warn level without per-attempt noise.
## Troubleshooting
- Validation fails at startup with a missing-stream error. The stream named in
jetstream_stream_name does not exist on the server, or is not reachable.
Provision it, or correct the name.
- Validation fails with a subject, no_ack, sealed, or duplicate-window error. The
pre-provisioned stream does not capture the publish subject, has no_ack set, is
sealed, or has a duplicate window below the worst-case retry span. Correct the
stream configuration to satisfy the pre-provisioning requirements above.
- Repeated stream-not-found errors in the log. The stream disappeared or was
never created. The publish is retryable, so the events stay on the live queue
and retry. Provision the stream so the queued events deliver on the next
replay.
- A warning that a publish was acknowledged by an unexpected stream. The
configured stream no longer captures the publish subject and another stream
does. The mismatched publish is rejected with a retryable error, the entry
stays on the live queue, and the mismatch resets the stream-validation
verdict. Later retries then fail validation before publishing, so the log
shows the validation failure rather than a repeating mismatch warning. The
mismatch warning fires again only after a validation pass lets another
publish through. The acknowledging stream named in the warning has persisted
a copy of each mismatched publish. Inspect that stream and remove the stray
copies. Delivery resumes once the configured stream captures the subject
again and validation passes.
- Health checks slow while a stream fails validation. Health reflects the last
stream-validation verdict rather than a live lookup on every check. The verdict
resets on a reconnect, a TLS rotation, an acknowledgment naming an unexpected
stream, or a stream-not-found publish outcome, and a reset forces a live stream
lookup on the next check or publish, bounded by the acknowledgement timeout. A
health snapshot taken right after a reset can take up to that timeout per
affected target until the stream is repaired. After a broker connection heals
on its own, publishes and health rely on the last validation verdict until an
error outcome resets it, a wrong-stream acknowledgment or a stream-not-found
publish outcome, so a stream reconfigured during a silent reconnect is
detected on the first publish evidence rather than immediately.
- Events stay in the queue and do not clear. The server is not acknowledging.
Check connectivity, that the stream subject filter captures the publish
subject, and that the server has capacity. Unacknowledged events are retained,
not lost.
- Duplicate deliveries observed. The duplicate window is shorter than the
worst-case retry span. Raise the stream duplicate window using the formula
above, especially after raising the acknowledgement timeout. An outage or a
restart that keeps an unacknowledged entry queued longer than the duplicate
window can also produce a duplicate on replay, consistent with at-least-once
delivery. Throttling RUSTFS_NOTIFY_TARGET_STREAM_CONCURRENCY below the number
of concurrently backlogged targets inserts untimed waits between the retries
of one entry and can push a late retry past the window. Keep the default or
add window margin when lowering it.
- Failed-store entries accumulating. A terminal rejection is recurring: a
message over the maximum payload size, a wrong expected last message identifier
or last sequence, or a sealed stream. Failed entries are on-disk diagnostic
records carrying the error class and a fixed diagnostic detail for inspection.
They are not re-published, and they are removed after the retention period. A
wrong subject or wrong credentials do not land here. They surface as retryable
and hold the events on the live queue until the configuration is fixed.
| Symptom | Cause | Action |
| --- | --- | --- |
| Validation fails at startup with a missing-stream error | The stream named in `jetstream_stream_name` does not exist or is unreachable | Provision it, or correct the name |
| Validation fails with a subject, `no_ack`, sealed, or duplicate-window error | The stream violates one of the requirements above | Correct the stream configuration |
| Repeated stream-not-found errors in the log | The stream disappeared or was never created; the publish is retryable | Provision the stream; queued events deliver on the next replay |
| Warning that a publish was acknowledged by an unexpected stream | The configured stream no longer captures the publish subject and another stream does. The publish is rejected as retryable and the validation verdict is reset, so later retries fail validation instead of repeating the warning | Inspect the acknowledging stream named in the warning and remove the stray copies; delivery resumes once the configured stream captures the subject and validation passes |
| Health checks slow while a stream fails validation | Health reflects the last validation verdict; a reset forces a live stream lookup on the next check or publish, bounded by the acknowledgement timeout, so a snapshot right after a reset can take up to that timeout per affected target | Repair the stream. After a silent reconnect, a reconfigured stream is detected on the first publish evidence (wrong-stream acknowledgement or stream-not-found), not immediately |
| Events stay in the queue and do not clear | The server is not acknowledging | Check connectivity, that the subject filter captures the publish subject, and server capacity. Unacknowledged events are retained, not lost |
| Duplicate deliveries observed | Duplicate window shorter than the worst-case retry span, an outage or restart that kept an entry queued longer than the window, or `RUSTFS_NOTIFY_TARGET_STREAM_CONCURRENCY` throttled below the backlogged-target count | Raise the stream duplicate window per the formula above, especially after raising the acknowledgement timeout |
| Failed-store entries accumulating | A terminal rejection is recurring (payload too large, wrong expected sequence or message id, sealed stream) | Fix the cause; the records are diagnostic, are not republished, and expire after `FAILED_STORE_TTL`. Wrong subject or credentials never land here — they hold events on the live queue |
## Disabling the Path
## Disabling the path
Set jetstream_enable off. The target reverts to the NATS Core path. Events
already in the queue are delivered by the standard replay. No failed-store
entries are created while the path is off.
Set `jetstream_enable` off. The target reverts to the NATS Core path; events already in the queue are delivered by the standard replay, and no failed-store entries are created while the path is off.
+25 -74
View File
@@ -1,98 +1,49 @@
# No-Parity Bitrot Recovery Guide
This guide covers historical objects written with erasure data shards but no
parity shards, for example an object whose `xl.meta` reports `EcM=1` and
`EcN=0`. PR #5179 prevents RustFS from committing new objects after it detects
this class of no-parity bitrot failure, but operators may still find already
committed objects on disk.
**Use this when:** a GET or deep heal reports `FileCorrupt` / `bitrot hash mismatch` on an object whose `xl.meta` shows `EcN=0` (data shards, no parity), and the raw `part.N` file is still readable from the filesystem.
**Source of truth:** `crates/ecstore/src/set_disk/ops/heal.rs` (heal returns `FileCorrupt` for confirmed no-parity bitrot, `ErasureReadQuorum` otherwise); `crates/filemeta/examples/dump_fileinfo.rs` (offline `xl.meta` decoder).
This guide covers historical objects written with erasure data shards but no parity shards (for example `EcM=1`, `EcN=0`). The write path now self-verifies no-parity writes and refuses to commit an object whose shard fails bitrot verification, so new objects of this class cannot be created, but already committed ones may still exist on disk.
## Symptom
An affected object can look surprising during incident response:
- The raw shard file (`part.1`) is visible and readable from the local filesystem.
- An S3 GET or deep heal reports an integrity failure: `FileCorrupt`, `bitrot hash mismatch`, an unrecoverable heal result, or a truncated streaming response.
- No parity shards exist to reconstruct the corrupted data shard.
- the raw shard file, such as `part.1`, is visible and readable from the local
filesystem;
- an S3 GET or deep heal reports an integrity failure, commonly through
`FileCorrupt`, `bitrot hash mismatch`, or an unrecoverable heal result;
- there are no parity shards available to reconstruct the corrupted data shard.
The filesystem-readable `part.N` is evidence, not trusted object data: the stored hash no longer matches the bytes on disk, and RustFS must not bypass bitrot validation to serve it.
The filesystem-readable `part.N` file is therefore evidence, not trusted object
data. RustFS must not bypass bitrot validation to serve it through S3, because
the stored hash no longer matches the bytes on disk.
## What To Capture
## What to capture
Before deleting or moving anything, capture:
- bucket name, object key, and version ID if versioning is enabled;
- the RustFS version and whether the deployment was running with no parity
(`EcN=0`) at the time the object was written;
- the heal or GET error, including any `FileCorrupt`, `bitrot hash mismatch`,
`ErasureReadQuorum`, or truncated streaming response message;
- `xl.meta` from every shard disk that still has the object;
- the raw `part.N` file from every shard disk that still has the object.
1. Bucket name, object key, and version ID if versioning is enabled.
2. The RustFS version, and whether the deployment ran with no parity (`EcN=0`) when the object was written.
3. The heal or GET error text.
4. `xl.meta` and the raw `part.N` file from every shard disk that still has the object.
For local inspection, decode metadata with:
Decode metadata locally and record the erasure geometry (`EcM`, `EcN`), object size, part number, part logical size, data directory, and checksum algorithm:
```bash
cargo run -p rustfs-filemeta --example dump_fileinfo -- /path/to/disk/bucket/object/xl.meta
```
Record the erasure geometry (`EcM`, `EcN`), object size, part number, part
logical size, data directory, and checksum algorithm from the decoded metadata.
## Size accounting
## Size Accounting
Erasure shard files include bitrot hash data in addition to object bytes: with the default `HighwayHash256S` checksum each protected block adds 32 bytes, so a raw `part.1` larger than the logical object size is normal (for example 8,250,370 logical bytes in 8 blocks → 8,250,626 raw bytes). The size relationship only shows the layout is plausible; the bitrot reader is the authority for integrity.
RustFS erasure shard files include bitrot hash data in addition to object bytes.
For the default HighwayHash256S checksum, each protected block adds 32 bytes of
hash data to the shard file. A raw `part.1` size can therefore be larger than
the object logical size and still be normal.
## Recovery boundary
Example:
If `EcN=0` and a data shard fails bitrot verification, RustFS cannot reconstruct the object from the erasure set. The valid options are:
```text
logical object bytes: 8,250,370
protected blocks: 8
hash overhead: 8 * 32 = 256 bytes
raw part.1 bytes: 8,250,626
```
- Restore the object from an external backup, replica, upstream source, or a known-good copy outside the affected erasure set.
- Preserve the affected `xl.meta` and `part.N` files as incident evidence, then delete the object through the normal S3/admin path when retention policy allows.
- Quarantine by copying evidence out of the live data path first; remove or isolate the live object path only after the incident owner confirms the evidence is no longer needed.
This size relationship only proves that the file layout is plausible. It does
not prove the bytes are valid. The bitrot reader is the authority for integrity.
Do not edit `xl.meta`, rewrite `part.N`, or serve raw shard bytes to clients as the object. Those actions hide evidence and convert a detected integrity failure into silent data corruption.
## Recovery Boundary
If `EcN>0`, this guide is not the primary recovery path: run normal heal first, since parity may allow RustFS to reconstruct the missing or corrupt shard.
If `EcN=0` and a data shard fails bitrot verification, RustFS cannot reconstruct
the object from the erasure set. The valid recovery options are:
## Expected diagnostics
- restore the object from an external backup, replica, upstream source, or a
known-good copy outside the affected erasure set;
- preserve the affected `xl.meta` and `part.N` files as incident evidence, then
delete the object through the normal S3/admin path when retention policy
allows it;
- quarantine by copying evidence out of the live data path first, then remove
or isolate the live object path only after the incident owner confirms the
evidence is no longer needed.
Do not edit `xl.meta`, rewrite `part.N`, or serve raw shard bytes to clients as
the object. Those actions hide evidence and can convert a detected integrity
failure into silent data corruption.
If `EcN>0`, this guide is not the primary recovery path. Use normal heal first;
parity may allow RustFS to reconstruct the missing or corrupt shard.
## Expected Diagnostics
Deep heal should report no-parity corruption as an unrecoverable integrity
failure rather than only a generic read-quorum problem. The diagnostic context
should include:
- bucket, object, and version ID;
- erasure data shard count and parity shard count;
- part number;
- whether the failing part had a bitrot failure;
- the number of missing or corrupt shards.
When the failure is confirmed bitrot on a no-parity object, the heal error is
reported as `FileCorrupt`, and the heal result `detail` states that the
no-parity object is unrecoverable.
Deep heal reports no-parity corruption as an unrecoverable integrity failure rather than a generic read-quorum problem: the heal error is `FileCorrupt`, and the heal result `detail` states that the no-parity object is unrecoverable. The diagnostic context includes bucket, object, and version ID; data and parity shard counts; part number; whether the failing part had a bitrot failure; and the number of missing or corrupt shards.
+42 -113
View File
@@ -1,50 +1,29 @@
# Object I/O (GET/PUT) tuning A/B matrix runbook
> Scope: **parameter tuning** measuring the effect of changing one
> `RUSTFS_*` runtime knob at a time, against a fixed binary. This is
> deliberately different from the code-change A/B gate in
> [`hotpath-warp-ab-runbook.md`](hotpath-warp-ab-runbook.md) and the formal
> ABBA validation in
> [`hotpath-warp-abba-runbook.md`](hotpath-warp-abba-runbook.md), which compare
> a baseline binary against a candidate binary.
>
> When a knob change turns out to need a code change, use those two runbooks
> for the code-level validation and come back here for the knob-level sweep.
**Use this when:** you want to measure the effect of changing one `RUSTFS_*` runtime knob at a time against a fixed binary, or you need the catalog of GET/PUT tuning knobs, their defaults, and the stage metric that validates each.
**Source of truth:** `crates/io-metrics/src/lib.rs` (stage histogram names and stage tokens); `crates/config/src/constants/object.rs` and `crates/ecstore/src/erasure/coding/encode.rs` (knob defaults); `crates/ecstore/src/set_disk/mod.rs` (codec-streaming rollout and engine defaults); `rustfs/src/app/object/get.rs` (GET experimental switches).
## 1. What this runbook answers
Scope: **parameter tuning** with a fixed binary. The code-change A/B gate that compares a baseline binary against a candidate binary is [`hotpath-warp-ab-runbook.md`](hotpath-warp-ab-runbook.md); when a knob change turns out to need a code change, validate it there and come back here for the knob-level sweep.
For each tuning knob it answers three questions:
## 1. Questions this runbook answers
1. Which stage is actually slow — `set_disk_encode`, `set_disk_rename`,
`metadata_fanout`, `bitrot_verify`, etc.?
1. Which stage is actually slow — `set_disk_encode`, `set_disk_rename`, `metadata_fanout`, `bitrot_verify`, ...?
2. Is the knob the real bottleneck, or is the stage slow for another reason?
3. Does widening/loosening the knob buy throughput without an unacceptable
memory (RSS) or tail-latency regression?
3. Does widening the knob buy throughput without an unacceptable memory (RSS) or tail-latency regression?
The core discipline is **one variable per A/B cell**. Never change two knobs in
the same cell, or the result is unexplainable.
The core discipline is **one variable per A/B cell**. A cell that changes two knobs is thrown away.
## 2. Prerequisites
- Linux bench host (or an ansible-managed cluster); a laptop smoke run is too
noisy to decide anything.
- `warp` on `PATH` (or pass `--warp-bin` to the driver).
- The observability metrics runtime **enabled**. The stage histograms below are
not emitted when `RUSTFS_OBS_METRICS_EXPORT_ENABLED=false` or the runtime is
otherwise off — see
[`hotpath-warp-ab-runbook.md`](hotpath-warp-ab-runbook.md) for the no-log /
no-monitor baseline env.
- A warm, disposable data set. Recreate the bucket per run; do not bench against
production data.
Host, `warp`, and the no-log / no-monitor baseline environment are as in [`hotpath-warp-ab-runbook.md`](hotpath-warp-ab-runbook.md). Two additions: the observability metrics runtime must be **enabled** (the stage histograms are not emitted when `RUSTFS_OBS_METRICS_EXPORT_ENABLED=false`), and the data set must be disposable — recreate the bucket per run.
Load driver and gate are reused, not reimplemented:
- `scripts/run_object_batch_bench_enhanced.sh` — warp driver with rounds,
median aggregation, `baseline_compare.csv`, and Prometheus service-metric
capture.
- `scripts/hotpath_warp_ab_gate.sh` relative budget gate over the deltas.
- `scripts/run_hotpath_warp_ab.sh` — optional orchestrator when a knob needs
the full baseline-vs-candidate treatment (e.g. two different defaults).
| Script | Role |
| --- | --- |
| `scripts/run_object_batch_bench_enhanced.sh` | warp driver with rounds, median aggregation, `baseline_compare.csv`, Prometheus service-metric capture |
| `scripts/hotpath_warp_ab_gate.sh` | relative budget gate over the deltas |
| `scripts/run_hotpath_warp_ab.sh` | orchestrator when a knob needs the full baseline-vs-candidate treatment (e.g. two different defaults) |
## 3. Fixed test conditions (lock before you start)
@@ -56,9 +35,6 @@ network, erasure_set_drive_count, endpoint_mode (direct|lb),
rustfs_commit_sha, warp --version, durability mode
```
Workload matrix (the same shapes the hotpath gate uses, expanded for the
stage-breakdown object sizes):
| Workload | mode | sizes |
| --- | --- | --- |
| small-fixed | put / get | 4KiB, 100KiB |
@@ -66,39 +42,19 @@ stage-breakdown object sizes):
| large-stream | put / get | 10MiB, 16MiB, 32MiB |
| mixed | mixed | 256KiB |
Concurrency ladder: `8, 16, 32, 64` (add `96, 128` on a bigger rig). Duration
`120s`, `--rounds >= 3`, cooldown `>= 30s`.
Isolate background noise before the sweep: scanner deep-verify, heal,
replication, lifecycle transition, periodic capacity refresh — record whether
each is on rather than silently assuming it is off.
Concurrency ladder `8, 16, 32, 64` (add `96, 128` on a bigger rig); duration `120s`, `--rounds >= 3`, cooldown `>= 30s`. Record whether scanner deep-verify, heal, replication, lifecycle transition, and periodic capacity refresh are on rather than assuming they are off.
## 4. Measurement stack
The code already instruments every stage below. Drive each A/B cell with these
histograms (names verified against `crates/io-metrics/src/lib.rs`):
Drive each cell with the stage histograms emitted by `crates/io-metrics/src/lib.rs`; the stage and path label tokens are defined there and are not repeated here.
- PUT stages: `rustfs_s3_put_object_stage_duration_ms{stage=...}` — compute
P50/P95/P99 per stage. Stages: `app_bucket_validate`, `app_sse_config_lookup`,
`app_object_lock_config_lookup`, `app_put_opts_build`, `app_prelookup`,
`ingress_prepare`, `app_encryption_prepare`, `app_replication_decision`,
`app_store_put`, `app_post_store_bookkeeping`, `app_capacity_update`,
`set_disk_writer_setup`, `set_disk_encode`, `set_disk_rename`,
`set_disk_old_data_cleanup`.
- GET stages: `rustfs_io_get_object_stage_duration_seconds{path=..., stage=...}`
the `path` label separates the read paths: `legacy_duplex`, `codec_streaming`,
`direct_memory`, `body_cache`, `inline_direct`, `internal_meta`,
`remote_transition`, `set_disk`, `empty`. Stages: `metadata`,
`metadata_cache_lookup`, `metadata_fanout`, `metadata_resolve`, `object_info`,
`path_decision`, `quorum_reached`, `range`, `reader_setup`,
`stripe_read`, `stripe_read_first_shard`, `stripe_read_quorum`, `decode`,
`reconstruct`, `emit`, `fill`, `output_poll`, `output_lock_wait`,
`bitrot_verify`, `first_byte`, `full_body`, `response_handoff`,
`lock_acquire`.
- EC memory pressure: `rustfs_ec_encode_inflight_bytes_current` and the
allocator reclaim gauge; plus node RSS and CPU.
| Metric | Labels | Use |
| --- | --- | --- |
| `rustfs_s3_put_object_stage_duration_ms` | `stage` | P50/P95/P99 per PUT stage (`app_*`, `ingress_prepare`, `set_disk_*`) |
| `rustfs_io_get_object_stage_duration_seconds` | `path`, `stage` | Per GET stage, split by read path (`legacy_duplex`, `codec_streaming`, ...) |
| `rustfs_ec_encode_inflight_bytes_current` | — | EC encode memory pressure; pair with node RSS and CPU |
Host telemetry (collect alongside every cell):
Host telemetry, collected alongside every cell:
```bash
pidstat -durh 5 > telemetry/pidstat.txt &
@@ -108,37 +64,34 @@ iostat -xz 5 > telemetry/iostat.txt &
## 5. Tuning knob catalog
Defaults are verified against `crates/config/src/constants/object.rs` and
`crates/ecstore/src/erasure/coding/encode.rs`.
### 5.1 PUT
| Knob | Default | Controls | Validating stage | Risk if widened |
| --- | --- | --- | --- | --- |
| `RUSTFS_ERASURE_ENCODE_MAX_INFLIGHT_BYTES` | 32MiB | EC encode producer/consumer memory budget (blocks queued between encode and shard write) | `set_disk_encode` P95 + `rustfs_ec_encode_inflight_bytes_current` | RSS growth under high concurrency |
| `RUSTFS_OBJECT_IO_BUFFER_SIZE` | 128KiB | Streaming read-in / write-out block size | `ingress_prepare`, `set_disk_encode` | Larger buffers = fewer polls, more resident memory |
| `RUSTFS_OBJECT_DUPLEX_BUFFER_SIZE` | 4MiB | duplex pipe capacity (shared, but PUT path uses it less than GET) | `set_disk_encode` feed smoothness | Memory per in-flight request |
| `RUSTFS_DURABILITY_MODE` / `RUSTFS_DRIVE_SYNC_ENABLE` | mode-dependent | per-shard fsync/sync discipline on commit | `set_disk_rename` P99 | Weakening it changes the durability contract — treat as a deliberate tradeoff, not a free win |
| `RUSTFS_RUNTIME_WORKER_THREADS` / `RUSTFS_RUNTIME_MAX_BLOCKING_THREADS` | Tokio defaults | async workers + `spawn_blocking` pool feeding per-block encode | `set_disk_encode` P95 + mpstat | Oversubscription |
| `RUSTFS_ERASURE_ENCODE_MAX_INFLIGHT_BYTES` | 32MiB (`crates/ecstore/src/erasure/coding/encode.rs`) | EC encode producer/consumer memory budget (blocks queued between encode and shard write) | `set_disk_encode` P95 + `rustfs_ec_encode_inflight_bytes_current` | RSS growth under high concurrency |
| `RUSTFS_OBJECT_IO_BUFFER_SIZE` | 128KiB (`crates/config/src/constants/object.rs`) | Streaming read-in / write-out block size | `ingress_prepare`, `set_disk_encode` | Larger buffers = fewer polls, more resident memory |
| `RUSTFS_OBJECT_DUPLEX_BUFFER_SIZE` | 4MiB (`crates/config/src/constants/object.rs`) | Duplex pipe capacity (shared; PUT uses it less than GET) | `set_disk_encode` feed smoothness | Memory per in-flight request |
| `RUSTFS_DURABILITY_MODE` / `RUSTFS_DRIVE_SYNC_ENABLE` | mode-dependent | Per-shard fsync/sync discipline on commit | `set_disk_rename` P99 | Weakening it changes the durability contract — a deliberate tradeoff, never a free win |
| `RUSTFS_RUNTIME_WORKER_THREADS` / `RUSTFS_RUNTIME_MAX_BLOCKING_THREADS` | Tokio defaults | Async workers + `spawn_blocking` pool feeding per-block encode | `set_disk_encode` P95 + mpstat | Oversubscription |
### 5.2 GET
| Knob | Default | Controls | Validating stage | Risk if enabled |
| --- | --- | --- | --- | --- |
| `RUSTFS_GET_CODEC_STREAMING_ROLLOUT` | `off` | switches the read path from `legacy_duplex` to the pull-based `ErasureDecodeReader` (`codec_streaming`) | compare `path="legacy_duplex"` vs `path="codec_streaming"` for `decode`/`emit`/`output_lock_wait`/`stripe_read` | behavioral change to the read path; rollout is `off` by default for a reason |
| `RUSTFS_GET_CODEC_STREAMING_ENGINE` | `legacy` | `legacy` vs `rustfs` decode engine under the streaming reader | `reconstruct`/`decode` per `path` | engine swap on a correctness-critical path |
| `RUSTFS_GET_CODEC_STREAMING_MULTIPART_ENABLE` | `false` | multipart objects on the streaming reader | same, multipart cells | wider format coverage |
| `RUSTFS_GET_CODEC_STREAMING_DATA_BLOCKS_FIRST_ENABLE` (+ `_MAX_SIZE`, `_FIRST_READER_SETUP`) | `false` / 512KiB | prefer data-shard readers before parity | `stripe_read_first_shard`/`stripe_read_quorum` | shard-selection order change |
| `RUSTFS_OBJECT_GET_SKIP_BITROT_VERIFY` | `false` | skip per-shard HighwayHash verify | `bitrot_verify` | **do not default on** — measures the theoretical ceiling only |
| `RUSTFS_OBJECT_DUPLEX_BUFFER_SIZE` | 4MiB | legacy GET in-process pipe capacity | `output_lock_wait`/`output_poll` | memory per in-flight GET |
| `RUSTFS_GET_SEEK_BUFFER_ENABLE` | `false` | in-memory seek buffer for small GET | `first_byte` | experimental, startup-latched — see [`get-path-experimental-switches.md`](get-path-experimental-switches.md) |
| `RUSTFS_GET_OUTPUT_HANDOFF_ATTRIBUTION_ENABLE` | `false` | adds `response_handoff` attribution (metrics only) | `response_handoff` | small per-request bookkeeping cost |
| `RUSTFS_GET_CODEC_STREAMING_ROLLOUT` | `off` (`crates/ecstore/src/set_disk/mod.rs`) | Switches the read path from `legacy_duplex` to the pull-based `ErasureDecodeReader` (`codec_streaming`) | `path="legacy_duplex"` vs `path="codec_streaming"` for `decode`/`emit`/`output_lock_wait`/`stripe_read` | Behavioral change to the read path |
| `RUSTFS_GET_CODEC_STREAMING_ENGINE` | `legacy` | `legacy` vs `rustfs` decode engine under the streaming reader | `reconstruct`/`decode` per `path` | Engine swap on a correctness-critical path |
| `RUSTFS_GET_CODEC_STREAMING_MULTIPART_ENABLE` | `false` | Multipart objects on the streaming reader | Same, multipart cells | Wider format coverage |
| `RUSTFS_GET_CODEC_STREAMING_DATA_BLOCKS_FIRST_ENABLE` (+ `_MAX_SIZE`, `_FIRST_READER_SETUP`) | `false` / 512KiB | Prefer data-shard readers before parity | `stripe_read_first_shard`/`stripe_read_quorum` | Shard-selection order change |
| `RUSTFS_OBJECT_GET_SKIP_BITROT_VERIFY` | `false` | Skip per-shard HighwayHash verify | `bitrot_verify` | **Never default on** — measures the theoretical ceiling only |
| `RUSTFS_OBJECT_DUPLEX_BUFFER_SIZE` | 4MiB | Legacy GET in-process pipe capacity | `output_lock_wait`/`output_poll` | Memory per in-flight GET |
| `RUSTFS_GET_SEEK_BUFFER_ENABLE` [^latched] | `false` | Serves small GETs through an in-memory seek buffer (seek support without re-reading the object); gates `should_buffer_get_object_in_memory_with_threshold` | `first_byte` | Experimental; read by `scripts/run_get_1mib_abba_stage_metrics.sh` |
| `RUSTFS_GET_OUTPUT_HANDOFF_ATTRIBUTION_ENABLE` [^latched] | `false` | Attributes output-handoff timing to the `response_handoff` stage (metrics only) | `response_handoff` | Small per-request bookkeeping cost; set by `scripts/run_get_codec_streaming_smoke.sh` and `scripts/test_get_1mib_abba_stage_metrics.sh` |
[^latched]: Both switches are startup-latched `OnceLock` booleans in `rustfs/src/app/object/get.rs` (`ENV_RUSTFS_GET_SEEK_BUFFER_ENABLE` / `is_get_seek_buffer_enabled`, `ENV_RUSTFS_GET_OUTPUT_HANDOFF_ATTRIBUTION_ENABLE` / `is_get_output_handoff_attribution_enabled`), read once via `get_env_bool(.., false)`; changing a value requires a process restart. They exist to support A/B runs, are referenced only by the harness scripts named above, and are kept deliberately — do not remove either switch or the seek-buffer path as dead code. Leave both unset in production. The codec-streaming switches above are startup-latched as well.
## 6. A/B matrix
Run each row as an independent cell. Baseline is the shipped default; candidate
is one knob moved. Everything else (topology, sizes, concurrency, rounds,
durability) stays fixed.
Run each row as an independent cell. Baseline is the shipped default; candidate is one knob moved. Everything else (topology, sizes, concurrency, rounds, durability) stays fixed.
### 6.1 PUT
@@ -162,17 +115,12 @@ durability) stays fixed.
| G5 | duplex buffer | 4MiB | 8MiB, 16MiB | `output_lock_wait`/`output_poll` (legacy path) | only if still on `legacy_duplex` |
| G6 | skip bitrot verify | `false` | `true` | `bitrot_verify` | **ceiling measurement only**; do not carry into production |
## 7. Execution sequence
## 7. Execution and interpretation
1. Freeze the conditions in §3 and record the provenance block.
2. Run the baseline cell (all defaults) and capture stage histograms + host
telemetry.
3. Pick the **one** most-likely knob from the analysis. For large-object PUT
that is almost always `set_disk_encode` → P1; for GET it is G1 (the
`legacy_duplex``codec_streaming` switch).
4. Sweep that knob's candidate column one value at a time, same workload.
5. Read the decision table in §8; if the stage did not move, the knob is not
the bottleneck — stop widening it and pick the next stage.
2. Run the baseline cell (all defaults) and capture stage histograms plus host telemetry.
3. Pick the **one** most-likely knob from the decision table below (large-object PUT is almost always `set_disk_encode` → P1; GET is G1), sweep its candidate column one value at a time on the same workload, and stop widening as soon as the stage stops moving — the knob is then not the bottleneck.
4. Keep `baseline_compare.csv`, `median_summary.csv`, stage histograms, and host telemetry per cell; the conclusion must trace back to them. Record results in the issue tracker, not the repo.
Driver invocation for one PUT cell:
@@ -185,8 +133,6 @@ scripts/run_object_batch_bench_enhanced.sh \
--out-dir target/bench/put-tuning-p1-64mib
```
## 8. Interpretation / decision table
| Stage high | Most likely cause | Next action |
| --- | --- | --- |
| `set_disk_encode` | per-block EC encode scheduling + in-flight budget | P1 → P4 → P2, in that order |
@@ -198,21 +144,4 @@ scripts/run_object_batch_bench_enhanced.sh \
| `output_lock_wait` / `output_poll` | legacy duplex backpressure | G1 (move off duplex) or G5 |
| `stripe_read*` | shard concurrency / selection | G4 shard-selection, disk/network tail |
## 9. Guardrails
- **Never weaken correctness for throughput**: read/write quorum, bitrot verify,
`xl.meta` validation, and durability (`RUSTFS_DURABILITY_MODE`) are integrity
contracts, not knobs. P5 and G6 are ceiling measurements and must be labelled
as such; do not carry their values into production without an explicit
durability/correctness decision.
- **One variable per cell.** A cell that changes two knobs is thrown away.
- **Memory is part of the result.** A throughput win with unbounded RSS growth
is a regression; record RSS and the EC in-flight gauge for every PUT cell.
- **Startup-latched knobs** (`RUSTFS_GET_SEEK_BUFFER_ENABLE`,
`RUSTFS_GET_OUTPUT_HANDOFF_ATTRIBUTION_ENABLE`, and the codec-streaming
switches) require a process restart to change — see
[`get-path-experimental-switches.md`](get-path-experimental-switches.md).
- **Archive the raw data.** Keep the `baseline_compare.csv`, `median_summary.csv`,
stage histograms, and host telemetry per cell; the conclusion must trace back
to them. Do not commit benchmark result snapshots to the repo — record them in
the issue tracker.
Read/write quorum, bitrot verify, `xl.meta` validation, and durability are integrity contracts, not knobs: P5 and G6 are ceiling measurements and must be labelled as such. A throughput win with unbounded RSS growth is a regression — RSS and the EC in-flight gauge are part of every PUT result.
+217
View File
@@ -0,0 +1,217 @@
# OIDC Console integration
**Use this when:** connecting RustFS Console login to an OpenID Connect provider (Keycloak, Authing, or any standards-compliant IdP), or debugging an OIDC redirect, token, or policy-mapping failure.
**Source of truth:** `crates/config/src/constants/oidc.rs` (provider keys and `RUSTFS_IDENTITY_OPENID_*`), `crates/iam/src/oidc.rs` (discovery, PKCE, token validation, per-provider env suffixes), `rustfs/src/admin/handlers/oidc.rs` (authorize/callback handlers), `crates/config/src/constants/app.rs` (`ENV_RUSTFS_BROWSER_REDIRECT_URL`), `crates/utils/src/egress.rs` (`ENV_OUTBOUND_ALLOW_ORIGINS`), `crates/policy/src/policy/policy.rs` (built-in policies).
The RustFS side is vendor-neutral and is described once; what RustFS requires from any provider is tabulated in [oidc-provider-requirements.md](oidc-provider-requirements.md). The [Keycloak](#keycloak) and [Authing](#authing) sections contain only IdP-side steps and vendor caveats. Examples use provider id `default` and public origin `https://rustfs.example.com`.
## Integration model
RustFS requires a standards-compliant OpenID Connect provider: discovery at `<issuer>/.well-known/openid-configuration`, authorization and token endpoints, a JWKS, and an authorization-code flow that returns an `id_token`. RustFS never calls a vendor's authorization API; access is decided by RustFS IAM policies after claim mapping. Protocol requirements for IdP vendors are collected in [oidc-provider-requirements.md](oidc-provider-requirements.md).
Login flow:
1. The browser opens `https://rustfs.example.com/rustfs/admin/v3/oidc/authorize/<provider_id>`.
2. RustFS creates `state`, `nonce`, and a PKCE S256 challenge and redirects to the IdP.
3. The IdP redirects back to `/rustfs/admin/v3/oidc/callback/<provider_id>?code=...&state=...`.
4. RustFS exchanges the code at the token endpoint, sending `client_id` and `client_secret` in the request body (`client_secret_post`) together with the PKCE verifier.
5. RustFS validates the ID token signature (JWKS), issuer, audience, expiry, and nonce.
6. RustFS maps claim values to policy names and issues one-hour STS credentials to the Console.
In-flight `state` and PKCE verifiers are node-local: the authorize and callback requests must reach the same RustFS node.
## Configuration keys
Every provider key can be set as `RUSTFS_IDENTITY_OPENID_<KEY>` in the process environment or as `identity_openid` `<key>=<value>` through `mc admin config set`. Names are constants in `crates/config/src/constants/oidc.rs`.
| Provider key | Environment variable | Purpose |
| --- | --- | --- |
| `enable` | `RUSTFS_IDENTITY_OPENID_ENABLE` | `on` loads the provider. |
| `config_url` | `RUSTFS_IDENTITY_OPENID_CONFIG_URL` | Issuer URL used for discovery. A trailing `/.well-known/openid-configuration` is stripped; any other `.well-known` path is rejected. |
| `issuer` | `RUSTFS_IDENTITY_OPENID_ISSUER` | Expected `iss` when it differs from `config_url` (internal discovery URL, public token issuer). |
| `client_id`, `client_secret` | `RUSTFS_IDENTITY_OPENID_CLIENT_ID`, `RUSTFS_IDENTITY_OPENID_CLIENT_SECRET` | Confidential client credentials. |
| `scopes` | `RUSTFS_IDENTITY_OPENID_SCOPES` | Comma-separated; `openid` is required. |
| `other_audiences` | `RUSTFS_IDENTITY_OPENID_OTHER_AUDIENCES` | Additional accepted `aud` values. |
| `redirect_uri` | `RUSTFS_IDENTITY_OPENID_REDIRECT_URI` | Callback URL sent to the IdP; must equal the URL registered there. |
| `redirect_uri_dynamic` | `RUSTFS_IDENTITY_OPENID_REDIRECT_URI_DYNAMIC` | `on` derives the callback from request headers. Keep `off` behind proxies. |
| `claim_name`, `claim_prefix` | `RUSTFS_IDENTITY_OPENID_CLAIM_NAME`, `RUSTFS_IDENTITY_OPENID_CLAIM_PREFIX` | Policy claim name and a fixed string prepended to each value. `claim_prefix` is not a mapping table. |
| `groups_claim`, `roles_claim` | `RUSTFS_IDENTITY_OPENID_GROUPS_CLAIM`, `RUSTFS_IDENTITY_OPENID_ROLES_CLAIM` | Flat top-level array claims whose values are RustFS policy names. |
| `email_claim`, `username_claim` | `RUSTFS_IDENTITY_OPENID_EMAIL_CLAIM`, `RUSTFS_IDENTITY_OPENID_USERNAME_CLAIM` | Identity claims shown in the Console. |
| `role_policy` | `RUSTFS_IDENTITY_OPENID_ROLE_POLICY` | One fixed policy for every login from this provider. Connectivity testing only. |
| `display_name` | `RUSTFS_IDENTITY_OPENID_DISPLAY_NAME` | Login button label. |
| `hide_from_ui` | `RUSTFS_IDENTITY_OPENID_HIDE_FROM_UI` | Hides the provider from `/oidc/providers`. |
Process-level settings (environment only, never suffixed per provider):
| Variable | Purpose |
| --- | --- |
| `RUSTFS_BROWSER_REDIRECT_URL` | Public browser origin used for callback generation, Console success redirects, and logout fallback. |
| `RUSTFS_OUTBOUND_ALLOW_ORIGINS` | Exact `scheme://host[:port]` origins RustFS may contact for discovery, JWKS, and token requests when the IdP resolves to a private, loopback, or container-network address. See [outbound-connection-policy.md](outbound-connection-policy.md). |
Named providers: to use provider id `<id>`, suffix every provider env var with `_<id>` (for example `RUSTFS_IDENTITY_OPENID_CLIENT_ID_keycloak`) and register the callback `/rustfs/admin/v3/oidc/callback/<id>`. Suffix scanning is `parse_single_provider` in `crates/iam/src/oidc.rs`.
Restart RustFS after changing any of these settings.
### Environment example
```bash
export RUSTFS_BROWSER_REDIRECT_URL="https://rustfs.example.com"
export RUSTFS_IDENTITY_OPENID_ENABLE=on
export RUSTFS_IDENTITY_OPENID_CONFIG_URL="<ISSUER>"
export RUSTFS_IDENTITY_OPENID_CLIENT_ID="<CLIENT_ID>"
export RUSTFS_IDENTITY_OPENID_CLIENT_SECRET="<CLIENT_SECRET>"
export RUSTFS_IDENTITY_OPENID_SCOPES="openid,profile,email"
export RUSTFS_IDENTITY_OPENID_REDIRECT_URI="https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default"
export RUSTFS_IDENTITY_OPENID_REDIRECT_URI_DYNAMIC=off
export RUSTFS_IDENTITY_OPENID_DISPLAY_NAME="<IdP name>"
export RUSTFS_IDENTITY_OPENID_GROUPS_CLAIM="groups"
export RUSTFS_IDENTITY_OPENID_ROLES_CLAIM="roles"
export RUSTFS_IDENTITY_OPENID_EMAIL_CLAIM="email"
export RUSTFS_IDENTITY_OPENID_USERNAME_CLAIM="preferred_username"
```
The same keys through admin config:
```bash
mc admin config set rustfs identity_openid \
enable=on config_url="<ISSUER>" client_id="<CLIENT_ID>" client_secret="<CLIENT_SECRET>" \
scopes="openid,profile,email" \
redirect_uri="https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default" \
redirect_uri_dynamic=off display_name="<IdP name>" \
groups_claim="groups" roles_claim="roles" email_claim="email" username_claim="preferred_username"
mc admin service restart rustfs
```
`RUSTFS_BROWSER_REDIRECT_URL` is not an `identity_openid` key; it must still be set in the process environment.
## Redirect URL priority
1. Provider `redirect_uri`, when set, is the callback URL sent to the IdP.
2. `RUSTFS_BROWSER_REDIRECT_URL`, when set, is the public origin for callback generation when no provider `redirect_uri` exists, and for Console success and logout fallback redirects.
3. Request headers (`Host`, `X-Forwarded-Proto`) are used only when `redirect_uri_dynamic=on` and no browser redirect URL is configured.
Behind a reverse proxy or load balancer, set `RUSTFS_BROWSER_REDIRECT_URL` and keep session affinity for the authorize and callback requests.
## Policy mapping
Claim values are used verbatim as policy names (after `claim_prefix`, if any). Names must satisfy `is_safe_claim_policy_name` in `crates/iam/src/sys.rs`: ASCII letters, digits, `_`, `-`, `:`, `.` only, so a value containing `/` (for example Keycloak's full group path `/consoleAdmin`) never matches. Built-in policies:
| Policy | Grants |
| --- | --- |
| `consoleAdmin` | Full Console, admin, KMS, and S3 access. |
| `readwrite` | S3 read/write. |
| `readonly` | S3 read-only. |
| `writeonly` | S3 write-only. |
| `diagnostics` | Diagnostic admin access. |
For first-contact testing only, `RUSTFS_IDENTITY_OPENID_ROLE_POLICY=consoleAdmin` grants every login full access; remove it before production.
## Validation
1. Discovery:
```bash
curl -fsS "<ISSUER>/.well-known/openid-configuration" | jq '{issuer, authorization_endpoint, token_endpoint, jwks_uri, code_challenge_methods_supported, token_endpoint_auth_methods_supported, scopes_supported}'
```
`issuer` must equal `RUSTFS_IDENTITY_OPENID_ISSUER` when set, otherwise the issuer derived from `RUSTFS_IDENTITY_OPENID_CONFIG_URL`; `code_challenge_methods_supported` must include `S256`; `token_endpoint_auth_methods_supported` must include `client_secret_post`; `scopes_supported` must include every configured scope.
2. Provider visibility: `curl -fsS https://rustfs.example.com/rustfs/admin/v3/oidc/providers | jq` lists the provider unless `hide_from_ui=on`.
3. Browser login: open `https://rustfs.example.com/rustfs/admin/v3/oidc/authorize/default`. Expect a redirect to the IdP, sign-in, a redirect to `/rustfs/admin/v3/oidc/callback/default?code=...&state=...`, and then the Console with the mapped permissions.
4. ID token claims (decode the token after a test login): `iss` matches the issuer, `aud` includes the client id, `email` and `preferred_username` are present when configured, `groups` or `roles` is a flat array of policy names.
## Troubleshooting
| Symptom | Common cause | Fix |
| --- | --- | --- |
| `/oidc/providers` does not list the provider | provider failed to load, or RustFS was not restarted | Check env/admin config and restart RustFS. |
| Provider or login button missing; startup logs `OIDC provider discovery blocked by outbound policy` | IdP origin is private/internal and not allowlisted | Add the exact origin to `RUSTFS_OUTBOUND_ALLOW_ORIGINS` on every node and restart. |
| IdP reports a redirect mismatch (`invalid redirect_uri`) | registered callback differs from RustFS `redirect_uri` | Use the exact `/rustfs/admin/v3/oidc/callback/<provider_id>` URL on both sides. |
| Callback reports missing `code` or `state` | proxy dropped the query string | Preserve the full callback URL and query string. |
| Token exchange fails | wrong secret, or the IdP rejects request-body client authentication | Confirm the client is confidential and accepts `client_secret_post`. |
| No `id_token` in the token response | `openid` scope missing or a non-OIDC OAuth flow | Add `openid`; use the authorization-code flow. |
| ID token verification fails | issuer, audience, algorithm, or JWKS mismatch | Compare discovery metadata with `CONFIG_URL`/`ISSUER`/`CLIENT_ID`; prefer `RS256`. |
| Login succeeds, access denied | no claim value matches a policy name | Emit `groups` or `roles` as a flat array equal to policy names; check for `/` prefixes. |
| Console redirects to an internal host | `RUSTFS_BROWSER_REDIRECT_URL` unset or proxy headers wrong | Set `RUSTFS_BROWSER_REDIRECT_URL` to the public origin. |
| Invalid or expired OIDC state | callback reached a different node | Configure load-balancer session affinity for authorize and callback. |
## Production checklist
- [ ] RustFS and the IdP use HTTPS.
- [ ] The IdP registers the exact callback URL (no wildcard) and `RUSTFS_IDENTITY_OPENID_REDIRECT_URI` matches it.
- [ ] `RUSTFS_BROWSER_REDIRECT_URL` is the public browser origin.
- [ ] PKCE S256 is allowed or required at the IdP.
- [ ] ID tokens carry `groups` or `roles` values equal to RustFS policy names.
- [ ] `role_policy` is not used as a permanent shortcut.
- [ ] The load balancer preserves query strings and pins authorize/callback to one node.
- [ ] Internal IdP origins are listed exactly in `RUSTFS_OUTBOUND_ALLOW_ORIGINS` on every node.
## Keycloak
| Value | Example |
| --- | --- |
| Realm | `rustfs` |
| Issuer (`config_url`) | `https://keycloak.example.com/realms/rustfs` |
| Discovery URL | `https://keycloak.example.com/realms/rustfs/.well-known/openid-configuration` |
| Client id | `rustfs-console` |
| Scopes | `openid,profile,email` |
| Groups claim | `groups` (flat array) |
Client setup in the Keycloak Admin Console:
1. Create or select the realm and confirm discovery returns `issuer` equal to `https://keycloak.example.com/realms/rustfs`.
2. `Clients` → create: `Client type` = `OpenID Connect`, `Client ID` = `rustfs-console`.
3. Enable `Client authentication` and `Standard flow`; disable `Implicit flow`, `Direct access grants`, and `Service accounts roles`.
4. `Valid redirect URIs` = `https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default`; `Web origins` = `https://rustfs.example.com`.
5. `Proof Key for Code Exchange Code Challenge Method` = `S256`.
6. Save and copy the secret from `Credentials`. Do not apply a client policy that disables `client_secret_post`.
Group mapper (a `Group Membership` mapper in the client's dedicated scope):
| Mapper field | Value |
| --- | --- |
| Name | `rustfs-groups` |
| Token Claim Name | `groups` |
| Full group path | `Off` (a leading `/` breaks policy matching) |
| Add to ID token / access token / userinfo | `On` |
| Multivalued | `On` |
Create Keycloak groups named after RustFS policies (`consoleAdmin`, `readonly`, ...) and add users to them.
Roles instead of groups: assign realm or client roles named after policies, add a `User Realm Role` or `User Client Role` mapper that emits a flat top-level `roles` claim, and set `RUSTFS_IDENTITY_OPENID_ROLES_CLAIM=roles`. RustFS does not read Keycloak's nested `realm_access.roles` claim.
Internal discovery URL with a public issuer (for example in-cluster Keycloak on Kubernetes):
```bash
export RUSTFS_IDENTITY_OPENID_CONFIG_URL="http://keycloak.keycloak.svc.cluster.local:8080/realms/rustfs"
export RUSTFS_IDENTITY_OPENID_ISSUER="https://keycloak.example.com/realms/rustfs"
export RUSTFS_OUTBOUND_ALLOW_ORIGINS="http://keycloak.keycloak.svc.cluster.local:8080"
```
Discovery and issuer-relative JWKS requests use the `CONFIG_URL` base; `iss` validation uses `ISSUER`. The allowlist entry is the origin only (no realm or discovery path) and is read at startup on every node. Prefer HTTPS with a trusted CA for the internal URL: discovery and JWKS define the token-signing trust root, so plain HTTP is acceptable only where DNS and traffic cannot be tampered with.
## Authing
| Value | Example | Note |
| --- | --- | --- |
| Application domain | `https://example.authing.cn` | From the Authing application page. |
| Issuer (`config_url`) | `https://example.authing.cn/oidc` | Tenants differ (`/oidc`, `/oauth/oidc`): copy the issuer from the console and confirm discovery returns the same `issuer`. |
| App ID / App Secret | `<AUTHING_APP_ID>` / `<AUTHING_APP_SECRET>` | RustFS `client_id` / `client_secret`. |
| Scopes | `openid,profile,email,roles` | `roles` is needed when Authing emits role claims. |
| Roles claim | `roles` | Set `RUSTFS_IDENTITY_OPENID_ROLES_CLAIM=roles`. |
Application settings in the Authing console:
| Setting | Value |
| --- | --- |
| Protocol | OpenID Connect |
| Grant type / response type | Authorization Code / `code` |
| Token endpoint authentication | `client_secret_post` |
| PKCE | allow or require `S256` |
| ID token signing algorithm | `RS256` |
| Redirect URL | `https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default` (exact) |
Assign Authing roles whose names equal RustFS policy names; a test user with role `consoleAdmin` should produce `"roles": ["consoleAdmin"]` in the ID token. `claim_prefix` only prepends a fixed string, so keep role values equal to policy names unless policies with that prefix already exist.
@@ -0,0 +1,26 @@
# OIDC provider requirements
**Use this when:** evaluating whether an identity provider (or a vendor's "OIDC-like" OAuth product) can back RustFS console SSO, or debugging why a provider fails discovery, login, or policy mapping.
**Source of truth:** `crates/iam/src/oidc.rs` (`OidcProviderConfig`, `discover_provider`, `trusted_aud`, `map_claims_to_policies`); `crates/config/src/constants/oidc.rs` (`OIDC_DEFAULT_*`, `ENV_IDENTITY_OPENID_*`); `rustfs/src/admin/handlers/oidc.rs` (`derive_callback_uri_with_provider_config`, `browser_redirect_url`).
RustFS is a standard OpenID Connect relying party using the authorization-code flow with PKCE. It reads every claim it uses from the ID token; it never calls the UserInfo endpoint. Provider-side setup steps for Keycloak and Authing, and the RustFS-side configuration keys, are in [oidc-console-integration.md](oidc-console-integration.md).
## Requirements
| # | Requirement | Details | Code anchor |
| --- | --- | --- | --- |
| 1 | Discovery document | `RUSTFS_IDENTITY_OPENID_CONFIG_URL` names the provider (issuer base or full discovery URL); RustFS fetches `{issuer}/.well-known/openid-configuration` and needs `issuer`, `authorization_endpoint`, `token_endpoint`, `jwks_uri`, and the standard `*_supported` arrays. When `RUSTFS_IDENTITY_OPENID_ISSUER` is set, the document's `issuer` must equal it exactly; otherwise RustFS tries the issuer candidates derived from the config URL. | `crates/iam/src/oidc.rs` `discover_provider`, `discover_provider_from_config_url` |
| 2 | Outbound reachability | Discovery, JWKS, and token requests go through the shared egress policy. A provider on a private or loopback address needs its exact origin in `RUSTFS_OUTBOUND_ALLOW_ORIGINS`; otherwise startup logs `OIDC provider discovery blocked by outbound policy`. | `crates/iam/src/oidc.rs` `OIDC_DISCOVERY_BLOCKED_BY_OUTBOUND_POLICY`; [Outbound connection policy](outbound-connection-policy.md) |
| 3 | Signed ID token, verifiable via JWKS | The ID token signature is verified against `jwks_uri`; the key set is refreshed after `OIDC_JWKS_REFRESH_INTERVAL` and once more on a verification failure. A token response without `id_token` fails login — an OAuth-only access token is not sufficient. | `crates/iam/src/oidc.rs` `OIDC_JWKS_REFRESH_INTERVAL`, the `no id_token in token response` error |
| 4 | Authorization-code flow with PKCE (S256) | The authorization request carries `response_type=code`, `scope`, `redirect_uri`, `state`, `nonce`, `code_challenge`, `code_challenge_method=S256`; the token request carries the `code_verifier`. Providers that ignore or reject PKCE, `state`, or `nonce` are not supported. | `crates/iam/src/oidc.rs` `PkceCodeChallenge::new_random_sha256`; `crates/iam/src/oidc_state.rs` `pkce_verifier` |
| 5 | Scopes | Default `openid,profile,email` (`OIDC_DEFAULT_SCOPES`); override with `RUSTFS_IDENTITY_OPENID_SCOPES` (comma-separated). The provider must accept every configured scope. | `crates/config/src/constants/oidc.rs` `OIDC_DEFAULT_SCOPES`; `crates/iam/src/oidc.rs` `OidcProviderConfig::scopes` |
| 6 | Callback returns `state` | The callback must carry both `code` and the original `state`; `state` locates the in-flight session that holds the nonce and PKCE verifier. A callback with only `code` fails. | `rustfs/src/admin/handlers/oidc.rs` `OIDC_CALLBACK_SUFFIX`; `crates/iam/src/oidc_state.rs` |
| 7 | ID token claims | `iss`, `aud`, `exp`, and `nonce` are verified; `sub` identifies the user. `aud` must contain the client id or one of `RUSTFS_IDENTITY_OPENID_OTHER_AUDIENCES`. | `crates/iam/src/oidc.rs` `trusted_aud`, `OidcClaims` |
| 8 | Authorization claims in the ID token | Policy mapping reads the claim named by `RUSTFS_IDENTITY_OPENID_CLAIM_NAME` (default `groups`, `OIDC_DEFAULT_CLAIM_NAME`) together with `RUSTFS_IDENTITY_OPENID_GROUPS_CLAIM` and the optional `RUSTFS_IDENTITY_OPENID_ROLES_CLAIM`, as a string or array of strings, optionally prefixed by `RUSTFS_IDENTITY_OPENID_CLAIM_PREFIX`. Values must equal RustFS policy names (for example `consoleAdmin`, `readwrite`, `readonly`). Email and username come from `RUSTFS_IDENTITY_OPENID_EMAIL_CLAIM` (default `email`) and `RUSTFS_IDENTITY_OPENID_USERNAME_CLAIM` (default `preferred_username`). A provider that returns only a user id can authenticate but cannot express RustFS authorization. | `crates/iam/src/oidc.rs` `map_claims_to_policies`, `extract_canonical_group_values` |
| 9 | Registered redirect URI | The provider must accept the callback `{public-origin}/rustfs/admin/v3/oidc/callback/{provider_id}`. RustFS picks the origin in this order: the provider's `redirect_uri` (`RUSTFS_IDENTITY_OPENID_REDIRECT_URI`), then `RUSTFS_BROWSER_REDIRECT_URL`, then the request's own scheme and host — the last only when `RUSTFS_IDENTITY_OPENID_REDIRECT_URI_DYNAMIC` is enabled. | `rustfs/src/admin/handlers/oidc.rs` `derive_callback_uri_with_provider_config`, `browser_redirect_url` |
| 10 | Logout endpoint (optional) | When discovery advertises `end_session_endpoint`, RustFS builds an RP-initiated logout URL with `id_token_hint`, `client_id`, and `post_logout_redirect_uri`. Without it, logout falls back to the console login page. | `crates/iam/src/oidc.rs` `build_logout_url` (reads `end_session_endpoint` from `ProviderMetadataWithLogout`) |
## Deployment notes
- Behind a load balancer, authorize and callback requests must reach the same RustFS node while the `state` is in flight, or set `RUSTFS_BROWSER_REDIRECT_URL` so the callback URL is stable; the callback error text names both remedies.
- Provider-specific setup (client registration, claim mappers, redirect-URL priority) is covered by the console integration guide for the provider in use; this page states only what any provider must offer.
@@ -1,134 +0,0 @@
# OIDC Vendor Compatibility Checklist
Use this checklist when a vendor provides an OAuth or SSO document that is described as OIDC but does not clearly expose the standard OpenID Connect contract required by RustFS.
## 1. Discovery Metadata
RustFS expects provider metadata at:
```text
GET {issuer}/.well-known/openid-configuration
```
Ask the vendor to provide the discovery URL and confirm that it returns at least:
- `issuer`
- `authorization_endpoint`
- `token_endpoint`
- `jwks_uri`
- `response_types_supported`
- `subject_types_supported`
- `id_token_signing_alg_values_supported`
The returned `issuer` must exactly match the issuer configured in RustFS.
## 2. JWKS and Token Signature Verification
RustFS must verify the ID token signature. Ask the vendor to provide:
- `jwks_uri`
- supported signing algorithms, such as `RS256`
- key rotation behavior
- how the token `kid` maps to the JWKS key set
Without a verifiable ID token signature, the provider is not suitable for RustFS OIDC login.
## 3. Authorization Request Parameters
The provider must accept the standard authorization-code request parameters:
- `scope=openid profile email`
- `response_type=code`
- `client_id`
- `redirect_uri`
- `state`
- `nonce`
- `code_challenge`
- `code_challenge_method=S256`
If the vendor example omits `state`, `nonce`, or PKCE, confirm whether those parameters are supported.
## 4. Callback State
The provider must return the original `state` value in the callback:
```text
...?code=xxx&state=yyy
```
RustFS uses `state` for CSRF protection and to find the in-flight OIDC session. A callback that only returns `code` is not enough.
## 5. Token Response
The token endpoint response must be JSON and include at least:
- `access_token`
- `token_type`, usually `Bearer`
- `expires_in`
- `id_token`
RustFS requires `id_token`; an OAuth-only access token is not sufficient for Console OIDC login.
## 6. ID Token Claims
The ID token must contain standard claims that RustFS can verify:
- `iss`
- `sub`
- `aud`
- `exp`
- `iat`
- `nonce` when the authorization request includes `nonce`
Ask the vendor for a sample ID token payload and claim documentation.
## 7. UserInfo Endpoint
Standard OIDC UserInfo normally uses:
```text
GET /userinfo
Authorization: Bearer <access_token>
```
If the vendor only documents a private profile endpoint such as `/oidc/profile?access_token=...`, ask whether a standard `userinfo_endpoint` is available and returned in discovery.
## 8. Logout Endpoint
Standard RP-initiated logout is normally exposed through an `end_session_endpoint` in discovery. If the vendor only documents a private token removal endpoint, ask whether standard OIDC logout is available.
RustFS can still fall back to the Console login page when the provider does not advertise an end-session endpoint.
## 9. Authorization Claims
OIDC primarily authenticates the user. RustFS authorization is still based on RustFS policies. The provider must emit claims that can be mapped to RustFS policies, for example:
- `groups`
- `roles`
- `policy`
- another agreed flat array or string claim
Ask the vendor to confirm:
- whether group, role, or policy claims can be included in the ID token
- whether those claims can be included in UserInfo
- the exact claim names and value formats
- whether the claim values can match RustFS policy names such as `consoleAdmin`, `readwrite`, or `readonly`
If the provider only returns a user id or token validity result, it can authenticate the user but cannot by itself express RustFS authorization.
## 10. RustFS Redirect Requirements
RustFS browser-facing redirect behavior depends on these values:
- provider `redirect_uri`, when explicitly configured, is the callback URL sent to the provider
- `RUSTFS_BROWSER_REDIRECT_URL` is the public RustFS browser origin used for callback generation when no provider `redirect_uri` exists, and for Console success and logout fallback redirects
- dynamic request-header redirects are used only when no configured redirect source exists and dynamic redirects are enabled
Ask the vendor to register the exact callback URL, for example:
```text
https://rustfs.example.com/rustfs/admin/v3/oidc/callback/default
```
For load-balanced RustFS deployments, ensure authorize and callback requests reach the same RustFS node while the OIDC `state` is in flight.
+38 -110
View File
@@ -1,120 +1,62 @@
# Outbound Connection Policy
This document describes the outbound connection policy that RustFS applies to
server-initiated HTTP(S) requests, and the `RUSTFS_OUTBOUND_ALLOW_ORIGINS`
allowlist operators can use to reach endpoints on private or container networks.
**Use this when:** a webhook, audit target, OIDC provider, or object-lambda endpoint on a private or container network (Compose service names, `host.docker.internal`, RFC 1918 addresses) is not being reached, or you need to know which server-initiated connections RustFS restricts and how to allowlist one.
**Source of truth:** `crates/utils/src/egress.rs` (`OutboundPolicy`, `OutboundDnsResolver`, `validate_outbound_url`, `ENV_OUTBOUND_ALLOW_ORIGINS`).
It is written for operators whose outbound integrations stopped reaching
endpoints after an upgrade — typically Docker Compose service names,
`host.docker.internal`, or RFC 1918 addresses. Webhook and audit clients adopted
this policy in `1.0.0-beta.11`; OIDC provider requests adopted it in
`1.0.0-beta.12`.
RustFS validates every operator-configured outbound destination to close a server-side request forgery (SSRF) class. Two layers exist:
## Background: what the policy protects
| Layer | What it checks | Escape hatch |
| --- | --- | --- |
| Literal URL check (`validate_outbound_url`) | Scheme is `http`/`https`; the host is not `localhost` or a loopback, private, shared, reserved, link-local, unspecified, or metadata address (IPv4-mapped and embedded IPv6 forms are classified by the embedded IPv4) | None |
| Full policy (`OutboundPolicy` + `OutboundDnsResolver`) | The literal check, plus re-validation of every address DNS returns on each new connection, so a hostname cannot be rebound to a restricted address after it was accepted | `RUSTFS_OUTBOUND_ALLOW_ORIGINS` for the loopback, private, shared, and reserved classes |
Several RustFS subsystems open connections to operator-configured URLs. To close
a server-side request forgery (SSRF) class of problem, RustFS validates every such
destination and re-checks the addresses returned by DNS on each new connection,
so a hostname cannot be rebound to a restricted address after it is first
accepted.
## Which subsystem uses which layer
The policy governs the outbound clients used by:
| Subsystem | Layer | Notes |
| --- | --- | --- |
| Event-notification webhooks (`RUSTFS_NOTIFY_WEBHOOK_*`) and audit webhooks (`RUSTFS_AUDIT_WEBHOOK_*`) | Full policy | Proxies disabled and redirects not followed, so the endpoint must be reachable directly (`crates/targets/src/target/webhook.rs`) |
| Target configuration validation (startup and admin API) | Full policy | `crates/targets/src/config/common.rs` `validate_outbound_http_url`; `rustfs/src/admin/handlers/target_descriptor.rs` |
| OIDC discovery, JWKS, and token requests | Full policy | A blocked provider logs `OIDC provider discovery blocked by outbound policy` naming the origin to allowlist (`crates/iam/src/oidc.rs`) |
| Object Lambda targets | Full policy | `rustfs/src/admin/router.rs` `outbound_policy` |
| Bucket replication targets | Literal check, relaxed | Private addresses are always allowed; loopback only with `RUSTFS_REPLICATION_ALLOW_LOOPBACK_TARGET=true` (`crates/ecstore/src/bucket/bucket_target_sys.rs` `validate_replication_target_endpoint`) |
| Site replication peers | Literal check | `rustfs/src/site_replication/mod.rs` |
| Tiering warm backends (S3, MinIO, RustFS, Azure, GCS, Aliyun, Tencent, Huawei, R2) | Literal check | `crates/ecstore/src/services/tier/warm_backend.rs` `validate_endpoint`; the RustFS provider adds a debug-only, env-gated loopback exception for e2e tests |
| Keystone `auth_url` | Literal check | `crates/keystone/src/config.rs` |
- event-notification webhooks (`RUSTFS_NOTIFY_WEBHOOK_*`);
- audit webhooks (`RUSTFS_AUDIT_WEBHOOK_*`);
- OIDC identity-provider discovery, JWKS, and token requests (since `1.0.0-beta.12`);
- S3 tiering (warm-backend) endpoints;
- Keystone auth URLs.
The webhook and audit outbound clients also **disable proxies and do not follow
redirects**, so the destination must be reachable directly at the configured URL.
## What changed in beta.11 (and for OIDC in beta.12)
For webhook and audit clients:
| | beta.10 | beta.11+ |
|---|---|---|
| Literal `localhost` / private / loopback IPs | Rejected | Rejected |
| Hostnames that resolve to private/loopback addresses (`logstash`, `host.docker.internal`, Compose service DNS, …) | Allowed | **Blocked at DNS/connect time** unless allowlisted |
| Escape hatch for private destinations | None | `RUSTFS_OUTBOUND_ALLOW_ORIGINS` |
| Proxies / redirects for outbound clients | Followed | Disabled |
Before beta.11 a webhook endpoint whose hostname happened to resolve to a
private address was accepted. Beta.11 fails that resolution check unless the
exact origin is on the allowlist. This is why a Compose setup that delivered
events on beta.10 can go silent after the upgrade even though the configuration
is unchanged.
OIDC joined the same policy in beta.12. An internal identity provider that
worked in beta.11 can therefore fail discovery after upgrading to beta.12 unless
its exact origin is allowlisted. The policy remains active for discovery, JWKS,
and token requests.
The allowlist affects only the "Full policy" rows. A literal-check subsystem rejects a hostname that is itself a restricted IP literal, does not re-check what a hostname resolves to, and cannot be widened by `RUSTFS_OUTBOUND_ALLOW_ORIGINS`.
## Symptoms
- Bucket event rules and webhook configuration look correct.
- Uploads and audited API calls succeed.
- No HTTP POST reaches the internal webhook receiver.
- The target may appear offline or fail activation when its endpoint resolves to
a loopback, private, shared, or reserved address.
- Startup or target validation reports `webhook endpoint is not allowed: ...`
with a reason such as `private address` or `loopback host`.
- An OIDC provider or login button is missing, and startup reports
`OIDC provider discovery blocked by outbound policy` with the exact origin to
allowlist.
- Bucket event rules and webhook configuration look correct and uploads succeed, but no POST reaches the receiver.
- Target validation reports `<field> is not allowed: ...` with a reason such as `private address` or `loopback host`; when an exact-origin allowlist entry would fix it, the message says so.
- An OIDC login button is missing and startup logs `OIDC provider discovery blocked by outbound policy`.
## `RUSTFS_OUTBOUND_ALLOW_ORIGINS`
`RUSTFS_OUTBOUND_ALLOW_ORIGINS` is a comma-separated list of exact HTTP(S)
origins that are permitted to resolve to otherwise-restricted addresses. It is an
operator-owned process setting read once at startup; individual target
configuration cannot extend it.
A comma-separated list of exact HTTP(S) origins permitted to resolve to otherwise-restricted addresses. It is a process-level setting read once at startup; individual target configuration cannot extend it.
```bash
# exact scheme://host:port — comma-separate multiple origins
RUSTFS_OUTBOUND_ALLOW_ORIGINS=http://logstash:8080,http://host.docker.internal:3020
```
### Origin format rules
Each entry is matched as an **exact origin** (`scheme://host:port`):
- The scheme must be `http` or `https`.
- The host and port must match the destination exactly. An allowlisted
`http://logstash:8080` does **not** authorize `http://logstash:9090` or
`https://logstash:8080`.
- If the port is omitted, the scheme's default is used (`80` for `http`, `443`
for `https`); the destination must then use that same default port.
- Entries must be origins only. A trailing `/` is accepted, but a path, query,
or fragment (for example `http://logstash:8080/events`) is **rejected** as an
invalid origin — the process fails closed rather than silently ignoring the
path.
- Userinfo (`http://user:pass@host`) is not allowed.
- An empty entry (for example a trailing or doubled comma) is rejected.
An invalid list fails closed: the affected subsystem reports an
`invalid outbound policy` / `invalid origin at position N` error instead of
starting with a partially applied allowlist.
| Rule | Detail |
| --- | --- |
| Exact origin | `scheme://host:port`. `http://logstash:8080` does not authorize `http://logstash:9090` or `https://logstash:8080` |
| Scheme | `http` or `https` only |
| Default port | If omitted, the scheme default (`80` / `443`) applies and the destination must use that port |
| Origin only | A trailing `/` is accepted; any path, query, or fragment (`http://logstash:8080/events`) is rejected |
| No userinfo | `http://user:pass@host` is rejected |
| No empty entries | A trailing or doubled comma is rejected |
| Fail closed | An invalid list yields `invalid outbound policy` / `invalid origin at position N` and the affected subsystem does not start with a partially applied allowlist |
### What stays blocked even when allowlisted
Allowlisting an origin only relaxes the loopback, private, shared, and reserved
address classes for that exact origin. The following remain forbidden for every
origin, allowlisted or not:
- Cloud metadata endpoints (`169.254.169.254` and the other well-known IMDS addresses).
- Link-local addresses (`169.254.0.0/16`, `fe80::/10`) and the unspecified address (`0.0.0.0`, `::`).
- IPv4-mapped, IPv4-compatible, and NAT64/6to4-embedded forms of the above; the embedded IPv4 address is what gets classified, so `::ffff:127.0.0.1` cannot bypass the policy.
- cloud metadata endpoints (for example `169.254.169.254` and the other
well-known IMDS addresses);
- link-local addresses (`169.254.0.0/16`, `fe80::/10`);
- the unspecified address (`0.0.0.0`, `::`);
- IPv4-mapped, IPv4-compatible, and NAT64/6to4-embedded forms of any of the
above (RustFS classifies the embedded IPv4 destination, so `::ffff:127.0.0.1`
and similar cannot be used to bypass the policy).
The allowlist authorizes only the exact host you name. A DNS answer for a
different hostname that points at a private address is still rejected, and each
new connection re-validates the resolved addresses so a rebinding answer fails
closed.
The allowlist authorizes only the exact host named. A DNS answer for a different hostname that points at a private address is still rejected, and each new connection re-validates the resolved addresses.
## Docker Compose example
@@ -128,25 +70,11 @@ services:
RUSTFS_NOTIFY_WEBHOOK_ENDPOINT_PRIMARY: "http://logstash:8080/events"
RUSTFS_NOTIFY_WEBHOOK_QUEUE_DIR_PRIMARY: "/tmp/rustfs-events"
# Allow the webhook host to resolve to the Compose private network.
# Note: the allowlist takes the origin only, without the /events path.
# The allowlist takes the origin only, without the /events path.
RUSTFS_OUTBOUND_ALLOW_ORIGINS: "http://logstash:8080"
logstash:
image: docker.elastic.co/logstash/logstash:8.15.0
# ...
```
The endpoint keeps its full path (`/events`); the allowlist entry is the origin
(`http://logstash:8080`) only.
## Upgrade checklist (beta.10 → beta.11+, or OIDC beta.11 → beta.12+)
1. List every outbound endpoint whose hostname resolves to a loopback, private,
shared, or reserved address: notification webhooks, audit webhooks, OIDC
providers, tiering endpoints, and Keystone auth URLs.
2. Add each one to `RUSTFS_OUTBOUND_ALLOW_ORIGINS` as an exact
`scheme://host:port` origin (no path).
3. Ensure the endpoint is reachable directly — for webhook and audit targets,
proxies are disabled and redirects are not followed.
4. Restart RustFS; the policy is read at startup.
5. Confirm delivery, and check the logs for `... is not allowed` messages if a
target still fails to activate.
The endpoint keeps its full path (`/events`); the allowlist entry is the origin only. Restart RustFS after changing the variable — the policy is read at startup — and check the logs for `is not allowed` messages if a target still fails to activate.
+15 -43
View File
@@ -1,8 +1,9 @@
# Pool metadata upgrade and recovery
`pool.bin` is cluster state. Do not delete or copy it independently on a live
node. Version 3 adds a deployment identity, epoch, durable generation, and a
recoverable prepare/commit record on every pool.
**Use this when:** upgrading a cluster to `pool.bin` V3, a node cannot rejoin after a metadata-drive replacement, or startup reports `pool.bin` as incompatible, corrupt, or recovery required.
**Source of truth:** `crates/ecstore/src/core/pools.rs` (`pool.bin` / `pool.bin.identity` reader and writer, the `RUSTFS_POOL_META_V3_WRITE` and `RUSTFS_POOL_META_V3_FLEET_CONFIRMED` gates).
`pool.bin` is cluster state. Do not delete or copy it independently on a live node. Version 3 adds a deployment identity, epoch, durable generation, and a recoverable prepare/commit record on every pool.
## Compatibility matrix
@@ -12,59 +13,30 @@ recoverable prepare/commit record on every pool.
| V2-capable binary | read/write while mixed | read/write after the V2 fleet gate | reject |
| V3-capable binary | read/migrate | read/migrate | read/write; never downgrade |
Leave `RUSTFS_POOL_META_V3_WRITE` or
`RUSTFS_POOL_META_V3_FLEET_CONFIRMED` disabled while any running process lacks
V3 support. Both must be `true` before an existing cluster migrates. A fresh
deployment can initialize directly at V3. Once a committed V3 generation is
observed, rollback to a V1/V2-only binary is not supported.
Repairing a missing identity on an existing V1/V2 snapshot does not cross the
V3 gate; the identity is committed as initialized while `pool.bin` stays on its
observed legacy version.
Leave `RUSTFS_POOL_META_V3_WRITE` and `RUSTFS_POOL_META_V3_FLEET_CONFIRMED` disabled while any running process lacks V3 support. Both must be `true` before an existing cluster migrates. A fresh deployment can initialize directly at V3. Once a committed V3 generation is observed, rollback to a V1/V2-only binary is not supported. Repairing a missing identity on an existing V1/V2 snapshot does not cross the V3 gate; the identity is committed as initialized while `pool.bin` stays on its observed legacy version.
Unknown fields are not ignored. An unsupported version or field layout is
reported as **incompatible** and is never overwritten. A truncated or invalid
payload is **corrupt** and may be repaired only from a verified committed
replica. Conflicting identities, epochs, or transactions at the same generation
are **recovery required** and need an operator-selected source.
| Startup verdict | Cause | Handling |
| --- | --- | --- |
| **incompatible** | Unsupported version or field layout (unknown fields are not ignored) | Never overwritten |
| **corrupt** | Truncated or invalid payload | Repaired only from a verified committed replica |
| **recovery required** | Conflicting identities, epochs, or transactions at the same generation | Needs an operator-selected source |
## Partial writes
A V3 update first conditionally writes a pending generation containing the last
committed snapshot, then conditionally replaces it with the committed record.
During initial bootstrap, `pool.bin.identity` remains `initialized=false` and
carries a unique fresh-bootstrap nonce until that committed V3 record is
verified. Restarting from an initial prepare record finishes generation 1; it
never rewrites the record as V1 or V2.
On restart:
A V3 update first conditionally writes a pending generation containing the last committed snapshot, then conditionally replaces it with the committed record. During initial bootstrap, `pool.bin.identity` remains `initialized=false` and carries a unique fresh-bootstrap nonce until that committed V3 record is verified. Restarting from an initial prepare record finishes generation 1; it never rewrites the record as V1 or V2. On restart:
- prepare-only replicas expose their previous committed snapshot;
- one committed replica makes that transaction authoritative;
- remaining pending or older replicas are repairable by the next fenced save;
- two different committed transactions at one generation stop startup.
Do not hand-edit a pending record or select a replica only because it is in pool
zero. Preserve all copies when escalating recovery.
Do not hand-edit a pending record or select a replica only because it is in pool zero. Preserve all copies when escalating recovery.
## Disk replacement and metadata erasure
1. Keep a quorum of nodes online and verify the cluster is ready.
2. Stop the lagging node before replacing or erasing its metadata drive.
3. Restore storage formats and the `pool.bin.identity` marker from the same
deployment before rejoining it.
4. Start the node and wait for it to load the verified committed generation and
repair its replicas before touching another node.
3. Restore storage formats and the `pool.bin.identity` marker from the same deployment before rejoining it.
4. Start the node and wait for it to load the verified committed generation and repair its replicas before touching another node.
An initialized identity with every `pool.bin` missing is recovery required.
Existing storage formats with neither identity nor `pool.bin` are also recovery
required. Format creation alone is not fresh-cluster proof. Only the elected
first topology node may create a durable `initialized=false` bootstrap identity
with a fresh-bootstrap nonce, and only after every configured disk explicitly
responds that it is unformatted.
An unreachable peer, a non-elected distributed node, or an existing format is
not sufficient proof. All-missing `pool.bin` replicas are accepted only by the
same startup that proved the fresh topology and persisted that pending identity.
When every `pool.bin` is missing, a later startup must recover even if the
pending identity survived. This prevents a wiped or lagging node from rebuilding
empty state and overwriting the cluster. Runtime reload, rebalance activation,
and rebalance worker admission all fail closed and latch the same recovery gate
until the node is restarted with readable metadata.
An initialized identity with every `pool.bin` missing is recovery required, as are existing storage formats with neither identity nor `pool.bin`. Format creation alone is not fresh-cluster proof: only the elected first topology node may create a durable `initialized=false` bootstrap identity with a fresh-bootstrap nonce, and only after every configured disk explicitly responds that it is unformatted. An unreachable peer, a non-elected distributed node, or an existing format is not sufficient proof. All-missing `pool.bin` replicas are accepted only by the same startup that proved the fresh topology and persisted that pending identity; when every `pool.bin` is missing, a later startup must recover even if the pending identity survived. This prevents a wiped or lagging node from rebuilding empty state and overwriting the cluster. Runtime reload, rebalance activation, and rebalance worker admission all fail closed and latch the same recovery gate until the node is restarted with readable metadata.
@@ -1,55 +0,0 @@
# Presigned multipart total-size limit
RustFS V2 supports an optional capability on a signed or SigV4-presigned
`CreateMultipartUpload` request:
```text
x-rustfs-max-total-object-size=<unsigned 64-bit integer>
```
The backend must include the parameter before calculating the SigV4
signature. It is part of the canonical query and cannot be added, removed, or
changed by the browser. RustFS stores the verified limit in the multipart
upload session and applies it to every `UploadPart` and to
`CompleteMultipartUpload`.
Backend pseudocode (the custom query must be present before signing):
```text
uri = "/photos/archive.zip?uploads"
uri += "&x-rustfs-max-total-object-size=104857600"
presigned_url = sigv4_presign("POST", uri, credentials)
# Return presigned_url to the browser. Never append the parameter afterwards.
```
The resulting flow is:
1. The backend signs `CreateMultipartUpload?...&x-rustfs-max-total-object-size=104857600`.
2. RustFS verifies the SigV4 request and persists the limit with the upload ID.
3. The browser uploads parts using the returned upload ID.
4. RustFS rejects a part whose declared logical size would exceed the remaining
budget and rejects completion if the server-side part metadata exceeds the
limit.
The limit is measured in logical object bytes (`actual_size`), not erasure,
encryption, or compression bytes. Replacing an existing part uses replacement
semantics: the old part size is removed before the new part size is admitted.
Unknown-length parts are rejected for capped sessions rather than buffered
without a bound. Capped parts are admitted under an upload-wide write lock
before temporary shards are created and use a per-upload staging permit to
bound local in-flight data. The distributed lock is released while the body is
read and reacquired for the final check/rename, so Complete and Abort are not
blocked behind a slow upload. The normal request-body stall timeout releases
the staging permit when a client stops sending.
The parameter is accepted only on `CreateMultipartUpload`. Supplying it on
`UploadPart`, `CompleteMultipartUpload`, `AbortMultipartUpload`, listing, or
copy operations returns `InvalidRequest`; those requests use the persisted
session state. A multipart upload created without this parameter remains
unlimited for backward compatibility. The V1 single-request capability
(`x-rustfs-max-content-length`) is independent and is not a multipart limit.
Because enforcement happens in the multipart data plane, every node that may
receive requests for a capped upload must run the V2 implementation. During a
rolling upgrade, route capped uploads only to upgraded nodes; older nodes treat
the internal metadata as unknown and cannot enforce the limit.
@@ -1,35 +0,0 @@
# Presigned PutObject size limit
RustFS V1 supports an optional, RustFS-specific capability on a SigV4
presigned `PutObject` URL:
```text
x-rustfs-max-content-length=<unsigned 64-bit integer>
```
The backend that creates the URL must add this query parameter to the request
URI before calculating the SigV4 presign. It is part of the canonical query;
adding, removing, or changing it after signing invalidates the signature. A
browser can then upload with a plain `PUT` and does not need a custom size
header.
RustFS validates the capability after SigV4 authentication and enforces it on
the decoded request body. A declared `Content-Length` above the limit is
rejected before storage. If the body produces more bytes than the limit while
streaming, RustFS returns `EntityTooLarge` and does not publish the object.
The V1 contract is deliberately narrow:
- The parameter is accepted only on a SigV4 presigned `PutObject` request.
- Duplicate, case-variant, malformed, negative, or overflowing values return
`InvalidRequest`.
- Requests without the parameter, including ordinary authenticated or
anonymous `PUT`, keep the existing behavior.
- The parameter on `CopyObject`, multipart, `GET`, `HEAD`, `DELETE`, bucket, or
other operations returns `InvalidRequest`.
- Unknown-length and SigV4 streaming-chunked uploads remain unsupported by the
existing PutObject admission contract and are not enabled by this feature.
This capability is per request; it is not a cumulative multipart-upload cap.
Multipart session limits are planned for V2 under a separate query/API
contract.
+57
View File
@@ -0,0 +1,57 @@
# Presigned upload size limits
**Use this when:** a backend issues SigV4-presigned upload URLs to browsers and must cap how much a client can upload with one URL — per request (`PutObject`) or per multipart upload.
**Source of truth:** `rustfs/src/auth.rs` (`RUSTFS_MAX_CONTENT_LENGTH_QUERY`, `RUSTFS_MAX_TOTAL_OBJECT_SIZE_QUERY`, `parse_presigned_put_max_content_length`, `parse_presigned_multipart_max_total_object_size`); enforcement in `rustfs/src/app/object/put.rs` (`MaxContentLengthStream`), `rustfs/src/app/multipart_usecase.rs` (`multipart_max_total_object_size`), and `crates/ecstore/src/set_disk/ops/multipart.rs` (`multipart_size_limit_from_metadata`, `admitted_multipart_size`).
Both limits are RustFS-specific query parameters carried inside the SigV4 canonical query.
## Shared signing rule
1. The backend appends the parameter to the request URI **before** computing the SigV4 presigned signature. It is part of the canonical query, so adding, removing, or changing it afterwards invalidates the signature; the browser cannot alter it.
2. RustFS parses the parameter only after the request has been accepted as SigV4-signed (the `VerifiedPresignedRequest` / `VerifiedSigV4Request` request markers). The same query string on an unsigned request is rejected with `InvalidRequest`.
3. The value is an unsigned 64-bit integer. Duplicate, case-variant, malformed, negative, or overflowing values return `InvalidRequest`.
4. Requests that do not carry the parameter, including ordinary authenticated or anonymous uploads, keep their existing behavior.
| | V1 per-request | V2 per-upload |
| --- | --- | --- |
| Query parameter | `x-rustfs-max-content-length=<u64>` | `x-rustfs-max-total-object-size=<u64>` |
| Accepted on | SigV4 presigned `PutObject` only | `CreateMultipartUpload` only (signed or presigned SigV4) |
| Any other operation carrying it | `InvalidRequest` (`CopyObject`, multipart, `GET`, `HEAD`, `DELETE`, bucket operations) | `InvalidRequest` (`UploadPart`, `CompleteMultipartUpload`, `AbortMultipartUpload`, listing, copy — these read the persisted session state instead) |
| What is measured | Decoded request body bytes of that one request | Logical object bytes (`actual_size`) summed across the upload's parts; not erasure, encryption, or compression bytes |
| Where the limit lives | The request only | Multipart session metadata (`SUFFIX_MAX_TOTAL_OBJECT_SIZE`), written at create time |
| Over-limit result | `EntityTooLarge`; the object is not published | `EntityTooLarge` on the offending `UploadPart`, and on `CompleteMultipartUpload` if the recorded parts exceed the limit |
## V1: `x-rustfs-max-content-length`
```text
uri = "/photos/avatar.png"
uri += "?x-rustfs-max-content-length=10485760"
presigned_url = sigv4_presign("PUT", uri, credentials)
# Return presigned_url to the browser. Never append the parameter afterwards.
```
- A declared `Content-Length` above the limit is rejected before storage. A body that streams more bytes than the limit is cut off with `EntityTooLarge` and nothing is published.
- Not combinable with archive auto-extraction (`x-amz-meta-snowball-auto-extract`): `InvalidRequest`.
- Unknown-length and SigV4 streaming-chunked uploads stay outside the existing PutObject admission contract; this parameter does not enable them.
- Per request only: it is neither a cumulative cap across several PUTs nor a multipart limit.
## V2: `x-rustfs-max-total-object-size`
```text
uri = "/photos/archive.zip?uploads"
uri += "&x-rustfs-max-total-object-size=104857600"
presigned_url = sigv4_presign("POST", uri, credentials)
# Return presigned_url to the browser. Never append the parameter afterwards.
```
1. RustFS verifies the SigV4 request and persists the limit with the upload ID.
2. The browser uploads parts with the returned upload ID; part requests carry no custom parameter.
3. Each `UploadPart` (and each `UploadPartCopy` into the upload) is admitted only if the upload's running logical total plus this part fits the budget. Replacing an existing part number uses replacement semantics: the old part's size is released before the new size is admitted.
4. `CompleteMultipartUpload` re-sums the recorded parts and rejects the completion if they exceed the limit.
Properties of a capped upload:
- Unknown-length or negative-length parts are rejected with `UnexpectedContent` rather than buffered without a bound.
- Capped parts are admitted under an upload-wide write lock before temporary shards are created, and hold a per-upload staging permit that bounds local in-flight data. The lock is released while the body is read and reacquired for the final check and rename, so `Complete` and `Abort` are not blocked behind a slow upload. The request-body stall timeout releases the staging permit when a client stops sending.
- An upload created without the parameter stays unlimited.
- Enforcement runs in the multipart data plane on every node. During a rolling upgrade, route capped uploads only to nodes that carry the V2 implementation; a node without it treats the internal metadata as unknown and cannot enforce the limit.
@@ -1,50 +1,35 @@
# Rebalance Stored-Representation Impact Guide
This guide covers the historical data-movement read defect tracked by
[`rustfs/backlog#1850`](https://github.com/rustfs/backlog/issues/1850). It is an
impact-assessment and read-only triage guide. It does not repair, rewrite,
migrate, delete, or quarantine any object.
**Use this when:** a deployment ran pool rebalance (or, in a narrower window, decommission) on an affected release and you must assess whether compressed or server-side-encrypted objects were copied as plaintext under their original metadata. Read-only triage only; this guide repairs nothing.
The defect affected data movement when the source reader returned logical
plaintext but the target writer preserved the source's stored-representation
metadata and sizes. Compressed objects could therefore be copied as plaintext
under compression metadata. Server-managed encrypted objects could be copied as
plaintext under encryption metadata. The forward rebalance fix reached `main`
in commit
[`e11fcfbd`](https://github.com/rustfs/rustfs/commit/e11fcfbd087f8a8dae2c0f2c62bc0f6e40e3f10a)
through [PR #6057](https://github.com/rustfs/rustfs/pull/6057).
**Source of truth:** `crates/ecstore/src/services/rebalance/migration.rs` and `crates/ecstore/src/core/pools.rs` (`raw_data_movement_read` on the source read options), `crates/ecstore/src/object_api/readers.rs` (raw stored-range read path), `crates/ecstore/src/data_movement/mod.rs` (metadata, part-size, ETag and index preservation), `crates/filemeta/examples/dump_fileinfo.rs` (evidence decoder). Tracked as `rustfs/backlog#1850`.
Upgrading prevents this defect in later rebalance runs. It does not validate or
repair copies produced by an earlier run.
## Mechanism
The migration pipeline is a stored-representation copier: it preserves the source ETag and internal metadata, divides the stream using stored `part.size` values, and carries the decoded compression index. The affected rebalance read options supplied only the version ID and lock setting, so the normal GET read plan decompressed or decrypted the stream first. The target write could therefore complete while its bytes no longer matched the metadata describing them: compressed objects became plaintext under compression metadata, SSE-S3 and SSE-KMS objects became plaintext under encryption metadata.
Historical rebalance cleanup deleted the source entry only after every version in it was reported moved. A target write accepted as a successful move could therefore be followed by source deletion even though a later GET of the target fails. Conversely, a source-read failure prevented the version from being counted as moved and blocked normal source cleanup.
The forward fix (commit [`e11fcfbd`](https://github.com/rustfs/rustfs/commit/e11fcfbd087f8a8dae2c0f2c62bc0f6e40e3f10a), [PR #6057](https://github.com/rustfs/rustfs/pull/6057)) sets `raw_data_movement_read: true` for rebalance source reads. Upgrading prevents the defect in later runs; it does not validate or repair copies produced by an earlier run.
## Immediate Operator Decision
Treat a deployment as exposed when both conditions are true:
1. it ran rebalance in an affected build, or decommission in the narrower
historical window described below; and
1. it ran rebalance in an affected build, or decommission in the narrower historical window below; and
2. the operation could have selected compressed, SSE-S3, or SSE-KMS objects.
For an exposed deployment:
- preserve old pool media, snapshots, replicas, and backups before any pool is
removed, reformatted, reused, or returned;
- stop destructive cleanup and do not use another rebalance or decommission run
as a repair mechanism;
- preserve old pool media, snapshots, replicas, and backups before any pool is removed, reformatted, reused, or returned;
- stop destructive cleanup and do not use another rebalance or decommission run as a repair mechanism;
- inventory and validate candidates with read-only operations;
- handle SSE-S3 and SSE-KMS candidates as a confidentiality incident as well as
a data-integrity incident;
- restore only from a separately verified source under an incident-specific
recovery plan.
- handle SSE-S3 and SSE-KMS candidates as a confidentiality incident as well as a data-integrity incident;
- restore only from a separately verified source under an incident-specific recovery plan.
## Affected Versions
The release boundaries below were verified by tag ancestry. Commit
[`a236b0d0`](https://github.com/rustfs/rustfs/commit/a236b0d01d40a152309446a553756ea991c9f901)
introduced the merged rebalance and decommission implementation. Commit
[`2f25cf60`](https://github.com/rustfs/rustfs/commit/2f25cf606e5ca814fe992be6327a91e31fe066b3)
introduced the raw stored-representation read mode and wired it into
decommission. Commit `e11fcfbd` wired the same mode into rebalance.
Boundaries were verified by tag ancestry. Commit [`a236b0d0`](https://github.com/rustfs/rustfs/commit/a236b0d01d40a152309446a553756ea991c9f901) introduced the merged rebalance and decommission implementation. Commit [`2f25cf60`](https://github.com/rustfs/rustfs/commit/2f25cf606e5ca814fe992be6327a91e31fe066b3) introduced the raw stored-representation read mode and wired it into decommission. Commit `e11fcfbd` wired the same mode into rebalance.
| Release or commit range | Rebalance | Decommission | Operator classification |
| --- | --- | --- | --- |
@@ -53,31 +38,9 @@ decommission. Commit `e11fcfbd` wired the same mode into rebalance.
| `1.0.0-beta.9` through `1.0.0-rc.1`, from `2f25cf60` up to but excluding `e11fcfbd` | Decoded read | Raw stored-representation read | Rebalance requires assessment; decommission is not affected by this defect |
| `1.0.0-rc.2` and later, at or after `e11fcfbd` | Raw stored-representation read | Raw stored-representation read | Forward-fixed; earlier copies still require assessment |
Preview tags follow the commit they reference. In particular, the `rc.1`
preview is affected and the `rc.2` preview contains the forward fix. For custom
or untagged builds, compare the deployed commit with the three commit boundaries
rather than inferring behavior from a version string.
Preview tags follow the commit they reference: the `rc.1` preview is affected and the `rc.2` preview contains the forward fix. For custom or untagged builds, compare the deployed commit with the three commit boundaries rather than inferring behavior from a version string.
The historical decommission result is narrower than the rebalance result but is
not empty. Before `2f25cf60`, decommission used the same ordinary decoded reader.
From `1.0.0-beta.9` onward it explicitly used `raw_data_movement_read: true`.
Any code change or automated remediation for the earlier decommission window is
outside this report and requires a separate issue.
## Why The Copy Could Be Accepted
The migration pipeline is a stored-representation copier. It preserves the
source ETag and internal metadata, uses stored `part.size` values to divide the
stream, and carries the decoded compression index. The affected rebalance read
options supplied only the version ID and lock setting, so the normal GET read
plan decompressed or decrypted the stream first. A target write could therefore
complete while its bytes no longer matched the metadata that described them.
Historical rebalance cleanup ran only after every version in an entry was
reported moved. It then deleted the source entry. A target write accepted as a
successful move could therefore be followed by source deletion even though a
later GET of the target would fail. Conversely, a source-read failure prevented
the version from being counted as moved and prevented normal source cleanup.
The decommission exposure is narrower than the rebalance exposure but not empty: before `2f25cf60` decommission used the same ordinary decoded reader. Any code change or automated remediation for that earlier decommission window is outside this guide and requires a separate issue.
## Object Classification
@@ -90,32 +53,21 @@ the version from being counted as moved and prevented normal source cleanup.
| SSE-C | The migration request did not have the customer key, so the normal read failed closed | Migration failure and possible incomplete progress; no successful corrupting copy is expected from this path | Medium; confirm the source was retained |
| Any compressed and encrypted combination | Multiple stored-representation assumptions were violated | Confidentiality exposure and data corruption | Critical |
The classification is specific to this defect. A low-risk classification does
not certify an object against unrelated corruption.
The classification is specific to this defect. A low-risk classification does not certify an object against unrelated corruption.
## Read-Only Assessment Workflow
### 1. Establish The Operation Window
### 1. Establish the operation window
Record the exact RustFS version and commit for every node that participated.
Collect the authenticated rebalance status response, decommission status when
applicable, service logs, deployment change records, and release history.
Record the exact RustFS version and commit for every node that participated. Collect the authenticated rebalance status response, decommission status when applicable, service logs, deployment change records, and release history.
Persisted rebalance metadata records the run ID, participating pools, start and
end state, bucket lists, counters, and the last bucket/object progress value. It
does not persist a complete per-object movement ledger. Status metadata can
prove that a run occurred and narrow time, pool, and bucket scope, but it cannot
by itself enumerate every moved object.
Persisted rebalance metadata records the run ID, participating pools, start and end state, bucket lists, counters, and the last bucket/object progress value. It does not persist a per-object movement ledger: status metadata can prove that a run occurred and narrow time, pool, and bucket scope, but it cannot enumerate every moved object.
If no reliable operation record remains, assume that every object version in a
bucket present during the affected deployment interval is a candidate until
other evidence narrows the set.
If no reliable operation record remains, assume that every object version in a bucket present during the affected deployment interval is a candidate until other evidence narrows the set.
### 2. Build A Candidate Inventory
### 2. Build a candidate inventory
Use read-only S3 list and list-object-versions operations for the buckets in
scope. Preserve bucket, key, version ID, last-modified time, size, ETag, storage
class, and any client-side content digest. Join that list with:
Use read-only S3 list and list-object-versions operations for the buckets in scope. Preserve bucket, key, version ID, last-modified time, size, ETag, storage class, and any client-side content digest. Join that list with:
- upload records that identify compression settings or SSE mode;
- KMS audit history and application catalogs;
@@ -123,20 +75,13 @@ class, and any client-side content digest. Join that list with:
- rebalance/decommission timestamps and source/target pool records;
- server access logs showing successful or failed reads after movement.
Do not use ETag equality as proof of content integrity. The migration writer
preserved the source ETag, including for a malformed target copy, and multipart
or encrypted ETags are not general-purpose content hashes.
Do not use ETag equality as proof of content integrity. The migration writer preserved the source ETag, including for a malformed target copy, and multipart or encrypted ETags are not general-purpose content hashes.
### 3. Classify Stored Metadata On Evidence Copies
### 3. Classify stored metadata on evidence copies
When API and application records cannot classify a candidate, copy `xl.meta`
from each relevant shard disk to a restricted evidence location and inspect the
copy on an offline host. Do not edit or decode metadata in place on a live data
path. Keep the evidence copies under the same access controls as the object.
When API and application records cannot classify a candidate, copy `xl.meta` from each relevant shard disk to a restricted evidence location and inspect the copy on an offline host. Do not edit or decode metadata in place on a live data path. Keep the evidence copies under the same access controls as the object.
The existing `rustfs-filemeta` example can decode an evidence copy. It prints
metadata values, some of which are sensitive encryption material, so redact
metadata values before they reach a terminal or report:
The `rustfs-filemeta` example decodes an evidence copy. It prints metadata values, some of which are sensitive encryption material, so redact values before they reach a terminal or report:
```bash
cargo run --quiet -p rustfs-filemeta --example dump_fileinfo -- /evidence/object/xl.meta |
@@ -145,119 +90,51 @@ cargo run --quiet -p rustfs-filemeta --example dump_fileinfo -- /evidence/object
Use the output only as a screen:
- either the `x-rustfs-internal-compression` or
`x-minio-internal-compression` key marks a compressed representation;
- `actual-size`, per-part `size`/`actual_size`, and compression-index totals
should be arithmetically consistent;
- either the `x-rustfs-internal-compression` or `x-minio-internal-compression` key marks a compressed representation;
- `actual-size`, per-part `size`/`actual_size`, and compression-index totals should be arithmetically consistent;
- SSE-C customer-algorithm/MD5 markers identify SSE-C;
- KMS key-ID/context markers identify SSE-KMS;
- a managed encryption envelope without SSE-C or KMS markers identifies an
SSE-S3 candidate.
- a managed encryption envelope without SSE-C or KMS markers identifies an SSE-S3 candidate.
Never include encryption metadata values in tickets, logs, chat, or assessment
reports. Metadata consistency is necessary but not sufficient: the defect
preserved metadata, so plausible sizes and a decodable index do not prove that
the stored bytes match it.
Never include encryption metadata values in tickets, logs, chat, or assessment reports. Metadata consistency is necessary but not sufficient: the defect preserved metadata, so plausible sizes and a decodable index do not prove that the stored bytes match it.
### 4. Validate Logical Content Without Mutation
### 4. Validate logical content without mutation
For each high- or critical-risk candidate, perform a complete authenticated GET
of the exact version into a restricted validation sink. Supply the customer key
only for an authorized SSE-C check. Record the status, byte count, and a
cryptographic digest calculated by the validation client. Compare it with a
digest from an independently trusted source, backup, replica, or application
record.
For each high- or critical-risk candidate, perform a complete authenticated GET of the exact version into a restricted validation sink. Supply the customer key only for an authorized SSE-C check. Record the status, byte count, and a cryptographic digest calculated by the validation client. Compare it with a digest from an independently trusted source, backup, replica, or application record.
Interpret the result conservatively:
- a GET decode/decrypt error, unexpected EOF, or short byte count is a strong
affected-copy signal, but may also have another corruption cause;
- a GET decode/decrypt error, unexpected EOF, or short byte count is a strong affected-copy signal, but may also have another corruption cause;
- a matching independent cryptographic digest validates that logical version;
- a successful GET without an independent digest proves readability, not
identity;
- a successful GET without an independent digest proves readability, not identity;
- a matching ETag alone is inconclusive;
- an SSE-S3/KMS candidate moved in the affected window remains a confidentiality
incident until storage-level review excludes plaintext target copies and
derivative snapshots or backups.
- an SSE-S3/KMS candidate moved in the affected window remains a confidentiality incident until storage-level review excludes plaintext target copies and derivative snapshots or backups.
Storage-level confirmation for managed-SSE candidates may expose plaintext and
sealed-key material. It must be performed only by the incident/security owner on
offline evidence copies. Do not print, upload, or serve raw shard bytes, and do
not bypass RustFS to return them to an application.
Storage-level confirmation for managed-SSE candidates may expose plaintext and sealed-key material. It must be performed only by the incident/security owner on offline evidence copies. Do not print, upload, or serve raw shard bytes, and do not bypass RustFS to return them to an application.
### 5. Record Confidence And Outcome
### 5. Record confidence and outcome
Record one result for every candidate version:
- `confirmed-good`: full logical bytes match an independent digest;
- `confirmed-affected`: target decode/decrypt/length evidence and a trusted
source establish the mismatch, or authorized storage review confirms
plaintext under managed-SSE metadata;
- `suspected`: the version and operation window match, but proof is incomplete;
- `not-applicable`: evidence proves the object was plain and uncompressed or was
never selected by an affected operation;
- `unrecoverable-pending-source`: affected or suspected, with no verified source
yet found.
| Result | Meaning |
| --- | --- |
| `confirmed-good` | Full logical bytes match an independent digest. |
| `confirmed-affected` | Target decode/decrypt/length evidence and a trusted source establish the mismatch, or authorized storage review confirms plaintext under managed-SSE metadata. |
| `suspected` | The version and operation window match, but proof is incomplete. |
| `not-applicable` | Evidence proves the object was plain and uncompressed or was never selected by an affected operation. |
| `unrecoverable-pending-source` | Affected or suspected, with no verified source yet found. |
Retain the evidence used for each decision. Do not collapse object versions with
the same key into one result.
Retain the evidence used for each decision. Do not collapse object versions with the same key into one result.
## Source Retention And Recovery Limits
## Source Retention and Recovery Limits
Successful historical migration could be followed by source-entry deletion.
Therefore, neither successful rebalance status nor absence from the old source
pool proves that the target bytes are sound. Recovery is possible only from a
separately verified source, such as:
Successful historical migration could be followed by source-entry deletion, so neither successful rebalance status nor absence from the old source pool proves that the target bytes are sound. Recovery is possible only from a separately verified source:
- retained source-pool media or a snapshot taken before cleanup;
- an independently validated replica;
- an external backup;
- the original application or upstream source with a trusted digest.
SSE-C normally failed before the target copy was accepted because the migration
read had no customer key. That failure prevented normal source cleanup, but
operators must verify the exact version on retained source media rather than
assuming it is present.
SSE-C normally failed before the target copy was accepted because the migration read had no customer key. That failure prevented normal source cleanup, but operators must verify the exact version on retained source media rather than assuming it is present.
If no verified source exists, mark the version unrecoverable for this incident.
Do not edit `xl.meta`, rewrite shard files, clear encryption/compression markers,
or overwrite the object in place. Those actions can destroy evidence, violate
retention/versioning policy, or turn a visible read failure into silent data
substitution. Any restoration or replacement procedure needs its own reviewed,
rollback-aware plan.
## Release Guidance
Release notes for `1.0.0-rc.2` and later should state:
> Rebalance now copies the stored object representation for compressed and
> encrypted objects. Deployments that ran rebalance on versions from
> `1.0.0-alpha.91` through `1.0.0-rc.1` should preserve old pool media and run
> the read-only assessment in this guide. Upgrading prevents new copies from
> this defect but does not repair historical copies. Deployments that ran
> decommission from `1.0.0-alpha.91` through `1.0.0-beta.8` require the same
> assessment. SSE-S3 and SSE-KMS candidates require security incident handling.
Do not recommend rerunning rebalance as remediation. Do not remove or repurpose
old pool media until high- and critical-risk candidates have a recorded outcome
and the incident owner has accepted the recovery limits.
## Evidence Audit
The conclusions above are grounded in these repository facts:
- `crates/ecstore/src/services/rebalance/migration.rs` now sets both
`data_movement` and `raw_data_movement_read` for rebalance source reads;
- `crates/ecstore/src/core/pools.rs` sets the same flags for decommission source
reads;
- `crates/ecstore/src/object_api/readers.rs` returns the stored byte range before
compression or encryption transforms when `raw_data_movement_read` is set;
- `crates/ecstore/src/data_movement/mod.rs` preserves stored part sizes, ETags,
indexes, and internal metadata during migration;
- the historical `a236b0d0` rebalance and decommission readers both used normal
read options, while `2f25cf60` changed only decommission to the raw mode;
- the historical rebalance entry deleted its source prefix only after all
versions were counted as moved;
- the tag ancestry boundaries are `1.0.0-alpha.91`, `1.0.0-beta.9`, and
`1.0.0-rc.2` for the implementation, decommission raw-read fix, and rebalance
raw-read fix respectively.
If no verified source exists, mark the version unrecoverable for this incident. Do not edit `xl.meta`, rewrite shard files, clear encryption/compression markers, or overwrite the object in place. Those actions can destroy evidence, violate retention/versioning policy, or turn a visible read failure into silent data substitution. Any restoration or replacement procedure needs its own reviewed, rollback-aware plan.
+20 -32
View File
@@ -1,29 +1,25 @@
# Replication target check
`GET /BUCKET?replication-check` is a signed S3 extension for validating every
replication target referenced by a bucket replication configuration.
**Use this when:** you are about to call, automate, or debug `GET /BUCKET?replication-check`, or need to explain why a `GET` wrote and deleted objects on a replication target.
**Source of truth:** `rustfs/src/admin/router.rs` (`REPLICATION_CHECK_PROBE_PREFIX`, `REPLICATION_CHECK_ERROR_MAX_BYTES`, the `replication-check` route handler).
`GET /BUCKET?replication-check` is a signed S3 extension that validates every replication target referenced by a bucket replication configuration.
## Active mutation warning
Despite using `GET`, this operation is **not read-only**. On each target it:
1. writes an 8-byte object under `.rustfs.sys/replication-check/<uuid>/<uuid>`;
1. writes an 8-byte object under `.rustfs.sys/replication-check/<uuid>/<uuid>` (`REPLICATION_CHECK_PROBE_PREFIX`);
2. creates a replicated delete marker;
3. permanently deletes the probe object version; and
4. enumerates that exact probe key and attempts to delete every remaining
object version and delete marker.
4. enumerates that exact probe key and attempts to delete every remaining object version and delete marker.
Callers should obtain operator confirmation before sending the request. Probe
keys use a reserved namespace and two independent random UUIDs. Before writing,
the server verifies that no version or delete marker exists at the exact key,
then uses an atomic `If-None-Match: *` write so it cannot overwrite a key created
concurrently by an application.
Obtain operator confirmation before sending the request. Probe keys use a reserved namespace and two independent random UUIDs. Before writing, the server verifies that no version or delete marker exists at the exact key, then uses an atomic `If-None-Match: *` write so it cannot overwrite a key created concurrently by an application.
## Response contract
The route returns HTTP 200 with JSON after all configured targets have been
checked. `Status` is `FAILED` when any target or cleanup phase failed; successful
target results remain present when another target fails.
The route returns HTTP 200 with JSON after all configured targets have been checked. `Status` is `FAILED` when any target or cleanup phase failed; successful target results remain present when another target fails.
```json
{
@@ -55,23 +51,15 @@ target results remain present when another target fails.
}
```
Phase states are `OK`, `FAILED`, or `SKIPPED`. Errors are single-line, bounded
to 512 bytes, and omit remote messages, endpoints, credentials, signatures, and
authorization material. A cleanup failure is always explicit; it is never
reported as a successful check.
| Field | Contract |
| --- | --- |
| `Phases.*.Status` | `OK`, `FAILED`, or `SKIPPED`. |
| `Error` | Single line, bounded to `REPLICATION_CHECK_ERROR_MAX_BYTES` (512 bytes); omits remote messages, endpoints, credentials, signatures, and authorization material. |
| `Cleanup` | A cleanup failure is always explicit; it is never reported as a successful check. |
| `Code` | Appears only on failures callers are expected to branch on (currently `BucketRemoteTargetVersionMismatch`). Go decoders ignore the unknown key. |
`VersionFidelity` pins the version-identity contract on **both** write paths:
the probe PUT carries a source version id (header plus `?versionId=` query,
the exact shape live replication uses) and the target must answer with the
same id, and a second probe repeats it through CreateMultipartUpload ->
UploadPart -> CompleteMultipartUpload, where the target fixes the version at
initiate and only reports it on completion. A target can adopt PutObject ids
and still mint its own for multipart, which would leave multipart deletes and
heals addressing a version that never existed; the failure message names the
path that drifted. Targets that
mint their own version ids break every version-addressed operation that
follows (version deletes, heal re-drives), so the phase fails with the
machine-readable extension key `"Code": "BucketRemoteTargetVersionMismatch"`,
the later mutation phases are skipped, and cleanup still removes the probe via
the version id the target actually assigned. `Code` only appears on failures
that callers are expected to branch on; Go decoders ignore the unknown key.
## VersionFidelity phase
`VersionFidelity` pins the version-identity contract on both write paths. The probe PUT carries a source version id (header plus `?versionId=` query, the exact shape live replication uses) and the target must answer with the same id; a second probe repeats the check through CreateMultipartUpload -> UploadPart -> CompleteMultipartUpload, where the target fixes the version at initiate and only reports it on completion. A target can adopt PutObject ids and still mint its own for multipart; the failure message names the path that drifted.
A target that mints its own version ids breaks every version-addressed operation that follows (version deletes, heal re-drives). The phase therefore fails with `"Code": "BucketRemoteTargetVersionMismatch"`, the later mutation phases are skipped, and cleanup still removes the probe via the version id the target actually assigned.
@@ -1,5 +1,8 @@
# Replication object size and shape limits (generic S3 targets)
**Use this when:** an object fails to replicate to an S3-compatible target with `EntityTooLarge`/`EntityTooSmall`, or you need to know whether a large or oddly-chunked object is replicable before relying on it.
**Source of truth:** `crates/ecstore/src/bucket/replication/` (transport selection and part replay), `crates/replication/` (target client), `crates/config/src/constants/` (`RUSTFS_OBS_LOGGER_LEVEL`).
What RustFS can and cannot replicate to a generic S3 target (AWS S3, Wasabi,
MinIO, or any other S3-compatible endpoint configured as a bucket replication
target), and how a rejected object shows up in the log.
@@ -97,5 +100,4 @@ underneath this summary.
- [Replication target check](replication-check.md) — validate a target's
configuration, versioning, and version fidelity before relying on it.
- [Presigned PUT size limit](presigned-put-size-limit.md)
- [Presigned multipart size limit](presigned-multipart-size-limit.md)
- [Presigned size limits](presigned-size-limits.md) — per-request and per-upload caps a backend can put on presigned uploads.
+34 -90
View File
@@ -1,70 +1,38 @@
# Running RustFS behind a reverse proxy
RustFS speaks plain S3 over HTTP/1.1 and HTTP/2 and works behind reverse
proxies (Caddy, Nginx, HAProxy) and CDNs (Cloudflare). Most proxy problems are
**not** RustFS storage bugs — the same request sent directly to `:9000`
succeeds, while the proxied request fails. This page documents the request
semantics RustFS expects from the proxy layer and gives known-good
configurations.
**Use this when:** a request succeeds against `http://<host>:9000` directly but fails, hangs, or resets through Caddy, Nginx, HAProxy, or Cloudflare.
> Rule of thumb: if a request works against `http://<host>:9000` directly but
> fails through the proxy, the fault is in the proxy/CDN request forwarding, not
> in RustFS object handling. Use the checklist below to find which forwarding
> behavior broke.
**Source of truth:** `crates/config/src/constants/tls.rs` (`DEFAULT_HTTP1_HEADER_READ_TIMEOUT`, `DEFAULT_HTTP_REQUEST_BODY_READ_TIMEOUT`); the `put_object_body_read_stalled` log event.
RustFS speaks plain S3 over HTTP/1.1 and HTTP/2. Most proxy problems are not RustFS storage bugs: if the same request works directly against `:9000`, the fault is in proxy/CDN request forwarding. Use the checklist below to find which forwarding behavior broke.
## What RustFS requires from the proxy
S3 clients sign requests with AWS SigV4. RustFS (via `s3s`) re-derives the
signature from the forwarded request, and streams the request body to storage.
For this to succeed the proxy must forward the request **byte-for-byte** with
respect to the signed material and the body:
S3 clients sign requests with AWS SigV4. RustFS (via `s3s`) re-derives the signature from the forwarded request and streams the request body to storage, so the proxy must forward the signed material and the body byte-for-byte:
1. **Do not alter the body.** No transparent compression, no re-encoding, no
truncation. If the client sent `Content-Length: N`, exactly `N` body bytes
must reach RustFS. If fewer bytes arrive, RustFS waits for the rest per the
HTTP spec and the request appears to hang until the client aborts.
2. **Do not rewrite signed headers.** `Host` and any `x-amz-*` / signed headers
must reach RustFS unchanged. Rewriting `Host` is fine only if the client
signed with that same host.
3. **Preserve `Content-Length`; avoid re-chunking large bodies.** Some CDNs
drop `Content-Length` and switch to `Transfer-Encoding: chunked`, or buffer
the whole request body before forwarding — both change the timing and
framing RustFS sees.
4. **Keep upstream idle keep-alive shorter than RustFS's, or vice-versa** (see
next section) so the proxy never reuses a connection RustFS has already
closed.
5. **Do not strip `ETag`** from responses (breaks multipart completion).
| Requirement | Why |
| --- | --- |
| Do not alter the body (no compression, re-encoding, truncation). | If the client sent `Content-Length: N`, exactly `N` body bytes must arrive; with fewer, RustFS waits for the rest and the request appears to hang until the client aborts. |
| Do not rewrite signed headers (`Host`, `x-amz-*`). | Rewriting `Host` is fine only if the client signed with that same host; otherwise `SignatureDoesNotMatch`. |
| Preserve `Content-Length`; do not re-chunk or buffer large bodies. | Switching to `Transfer-Encoding: chunked` or buffering the whole body changes the framing and timing RustFS sees. |
| Keep the proxy's upstream idle keep-alive shorter than RustFS's timeout (next section). | Otherwise the proxy reuses a connection RustFS has already closed. |
| Do not strip `ETag` from responses. | Breaks multipart completion. |
## Idle keep-alive: the #1 cause of `socket hang up` on writes
## Idle keep-alive: the main cause of `socket hang up` on writes
RustFS closes **idle** upstream HTTP/1.1 keep-alive connections after
`RUSTFS_HTTP1_HEADER_READ_TIMEOUT` seconds (default **75s**; see
`crates/config/src/constants/tls.rs`). Reverse proxies keep a pool of upstream
connections and reuse them. If the proxy's upstream idle-keepalive window is
**longer** than RustFS's timeout, the proxy can pick a connection that RustFS
has already FIN'd, write a request onto the dead socket, and the client sees:
RustFS closes idle upstream HTTP/1.1 keep-alive connections after `RUSTFS_HTTP1_HEADER_READ_TIMEOUT` seconds (`DEFAULT_HTTP1_HEADER_READ_TIMEOUT`, 75). Reverse proxies pool and reuse upstream connections. If the proxy's upstream idle-keepalive window is longer than RustFS's timeout, the proxy can pick a connection RustFS has already FIN'd, write a request onto the dead socket, and the client sees:
```
```text
TimeoutError: socket hang up # ECONNRESET
AbortError: Request aborted
```
This is most visible on large `PutObject` uploads because:
This is most visible on large `PutObject` uploads: `PUT` is non-idempotent, so proxies will not transparently retry it, and a larger body keeps the connection in use longer, widening the race window, so small uploads on the same path often succeed.
- `PUT` is non-idempotent, so proxies will **not** transparently retry it; and
- a larger body keeps the connection in use longer, widening the race window,
so small uploads on the same path often succeed.
Fix by making the two windows agree (doing both is safest):
### Fix — make the two windows agree
Pick **either** side; doing both is safest:
- **RustFS side:** keep `RUSTFS_HTTP1_HEADER_READ_TIMEOUT` (default 75s) *above*
the proxy's upstream idle-keepalive. To harden slowloris protection on a
directly-exposed node instead, lower it — but then also lower the proxy
keepalive below it.
- **Proxy side:** lower the proxy's upstream idle-keepalive below RustFS's
timeout, or disable upstream keep-alive entirely.
1. RustFS side: keep `RUSTFS_HTTP1_HEADER_READ_TIMEOUT` above the proxy's upstream idle-keepalive. To harden slowloris protection on a directly exposed node instead, lower it, and then also lower the proxy keepalive below it.
2. Proxy side: lower the proxy's upstream idle-keepalive below RustFS's timeout, or disable upstream keep-alive entirely.
## Known-good Caddy configuration
@@ -127,47 +95,23 @@ location / {
## Cloudflare (orange-cloud) caveats
Cloudflare's proxy (orange cloud) may **buffer the entire request body** before
forwarding, and can rewrite requests to `Transfer-Encoding: chunked`, dropping
the client's `Content-Length`. Symptoms match this pattern exactly: tiny uploads
succeed, larger uploads fail with `socket hang up`.
- For large object writes, prefer **DNS-only (grey cloud)** for the S3 endpoint,
or a Cloudflare plan/tunnel configuration that does not buffer/re-chunk the
request body.
- Force `Accept-Encoding: identity` so nothing in the path negotiates
compression (see issues #609, #1492).
- Ensure `Content-Length` reaches RustFS; disable chunked re-encoding in tunnel
settings (see issue #934).
Cloudflare's proxy may buffer the entire request body before forwarding and can rewrite requests to `Transfer-Encoding: chunked`, dropping the client's `Content-Length`. The symptom is exactly the pattern above: tiny uploads succeed, larger uploads fail with `socket hang up`. For large object writes prefer DNS-only (grey cloud) for the S3 endpoint, or a plan/tunnel configuration that does not buffer or re-chunk the body. The `Accept-Encoding` and `Content-Length` rows in the issue table below are the Cloudflare-specific failures seen so far.
## Diagnosis checklist
Run each step and note where behavior diverges:
1. Bypass the proxy. Send the failing request to `http://<host>:9000` directly. Success confirms the fault is in the proxy/CDN path.
2. Bypass the CDN, keep the proxy. Point the proxy straight at the origin (Cloudflare grey cloud / direct DNS). If it now works, the CDN was buffering or re-chunking the body.
3. Check idle reuse. Intermittent failures that correlate with upload size are almost always the keep-alive mismatch. Lower the proxy keepalive (or disable it) and retry.
4. Check for a truncated body. If the upload hangs indefinitely rather than resetting, the proxy is forwarding a partial body and then going silent without closing the connection. RustFS bounds this wait with `RUSTFS_HTTP_REQUEST_BODY_READ_TIMEOUT` (`DEFAULT_HTTP_REQUEST_BODY_READ_TIMEOUT`, 300; `0` disables) and on timeout logs `put_object_body_read_stalled` with the received/expected byte counts.
5. Compare bytes. Confirm the proxy forwards exactly `Content-Length` body bytes with no compression or transformation.
6. Confirm signed headers survive. `Host` and `x-amz-*` must reach RustFS unchanged; a `SignatureDoesNotMatch` (rather than a hang) points here.
1. **Bypass the proxy.** Send the failing request to `http://<host>:9000`
directly. Success here confirms the fault is in the proxy/CDN path.
2. **Bypass the CDN, keep the proxy.** Point the proxy straight at the origin
(Cloudflare grey cloud / direct DNS). If it now works, the CDN was
buffering/re-chunking the body.
3. **Check idle reuse.** If failures are intermittent and correlate with upload
size, it is almost always the keep-alive mismatch above. Lower the proxy
keepalive (or disable it) and retry.
- If instead the upload **hangs indefinitely** (rather than resetting), the
proxy is likely forwarding a *partial* body and then going silent without
closing the connection. RustFS bounds this wait with
`RUSTFS_HTTP_REQUEST_BODY_READ_TIMEOUT` (default 300s; `0` disables) and, on
timeout, logs a `put_object_body_read_stalled` event with the
received/expected byte counts — grep the server log for it to confirm a
truncated-body forwarding problem.
4. **Compare bytes.** Confirm the proxy forwards exactly `Content-Length` body
bytes with no compression/transformation.
5. **Confirm signed headers survive.** `Host` and `x-amz-*` headers must reach
RustFS unchanged; a `SignatureDoesNotMatch` (rather than a hang) points here.
## Known failure signatures
## Related issues
- #3076 Large single-request PutObject fails behind Caddy (this document)
- #609 Bucket inaccessible via Cloudflare proxied DNS (`Accept-Encoding`)
- #1492 SigV4 `SignatureDoesNotMatch` on Cloudflare tunnel (`Accept-Encoding`)
- #934 Console fails behind Cloudflare tunnels (chunked / `Content-Length`)
- #1766 Large multipart upload fails through Nginx (`ETag` stripping)
| Symptom | Forwarding fault | Issue |
| --- | --- | --- |
| Large single-request PutObject fails behind Caddy | Upstream idle keep-alive longer than RustFS's timeout | #3076 |
| Bucket inaccessible via Cloudflare proxied DNS | `Accept-Encoding` negotiation / body transformation | #609 |
| SigV4 `SignatureDoesNotMatch` on Cloudflare tunnel | `Accept-Encoding` header rewritten | #1492 |
| Console fails behind Cloudflare tunnels | Chunked re-encoding drops `Content-Length` | #934 |
| Large multipart upload fails through Nginx | `ETag` stripped from responses | #1766 |
+35 -112
View File
@@ -1,153 +1,76 @@
# Restarting a multi-node RustFS cluster
How to restart nodes of an erasure-coded multi-node cluster without losing
availability, what to expect when several nodes are down at once (sequential
cold start), and how to read the degraded-mode signals. Written for the
failure pattern reported in rustfs/rustfs#4304.
**Use this when:** restarting or upgrading nodes of an erasure-coded multi-node cluster, bringing a cluster back after several nodes were down at once, or interpreting `503` degraded-mode responses during startup.
> Upgrading the binary or container image does not change the on-disk data
> format unless an explicitly enabled feature documents a version floor.
> Replacing the executable and restarting does not run a migration step on
> startup.
**Source of truth:** `crates/config/src/constants/health.rs` (`DEFAULT_STARTUP_READINESS_MAX_WAIT_SECS`), `rustfs/src/server/readiness.rs` (readiness responses and `Retry-After`), the `iam_bootstrap_retry_failed` log event.
> [!WARNING]
> The release that switches local SSE wrapped DEKs from the legacy
> `base64(nonce):base64(ciphertext)` representation to the versioned JSON
> envelope is a deliberate exception. Do not run that release together with
> an older RustFS version: older nodes cannot read objects written with the
> JSON envelope. Freeze every source of object mutation, including client
> writes and background lifecycle or replication work, upgrade every node,
> and then resume traffic. Downgrading or rolling back after new encrypted
> objects are written is not supported.
Upgrading the binary or container image does not change the on-disk data format unless an explicitly enabled feature documents a version floor. Replacing the executable and restarting does not run a migration step on startup.
> [!WARNING]
> `RUSTFS_DATA_MOVEMENT_PART_CHECKSUMS_WRITE` remains inactive unless
> `RUSTFS_DATA_MOVEMENT_PART_CHECKSUMS_FLEET_CONFIRMED` is also `true`. Enable
> both only after every node that can read or write object metadata supports
> the `part-checksums` sidecar and the fleet has adopted that version as its
> rollback floor. Leave either setting disabled throughout a mixed-version
> rolling upgrade. Once rebalance or decommission has migrated a legacy
> checksummed multipart object with both settings enabled, rolling back to an
> older build is not supported: older readers ignore the sidecar and can
> report an object checksum in place of the requested part checksum.
## Version floors that break mixed-version fleets
> [!WARNING]
> Writing pool metadata version 2 remains inactive unless both
> `RUSTFS_POOL_META_V2_WRITE=true` and
> `RUSTFS_POOL_META_V2_FLEET_CONFIRMED=true`. Leave either setting disabled
> until every node that can read or write `pool.bin` supports version 2. Once a node
> observes or writes version 2 it will not downgrade the file, and older
> binaries or rollback builds cannot read it. Unresolved decommission entries
> fail closed instead of being written in the version 1 format.
Each row is a feature whose activation makes older binaries unable to read what newer ones write. Keep every gate in its inactive state throughout a mixed-version rolling upgrade.
> [!WARNING]
> Pool metadata version 3 remains inactive on an existing cluster unless both
> `RUSTFS_POOL_META_V3_WRITE=true` and
> `RUSTFS_POOL_META_V3_FLEET_CONFIRMED=true`. V3 adds durable generations and a
> recoverable cross-pool commit protocol. Once committed, V1/V2-only binaries
> cannot rejoin. Follow [Pool metadata upgrade and recovery](pool-metadata-recovery.md)
> for the compatibility matrix and disk-replacement order.
| Feature gate | Activation | Consequence once active | Owning doc |
| --- | --- | --- | --- |
| Local SSE wrapped-DEK JSON envelope | The release that replaces the legacy `base64(nonce):base64(ciphertext)` representation with the versioned JSON envelope | Older nodes cannot read objects written with the JSON envelope. Freeze every source of object mutation (client writes, lifecycle, replication), upgrade every node, then resume traffic. Downgrading after new encrypted objects are written is not supported. | [compat-cleanup-register.md](../architecture/compat-cleanup-register.md) (`sse-local-dek-json-v1`), [minio-file-format-compat.md](../architecture/minio-file-format-compat.md) |
| Data-movement part checksums sidecar | `RUSTFS_DATA_MOVEMENT_PART_CHECKSUMS_WRITE=true` and `RUSTFS_DATA_MOVEMENT_PART_CHECKSUMS_FLEET_CONFIRMED=true` (inactive unless both) | Enable only after every node that reads or writes object metadata supports the `part-checksums` sidecar and the fleet has adopted that version as its rollback floor. Once rebalance or decommission has migrated a legacy checksummed multipart object with both enabled, rollback is not supported: older readers ignore the sidecar and can report an object checksum in place of the requested part checksum. | This page |
| Pool metadata version 2 | `RUSTFS_POOL_META_V2_WRITE=true` and `RUSTFS_POOL_META_V2_FLEET_CONFIRMED=true` (inactive unless both) | Once a node observes or writes version 2 it never downgrades `pool.bin`; older binaries and rollback builds cannot read it. Unresolved decommission entries fail closed instead of being written in the version 1 format. | [pool-metadata-recovery.md](pool-metadata-recovery.md) |
| Pool metadata version 3 | `RUSTFS_POOL_META_V3_WRITE=true` and `RUSTFS_POOL_META_V3_FLEET_CONFIRMED=true` (inactive on an existing cluster unless both) | Adds durable generations and a recoverable cross-pool commit protocol. Once committed, V1/V2-only binaries cannot rejoin. | [pool-metadata-recovery.md](pool-metadata-recovery.md) (compatibility matrix, disk-replacement order) |
## TL;DR
- **Rolling restart (no downtime):** restart **one node at a time**, and wait
for the restarted node to report `200` on `/health/ready` before touching
the next one. The remaining nodes keep serving traffic.
- **Sequential cold start (several nodes down):** nodes started before the
cluster has quorum come up in **degraded mode** — the process stays alive,
answers `503` with the blocking reason, and recovers **automatically** as
soon as enough peers are online. Do not restart-loop them; just keep
starting the remaining nodes.
- Rolling restart (no downtime): restart one node at a time and wait for the restarted node to report `200` on `/health/ready` before touching the next one. The remaining nodes keep serving traffic.
- Sequential cold start (several nodes down): nodes started before the cluster has quorum come up in degraded mode. The process stays alive, answers `503` with the blocking reason, and recovers automatically as soon as enough peers are online. Do not restart-loop them; keep starting the remaining nodes.
## Why a single node cannot serve alone
Erasure coding shards every object (including internal metadata such as IAM
users, groups, and policies under `.rustfs.sys`) across the drives of a set.
Reading an object back needs a **read quorum** of shards online. With the
drives of one set spread over several nodes, one node alone can never satisfy
the read quorum — this is a mathematical property of erasure coding, not a
bug. The cluster becomes readable once enough nodes are up (for internal
configuration objects, which are written with maximum parity, that is
typically about half the nodes of a set).
Erasure coding shards every object, including internal metadata such as IAM users, groups, and policies under `.rustfs.sys`, across the drives of a set. Reading an object back needs a read quorum of shards online. With the drives of one set spread over several nodes, one node alone can never satisfy the read quorum; this is a property of erasure coding, not a bug. The cluster becomes readable once enough nodes are up (for internal configuration objects, written with maximum parity, typically about half the nodes of a set).
Distributed locking similarly needs a majority of nodes' lock RPC endpoints.
The startup path no longer takes namespace locks while loading IAM
(rustfs/rustfs#4363), so IAM recovery depends only on the storage read
quorum.
Distributed locking similarly needs a majority of nodes' lock RPC endpoints. The startup path does not take namespace locks while loading IAM, so IAM recovery depends only on the storage read quorum.
## Rolling restart procedure
For each node, in any order, **one at a time**:
For each node, in any order, one at a time:
1. Restart the node (upgrade the binary/image first if this is an upgrade).
2. Wait until the node reports ready:
2. Wait until the node reports ready; a ready node returns `200` with `"ready": true` in the JSON body:
```bash
curl -fsS http://<node>:9000/health/ready
```
A ready node returns `200` with `"ready": true` in the JSON body.
3. Only then move on to the next node.
While one node is down, the rest of the cluster keeps quorum and serves all
traffic. If you take a second node down before the first is back, some
erasure sets may lose write or even read quorum and requests start failing —
this is the situation to avoid.
While one node is down, the rest of the cluster keeps quorum and serves all traffic. Taking a second node down before the first is back can cost some erasure sets their write or even read quorum; that is the situation to avoid.
## Sequential cold start (multiple nodes down)
When the whole cluster (or several nodes) went down — power loss, host
maintenance, crash-looping deployment — and nodes are brought back one at a
time:
When the whole cluster (or several nodes) went down and nodes are brought back one at a time:
1. **Early nodes come up degraded.** The process does not exit. S3 requests
receive `503 Service Unavailable` with a `Retry-After: 5` header, an
`x-rustfs-readiness-pending` header, and a body naming the blocking
dependency:
1. Early nodes come up degraded. The process does not exit. S3 requests receive `503 Service Unavailable` with a `Retry-After: 5` header, an `x-rustfs-readiness-pending` header, and a body naming the blocking dependency:
- `storage_quorum` — waiting for enough nodes/disks for the erasure read
quorum;
- `iam` — storage is up, IAM cache is still loading;
- `startup_finalization` — last startup steps are being published.
| Blocking dependency | Meaning |
| --- | --- |
| `storage_quorum` | Waiting for enough nodes/disks for the erasure read quorum. |
| `iam` | Storage is up; the IAM cache is still loading. |
| `startup_finalization` | Last startup steps are being published. |
2. **Logs say what the node waits for.** The IAM recovery loop retries with
backoff and logs `event="iam_bootstrap_retry_failed"` with an actionable
`hint` field (for example, "storage read quorum not met yet; waiting for
enough cluster nodes/disks to come online"). After repeated failures the
log level escalates from WARN to ERROR — this still does not kill the
process.
3. **Recovery is automatic.** As soon as enough peers are online for the
storage read quorum, the pending nodes finish IAM bootstrap on the next
retry and flip `/health/ready` to `200` on their own. No manual restart is
needed, and restarting them does not speed anything up.
4. **Check readiness detail while waiting.** `/health/ready` (and
`/minio/health/ready`) return per-dependency detail during degradation:
2. Logs say what the node waits for. The IAM recovery loop retries with backoff and logs `event="iam_bootstrap_retry_failed"` with an actionable `hint` field (for example, "storage read quorum not met yet; waiting for enough cluster nodes/disks to come online"). After repeated failures the level escalates from WARN to ERROR; this still does not kill the process.
3. Recovery is automatic. As soon as enough peers are online for the storage read quorum, the pending nodes finish IAM bootstrap on the next retry and flip `/health/ready` to `200` on their own. Restarting them does not speed anything up.
4. Check readiness detail while waiting. `/health/ready` (and `/minio/health/ready`) return per-dependency detail during degradation; the `details` object shows `storage` / `iam` / `lock` readiness and `degradedReasons` lists machine-readable causes such as `storage_quorum_unavailable` or `lock_quorum_unavailable`:
```bash
curl -s http://<node>:9000/health/ready | jq
```
The `details` object shows `storage` / `iam` / `lock` readiness, and
`degradedReasons` lists machine-readable causes such as
`storage_quorum_unavailable` or `lock_quorum_unavailable`.
## Tuning
- `RUSTFS_STARTUP_READINESS_MAX_WAIT_SECS` (default `120`): how long startup
waits for full readiness before continuing in degraded mode with background
recovery. Raising it delays the listener during genuinely slow starts;
lowering it surfaces degraded mode sooner. Recovery retries continue
regardless of this limit.
| Variable | Default | Effect |
| --- | --- | --- |
| `RUSTFS_STARTUP_READINESS_MAX_WAIT_SECS` | `120` (`DEFAULT_STARTUP_READINESS_MAX_WAIT_SECS`) | How long startup waits for full readiness before continuing in degraded mode with background recovery. Raising it delays the listener during genuinely slow starts; lowering it surfaces degraded mode sooner. Recovery retries continue regardless of this limit. |
## What is *not* normal
## What is not normal
- A node process **exiting** with a fatal IAM/lock error during startup
that fatal path was removed after v1.0.0-beta.5 (rustfs/rustfs#4304);
upgrade if you still see it.
- A node stuck degraded **after** the whole cluster is back: check network
reachability between nodes (peer RPC ports) and per-node clocks, then
inspect `degradedReasons` and the `hint` field of the IAM retry logs.
- A node shown offline in the console with no log output — tracked
separately, see rustfs/backlog#888.
- A node process exiting with a fatal IAM/lock error during startup. That fatal path was removed after v1.0.0-beta.5 (rustfs/rustfs#4304); upgrade if you still see it.
- A node stuck degraded after the whole cluster is back: check network reachability between nodes (peer RPC ports) and per-node clocks, then inspect `degradedReasons` and the `hint` field of the IAM retry logs.
- A node shown offline in the console with no log output is tracked separately (rustfs/backlog#888).
@@ -1,568 +0,0 @@
# RustFS heal & scanner vs MinIO — comprehensive parity analysis (v2, 2026-08-16)
> English | [中文版](rustfs-heal-scanner-vs-minio-comprehensive-analysis-2026-08-16_zh.md)
- Date: 2026-08-16 (based on that day's `main` code; audit HEAD ≈ `a118d7e4f`)
- Scope: `crates/heal` (src 19,560 lines + tests 2,274 lines), `crates/scanner` (src ~26,000 lines + tests), `crates/data-usage`, the heal/heal_walk/bitrot_self_verify and config parts of `crates/ecstore`, `crates/heal-contracts/src/heal_channel.rs`, `crates/madmin` (heal/scanner wire types), `rustfs/src` (startup wiring, admin handlers, cluster RPC)
- Parity baseline: minio/minio master (HEAD `7aac2a2c5b`; the repo has entered maintenance mode with master frozen, i.e. its final state)
- Method: four parallel audit tracks (heal crate / scanner crate / ecstore integration layer / MinIO source study), with key conclusions verified by hand one by one (points marked "verified first-hand" below were checked against the source directly)
- This document supersedes `docs/rustfs-heal-scanner-vs-minio-parity-assessment.md` (2026-06-15, v1). Since v1 there have been more than 80 heal/scanner commits (the full automatic drive-replacement healing chain, the resume state machine, making usage convergence authoritative, cluster-level heal coordination, ILM restore semantics, etc.), so v1's feature inventory and gap judgments are comprehensively outdated; v1 conclusions such as "bloom filter missing" were verified this round to be **misjudgments** (see §5.4).
---
## 0. Conclusion summary
1. **Overall verdict: the core functional chains of heal and scanner are complete.** Object-level heal (quorum arbitration + ETag fallback + bitrot Deep verification + dangling handling), erasure set deep scans (per-set disk-walk union enumeration), per-version resumable scans (schema'd persistence layer + CAS atomic publish + crash-window backfill), automatic drive-replacement healing (readiness validation + identity fencing + durable intent + completion proof), the scanner cycle loop (leader lock + persisted leader-epoch fence), data usage statistics (bucket-level/cluster-level, primary + backup + observed snapshots, epoch/cycle anti-rollback), the full ILM action set (expiry/transition/noncurrent/free-version/delete-marker cleanup), and the admin Start/Query/Cancel protocol (clientToken semantics aligned with madmin) — all of these are implemented and carry regression tests. There are **no empty implementations / early-return stubs** inside the two crates; every exceptional path has logs + metrics + error semantics.
2. **The main gaps concentrate on "entry points and the observability surface", not on the repair algorithms themselves**: the MRF/ECDecode/Metadata task executors are implemented but have no production trigger entry (`HealEvent` is entirely unwired); `CheckAbandonedParts` is `NotImplemented` at all three ecstore layers; the heal/scanner trace channels are missing; scanner excess S3 events are missing; madmin client methods are missing (only wire types exist); heal byte-level progress/ETA is not implemented.
3. **Important corrections to the v1 understanding**: the bloom filter has been **removed** from current MinIO master (`.bloomcycle.bin` stores only a cycle count), so RustFS's current state matches MinIO; the MinIO scanner is likewise a **cluster-level leader singleton**, and RustFS's leader.lock model is the same shape as MinIO's; RustFS's ETag majority-fallback arbitration is already implemented (`crates/ecstore/src/set_disk/ops/heal.rs:525-567,679`, verified first-hand) — the arbitration gap v1 worried about does not exist.
4. **RustFS exceeds MinIO in several places**: the remote_scanner RPC protocol (remote peers scan locally instead of the leader reading remote drives across the network), the persisted leader-epoch CAS fence, cycle budgets and per-set/per-disk concurrency gates, the pending-heal ledger, the durable replacement intent + completion proof state machine, foreground pressure gating (mainline throttle), and the cluster heal control coordinator + envelope replay protection.
5. Gap severity tally: 8 P1 items (behavioral/operational alignment gaps), 9 P2 items (completeness), 3 P3 items (cleanup/low risk), and 7 items of "not pursuing parity by design". Full list in §6.
---
## 1. Architecture overview
### 1.1 RustFS's three-layer architecture
RustFS splits the heal/scanner functionality that MinIO keeps inside the `cmd/` monolith into three layers plus two standalone crates:
| Layer | Location | Responsibilities |
|---|---|---|
| Primitives layer | `crates/ecstore/src/set_disk/ops/heal.rs` (~3,240 lines), `ops/heal_walk.rs`, `ops/bitrot_self_verify.rs`; upper wrappers `store/heal.rs`, `store/heal_walk.rs`, `core/sets.rs` | Object/bucket/format/replacement-drive format repair, disk-walk union enumeration, write-path bitrot self-verification; the `rustfs_storage_api::HealOperations` contract is implemented by `SetDisks`/`Sets`/`ECStore` (`crates/storage-api/src/object.rs:503-519`) |
| heal runtime | `crates/heal` | Process-level HealManager (priority queue/scheduler/auto disk scanner/resumable resume), HealChannelProcessor (consumes the global heal channel), drive-replacement recovery state machine |
| scanner runtime | `crates/scanner` | Data usage scanning, ILM evaluation and enqueueing, heal candidate production, replication usage statistics, remote scanner RPC |
| Shared protocol | `crates/heal-contracts/src/heal_channel.rs` (~776 lines) | Start/Query/Cancel command channel, `HealOpts`/`HealScanMode`/`HealRequestSource`/`HealAdmission*` shared types, `HealResultItem` (madmin) |
| Shared data | `crates/data-usage` | `DataUsageEntry/Info`, histograms, `hash_path`; produced by the scanner, consumed by ecstore/admin |
Startup chain (wiring verified first-hand):
1. `rustfs/src/startup_services.rs:93``init_background_service_runtime(store)`.
2. `rustfs/src/startup_background.rs:41-81`: create the global heal service cancel token; read `RUSTFS_SCANNER_ENABLED` (alias `RUSTFS_ENABLE_SCANNER`, default true) and `RUSTFS_HEAL_ENABLED` (alias `RUSTFS_ENABLE_HEAL`, default true); **the heal manager is initialized whenever either heal or scanner is enabled** (heal candidates produced by the scanner need a consumer; with both off, the heal channel is not initialized and `send_heal_request` reports "Heal channel not initialized").
3. `crates/heal/src/lib.rs:142-216`: atomic initialization inside an owned task (a caller cancel cannot leave a half-initialized manager behind, `lib.rs:123-131`; `GLOBAL_HEAL_RUNTIME_INIT` mutex single-flight) → `HealManager::start()``rustfs_common::heal_channel::init_heal_channels()` → spawn `HealChannelProcessor::start_with_receipts`.
4. `crates/heal/src/heal/manager.rs:1301-1356` `HealManager::start`: `start_scheduler()` (`manager.rs:2394-2461`, interval default 10s + `Notify` event-driven wakeup) → `process_unclean_shutdown()` (`manager.rs:1362-1695`) → when `enable_auto_heal` (default true), `start_auto_disk_scanner()` (`manager.rs:2464-2999`).
5. After the server is ready, `rustfs/src/startup_lifecycle.rs:150-152`: when `enable_scanner`, `init_data_scanner(token, store)` (`crates/scanner/src/scanner.rs:1293-1372`).
6. Graceful shutdown: `rustfs/src/startup_shutdown.rs:308` `shutdown_ahm_services()` (cancel token); `:414` `clear_unclean_shutdown_markers()`.
### 1.2 MinIO's corresponding structure (final master state)
| MinIO file | Responsibilities |
|---|---|
| `cmd/admin-heal-ops.go` | Manual admin heal sequence (healSequence, clientToken/forceStart/forceStop) |
| `cmd/global-heal.go` | Resident background heal queue (newBgHealSequence, token fixed `0000-…`, never ends) + `healErasureSet` (full-object heal per set) |
| `cmd/background-heal-ops.go` | healRoutine worker pool (`_MINIO_HEAL_WORKERS`, default GOMAXPROCS/2) consuming healTask |
| `cmd/mrf.go` | MRF (Most Recent Fail) queue (capacity 100,000), persisted at process exit to `.minio.sys/buckets/.heal/mrf/list.bin` with startup replay |
| `cmd/background-newdisks-heal-ops.go` | Automatic resync for new/replaced drives (monitorLocalDisksAndHeal 10s polling + healFreshDisk + healingTracker) |
| `cmd/erasure-healing.go` / `erasure-healing-common.go` | Object-level heal core (~800 lines), listAndHeal |
| `cmd/data-scanner.go` | Scanner loop (globalLeaderLock cluster singleton) + folderScanner + applyActions |
| `cmd/erasure.go` (nsScanner) / `erasure-server-pool.go` | NSScanner three-layer structure |
| `cmd/bucket-lifecycle.go` | ILM executor (expiry/transition worker pools) |
| `cmd/xl-storage.go` | DiskInfo.Healing, CheckParts/VerifyFile, CleanAbandonedData, RenameData healing branch |
| `cmd/prepare-storage.go` | waitForFormatErasure new-drive startup handshake |
### 1.3 Architecture-level differences (design trade-offs, not defects)
1. **heal queue model**: MinIO funnels every heal (scanner sampling/MRF/admin/new-disk resync) into a single channel + a fixed worker pool (new-disk resync additionally has a per-drive worker pool); RustFS is a multi-policy scheduler built from a priority heap + dedup-merge + capacity-tiered dropping + per-set bulkhead + foreground pressure gating (`manager.rs:3003-3420`). RustFS is more expressive, at the cost of an observability question around "duplicate requests being merged" (already pointed out in v1; the current `HealAdmissionReceipt` canonical task_id + alias mechanism answers it, `manager.rs:1759-1846`).
2. **scanner remote-drive access**: the MinIO leader transparently reads and writes remote-node drives through the disk abstraction layer; the RustFS leader pushes scan execution down to the remote peer to run locally via the remote_scanner RPC (`crates/scanner/src/remote_scanner.rs`), with only results and progress heartbeats sent back. Both are cluster single-leader. RustFS's approach saves the leader↔remote metadata read amplification, at the cost of maintaining a separate RPC protocol (HMAC per-frame authentication, session replay cache, fence re-validation, `remote_scanner.rs:52-61,405-496,1024-1065`).
3. **heal state persistence**: MinIO uses a single file `.healing.bin` (msgp healingTracker, reset whenever the diskID mismatches); RustFS uses a schema'd multi-file layout (resume/checkpoint/intent/seal/proof, each CAS-published, `resume.rs:38-61`), with the crash window explicitly backfilled (`erasure_healer.rs:389-402`, `resume.rs:1027-1057`).
4. **write-path self-protection**: MinIO relies on background heal to converge after writes; RustFS, after the commit rename in PutObject/CompleteMultipartUpload, actively checks `convergence.needs_heal()` and immediately enqueues an object heal (`set_disk/ops/object.rs:2291-2306`, `ops/multipart.rs:2574-2589`), and additionally has read repair (`io_primitives.rs:1040-1160`).
---
## 2. Heal implemented-feature panorama
### 2.1 Task types (`HealType`, `crates/heal/src/heal/task.rs:85-111`)
| Type | Semantics | Executor | Production trigger |
|---|---|---|---|
| `Cluster` | all buckets healed in turn (structure + optional recursive objects), in-batch retry ≤3 | `heal_cluster` task.rs:1420-1490 | channel: empty bucket means Cluster (channel.rs:576-577) |
| `Object{bucket,object,version_id}` | single object/version; when absent, rebuild per `recreate_missing` or error out | `heal_object` task.rs:855-1146 | admin, scanner, read-repair, write-path convergence, add_partial |
| `Bucket{bucket}` | bucket metadata/structure; `recursive` additionally walks all object versions | `heal_bucket` task.rs:1284-1418 + `heal_bucket_objects` task.rs:1508-1698 | admin (POST /v3/heal/{bucket}), scanner `build_bucket_heal_request` |
| `Prefix{bucket,prefix}` | recursive by prefix | `heal_prefix` task.rs:1492-1506 | channel: `recursive && prefix` non-empty (channel.rs:578-585) |
| `ErasureSet{buckets,set_disk_id}` | format repair + healing marker + per-bucket preprocessing + resumable per-version deep scan | `heal_erasure_set` task.rs:2158-2642 | admin (pool/set params), auto disk scanner, unclean shutdown, renew_disk, durable replacement recovery |
| `Metadata{bucket,object}` | metadata only (Deep, does not rebuild data) | `heal_metadata` task.rs:1700-1859 | **no production trigger** (§6 HS-01) |
| `MRF{meta_path}` | failure-path-driven Deep repair (recursive+update_parity) | `heal_mrf` task.rs:1861-1992 | **no production trigger** (only `HealEvent` can generate it, unwired) |
| `ECDecode{bucket,object,version_id}` | EC decode rebuild (Deep+recreate+update_parity), Urgent priority | `heal_ec_decode` task.rs:1994-2156 | **no production trigger** (only `HealEvent` can generate it, unwired) |
Priorities `Low/Normal/High/Urgent` (task.rs:168-179); state machine `Pending/Running/Retrying/Completed/Failed/Cancelled/Timeout` (task.rs:225-241).
### 2.2 Trigger-path panorama (beyond admin)
| Channel | source | Priority | Evidence |
|---|---|---|---|
| Scanner periodic sampling (1/1024, `RUSTFS_HEAL_OBJECT_SELECT_PROB`) | Scanner | Low | `scanner_folder.rs:2117-2136`, `:1150`; `remove_corrupted=HEAL_DELETE_DANGLING(true)`, `recreate_missing=false` (`common/heal_channel.rs:24`, `scanner_folder.rs:510-511`) |
| Scanner metadata corruption (get_size failure classified HealMetadata) | Scanner | High | `scanner_folder.rs:2147-2208`, `:1244-1260` |
| Scanner abandoned children (present in cache, absent on disk, list_path_raw quorum verification) | Scanner | High (bucket-level + object-level) | `scanner_folder.rs:2528-2792` |
| Scanner pending-heal ledger retry (persisted after rejection by a full heal channel, ≤128 per bucket per round, 10k cap) | Scanner | original priority | `scanner_folder.rs:1721-1763`, `:99-100` |
| auto disk scanner (unformatted drive confirmed via replacement_readiness / `runtime_state=="returning"` drive / durable-intent re-entry) | AutoHeal | Low | `manager.rs:2464-2999` |
| unclean shutdown recovery (startup reads the `unclean-shutdown` marker → ErasureSet heal for all local sets) | AutoHeal | Low | `manager.rs:1362-1695` |
| write-path convergence (after PutObject/CompleteMultipartUpload, `convergence.needs_heal()`) | Internal | Normal | `set_disk/ops/object.rs:2291-2306`, `ops/multipart.rs:2574-2589` |
| partial-object heal (add_partial) | Internal | Normal | `set_disk/ops/object.rs:5808-5825` |
| stale data-directory cleanup leftover enqueue | Internal | Normal | `set_disk/core/io_primitives.rs:3880-3907` |
| read repair (metadata_read_error / missing_shards / decode_error, TTL dedup cache) | ReadRepair | Low | `set_disk/read.rs:407,995,1079``submit_read_repair_heal` (`io_primitives.rs:1105-1160`), `recreate_missing=true` |
| drive reconnect hits UnformattedDisk → send_heal_disk | AutoHeal | Normal | `set_disk/ops/locking.rs:339-347` |
| Admin API (incl. cluster coordinator routing) | Admin | High | `rustfs/src/admin/handlers/heal.rs:174-212`, `:771-930` |
| cluster RPC heal (peer invocation) | — | — | `rustfs/src/storage/rpc/node_service/heal.rs`, `ecstore/src/cluster/rpc/peer_s3_client.rs:296,1209` |
Note: MinIO's MRF channel (read-path immediate delivery on missing/corrupt parts + queue persistence + shutdown replay, `cmd/mrf.go`, `erasure-object.go:395-410,800-812`) is **partially replaced** in RustFS by read-repair + write-path convergence; the three executors `HealType::MRF`/`ECDecode`/`Metadata` have no production entry (see §6 HS-01 for details).
### 2.3 Object-level heal semantics (ecstore `set_disk/ops/heal.rs`)
Flow (`heal_object_with_explicit_version_regen` from :426):
1. Take the object write lock (unless `no_lock`); an `object` ending with `/` goes through object-directory heal (`heal_object_dir_locked` :1587-1717: dangling determination + `remove` deletion + missing-volume rebuild).
2. `read_all_fileinfo` reads xl.meta from all disks; all-not-found is treated as already deleted and returns.
3. **quorum arbitration + ETag fallback** (verified first-hand): `list_online_disks` treats the mod-time quorum as authoritative; when quorum fails it falls back to ETag majority arbitration (`:525-567` `filter_by_etag`/`quorum_etag`); `pick_valid_fileinfo` picks the canonical metadata; the cannotHeal determination for "number of bad-meta disks > parity" is waived when the ETag agrees across all disks (`:679`). Matches MinIO's dual arbitration in `filterDisksByETag`.
4. `disks_with_all_parts` (:562-572) validates parts per `scan_mode`: **Normal only stats (CheckParts semantics), Deep does full bitrot verification (VerifyFile semantics)**; when a Normal scan detects `FileCorrupt` it automatically escalates to Deep and retries once (`:2022-2031`, same shape as MinIO erasure-healing.go:1101-1106); a no-parity object (EC:0) with a bitrot failure is judged unrecoverable (`:700-726`).
5. `should_heal_object_on_disk` (:606-650) classifies each disk as missing/corrupt/offline/outdated → rebuild: per-part bitrot reader/writer (using per-part checksum + algorithm), write into a temporary volume then rename to commit (`HEAL_RENAME_INCOMPLETE` retry semantics :24); dangling-deletion safety check `dangling_delete_safety` (:1488); **orphan data-directory reclamation `reclaim_orphan_data_dirs_best_effort` (:1428)** — this part covers the main scenarios of MinIO's `CleanAbandonedData` (but there is no standalone `CheckAbandonedParts` API, see §6 HS-02).
6. Versioned objects: enumerate "every version" (`storage.rs:1494-1530`); the delete-marker path is decided by `latest_meta.deleted` (`storage.rs:262-277` comment); regression tests `tests/heal_b5_versioned_regression_test.rs:282,334`.
7. Explicit-version rebuild `try_regenerate_explicit_version_meta` (:1318); cleanup of local leftovers of transitioned objects.
8. The write path additionally has shard-level bitrot self-verification `verify_written_bitrot_shards` (`ops/bitrot_self_verify.rs:45-129`, HighwayHash256S, verifying freshly written shards right before the final rename, serving the EC:0 no-parity case) — **note this is not background bitrot patrol**; background patrol is carried by scanner bitrot_cycle-driven Deep heal.
heal-crate-side wrapper (`task.rs:855-1146`): existence check (transient errors become `TransientSkip` to avoid false failures :551-569); scanner synthetic-directory normalization (:1148-1180); `recreate_missing` rebuild (:1183-1282); data-usage-cache object-lock timeout exemption (:571-653); not-found → treated_as_deleted success (:1012-1029); results `HealResultItem` keep at most 1024 entries + truncated flag (:50,845-852).
Recursive walk (`heal_bucket_objects` task.rs:1508-1698): paginated enumeration of all versions including delete markers, transient-error exponential-backoff retry ≤3 (2^n + jitter :620-627), failure-sample log truncation ≤5 entries, aggregated `BatchHealFailure`.
### 2.4 erasure set heal and resumable scans
`heal_erasure_set` (task.rs:2158-2642) runs in four phases (4-step progress tracking):
1. **Replacement intent and recovery-drive selection** (AutoHeal only + non-empty heal_endpoints): reuse the drive holding the durable intent / exclude the target endpoints and pick surviving drives; already-completed generations get an idempotent CleanupPending wrap-up.
2. **Format repair**: `heal_replacement_format(dry_run, pool, set, targets)` (`storage.rs:1372-1384`, trait default fail-closed); per-target-drive results must all be ok (`erasure_healer.rs:97-102`) + identity-fence re-check (task.rs:2410-2420).
3. **healing marker**: write an owner CAS marker `{set_disk_id}:{task_id}` to the target drive (`mod.rs:80-229`, CAS + rollback + unique concurrent owner), which makes `DiskInfo.healing` true (assignment chain verified first-hand `set_disk/mod.rs:4988`).
4. **Per-bucket preprocessing + resumable deep scan**: `ErasureSetHealer::heal_erasure_set` (`erasure_healer.rs:242-278`).
`ErasureSetHealer` scan details (benchmarked against MinIO `healErasureSet`; the `heal_walk.rs:15-23` module comment explicitly cites MinIO `global-heal.go`'s listPathRaw + objQuorum=1 + mergeXLV2Versions):
- **Enumerator choice (backlog#920)**: Deep or AutoHeal → per-set **disk-walk union enumeration** `list_versions_for_heal_page_disk_walk` ("exists on any drive" means sub-quorum reconstructible; `storage.rs:1559-1644`, page bounds 1,000 objects/10,000 versions, `dw1:` cursor); ordinary requests go through read-quorum `list_object_versions`.
- **Resume cursor**: the authoritative cursor is an opaque continuation token (`v1:` = marker JSON, `dw1:` = disk-walk key; the two namespaces are mutually exclusive against misreads, `storage.rs:81-260`); after each completed page, persist the cursor first, then clear the dedup set (`erasure_healer.rs:922-927`).
- **In-page concurrency**: FuturesUnordered + Semaphore, default `RUSTFS_HEAL_PAGE_OBJECT_CONCURRENCY=8`, Deep/AutoHeal forces 1 (`erasure_healer.rs:105-142`).
- **per-version dedup**: `compose_key` length-prefix injection encoding (`resume.rs:281-288`).
- **Error classification**: truly absent (FileNotFound etc.) → Absent (counted as success); infrastructure-transient (quorum/DiskNotFound/SlowDown etc.) → Transient (counted as skipped); everything else Failed (`erasure_healer.rs:148-182`; the comment cites backlog#856/#799 B7: offline drives must not be recorded healed/absent).
- **Loop protection**: abort when an empty page is truncated or the page-tail version identity does not advance (:933-949).
- **Completion determination**: if any of failed/skipped/failed_buckets is >0, do not mark complete; `schedule_retry()` resets both the resume and checkpoint layers (:561-626; backlog#855/B6/#1033: a skip round must not be marked complete).
- **Replacement-drive commit proof**: physical read-back on the target endpoints `replacement_targets_have_version` (`ops/heal.rs:340-412`); unconfirmed → transient skip.
### 2.5 Automatic drive-replacement healing (replacement recovery)
- **Identification** (`replacement_readiness.rs:25-73`): `replacement_mount_lease_root()` exists, canonicalize succeeds, is a mount point, the physical device id is non-empty, disjoint from the root device, and shares no physical device with sibling drives (Linux uses /proc/self/mountinfo mount-id+dev+ino). The non-root mount check has a regression test (`manager.rs:3549`).
- **State machine** (`resume.rs:63-73`): `Intent → Rebuilding → (write proof) Verified → CleanupPending → cleanup`; `Abandoned` is a terminal state; state transitions write the persistence layer first, then mutate (`save_state_strict`).
- **Persistence** (`resume.rs:38-61`, schema ResumeState=5/Checkpoint=5/proof=1): `{task_id}_ahm_resume_state.json`, `_ahm_checkpoint.json`, and intent/seal/completion_proof under the `buckets/ahm-replacement/` namespace; torn write + no seal is recognizable and rebuilt atomically (:1316-1338); CAS publish, refuses to overwrite a concurrently valid proof (:1512-1585).
- **Recovery**: both unclean shutdown and the periodic scan recover unfinished/pending-cleanup replacement generations from surviving drives (`manager.rs:1435-1640,2663-2815`); multi-generation conflict / validation failure → freeze that set (`replacement_recovery_blocked_sets`, `manager.rs:69-87,2782-2815`).
- **External snapshot**: `current_replacement_recovery_snapshot` (`lib.rs:262-333`) merges local surviving-drive records; conflict → Unknown / non-definitive; admin `GET /v4/heal/replacement-recovery`.
### 2.6 Scheduler (manager.rs)
- Priority heap + FIFO within the same priority (:148-191,330-347); dedup key per type (:469-506); enqueue three-state dedup active→queued→retrying (:1759-1785); duplicates default to Merged and return the canonical task_id (`HealAdmissionReceipt`, :1821-1846) + client token alias (:1219-1246).
- Capacity: when the queue is full, best-effort sources (Scanner/AutoHeal/ReadRepair) or low-priority items get Dropped(QueueFull); Admin/Internal may evict queued lower-priority items (`push_displacing_lower_priority` :353-396); 80%/95% tiered pressure handling (:885-909).
- Concurrency: global `max_concurrent_heals` (default 4) + per-set bulkhead `max_concurrent_per_set` (default 1) (:3040-3073,3434-3447).
- Foreground pressure gating, mainline throttle: delay best-effort tasks when foreground read/write permit utilization is ≥80% (:919-1009,2999-3020).
- Timeout: task-level aggregate timeout (default 300s), remaining budget preserved across retries (task.rs:444-451, PR #6101).
- Recoverable retry: `is_recoverable_heal()` (error.rs:83-136) ≤3 attempts, 2^n backoff capped at 30s; retries hold ownership inside a standalone backoff task (:3235-3382).
- Completion states are retained for 10 minutes for querying (:42).
### 2.7 Admin API and cluster coordination
- Routes (`rustfs/src/admin/handlers/heal.rs:174-212`): `POST /rustfs/admin/v3/heal/`, `/heal/{bucket}`, `/heal/{bucket}/{prefix}` (the same POST distinguishes start/query/cancel by the query `clientToken/forceStart/forceStop`, aligned with mc admin heal semantics); `POST /v3/background-heal/status`; `GET /v4/heal/replacement-recovery`. Permission `HealAdminAction` (route_policy.rs:334-341).
- Cluster coordination (heal.rs:771-930 + `node_service.rs:514-606`): `heal_topology_fingerprint` + deterministic-by-topology coordinator-node selection + coordinator epoch; envelope validation + SHA256 digest replay protection; when the coordinator is not local, go through peer gRPC `heal_control`; `probe_heal_control` capability probe (rolling-upgrade scenario).
- Request: the body is `HealOpts` (`recursive/dryRun/remove/recreate/scanMode(0/1/2)/updateParity/nolock/pool/set`, serde camelCase, fields aligned with madmin.HealOpts); a root heal start requires `recursive=true` or a `pool+set` pair; body cap 1MB.
- Response: `HealStartSuccess{clientToken, clientAddress, startTime}`; `HealTaskStatus{summary, detail, startTime, settings, items, truncated, progress}` (summary ∈ running/finished/stopped/notFound); `BackgroundHealStatus` (bitrot start time/cycle/current mode + `disabled/uninitialized/idle/active/degraded` states — an unreachable peer is explicitly degraded rather than impersonating idle, issue #5850) + `healOperations` as a priority×source matrix + cluster progress.
- `HealResultItem`/`HealDriveInfo`/`HealItemType`/DriveState enums are JSON-compatible with madmin (`crates/madmin/src/heal_commands.rs:19-65`).
- A status payload over 8MiB is truncated by halving (channel.rs:37,73-104); path-token validation (wrong token rejected; an empty path matches Cluster only).
### 2.8 heal metrics and logs
Metrics: `rustfs_heal_admission_total{source,result,reason,context}`, `rustfs_heal_task_start_total`, `rustfs_heal_task_running{type,set}`, `rustfs_heal_queue_delay_seconds`, `rustfs_heal_scheduler_skip_total`, `rustfs_heal_mainline_throttle_total`, `rustfs_heal_page_concurrency_current{set}`, `rustfs_heal_candidate_enqueue/merge/drop/priority_reject_total`, `rustfs_heal_read_repair_dedup_total{reason}`, etc. All logs are structured event style (PR #5720); per-object logs are demoted to prevent storms (`demote_to_debug_when!`, #5716/#5719/#5727).
---
## 3. Scanner implemented-feature panorama
### 3.1 Loop, leader, immediate triggering
- **Cluster single leader**: distributed ns write lock `leader.lock` (`scanner.rs:3156-3207`, timeout default 5s) + **persisted leader-epoch CAS fence**: the leader writes (cycle, leader_epoch) encoded as `RSCYC001` into `.bloomcycle.bin` using an ETag precondition (`scanner.rs:118,1850-1861,2177-2334`); usage snapshots additionally carry an epoch fence (:2087-2153). Lock lost → cancel the current cycle, converging within 30s (:108-111,2623-2642).
- One round executes immediately after the lock is acquired; cycle = `RUSTFS_SCANNER_CYCLE` > config cycle > start_delay > deployment default > speed tier (±10% jitter, floor 1s).
- **clean-idle exponential backoff**: consecutive fully-clean idle intervals double (capped at 24h; bitrot-cycle compression cap; disabled when a bucket has active lifecycle/replication rules, :383-456,1382-1512).
- **superseded/deferred backoff**: exponential backoff from 5s capped at 30min (:105-106,3432-3438); maintenance probing failures get an independent backoff (:459-505).
- **Immediate wakeup**: ① dirty-usage fast path — write-path put/delete/multipart/bucket operations call `record_dirty_usage_bucket` (`scanner_io.rs:222-235`; call sites include `rustfs/src/app/object_usecase.rs:6221`), bump the generation and Notify-wake the leader; dirty buckets are queued first (`scanner_io.rs:462-488`); ② maintenance-config changes (lifecycle/replication settings call `record_scanner_maintenance_change`); ③ runtime-config hot updates generation+Notify; ④ cluster activity snapshot changes.
- **Cluster coordination**: `probe_scanner_activity` gathers this node's and peers' `ScannerNodeActivity` (instance_id/namespace_generation/maintenance_generation/protocol_version/topology_digest/data_movement_active/dirty usage); the topology digest covers pools/sets/drives URLs; a mismatched protocol version refuses to share the cache lock (`scanner.rs:970-1068`); **cycles are deferred during data movement (rebalance/decommission)** (`scanner_io.rs:2226-2374`); at cycle end, per-peer RPC confirms the dirty-usage ack (`scanner.rs:2925-2952`).
### 3.2 Traversal model
- The main traversal is a **full directory walk** (tokio::fs::read_dir recursion, `scanner_folder.rs:1915-2234`), not via metacache; metacache/`list_path_raw` is used only for the abandoned-children cross-drive verification (:2528-2792).
- Three-level concurrency: leader → per-set (semaphore default 4) → per-disk bucket scans (default 4) → single-drive recursion; a cache lock per bucket per set `.scanner-cycle.lock.pool-N.set-M` (losing the lock cancels that bucket's scan; lock contention re-queues); single-scan admission per drive (local drives also go through the semaphore, `scanner_io.rs:3246-3274`).
- Bucket ordering: after shuffle, re-ordered as dirty → uncached → cached (`scanner_io.rs:2947-2949,462-488`); entries within a directory sorted by name + resume-hint rotation (`scanner_folder.rs:333-359`).
- **Resumable scanning**: `DataUsageScanCheckpoint{version,resume_after,reason}` persisted in the cache info (`data_usage_define.rs:68,293-307`); written on budget exhaustion/cancel; resumption has Used/Stale/NoHint metrics; the resume unit is a directory (no cross-cycle object-level pagination).
- Erasure semantics: finding `xl.meta` marks an object boundary with no descent; at most 64 UUID data-dir candidate entries probed; data without metadata → record failed + high-priority heal; symlink directories ignored / cycles skipped.
- Cooperative yielding: `yield_now` every N objects (default 128).
### 3.3 Large-bucket skip strategy (benchmarked against MinIO compaction)
1. Cache-currency reuse: if the bucket and scan plan are unchanged (name/source/snapshot_complete/plan digest/next_cycle/leader_epoch/cache_key_format all match), the whole bucket is skipped (`scanner_io.rs:1062-1109`).
2. compacted-directory 16-cycle rotation window: rescan only when `hash mod (next_cycle, 16)` hits, otherwise copy from the old cache (`scanner_folder.rs:74,2429-2442`).
3. compaction thresholds: children <500 or pure-object leaves compress into a single entry; subfolders ≥2500 (root 10000) pre-compressed; children ≥10000 reduced (:75-78,2314-2340,2846-2887).
4. failed-object TTL skip: 86400s / at most 10,000 entries (:88-91,1354-1381).
Compared with MinIO master: MinIO's skip strategy is likewise hash-mod-16 cycles + a compaction threshold tree (500/10000/2500), and the **bloom filter has been removed from master**. RustFS's constants and structure share the same origin as MinIO's current state (MinIO does not adopt cross-drive dirty-generation prioritization; RustFS additionally has two more skip layers — plan digest and cache-currency validation).
### 3.4 data usage statistics
- Dimensions: per-directory entry (size/objects/versions/delete_markers/size histogram/version histogram/replication stats/failed_objects/per-tier stats/children/compacted, `data-usage/src/data_usage.rs:661-679`); per-object SizeSummary (incl. per-ARN replication-target stats and tier stats; tier classification: fully transitioned counts toward its tier, otherwise by storage class; free versions not counted); bucket-level `BucketUsageInfo`; cluster-level `DataUsageInfo` (incl. scanner_cycle/scanner_epoch fence + usage_snapshot_complete).
- Storage: per bucket per set `{bucket}/.usage-cache.bin` (primary + `.bkp` backup + CAS retry); the authoritative cluster snapshot `buckets/data-usage/data-usage.json` (`.bkp` synced every 10 cycles, legacy path compatible); stale snapshots rejected on write (triple epoch/cycle/last_update determination); observation snapshots superseded by a race are stored separately as `data-usage-observed.json`.
- Consumption: `replace_bucket_usage_memory_from_info` refreshes bucket-usage memory + two-level cache invalidation (`scanner.rs:4142-4152`) → bucket stats/quota/admin account_info/system; the write path overlays memory in real time; at startup, reading the snapshot detects a cold cache and skips startup delay.
- Incomplete multipart uploads are not counted (consistent with MinIO, which also does not scan the multipart bucket).
### 3.5 ILM integration
- Per object `ScannerItem::apply_actions` (`scanner_folder.rs:747-1032`): `Evaluator::new(lifecycle).with_lock_retention(...).with_replication_config(...).eval()` batch evaluation.
- Implemented actions (the full IlmAction set, `scanner-contracts/src/metrics.rs:34-45`): expiry deletes (Delete/DeleteRestored/DeleteRestoredVersion), all-versions deletes (DeleteAllVersions/DelMarkerDeleteAllVersions, stop further versions after handling), transition (Transition/TransitionVersion, tier list read at runtime), noncurrent batches (DeleteVersionAction → `enqueue_by_newer_noncurrent`), free-version cleanup (`enqueue_free_version`), object-lock retention constraints. **A one-to-one mapping onto MinIO's 9 ILM actions.**
- Execution model: the scanner is the "discover and enqueue" role (the expiry/transition queues live in ecstore `bucket_lifecycle_ops.rs`); actions are consumed by worker pools — the same shape as MinIO's globalExpiryState/globalTransitionState.
- AbortIncompleteMultipartUpload is not executed inside scanner/ILM (MinIO likewise: `internal/bucket/lifecycle/rule.go` has a FIXME, and it is actually carried by the `erasureSets.cleanupStaleUploads` global routine); in RustFS it is an independent ecstore background task `init_background_stale_multipart_upload_cleanup` (`bucket_lifecycle_ops.rs:3289-3320`) + on-demand at bucket deletion.
- Integration-test coverage: transition+restore, free-version, noncurrent, delete-marker, 0-day, background-scan expiry (`scanner/tests/lifecycle_integration_test.rs:1071-2095`).
### 3.6 heal candidate production (scanner side)
- Sampling: `hash mod_alt(next_cycle/prob_div, 1024/prob_div)`; when rescanning via the compacted branch, prob_div=16 gives an equivalent ×16 probability (the same compensation as MinIO, `scanner_folder.rs:125-127,2117-2122`).
- deep/normal: cycle-level `get_cycle_scan_mode` (bitrot_cycle default 30d, `scanner.rs:1626-1657`) → object-level with `HealScanMode::Deep`; fresh objects (modified within 60s) are demoted to Normal (:146-155); state persisted in `.background-heal.json` (`BackgroundHealInfo{bitrot_start_time,bitrot_start_cycle,current_scan_mode}`, same path and structure as MinIO).
- The scanner only enqueues, never executes inline (inline heal was removed; the compat flag only warns, `scanner_folder.rs:411-427`); `HealScanMode::Deep` is just a marker — the bitrot-verification read happens at the heal consumer (the ecstore Deep path).
- Metadata corruption → high-priority heal (`classify_get_size_failure` → HealMetadata); abandoned children → list_path_raw quorum verification + bucket-level/object-level high-priority heal; healing drives get sticky skipping (`should_heal` :1628-1648).
- pending-heal ledger: candidates rejected by a full heal channel are persisted into the cache info and retried next round.
- Replication heal: `queue_replication_heal` → the replication queue (going through the replication channel, not the heal channel); per-ARN replication usage statistics.
### 3.7 remote_scanner RPC protocol (RustFS-specific)
Requests ≤16KB msgpack (version/request_id/server_epoch/session_id/session_sequence/bucket/next_cycle/leader_epoch/scan_plan_digest/skip_healing/scan_mode/budget); frames ≤2MB, HMAC-SHA256 per-frame authentication (domain `rustfs-ns-scanner-frame-v3`); progress heartbeats 1s (250ms in budget mode); phase announcements Scanning→Persisting; RPC lifetime cap 24h, disconnect grace 2min; anti-replay session+sequence cache (capacity 65536); the server validates leader-fence and persisted-cycle consistency + fence re-validation every 5s; results Complete/Partial/NamespaceNotFound/CycleAhead; remote drives without v4-protocol support fall back to the leader scanning locally (`remote_scanner.rs` whole file; `scanner_io.rs:2750-2812`).
### 3.8 Rate limiting / budgets / hot updates / observability
- DynamicSleeper proportional backoff (speed tiers fastest/fast/default/slow/slowest, same five-tier parameters as MinIO); idle_mode master switch; an extra backoff capped at 250ms per request (10ms base) driven by foreground S3 read traffic.
- Cycle budget ScannerCycleBudget: max_duration/max_objects/max_directories (default 0 = unlimited); partial cycles still advance the cycle count.
- runtime_config with three-layer sources (env > config > default) and per-field source markers (Env/Config/ScannerCompatConfig/Default); admin `PUT /v3/config` hot update → generation+Notify takes effect immediately; `GET /v3/scanner/status` returns enabled/freshness(fresh/stale/unknown)/metrics/cycle_schedule/runtime_config; `GET /v3/ilm/expiry/status` returns expiry queue/workers/missed/blocked.
- Metrics: leader lock; cycle complete/partial/deferred/superseded; versions scanned; per-source (Usage/Lifecycle/BucketReplication/SiteReplication/Heal/Bitrot/Alerts) checked/executed/queued/missed; checkpoint set/used/stale; current path (per-disk+bucket in real time); cache save series; concurrency series; alerts (excess versions/version size/folders).
---
## 4. Item-by-item parity versus MinIO
### 4.1 heal trigger-channel comparison
| MinIO channel | RustFS counterpart | Status |
|---|---|---|
| A. Manual admin heal (healSequence, clientToken/forceStart/forceStop) | heal channel Start/Query/Cancel + cluster coordinator + envelope replay protection | ✅ equivalent and enhanced (cluster routing); sequence-semantics differences in §6 HS-06 |
| B. Resident background heal queue (newBgHealSequence + healRoutine worker pool) | HealManager resident scheduler + priority queue + bulkhead | ✅ equivalent and enhanced |
| C. Automatic new/replaced-drive resync (monitorLocalDisksAndHeal 10s + healFreshDisk + healingTracker + waitForFormatErasure handshake) | auto disk scanner (10s) + replacement_readiness + durable intent/proof state machine + heal_replacement_format | ✅ equivalent and enhanced (identity fence + completion proof; MinIO's tracker is stronger on external visibility, see §6 HS-07) |
| D. MRF (100k queue + persisted list.bin + shutdown replay + read-path corrupt delivery) | read-repair (Low + TTL dedup) + write-path convergence heal carry it partially; the `HealType::MRF` executor has no production entry | ⚠️ partially equivalent (§6 HS-01) |
| E. Scanner sampled heal (1/1024 + compacted ×16 compensation) + abandoned children | the same sampling + ×16 compensation + abandoned children + pending-heal ledger | ✅ equivalent and enhanced (the ledger) |
| F. Read-path inline trigger → MRF (GetObject part missing/corrupt, metadata rebuild missingBlocks>0) | read repair (three entries: missing_shards/decode_error/metadata_read_error) | ✅ equivalent (enqueued into the heal queue rather than the MRF queue) |
### 4.2 Object-level heal semantics comparison
| Feature | MinIO | RustFS | Status |
|---|---|---|---|
| mod-time quorum arbitration | listOnlineDisks | same | ✅ |
| ETag majority fallback (clock drift) | filterDisksByETag | `filter_by_etag`/`quorum_etag` (heal.rs:525-567) | ✅ verified first-hand |
| cannotHeal ETag waiver | waived on all-consistent ETag retry | heal.rs:679 | ✅ |
| Normal=CheckParts (stat) / Deep=VerifyFile (bitrot) | yes | `disks_with_all_parts` by scan_mode (ops/heal.rs:562-572,978-1024) | ✅ |
| Normal detecting corrupt auto-escalates to one Deep retry | erasure-healing.go:1101-1106 | ops/heal.rs:2022-2031 | ✅ |
| dangling determination (not-found > parity) + deletion auditing | isObjectDangling/deleteIfDangling | `dangling_delete_safety` (:1488) + scanner HEAL_DELETE_DANGLING | ✅ (audit-tags details differ) |
| Orphan data-dir/inline cleanup (CleanAbandonedData) | CheckAbandonedParts (invoked explicitly on scanner sampling + admin Remove) | in-heal-path `reclaim_orphan_data_dirs_best_effort` (:1428); standalone API NotImplemented at all three layers | ⚠️ partially equivalent (§6 HS-02) |
| Versioned/delete-marker heal | HealObject versionID; nullVersionID special case | per-version enumeration + delete-marker latest heal (B5 regression) | ✅ |
| Object-level healing metadata marker (x-minio-healing, RenameData skips version cleanup) | yes | no object-level marker; relies on drive-level healing.bin + NSLock + rename semantics | ⚠️ evaluation item (§6 HS-12) |
| Distribution/Index consistency, three lines of defense | yes (manual modification rejected) | target-drive format results all-ok check + identity fence | ✅ (different granularity) |
| no-parity (EC:0) objects | bitrot treated as unrecoverable | judged unrecoverable (:700-726) + write self-verification | ✅ enhanced (write-path self-verification) |
| three-layer distribution inconsistency refuses heal | yes | heal_walk normalization + page-bound defense | ✅ (different implementation approach) |
| multipart orphan reconciliation | carried by CheckAbandonedParts | explicitly NotImplemented (carried by lifecycle cleanup) | ⚠️ §6 HS-02 |
| suspended/decommissioned pool handling | skipped via IsSuspended | deferral semantics (store/heal.rs:192-207, PR #5876) | ✅ |
| heal mutually exclusive with concurrent deletes | NSLock + healing marker | NSLock + write lock | ✅ |
### 4.3 new-drive resync comparison
| MinIO | RustFS | Status |
|---|---|---|
| waitForFormatErasure handshake waiting indefinitely on four classes of recoverable errors | startup drive resolution + renew_disk reconnect path | ✅ (different model: RustFS does not block at startup waiting for format) |
| HealFormat NSLock + errNoHealRequired + refFormat-mismatch rejection | `heal_format`/`heal_replacement_format` fail-closed + target-slot restriction (PR #1787 semantics) | ✅ enhanced |
| per (pool,set) distributed lock preventing concurrent resync | set-level queue dedup + bulkhead (manager.rs:2854-2889) | ✅ |
| brand-new-cluster detection (drives-to-heal == total drives does not trigger) | replacement_readiness (independent mount point / physical-device validation, non-root) | ✅ enhanced |
| healingTracker (.healing.bin: Bytes/Items counters, QueuedBuckets/HealedBuckets, Resume snapshot, RetryAttempts ≤4, HealID linkage, diskID-change reset) | resume/checkpoint schema'd persistence + durable intent/proof (per-task files, CAS) | ✅ equivalent and enhanced (crash-window backfill); but **external snapshot visibility** is weaker than MinIO's (§6 HS-07) |
| skip versions written after heal start (ModTime > Started) | no such filter | ⚠️ §6 HS-13 |
| skip ILM-expired versions (filterLifecycle) | no such filter | ⚠️ §6 HS-13 |
| worker count max(GOMAXPROCS,NR)/4 floor 4, heal:drive_workers override | in-page concurrency 8 (Deep/AutoHeal forced to 1) + per-set bulkhead | ✅ (different parameter model) |
| waitForLowHTTPReq yield per entry | mainline throttle (foreground-utilization gating) | ✅ enhanced |
| heal scope includes the two pseudo-buckets `.minio.sys/config` and `.minio.sys/buckets`; newest bucket first | ErasureSet task pre-processes per bucket (meta-bucket semantics carried by heal_bucket) | ✅ (no "newest first" ordering) |
| whole-failure retry ≤4 (resetHealing + errRetryHealing) | schedule_retry resets both layers + recoverable retry ≤3 | ✅ |
### 4.4 scanner comparison
| MinIO | RustFS | Status |
|---|---|---|
| cluster single leader (globalLeaderLock) | leader.lock + persisted leader-epoch CAS fence | ✅ enhanced (epoch fence against split-brain; MinIO has no persisted epoch) |
| `.bloomcycle.bin` stores only the cycle (bloom removed) | same path stores cycle+leader_epoch (RSCYC001) | ✅ aligned (v1 misjudgment corrected) |
| folderScanner hash-mod-16 + compaction (500/10000/2500) | same constants + plan digest + cache-currency validation + dirty-first | ✅ enhanced |
| ≤GOMAXPROCS parallel scans per drive; healing drives excluded | per-set/per-disk semaphores + sticky skip of healing drives | ✅ |
| scannerSleeper (factor 2/max 1s, speed tiers hot-swapped) | DynamicSleeper same + idle_mode + foreground-read backoff | ✅ enhanced |
| idle semantics: `scanner:idle_speed=on` (throttle only in idle windows, full speed when busy) | `RUSTFS_SCANNER_IDLE_MODE=true` (master switch for rate limiting) | ⚠️ opposite semantic direction, §6 HS-14 |
| applyActions order (heal→ILM→replication→alerts) | apply_actions same order (heal candidates→ILM→replication heal→alerts) | ✅ |
| ILM 9 actions + batch evaluation + DeletePrefixObject optimization | same 9 actions + batch evaluation + expiry queue | ✅ (whether DeleteAllVersions has the single-call optimization was not checked line by line) |
| abandoned children (listPathRaw minDisks=N/2 detects under-written drives) | list_path_raw + quorum verification + high-priority heal | ✅ |
| incomplete multipart independent routine (6h interval/24h expiry, rename into .trash) | ecstore independent background task (configurable interval/expiry) | ✅ (trash two-stage cleanup detail differences, §6 HS-18) |
| usage dimensions (size/objects/versions/DM/histograms/replication/tier/bucket level) | full coverage + cluster snapshot with triple anti-rollback | ✅ enhanced |
| prefix-level usage (loadPrefixUsageFromBackend, consumed by console) | the cache holds the directory tree but flattens only to bucket level | ❌ §6 HS-08 |
| excess events s3:ObjectManyVersions/LargeVersions/PrefixManyFolders + auditing | metrics alert_excess_* only (defaults 100/1TiB/65538 vs MinIO 100/1TB/50000) | ⚠️ §6 HS-04/HS-17 |
| scanner metrics v3 (bucket_scans/directories/objects/versions/last_activity) | full rustfs_scanner_* suite + freshness | ✅ (different naming scheme) |
| TraceScanner / realtime metrics (mc admin scanner status/trace) | no trace channel; /v3/scanner/status has its own structure | ⚠️ §6 HS-03 |
### 4.5 admin/CLI/API surface comparison
| MinIO | RustFS | Status |
|---|---|---|
| `POST /minio/admin/v3/heal/...` start/status/cancel | `POST /rustfs/admin/v3/heal/...` same three states | ✅ (different path prefix is expected) |
| `HealStartSuccess`/`HealTaskStatus`/`HealResultItem`/DriveState | same-named fields JSON-compatible | ✅ |
| `POST /v3/background-heal/status` (BgHealState aggregate) | same path + degraded semantics + operations matrix | ✅ enhanced (no MRF per-endpoint sub-state, because there is no MRF) |
| `GET /v3/healthinfo` per-drive `HealInfo *HealingDisk` | no equivalent healthinfo heal field (replacement-recovery v4 covers part of it) | ⚠️ §6 HS-07 |
| madmin client HealStart/HealStatus/BackgroundHealStatus/ScannerStatus methods | wire types only, no client methods | ❌ §6 HS-05 |
| mc admin heal --pool/--set, --scan-mode, --force-start/stop | HealOpts full field support (pool/set/scanMode/forceStart/forceStop) | ✅ (server-side ready; missing the mc-side entry, HS-05) |
| ErrHealAlreadyRunning / ErrHealOverlappingPaths typed errors | dedup-merge + eviction semantics; no typed overlap rejection | ⚠️ §6 HS-06 |
| result backpressure (maxUnconsumedItems=1000, 10s keep-alive streaming, 24h unconsumed abort) | snapshot-style query (1024 entries + 8MiB truncation + 10min retention) | ⚠️ §6 HS-06 |
| `mc support inspect`/healing-bin offline dump | none (inspect.rs exists but the healing dump is unconfirmed) | ⚠️ P3 |
### 4.6 observability surface comparison
| Dimension | MinIO | RustFS | Status |
|---|---|---|---|
| heal metrics | minio_heal_objects_total/heal_total/errors_total/time_last_activity + v3 drive_health 2=healing | full rustfs_heal_* suite (admission/queue delay/running/throttle/page concurrency) | ✅ (RustFS lacks an equivalent of the single drive_health=healing gauge; DiskInfo.healing is already assigned) |
| scanner metrics | v3 6 + realtime 18 items | full rustfs_scanner_* suite + per-source dimensions | ✅ |
| ILM metrics | v3 5 (expiry/transition pending/active/missed + versions_scanned) | ilm expiry status API + scanner per-source | ✅ (different metrics and API shape) |
| trace | TraceHealing/TraceScanner channels | none | ❌ §6 HS-03 |
| auditing | HealObject events, dangling-deletion audit, scanner:manyversions etc. | structured logs (event style) + metrics; no audit-log events | ⚠️ §6 HS-04 |
| progress | healingTracker Bytes/Items/QueuedBuckets/current object + usage-cache total baseline | HealProgress{scanned/healed/failed/bytes/current_object/percentage}; bytes_processed annotated as 0, estimated_completion_time always None | ⚠️ §6 HS-07 |
### 4.7 configuration surface comparison (defaults)
| MinIO | RustFS | Notes |
|---|---|---|
| `heal:bitrotscan` (default off; on=every cycle; Nm=N×30×24h) | `heal.bitrot_cycle` / `RUSTFS_SCANNER_BITROT_CYCLE_SECS` (default 30d=2592000s; 0/on=Deep every cycle, off=disabled) | ✅ same semantics (RustFS default 30d, MinIO default off — **different defaults**, RustFS more aggressive) |
| `heal:max_io=100`/`max_sleep=250ms` (waitForLowIO) | mainline throttle thresholds 80%/80%, max_sleep 250ms | ✅ same shape (different threshold model) |
| `heal:drive_workers` (default -1 auto) | in-page concurrency 8 + per-set 1 | ✅ same shape |
| `_MINIO_HEAL_WORKERS` (GOMAXPROCS/2) | `RUSTFS_HEAL_MAX_CONCURRENT_HEALS=4` + `_MAX_CONCURRENT_PER_SET=1` | ✅ |
| `_MINIO_AUTO_DRIVE_HEALING` (on) | `RUSTFS_HEAL_AUTO_HEAL_ENABLE=true` | ✅ |
| `_MINIO_SCANNER` (on) | `RUSTFS_SCANNER_ENABLED=true` | ✅ |
| `scanner:speed` five tiers (default=2x/1s/1m) | same five tiers, same names, same parameters | ✅ |
| `scanner:idle_speed` (on) | `RUSTFS_SCANNER_IDLE_MODE` (true) | ⚠️ semantic direction (HS-14) |
| `scanner:alert_excess_versions=100` | 100 | ✅ |
| `scanner:alert_excess_folders=50000` | 65538 (compatible with the PBS layout) | ⚠️ HS-17 |
| `ilm:expiration_workers=100`/`transition_workers=100` | ecstore expiry/transition worker pools (keys under the ilm subsystem) | ✅ (defaults not checked item by item) |
| `api:stale_upload_cleanup_interval=6h`/`expiry=24h` | ecstore background task, configurable via env | ✅ (defaults not checked item by item) |
| — (none) | `RUSTFS_HEAL_QUEUE_SIZE=10000`, `_TASK_TIMEOUT_SECS=300`, `_INTERVAL_SECS=10`, `_LOW_PRIORITY_MERGE/DROP`, `_PAGE_*`, `_SET_BULKHEAD`, `_MAINLINE_*`, `RUSTFS_SCANNER_CYCLE_MAX_*` budgets, `_MAX_CONCURRENT_SET/DISK_SCANS=4`, `_YIELD_EVERY_N_OBJECTS=128`, etc. | RustFS-specific (finer-grained) |
### 4.8 Where RustFS exceeds MinIO
1. remote_scanner RPC (scan execution pushed down to the remote peer locally, with HMAC authentication/replay cache/fence re-validation/disconnect grace).
2. Persisted leader-epoch CAS fence + usage-snapshot epoch/cycle anti-rollback (MinIO has only the lock, no persisted epoch).
3. Cycle budgets (max_duration/objects/directories) + partial-cycle advancement semantics.
4. per-set/per-disk scan concurrency gates + a cache lock per bucket per set.
5. pending-heal ledger (heal candidates are not lost when the heal channel is full).
6. Drive-replacement durable intent + completion proof state machine + identity fence (MinIO's healingTracker has no proof).
7. mainline throttle foreground pressure gating (driven by permit utilization).
8. Cluster heal control coordinator + envelope replay protection + explicit degraded fallback.
9. Write-path shard bitrot self-verification (the EC:0 case).
10. dirty-usage fast-path wakeup (immediate write-path notification + dirty buckets first).
11. heal runtime observability matrix (priority×source operations snapshot).
12. workload admission integration (the heal scheduler reads the foreground pressure snapshot).
---
## 5. Gap and improvement list
Severity definitions: P1 = behavioral/operational alignment gap (affects production operations or toolchain compatibility); P2 = completeness (the feature exists but is missing a corner); P3 = cleanup/low risk. Each item includes current-state evidence, MinIO behavior, impact, recommendation, and acceptance.
### P1 (8 items)
**HS-01 The MRF/ECDecode/Metadata heal task types have no production trigger; HealEvent unwired**
- Current state: the `HealType::MRF/ECDecode/Metadata` executors are complete (task.rs:1700-2156) but have no production trigger anywhere in the repo; `HealEvent`/`HealEventHandler` (event.rs:50-367) has zero references outside the crate (verified first-hand by grep); channel conversion produces only Cluster/Object/Bucket/Prefix/ErasureSet (channel.rs:566-601).
- MinIO: mrf.go has a standalone MRF queue (capacity 100k, drop-and-count when full), msgp persistence to `.heal/mrf/list.bin` at process exit + startup replay, 1s delay for enqueues <1s (waiting for network recovery), healSleeper rate limiting; on the read path, GetObject part missing/corrupt, metadata rebuild missingBlocks>0, partial Put success, DeleteObject, multipart, and the peer client add up to 7+ delivery points.
- Impact: RustFS's read-repair + write-path convergence covers the main scenarios, but lacks: ① an event-driven Urgent ECDecode rebuild entry (on ecstore decode failure there is currently only Low read-repair); ② a metadata-only heal entry (the scanner's HealMetadata classification exists but goes through ordinary object heal); ③ MRF queue persistence (unconsumed repair intents are lost on restart — partially mitigated by the scanner's pending-heal ledger).
- Recommendation: a pick-one-of-three decision — (a) wire HealEvent (emit events at ecstore decode-failure/metadata-corruption points) + implement a persistent retry ledger; (b) delete the MRF/ECDecode/Metadata dead code and keep only a documentation note; (c) keep the executors and demote HealEvent to an internal API. (a) is recommended, but first quantify whether read-repair already meets the response-time requirements for decode-failure scenarios.
- Acceptance: an e2e decode-failure → Urgent heal-request chain; replay of pending repair intents after restart; HealEvent ring-buffer metrics.
**HS-02 CheckAbandonedParts NotImplemented at all three layers (missing standalone abandoned-data reconciliation entry)**
- Current state: `set_disk/ops/heal.rs:2052-2056`, `core/sets.rs:1144-1148`, `store/heal.rs:258-266` explicitly return `Err(NotImplemented)` at all three layers (verified first-hand); the comment reads "intentionally retained above the set layer until there is a concrete caller".
- MinIO: `CheckAbandonedParts` → per-drive `CleanAbandonedData`: read xl.meta → list UUID data-dirs + inline entries → diff against getDataDirs → delete surplus data-dirs/inline entries and rewrite xl.meta; invoked explicitly on scanner-sampled heals and admin heal Remove.
- Impact: RustFS's in-heal-path `reclaim_orphan_data_dirs_best_effort` (:1428) covers "reclaim orphan directories while healing", but ① there is no standalone trigger point (MinIO can also clean abandoned data before an object reaches the heal threshold); ② orphan inline-data entry cleanup is unconfirmed; ③ multipart orphan reconciliation is explicitly out of scope (a design decision, carried by lifecycle).
- Recommendation: evaluate promoting `reclaim_orphan_data_dirs_best_effort` to a fixed step of heal_object (if it is not already) + implement a real HealOperations::check_abandoned_parts (calling the same reclamation logic), or explicitly document "carried by lifecycle" and close the API surface.
- Acceptance: construct data-dir/inline orphans → cleaned after scanner sampling/admin heal; the three-layer API returns success or an explicitly documented NotSupported.
**HS-03 heal/scanner trace channels missing**
- Current state: zero hits for TraceHealing/TraceScanner (verified first-hand by grepping the whole repo).
- MinIO: `madmin.TraceHealing` (mc admin trace --healing, FuncName=heal.Bucket/heal.Object/heal.CheckAbandonedParts, with dry/remove/mode/version-id/disks/bytes), `TraceScanner` (mc admin scanner trace, supports --filter-size/--response-duration).
- Impact: no way to observe in real time the latency and parameters of individual heal/scanner actions; troubleshooting can rely only on aggregated metrics and logs.
- Recommendation: instrument heal-channel execution and scanner folder/item handling, and hook them into the existing admin trace subscription surface (reuse the rustfs trace infrastructure if it exists; otherwise extend it per madmin TraceType).
- Acceptance: an mc-equivalent tool can subscribe to the heal/scanner trace stream.
**HS-04 Scanner excess S3 events and auditing missing**
- Current state: only `rustfs_scanner_excess_*_total` metrics (versions 100 / version size 1TiB / folders 65538).
- MinIO: emits `s3:ObjectManyVersions` (>100 versions), `s3:ObjectLargeVersions` (cumulative >1TB), `s3:PrefixManyFolders` (>50000 subdirectories) events (UserAgent: Scanner) + scanner:manyversions/largeversions/manyprefixes auditing.
- Impact: users relying on event subscriptions for capacity governance (console/external auditing) receive no alerts.
- Recommendation: hook the scanner_folder alert points into notify event publishing (reusing the lifecycle event-channel semantics).
- Acceptance: after configuring bucket notifications, an over-threshold object triggers an event.
**HS-05 madmin client methods missing**
- Current state: `crates/madmin/src/heal_commands.rs` has only wire types (HealDriveInfo/Infos/HealResultItem); no HealStart/HealStatus/BackgroundHealStatus/ScannerStatus client methods.
- MinIO: madmin-go provides the full client; mc admin heal/scanner/status/trace are all built on it.
- Impact: admin tools like mc cannot directly drive the RustFS heal/scanner admin surface; automated operations must hand-write HTTP.
- Recommendation: add the client following the madmin-go interface shape (the server side is ready; this is pure client work).
- Acceptance: complete the start→query→cancel flow with the madmin client.
**HS-06 admin heal sequence semantics differ from MinIO**
- Current state: duplicate/overlapping requests are dedup-merged (returning the canonical task_id) or evicted; no ErrHealAlreadyRunning/ErrHealOverlappingPaths typed errors (verified first-hand: manager.rs:1309's already_running is an idempotent-startup guard, not an admin semantic); results are snapshot-style queries (1024 entries/8MiB truncation/10min retention), not MinIO's streaming increments (clientToken pulls increments + maxUnconsumedItems=1000 backpressure + 10s keep-alive + 24h unconsumed abort).
- Impact: mc admin heal's interaction model (long connection pulling increments) behaves against RustFS as multiple snapshot polls; automation scripts cannot easily distinguish "merged" from "newly started".
- Recommendation: ① incremental semantics: channel query supports item increments since the last clientToken (or a cursor); ② overlapping requests return a typed error code (or an explicit merged_into field in the receipt — the existing alias mechanism already provides the base); ③ verify forceStart's stop-old-then-start-new semantics.
- Acceptance: an madmin-compatible client polling in the MinIO style can retrieve the full item set.
**HS-07 healing progress and drive-level healing state insufficiently visible externally**
- Current state: byte-recovery progress `progress.bytes_processed = 0 // set to 0 for now` (erasure_healer.rs:967); `HealProgress::estimated_completion_time` is always None and `HealStatistics::add_healed_objects` is never written (progress.rs:38,135-139 zero calls); healthinfo has no per-drive HealInfo equivalent (MinIO HealingDisk: BytesDone/Failed/Skipped, ObjectsTotal baseline, QueuedBuckets/HealedBuckets, Resume snapshot, current object); v3 metrics lack an equivalent of the single drive_health=2 (healing) gauge.
- Impact: during a drive rebuild (potentially hours to days) operations cannot answer "where are we / how much is left / when will it finish".
- Recommendation: ① accumulate bytes in erasure set heal (heal_object already yields the object size); ② read the object-total baseline from usage-cache (the same approach as MinIO); ③ expose a per-drive healing snapshot in admin healthinfo/background status (DiskInfo.healing already exists; add the aggregated exposure); ④ derive the ETA from baseline + rate.
- Acceptance: during a drive rebuild, admin shows byte progress and ETA; an mc info-equivalent output shows the Healing flag.
**HS-08 prefix-level usage not exposed**
- Current state: the DataUsageCache holds the directory-tree entries (organized by hash_path), but `dui()` flattens only to the bucket name (data_usage_define.rs:858-915).
- MinIO: `loadPrefixUsageFromBackend` (30s cache) aggregates prefix usage from each set's `.usage-cache.bin`, consumed by console bucket-prefix statistics.
- Impact: console/front ends cannot show prefix-level usage; there is no API to locate "which prefix is using the space" in a large bucket.
- Recommendation: implement a prefix-flattening query API (the data is already in the cache; this is pure aggregation and exposure work).
- Acceptance: a ListBuckets/PrefixUsage API returns statistics matching the prefix filter.
### P2 (9 items)
**HS-09 get_disk_status always returns Ok (the only TODO)**: `crates/heal/src/heal/storage.rs:930-943` (verified first-hand). Currently no production caller (low risk). Recommendation: delete the method or wire it to the real ecstore disk status (the DiskStatus enum is already defined).
**HS-10 About 1/3 of HealStorageAPI methods are dead code**: get_object_meta/get_object_data/put_object_data/delete_object/verify_object_integrity/ec_decode_rebuild/get_disk_status/format_disk/heal_bucket_metadata/get_object_size/get_object_checksum/list_objects_for_heal (the non-paginated version, with its own memory_heavy warning) all have 0 callers. Recommendation: clean up or wire them together with the HS-01 decision (dead interfaces mislead future maintainers into thinking a call path exists).
**HS-11 bitrot self-test missing**: MinIO at startup runs bitrotSelfTest over known vectors for the four algorithms and exits Fatal on failure (guarding against silent data corruption). RustFS has no equivalent (verified first-hand by grep). Recommendation: at startup, run known-vector self-tests for HighwayHash256S and the other algorithms in use (low cost, high value).
**HS-12 object-level healing metadata marker evaluation**: during heal, MinIO tags objects with `x-minio-healing:true`, and RenameData uses it to skip version cleanup/legacy purge (missing it lets heal and concurrent deletes destroy each other). RustFS has no object-level marker (verified first-hand by grep; object.rs has no healing branch) and relies on NSLock + rename semantics. Recommendation: audit whether the RustFS rename-commit path has a "heal commit racing concurrent delete/version cleanup" window; if not, document the difference, and if so, add a marker-equivalent mechanism.
**HS-13 erasure set heal lacks "skip newly written / ILM-expired versions" filters**: MinIO resync skips versions with ModTime>tracker.Started (so heal does not chase the tail of new writes) and ILM-expired versions (so work is not wasted). RustFS's erasure_healer does not implement such filters (per-version dedup exists; time/ILM filters do not). Impact: a long tail on rebuild completion (the completion decision for a continuously written bucket is pushed out by new versions) and wasted heal work. Recommendation: add a started_at time filter at the disk-walk enumeration point + an evaluator pre-check.
**HS-14 scanner idle semantics point the opposite way from MinIO**: MinIO `scanner:idle_speed=on` (default) means "throttle only when the cluster is idle, full speed when busy"; RustFS `RUSTFS_SCANNER_IDLE_MODE=true` (default) is a master switch for rate limiting (false = never sleep at all). The default behaviors may end up similar (both throttle), but the parameter semantics are not interchangeable; migration docs must state this explicitly; if mc config compatibility is the goal, a rename/re-semantization is needed. Recommendation: document the difference first, then evaluate aligning the semantics.
**HS-15 alert_excess_folders default differs**: RustFS 65538 (compatible with the PBS/Proxmox layout, scanner_folder.rs:79) vs MinIO 50000. The behavioral difference is that the trigger threshold differs out of the box. Recommendation: document it (keeping 65538 has local rationale).
**HS-16 single-node default-cycle hook not enabled**: `single_disk_default_cycle_secs(_features) -> None` is always empty (scanner.rs:1428-1430); single-node deployments get no dedicated default-cycle override. Recommendation: after deciding the single-node default-cycle policy, enable or delete the hook.
**HS-17 DeleteAllVersions batch-optimization check**: MinIO uses the single DeletePrefix+DeletePrefixObject call instead of per-version fan-out. Whether RustFS's expiry-queue path has the same optimization was not verified line by line (integration tests cover behavioral correctness). Recommendation: check the `apply_expiry_rule` all-versions delete path; if there is no prefix single-call optimization, evaluate adding it.
### P3 (3 items)
**HS-18 trash/temp-directory two-stage cleanup detail check**: MinIO cleans `.minio.sys/tmp/.trash` (delete_cleanup_interval default 5m + deleteCleanupSleeper) and stale uploads are renamed into trash in two stages. RustFS has delete_tail_activity.rs and the stale multipart task; whether the two-stage semantics are fully aligned was not verified line by line. Recommendation: align or document.
**HS-19 root-heal direct path is dead code**: `should_handle_root_heal_directly` is always false (admin/handlers/heal.rs:1200-1202, locked by a test); the store.heal_format direct branch is unreachable. Recommendation: delete the dead branch or restore the direct path as a fallback for cluster-coordination failure.
**HS-20 compat flags and dead metrics cleanup**: `RUSTFS_SCANNER_INLINE_HEAL_ENABLE` (enabling only warns) + the dead `rustfs_scanner_inline_heal_total` metric + the scanner-domain code in `rustfs_common::metrics` awaiting layering migration (backlog #1843 already filed). Recommendation: clean up along with the layering migration.
### Not pursuing parity by design (7 items, recorded to prevent later misreading as gaps)
1. **bloom filter**: removed from MinIO master; RustFS reuses `.bloomcycle.bin` as the cycle/epoch fence, consistent with MinIO's current state.
2. **scanner cluster single leader**: both sides agree; RustFS additionally has the epoch fence.
3. **heal emits no S3 bucket notification**: both sides agree (heal results go through admin status).
4. **incomplete multipart not executed inside scanner/ILM**: both sides agree (independent background routine).
5. **inline heal removal**: a deliberate RustFS choice (the scanner only enqueues); MinIO's applyHealing inline path is not a parity target.
6. **heal-sequence resident keep-alive (10s blank write-back)**: RustFS's snapshot-query model differs; handling incremental semantics per HS-06 is enough — do not copy the streaming keep-alive.
7. **`.trash`/`tmp-old` path-name compatibility**: RustFS's layout constants are independent; no literal alignment with MinIO paths.
---
## 6. Configuration defaults master table (RustFS)
heal (env prefix `RUSTFS_HEAL_`, `crates/config/src/constants/heal.rs`, consumed at `manager.rs:724-800`):
| Setting | Default | Hot update |
|---|---|---|
| AUTO_HEAL_ENABLE | true | no |
| QUEUE_SIZE | 10000 | no |
| INTERVAL_SECS | 10 | no (fixed at startup) |
| TASK_TIMEOUT_SECS | 300 | no |
| MAX_CONCURRENT_HEALS | 4 | no |
| MAX_CONCURRENT_PER_SET | 1 (≤min(global, value)) | no |
| LOW_PRIORITY_MERGE_ENABLE | true | no |
| LOW_PRIORITY_DROP_WHEN_FULL | true | no |
| PAGE_OBJECT_CONCURRENCY | 8 (Deep/AutoHeal forced to 1) | no |
| EVENT_DRIVEN_SCHEDULER_ENABLE | true | no |
| SET_BULKHEAD_ENABLE | true | no |
| PAGE_PARALLEL_ENABLE | true | no |
| MAINLINE_THROTTLE_ENABLE | true | no |
| MAINLINE_READ/WRITE_UTILIZATION_HIGH_PERCENT | 80/80 | no |
| MAINLINE_MAX_SLEEP_MS | 250 | no |
| (master switch) RUSTFS_HEAL_ENABLED | true | no |
| admin subsystem heal.bitrot_cycle | 30d | yes (via scanner runtime config) |
scanner (admin subsystem `scanner`, `crates/config/src/constants/scanner.rs` + `ecstore/src/config/scanner.rs` + `runtime_config.rs:527-673`):
| Key | env | Default |
|---|---|---|
| speed | RUSTFS_SCANNER_SPEED | default (2x/1s/60s) |
| delay / max_wait / cycle / start_delay | RUSTFS_SCANNER_* | derived/empty |
| cycle_max_duration/objects/directories | …_MAX_* | 0 (unlimited) |
| bitrot_cycle | …_BITROT_CYCLE_SECS | 2592000 (30d; 0/on=every cycle, off=disabled) |
| idle_mode | …_IDLE_MODE | true |
| cache_save_timeout | …_CACHE_SAVE_TIMEOUT_SECS | 14s |
| max_concurrent_set_scans / disk_scans | …_MAX_CONCURRENT_* | 4/4 |
| yield_every_n_objects | …_YIELD_EVERY_N_OBJECTS | 128 |
| alert_excess_versions / version_size / folders | …_ALERT_* | 100 / 1TiB / 65538 |
scanner-internal env: `RUSTFS_DATA_USAGE_UPDATE_DIR_CYCLES=16`, `RUSTFS_HEAL_OBJECT_SELECT_PROB=1024`, `RUSTFS_SCANNER_DEEP_VERIFY_COOLDOWN_SECS=60`, `RUSTFS_DATA_USAGE_FAILED_OBJECT_TTL_SECS=86400`/`_MAX=10000`, `RUSTFS_LOCK_ACQUIRE_TIMEOUT=5s`, `RUSTFS_SCANNER_ENABLED=true`, `RUSTFS_SCANNER_INLINE_HEAL_ENABLE=false` (compat warning).
All 17 scanner keys support the env > config dual channel + admin PUT hot update (generation+Notify takes effect immediately); heal runtime parameters are currently env-only (no admin hot-update entry; the `Arc<RwLock<HealConfig>>` structure is already reserved).
---
## 7. Related backlog / history index
- Automatic drive-replacement healing series (closed loop): backlog #1786 (redundant false-green algorithm), #1787 (target-slot restriction), #1789 (binding resume and the healing marker to the replacement instance), #1791 (black-box/white-box acceptance matrix).
- #801 DiskInfo.healing never assigned (fixed and closed; the assignment chain now lives at `set_disk/mod.rs:4988`).
- #1651 Scanner metrics node/source/bucket-drive dimensions (OPEN; related to §3.8/§4.6 of this analysis).
- #1843 crates/common 83% scanner/heal domain code layering migration (OPEN; includes HS-20).
- Historical defects cited in code comments (now guarded with regression tests): #856/#799 B7 (offline drive falsely recorded healed), #855/B6/#1033 (a skip round must not be marked complete), #920 (sub-quorum union enumeration), #856 B5 (per-version resume), #5173 (bitrot trailing bytes), #5029 (stale-version merge at regression nodes).
- v1 parity document: `docs/rustfs-heal-scanner-vs-minio-parity-assessment.md` (superseded by this document); the landing playbook `docs/rustfs-heal-scanner-vs-minio-improvement-playbook.md` (some entries have since been overtaken by implementation).
- Drive-replacement deep analyses: `docs/new-disk-replacement-and-healing-deep-analysis-zh.md`, `docs/node-disk-identity-and-healing-analysis-zh.md`.
## 8. Audit method and limitations
- Four parallel audit tracks (heal crate file by file, scanner crate file by file, ecstore integration-layer wiring, MinIO master source study) + the main session verifying each key "missing" conclusion first-hand (the get_disk_status TODO, HealEvent's zero external references, .bloomcycle.bin having no bloom implementation, check_abandoned_parts NotImplemented at all three layers, the ETag fallback being implemented, zero trace-channel hits, the already_running semantics).
- Points not verified line by line (marked "unconfirmed / not checked line by line" in the text): the DeleteAllVersions prefix single-call optimization (HS-17), trash two-stage cleanup details (HS-18), ilm worker default comparisons, stale multipart default comparisons, mc CLI flag spellings (MinIO side). Of these, HS-17 and HS-18 completed line-by-line verification on 2026-08-19; conclusions in §9.2/§9.3.
- MinIO-side references follow its master `7aac2a2c5b`; RustFS-side line numbers follow the 2026-08-16 workspace — for later evolution, search by symbol name instead.
## 9. Landing results (updated 2026-08-19)
All 14 sub-issues derived from this audit (backlog #1865~#1878) are closed. This section is the final disposition record for the gap list HS-01~HS-20, and also the incremental baseline for the next parity re-audit.
### 9.1 Landed (all PRs merged to main)
- HS-01 MRF wiring + persistent repair ledger (#1865, PR #6189): decision (a) chosen. common MRF channel (bounded 8192, try_send never blocks) + heal mrf_queue (100k entries / 8MiB dual-capacity ring) + `buckets/.heal/mrf/journal.bin` CRC-persisted replay (torn tail truncated, deleted after replay) + three delivery points (read decode_error→Urgent ECDecode, scanner metadata corruption→High Metadata, add_partial→Normal) + `RUSTFS_HEAL_MRF_ENABLE` one-switch rollback.
- HS-02 abandoned parts/data-dir reconciliation (#1866, PR #6179): wired up the abandoned-check entry, retaining dry-run / reclaim counters.
- HS-03 heal/scanner trace channels (#1867, PR #6179): in-process trace bus + `/v3/trace` admin streaming subscription + heal task / abandoned-parts / scanner folder / ILM / heal-candidate trace producers.
- HS-04 scanner excess S3 events (#1868, PR #6176): the three events `s3:Scanner:ManyVersions/LargeVersions/BigPrefix` + 24h edge cooldown; the HS-15 threshold delta documented (`docs/operations/scanner-excess-alerts.md`).
- HS-05 madmin client phase 1 (#1869, PR #6166): SigV4 admin client heal/scanner methods; incremental-consumption methods await a follow-up (the protocol was already folded in by HS-06).
- HS-06 admin heal incremental semantics and typed overlap (#1870, PR #6206): `sinceSeq/nextSeq/minSeq` incremental cursor (wire additive; absent = full snapshot) + `RUSTFS_HEAL_OVERLAP_POLICY` (default merge unchanged; under minio_error, typed AlreadyRunning/OverlappingPaths rejections) + forceStart stops the old sequence before starting the new one.
- HS-07 healing progress visibility (#1871, PR #6179): data-usage total baseline + baseline/current/healed counters.
- HS-08 prefix usage (#1872, PR #6171): `GET /v3/usage/{bucket}`.
- HS-11 bitrot startup self-test (#1873, PR #6165).
- HS-13 heal skip filters (#1875, PR #6179): filter-hit versions are no longer counted as failures.
- HS-16 single-node cycle hook (#1878, PR #6250): removed the always-None hook; the decision record is in `docs/operations/heal-scanner-parity-notes-zh.md`.
- HS-09/10/19/20 dead-code cleanup batch (#1877, PR #6256): net 911 lines, zero behavior change; the `get_disk_status` TODO (the repo's only product TODO) cleared to zero; `ec_decode_rebuild`/`get_object_meta`, kept due to the HS-01 linkage, are retained with Reserved annotations (MRF currently executes via `heal_object`).
### 9.2 Confirmed "already implemented / not a gap" after verification (audit-period misjudgment corrections, four in total)
- bloom filter (corrected in §0): removed from MinIO master; both sides now agree.
- ETag fallback arbitration (corrected in §0): RustFS already has the implementation (`set_disk/ops/heal.rs`).
- HS-17 (#1876, closed after line-by-line verification on 2026-08-19): the DeleteAllVersions prefix single-call optimization is fully implemented in RustFS — `apply_expiry_on_non_transitioned_objects` sets `delete_prefix + delete_prefix_object` for the two `delete_all()` actions and then performs a single `delete_object` call (`bucket_lifecycle_ops.rs:5047-5056`); the SetDisks branch takes one write lock + one all-version quorum read + inline per-version object-lock checks (`set_disk/ops/object.rs:5566-5612`), aligned line by line with MinIO `expire.go`'s `applyExpiryOnNonTransitionedObjects`. The item §8 listed as "not verified line by line" now has a conclusion: the current state is already the optimized path; nothing to implement.
- HS-14 (#1878, checked alongside PR #6250): MinIO's "idle = throttle only when idle" was the behavior before 2024-01 minio/minio#18734 (`scannerIdleMode` is now a static config; `idle_speed=on` by default means always throttling per the speed tier — the "idle" naming is a historical leftover); RustFS's `RUSTFS_SCANNER_IDLE_MODE` points the same way as MinIO's current semantics, and additionally has a foreground-read backoff floor that MinIO lacks. The real migration traps (the variable must carry the `RUSTFS_` prefix, the `on/off` vs `true/false` vocabulary, `false` also turning off foreground protection) are documented in `docs/operations/heal-scanner-parity-notes-zh.md`.
### 9.3 Audit-style conclusions (no code change needed)
- HS-12 (#1874, PR #6183): the class of race MinIO defends against with `x-minio-healing` does not exist — every commit surface for the same (bucket, object) is mutually exclusive under the same object-level ns write lock, and the heal lock guard covers the whole rename commit; delivered 2 concurrency-invariant regression tests + the intersection matrix in `docs/operations/heal-concurrency-safety-notes-zh.md`.
- HS-18 (#1878, line-by-line verification on 2026-08-19): trash/tmp three-stage cleanup fully aligned — stale multipart isolation-cleanup is equivalent and safer (`delete_all_with_quorum` recursively deletes per drive, i.e. the `move_to_trash` rename into `.rustfs.sys/tmp/.trash`, plus lock + fence); trash draining is essentially equivalent (no per-entry sleeper throttling; the 5m cycle naturally rate-limits); tmp non-trash 24h reclamation is equivalent (RustFS's 5m is more timely than MinIO's 6h); the three cycle defaults 24h/6h/5m all align. The item §8 listed as "not verified line by line" now has a conclusion.
### 9.4 Handed over to follow-ups (summarized in the backlog#1862 comment thread)
HS-01 bitrot GET→MRF full-chain e2e, kill -9 journal replay e2e, queue-full RSS stress test (≤ budget+10%); HS-05/06 madmin incremental-consumption methods + single-source wire + embedded e2e + multi-round polling soak; HS-08 multi-drive scanner cycle e2e; HS-04 excess audit entries; HS-18 the stale-multipart crash-residue window below quorum (crashing mid-fan-out with already-cleaned drives > parity means FileNotFound is not in the ignore set, so convergence is unnatural; the fix needs a dedicated quorum variant).
Recommendation for the next re-audit: trigger it after the next big heal/scanner feature lands, using this section as the incremental baseline.
@@ -1,568 +0,0 @@
# RustFS heal / scanner 全量功能分析与 MinIO 对标(v2)
> English version: [rustfs-heal-scanner-vs-minio-comprehensive-analysis-2026-08-16.md](rustfs-heal-scanner-vs-minio-comprehensive-analysis-2026-08-16.md)
- 日期:2026-08-16(基于 main 分支当日代码,审计时 HEAD ≈ `a118d7e4f`
- 范围:`crates/heal`src 19,560 行 + tests 2,274 行)、`crates/scanner`src 约 26,000 行 + tests)、`crates/data-usage``crates/ecstore` 中 heal/heal_walk/bitrot_self_verify 与 config、`crates/heal-contracts/src/heal_channel.rs``crates/madmin`heal/scanner wire 类型)、`rustfs/src`startup wiring、admin handlers、集群 RPC
- 对标基线:minio/minio masterHEAD `7aac2a2c5b`,仓库已进入维护模式,master 冻结,即最终态)
- 方法:四路并行审计(heal crate / scanner crate / ecstore 集成层 / MinIO 源码研究),关键结论逐条人工抽验(文内标注"已亲验"处为一手验证)
- 本文档取代 `docs/rustfs-heal-scanner-vs-minio-parity-assessment.md`2026-06-15 v1)。v1 之后 heal/scanner 相关提交超过 80 个(换盘自动修复全链路、resume 状态机、usage 收敛权威化、集群级 heal 协调、ILM restore 语义等),v1 的功能清单与差距判断已全面过时;v1 中"bloom filter 缺失"等结论经本次核实为**误判**(详见 §5.4)。
---
## 0. 结论摘要
1. **总体判断:heal 与 scanner 的核心功能链路已经完整**。对象级 healquorum 仲裁 + ETag 兜底 + bitrot Deep 校验 + dangling 处理)、erasure set 深扫(per-set disk-walk 并集枚举)、按版本断点续扫(schema 化持久层 + CAS 原子发布 + 崩溃窗口补齐)、换盘自动修复(readiness 校验 + 身份围栏 + durable intent + completion proof)、scanner 周期循环(leader lock + 持久化 leader-epoch 围栏)、data usage 统计(桶级/集群级、主+备+观测快照、epoch/cycle 防回退)、ILM 全动作(expiry/transition/noncurrent/free-version/delete-marker 清理)、admin Start/Query/Cancel 协议(clientToken 语义对齐 madmin)——以上均有实现且带回归测试。两个 crate 内**没有空实现/早退桩**,异常路径全部有日志 + 指标 + 错误语义。
2. **主要缺口集中在"入口与观测面",而不是修复算法本身**MRF/ECDecode/Metadata 三类任务执行体已实现但无生产触发入口(`HealEvent` 完全未接线);`CheckAbandonedParts` 在 ecstore 三层全部 `NotImplemented`heal/scanner trace 通道缺失;scanner 超限 S3 事件缺失;madmin 客户端方法缺失(只有 wire 类型);heal 字节级进度/ETA 未实现。
3. **与 v1 认知的重要修正**bloom filter 在 MinIO 当前 master **已删除**`.bloomcycle.bin` 只存 cycle 计数),RustFS 现状与 MinIO 一致;MinIO scanner 同样是**集群级 leader 单例**RustFS 的 leader.lock 模型与 MinIO 同型;RustFS 的 ETag 多数派兜底仲裁已实现(`crates/ecstore/src/set_disk/ops/heal.rs:525-567,679`,已亲验),v1 担心的仲裁缺口不存在。
4. **RustFS 在多处超出 MinIO**remote_scanner RPC 协议(远端 peer 本地扫描而非 leader 跨网读远盘)、持久化 leader-epoch CAS 围栏、周期预算与 per-set/per-disk 并发闸、pending-heal 账本、durable replacement intent + completion proof 状态机、前台压力门控(mainline throttle)、集群 heal control coordinator + envelope 重放防护。
5. 差距分级统计:P1(行为/运维对齐缺口)8 项,P2(完善性)9 项,P3(清理/低风险)3 项,"按设计不追平"7 项。完整清单见 §6。
---
## 1. 架构总览
### 1.1 RustFS 三层架构
RustFS 把 MinIO 在 `cmd/` 内单体的 heal/scanner 拆成三层 + 两个独立 crate:
| 层 | 位置 | 职责 |
|---|---|---|
| 原语层 | `crates/ecstore/src/set_disk/ops/heal.rs`~3,240 行)、`ops/heal_walk.rs``ops/bitrot_self_verify.rs`;上层封装 `store/heal.rs``store/heal_walk.rs``core/sets.rs` | 对象/桶/format/替换盘格式修复、disk-walk 并集枚举、写入路径 bitrot 自校验;由 `SetDisks`/`Sets`/`ECStore` 实现 `rustfs_storage_api::HealOperations` 契约(`crates/storage-api/src/object.rs:503-519` |
| heal 运行时 | `crates/heal` | 进程级 HealManager(优先级队列/调度器/auto disk scanner/断点续传 resume)、HealChannelProcessor(消费全局 heal channel)、换盘替换恢复状态机 |
| scanner 运行时 | `crates/scanner` | 数据使用扫描、ILM 评估与入队、heal 候选生产、复制用量统计、remote scanner RPC |
| 共享协议 | `crates/heal-contracts/src/heal_channel.rs`~776 行) | Start/Query/Cancel 命令通道、`HealOpts`/`HealScanMode`/`HealRequestSource`/`HealAdmission*` 共享类型、`HealResultItem`madmin |
| 共享数据 | `crates/data-usage` | `DataUsageEntry/Info`、直方图、`hash_path`scanner 产生、ecstore/admin 消费 |
启动链路(已亲验 wiring):
1. `rustfs/src/startup_services.rs:93``init_background_service_runtime(store)`
2. `rustfs/src/startup_background.rs:41-81`:创建全局 heal 服务取消令牌;读 `RUSTFS_SCANNER_ENABLED`(别名 `RUSTFS_ENABLE_SCANNER`,默认 true)与 `RUSTFS_HEAL_ENABLED`(别名 `RUSTFS_ENABLE_HEAL`,默认 true);**只要 heal 或 scanner 任一开启就初始化 heal manager**scanner 产生的 heal 候选需要消费端;两者都关时 heal channel 不初始化,`send_heal_request` 报 "Heal channel not initialized")。
3. `crates/heal/src/lib.rs:142-216`owned task 内原子初始化(caller 取消不会遗留半初始化 manager,`lib.rs:123-131``GLOBAL_HEAL_RUNTIME_INIT` 互斥单飞)→ `HealManager::start()``rustfs_common::heal_channel::init_heal_channels()` → spawn `HealChannelProcessor::start_with_receipts`
4. `crates/heal/src/heal/manager.rs:1301-1356` `HealManager::start``start_scheduler()``manager.rs:2394-2461`interval 默认 10s + `Notify` 事件驱动唤醒)→ `process_unclean_shutdown()``manager.rs:1362-1695`)→ `enable_auto_heal`(默认 true)时 `start_auto_disk_scanner()``manager.rs:2464-2999`)。
5. server ready 后 `rustfs/src/startup_lifecycle.rs:150-152``enable_scanner``init_data_scanner(token, store)``crates/scanner/src/scanner.rs:1293-1372`)。
6. 优雅停机:`rustfs/src/startup_shutdown.rs:308` `shutdown_ahm_services()`(取消令牌);`:414` `clear_unclean_shutdown_markers()`
### 1.2 MinIO 对应结构(master 最终态)
| MinIO 文件 | 职责 |
|---|---|
| `cmd/admin-heal-ops.go` | 手动 admin heal 序列(healSequence、clientToken/forceStart/forceStop |
| `cmd/global-heal.go` | 常驻后台 heal 队列(newBgHealSequencetoken 固定 `0000-…`,永不结束)+ `healErasureSet`(逐 set 全量对象 heal |
| `cmd/background-heal-ops.go` | healRoutine worker 池(`_MINIO_HEAL_WORKERS`,默认 GOMAXPROCS/2)消费 healTask |
| `cmd/mrf.go` | MRFMost Recent Fail)队列(容量 100,000),进程退出时持久化 `.minio.sys/buckets/.heal/mrf/list.bin` 并启动回放 |
| `cmd/background-newdisks-heal-ops.go` | 新盘/换盘自动 resyncmonitorLocalDisksAndHeal 10s 轮询 + healFreshDisk + healingTracker |
| `cmd/erasure-healing.go` / `erasure-healing-common.go` | 对象级 heal 核心(~800 行)、listAndHeal |
| `cmd/data-scanner.go` | scanner 循环(globalLeaderLock 集群单例)+ folderScanner + applyActions |
| `cmd/erasure.go`nsScanner/ `erasure-server-pool.go` | NSScanner 三层结构 |
| `cmd/bucket-lifecycle.go` | ILM 执行器(expiry/transition worker 池) |
| `cmd/xl-storage.go` | DiskInfo.Healing、CheckParts/VerifyFile、CleanAbandonedData、RenameData healing 分支 |
| `cmd/prepare-storage.go` | waitForFormatErasure 新盘启动握手 |
### 1.3 架构级差异(设计取舍,非缺陷)
1. **heal 队列模型**MinIO 所有 healscanner 抽样/MRF/admin/新盘 resync)汇入单 channel + 固定 worker 池(新盘 resync 另有 per-drive worker 池);RustFS 是优先级堆 + 去重合并 + 容量分级丢弃 + per-set bulkhead + 前台压力门控的多策略调度器(`manager.rs:3003-3420`)。RustFS 表达力更强,代价是"重复请求被合并"的可观测性问题(v1 已指出,现有 `HealAdmissionReceipt` canonical task_id + alias 机制回应了它,`manager.rs:1759-1846`)。
2. **scanner 远端盘访问**MinIO leader 通过磁盘抽象层透明读写远端节点磁盘;RustFS leader 通过 remote_scanner RPC 把扫描执行下放到远端 peer 本地进行(`crates/scanner/src/remote_scanner.rs`),只回传结果与进度心跳。两者都是集群单 leader。RustFS 方案省 leader↔远端的元数据读放大,代价是需要维护独立 RPC 协议(HMAC 逐帧认证、会话重放缓存、fence 复验,`remote_scanner.rs:52-61,405-496,1024-1065`)。
3. **heal 状态持久化**MinIO 用单文件 `.healing.bin`msgp healingTrackerdiskID 不匹配即重置);RustFS 用 schema 化多文件(resume/checkpoint/intent/seal/proof 各自 CAS 发布,`resume.rs:38-61`),崩溃窗口显式补齐(`erasure_healer.rs:389-402``resume.rs:1027-1057`)。
4. **写路径自保护**MinIO 写入后靠后台 heal 收敛;RustFS 在 PutObject/CompleteMultipartUpload 提交 rename 后主动检查 `convergence.needs_heal()` 并立即入队对象 heal`set_disk/ops/object.rs:2291-2306``ops/multipart.rs:2574-2589`),另有读修复 read repair`io_primitives.rs:1040-1160`)。
---
## 2. Heal 已实现功能全景
### 2.1 任务类型(`HealType``crates/heal/src/heal/task.rs:85-111`
| 类型 | 语义 | 执行体 | 生产触发方 |
|---|---|---|---|
| `Cluster` | 所有 bucket 依次 heal(结构 + 可选递归对象),批内重试 ≤3 | `heal_cluster` task.rs:1420-1490 | channelbucket 为空即 Clusterchannel.rs:576-577 |
| `Object{bucket,object,version_id}` | 单对象/版本;不存在时按 `recreate_missing` 重建或报错 | `heal_object` task.rs:855-1146 | admin、scanner、read-repair、写路径收敛、add_partial |
| `Bucket{bucket}` | 桶元数据/结构;`recursive` 再遍历全部对象版本 | `heal_bucket` task.rs:1284-1418 + `heal_bucket_objects` task.rs:1508-1698 | adminPOST /v3/heal/{bucket})、scanner `build_bucket_heal_request` |
| `Prefix{bucket,prefix}` | 按前缀递归 | `heal_prefix` task.rs:1492-1506 | channel`recursive && prefix` 非空(channel.rs:578-585 |
| `ErasureSet{buckets,set_disk_id}` | format 修复 + healing 标记 + 逐桶预处理 + 可恢复逐版本深扫 | `heal_erasure_set` task.rs:2158-2642 | adminpool/set 参数)、auto disk scanner、unclean shutdown、renew_disk、durable replacement 恢复 |
| `Metadata{bucket,object}` | 仅元数据(Deep、不重建数据) | `heal_metadata` task.rs:1700-1859 | **无生产触发方**(§6 HS-01 |
| `MRF{meta_path}` | 失败路径驱动的 Deep 修复(recursive+update_parity | `heal_mrf` task.rs:1861-1992 | **无生产触发方**(仅 `HealEvent` 可生成,未接线) |
| `ECDecode{bucket,object,version_id}` | EC 解码重建(Deep+recreate+update_parity),Urgent 优先级 | `heal_ec_decode` task.rs:1994-2156 | **无生产触发方**(仅 `HealEvent` 可生成,未接线) |
优先级 `Low/Normal/High/Urgent`task.rs:168-179);状态机 `Pending/Running/Retrying/Completed/Failed/Cancelled/Timeout`task.rs:225-241)。
### 2.2 触发路径全景(admin 之外)
| 通道 | source | 优先级 | 证据 |
|---|---|---|---|
| Scanner 周期抽样(1/1024`RUSTFS_HEAL_OBJECT_SELECT_PROB` | Scanner | Low | `scanner_folder.rs:2117-2136``:1150``remove_corrupted=HEAL_DELETE_DANGLING(true)``recreate_missing=false``common/heal_channel.rs:24``scanner_folder.rs:510-511` |
| Scanner 元数据损坏(get_size 失败分类 HealMetadata | Scanner | High | `scanner_folder.rs:2147-2208``:1244-1260` |
| Scanner abandoned children(缓存有、盘上无,list_path_raw quorum 核查) | Scanner | High(桶级+对象级) | `scanner_folder.rs:2528-2792` |
| Scanner pending-heal 账本重试(heal 通道满被拒后持久化,每桶每轮 ≤128 条、上限 10k) | Scanner | 原优先级 | `scanner_folder.rs:1721-1763``:99-100` |
| auto disk scannerunformatted 盘经 replacement_readiness 确认 / `runtime_state=="returning"` 盘 / durable intent 重入) | AutoHeal | Low | `manager.rs:2464-2999` |
| unclean shutdown 恢复(启动读 `unclean-shutdown` 标记 → 全部本地 set ErasureSet heal | AutoHeal | Low | `manager.rs:1362-1695` |
| 写路径收敛(PutObject/CompleteMultipartUpload 后 `convergence.needs_heal()` | Internal | Normal | `set_disk/ops/object.rs:2291-2306``ops/multipart.rs:2574-2589` |
| 部分对象 healadd_partial | Internal | Normal | `set_disk/ops/object.rs:5808-5825` |
| 旧数据目录清理残留 enqueue | Internal | Normal | `set_disk/core/io_primitives.rs:3880-3907` |
| 读修复(metadata_read_error / missing_shards / decode_errorTTL 去重缓存) | ReadRepair | Low | `set_disk/read.rs:407,995,1079``submit_read_repair_heal``io_primitives.rs:1105-1160`),`recreate_missing=true` |
| 盘重连遇 UnformattedDisk → send_heal_disk | AutoHeal | Normal | `set_disk/ops/locking.rs:339-347` |
| Admin API(含集群 coordinator 路由) | Admin | High | `rustfs/src/admin/handlers/heal.rs:174-212``:771-930` |
| 集群 RPC healpeer 调用) | — | — | `rustfs/src/storage/rpc/node_service/heal.rs``ecstore/src/cluster/rpc/peer_s3_client.rs:296,1209` |
注意:MinIO 的 MRF 通道(读路径检出 part 缺失/损坏即时投递 + 队列持久化 + shutdown 回放,`cmd/mrf.go``erasure-object.go:395-410,800-812`)在 RustFS 由 read-repair + 写路径收敛**部分替代**;`HealType::MRF`/`ECDecode`/`Metadata` 三个执行体没有生产入口(详见 §6 HS-01)。
### 2.3 对象级 heal 语义(ecstore `set_disk/ops/heal.rs`
流程(`heal_object_with_explicit_version_regen` :426 起):
1. 取对象写锁(除非 `no_lock`);`object``/` 结尾走对象目录 heal`heal_object_dir_locked` :1587-1717dangling 判定 + `remove` 删除 + 缺 volume 重建)。
2. `read_all_fileinfo` 全盘读 xl.meta,全部 not-found 视为已删除返回。
3. **quorum 仲裁 + ETag 兜底**(已亲验):`list_online_disks` 以 mod-time quorum 为准;quorum 失效时回退 ETag 多数派仲裁(`:525-567` `filter_by_etag`/`quorum_etag`);`pick_valid_fileinfo` 选 canonical 元数据;"meta 坏盘数 > parity" 的 cannotHeal 判定在 ETag 全盘一致时豁免(`:679`)。与 MinIO `filterDisksByETag` 双仲裁一致。
4. `disks_with_all_parts`:562-572)按 `scan_mode` 校验 part**Normal 仅 statCheckParts 语义),Deep 做全量 bitrot 校验(VerifyFile 语义)**Normal 扫描检出 `FileCorrupt` 自动升级 Deep 重试一次(`:2022-2031`,与 MinIO erasure-healing.go:1101-1106 同型);无 parity 对象(EC:0)bitrot 失败判不可恢复(`:700-726`)。
5. `should_heal_object_on_disk`:606-650)逐盘分类 missing/corrupt/offline/outdated → 重建:per-part bitrot reader/writer(用 per-part checksum + 算法)、写临时卷后 rename 提交(`HEAL_RENAME_INCOMPLETE` 重试语义 :24);dangling 删除安全检查 `dangling_delete_safety`:1488);**孤儿数据目录回收 `reclaim_orphan_data_dirs_best_effort`(:1428**——这部分覆盖了 MinIO `CleanAbandonedData` 的主场景(但无独立 `CheckAbandonedParts` API,见 §6 HS-02)。
6. 版本化对象:枚举"每个版本"(`storage.rs:1494-1530`);delete-marker 路径由 `latest_meta.deleted` 决定(`storage.rs:262-277` 注释);回归测试 `tests/heal_b5_versioned_regression_test.rs:282,334`
7. 显式版本重建 `try_regenerate_explicit_version_meta`:1318);transitioned 对象本地残留清理。
8. 写入路径另有 shard 级 bitrot 自校验 `verify_written_bitrot_shards``ops/bitrot_self_verify.rs:45-129`HighwayHash256S,最终 rename 前校验刚写出的 shard,服务 EC:0 无 parity 场景)——**注意这不是后台 bitrot 巡检**;后台巡检由 scanner bitrot_cycle 驱动 Deep heal 承担。
heal crate 侧包装(`task.rs:855-1146`):存在性检查(瞬时错误转 `TransientSkip` 不误判失败 :551-569);scanner 合成目录规范化(:1148-1180);`recreate_missing` 重建(:1183-1282);data-usage-cache 对象锁超时豁免(:571-653);not-found → treated_as_deleted 成功(:1012-1029);结果 `HealResultItem` 保留至多 1024 条 + truncated 标志(:50,845-852)。
递归遍历(`heal_bucket_objects` task.rs:1508-1698):分页枚举全部版本含 delete marker、瞬时错误指数退避重试 ≤32^n + 抖动 :620-627)、失败样本日志截断 ≤5 条、聚合 `BatchHealFailure`
### 2.4 erasure set heal 与断点续扫
`heal_erasure_set`task.rs:2158-2642)四阶段(4 步进度跟踪):
1. **替换意图与恢复盘选择**(仅 AutoHeal + heal_endpoints 非空):复用 durable intent 所在盘 / 排除目标端点选幸存盘;已完成代(CleanupPending)幂等收尾。
2. **格式修复**`heal_replacement_format(dry_run, pool, set, targets)``storage.rs:1372-1384`trait 默认实现 fail-closed);逐目标盘结果必须全 ok(`erasure_healer.rs:97-102`+ 身份围栏复核(task.rs:2410-2420)。
3. **healing 标记**:对目标盘写 owner CAS 标记 `{set_disk_id}:{task_id}``mod.rs:80-229`CAS + 回滚 + 并发唯一 owner),使 `DiskInfo.healing` 为真(已亲验赋值链 `set_disk/mod.rs:4988`)。
4. **逐桶预处理 + 可恢复深扫**`ErasureSetHealer::heal_erasure_set``erasure_healer.rs:242-278`)。
`ErasureSetHealer` 扫描细节(对标 MinIO `healErasureSet``heal_walk.rs:15-23` 模块注释明确引用 MinIO `global-heal.go` 的 listPathRaw + objQuorum=1 + mergeXLV2Versions):
- **枚举器选择(backlog#920**Deep 或 AutoHeal → per-set **disk-walk 并集枚举** `list_versions_for_heal_page_disk_walk`"任意盘上存在"即 sub-quorum 可重建;`storage.rs:1559-1644`,页界 1000 对象/10,000 版本,`dw1:` cursor);普通请求走 read-quorum `list_object_versions`
- **续扫游标**:权威 cursor 为 opaque continuation token`v1:`=marker JSON、`dw1:`=disk-walk key,两命名空间互斥防误读,`storage.rs:81-260`);每完成一页先持久化 cursor 再清 dedup 集合(`erasure_healer.rs:922-927`)。
- **页内并发**FuturesUnordered + Semaphore,默认 `RUSTFS_HEAL_PAGE_OBJECT_CONCURRENCY=8`Deep/AutoHeal 强制 1`erasure_healer.rs:105-142`)。
- **per-version dedup**`compose_key` 长度前缀注入编码(`resume.rs:281-288`)。
- **错误分类**:真缺席(FileNotFound 等)→ Absent(计成功);基础设施瞬时(quorum/DiskNotFound/SlowDown 等)→ Transient(计 skipped);其余 Failed`erasure_healer.rs:148-182`,注释引 backlog#856/#799 B7:离线盘不得记 healed/absent)。
- **防死循环**:空页 truncated 或页尾版本身份不前进即中止(:933-949)。
- **完成判定**failed/skipped/failed_buckets 任一 >0 不标记完成,`schedule_retry()` 复位 resume+checkpoint 两层(:561-626backlog#855/B6/#1033skip 轮不得标记完成)。
- **替换盘提交证据**:目标端点物理回读 `replacement_targets_have_version``ops/heal.rs:340-412`),未确认 → transient skip。
### 2.5 换盘自动修复(replacement recovery
- **识别**`replacement_readiness.rs:25-73`):`replacement_mount_lease_root()` 存在、canonicalize 成功、是挂载点、物理设备 id 非空、与根设备不相交、不与兄弟盘共享物理设备(Linux 用 /proc/self/mountinfo mount-id+dev+ino)。非 root 挂载检查有回归测试(`manager.rs:3549`)。
- **状态机**`resume.rs:63-73`):`Intent → Rebuilding →(写 proofVerified → CleanupPending → 清理``Abandoned` 终态;跨状态迁移先写持久层再变更(`save_state_strict`)。
- **持久化**`resume.rs:38-61`schema ResumeState=5/Checkpoint=5/proof=1):`{task_id}_ahm_resume_state.json``_ahm_checkpoint.json``buckets/ahm-replacement/` 命名空间下 intent/seal/completion_prooftorn write + 无 seal 可识别并原子重建(:1316-1338);CAS 发布、拒绝覆盖并发有效 proof:1512-1585)。
- **恢复**unclean shutdown 与周期扫描都从幸存盘恢复未完成/待清理替换代(`manager.rs:1435-1640,2663-2815`);多代冲突/校验失败 → 冻结该 set(`replacement_recovery_blocked_sets``manager.rs:69-87,2782-2815`)。
- **对外快照**`current_replacement_recovery_snapshot``lib.rs:262-333`)合并本地幸存盘记录,冲突 → Unknown/非 definitiveadmin `GET /v4/heal/replacement-recovery`
### 2.6 调度器(manager.rs
- 优先级堆 + 同优先级 FIFO:148-191,330-347);dedup key 按类型(:469-506);入队三态查重 active→queued→retrying:1759-1785);重复默认 Merged 并返回 canonical task_id`HealAdmissionReceipt`:1821-1846+ client token alias:1219-1246)。
- 容量:队列满时 best-effort 来源(Scanner/AutoHeal/ReadRepair)或低优先级被 Dropped(QueueFull)Admin/Internal 可驱逐低优先级排队项(`push_displacing_lower_priority` :353-396);80%/95% 压力分级(:885-909)。
- 并发:全局 `max_concurrent_heals`(默认 4+ per-set bulkhead `max_concurrent_per_set`(默认 1)(:3040-3073,3434-3447)。
- 前台压力门控 mainline throttle:前台读/写 permit 利用率 ≥80% 时延迟 best-effort 任务(:919-1009,2999-3020)。
- 超时:任务级聚合超时(默认 300s),跨重试保留剩余预算(task.rs:444-451PR #6101)。
- 可恢复重试:`is_recoverable_heal()`error.rs:83-136)≤3 次、2^n 退避封顶 30sretry 在独立 backoff task 中持有所有权(:3235-3382)。
- 完成态保留 10 分钟供查询(:42)。
### 2.7 Admin API 与集群协调
- 路由(`rustfs/src/admin/handlers/heal.rs:174-212`):`POST /rustfs/admin/v3/heal/``/heal/{bucket}``/heal/{bucket}/{prefix}`(同一 POST 按 query `clientToken/forceStart/forceStop` 区分 start/query/cancel,与 mc admin heal 语义对齐);`POST /v3/background-heal/status``GET /v4/heal/replacement-recovery`。权限 `HealAdminAction`route_policy.rs:334-341)。
- 集群协调(heal.rs:771-930 + `node_service.rs:514-606`):`heal_topology_fingerprint` + 按拓扑确定性选 coordinator 节点 + coordinator epochenvelope 校验 + SHA256 digest 重放缓防重放;coordinator 非本机走 peer gRPC `heal_control``probe_heal_control` 能力探测(滚动升级场景)。
- 请求:body 为 `HealOpts``recursive/dryRun/remove/recreate/scanMode(0/1/2)/updateParity/nolock/pool/set`serde camelCase,与 madmin.HealOpts 字段对齐);根 heal start 需 `recursive=true``pool+set` 成对;body 上限 1MB。
- 响应:`HealStartSuccess{clientToken, clientAddress, startTime}``HealTaskStatus{summary, detail, startTime, settings, items, truncated, progress}`summary ∈ running/finished/stopped/notFound);`BackgroundHealStatus`(bitrot 起始时间/周期/当前模式 + `disabled/uninitialized/idle/active/degraded` 状态——peer 不可达显式 degraded 不冒充 idleissue #5850 + `healOperations` 按优先级×来源矩阵 + 集群进度)。
- `HealResultItem`/`HealDriveInfo`/`HealItemType`/DriveState 枚举与 madmin JSON 兼容(`crates/madmin/src/heal_commands.rs:19-65`)。
- 状态 payload 超 8MiB 对折截断(channel.rs:37,73-104);path-token 校验(错误 token 拒绝,空 path 仅匹配 Cluster)。
### 2.8 heal 指标与日志
指标:`rustfs_heal_admission_total{source,result,reason,context}``rustfs_heal_task_start_total``rustfs_heal_task_running{type,set}``rustfs_heal_queue_delay_seconds``rustfs_heal_scheduler_skip_total``rustfs_heal_mainline_throttle_total``rustfs_heal_page_concurrency_current{set}``rustfs_heal_candidate_enqueue/merge/drop/priority_reject_total``rustfs_heal_read_repair_dedup_total{reason}` 等。日志全部结构化 event stylePR #5720);per-object 日志降级防风暴(`demote_to_debug_when!`#5716/#5719/#5727)。
---
## 3. Scanner 已实现功能全景
### 3.1 循环、leader、立即触发
- **集群单 leader**:分布式 ns 写锁 `leader.lock``scanner.rs:3156-3207`,超时默认 5s+ **持久化 leader-epoch CAS 围栏**leader 用 ETag 前置条件向 `.bloomcycle.bin``RSCYC001` 编码的 (cycle, leader_epoch)`scanner.rs:118,1850-1861,2177-2334`);usage 快照再打 epoch fence:2087-2153)。锁丢失 → 取消当前周期,30s 收敛(:108-111,2623-2642)。
- 抢锁后立即执行一轮;周期 = `RUSTFS_SCANNER_CYCLE` > config cycle > start_delay > 部署默认 > 速度档位(±10% 抖动、下限 1s)。
- **clean-idle 指数退避**:连续完整无脏周期间隔 ×2(封顶 24h;bitrot 周期压缩上限;桶有 lifecycle/replication 活动规则禁用,:383-456,1382-1512)。
- **superseded/deferred 退避**:5s 起指数退避封顶 30min:105-106,3432-3438);维护探测失败独立退避(:459-505)。
- **立即唤醒**:① dirty-usage 快路径——写路径 put/delete/multipart/bucket 操作调用 `record_dirty_usage_bucket``scanner_io.rs:222-235`;调用点 `rustfs/src/app/object_usecase.rs:6221` 等),自增 generation 并 Notify 唤醒 leader,脏桶优先排队(`scanner_io.rs:462-488`);② 维护配置变更(lifecycle/replication 设置时 `record_scanner_maintenance_change`);③ 运行时配置热更 generation+Notify;④ 集群活动快照变化。
- **集群协调**`probe_scanner_activity` 汇集本机+peer 的 `ScannerNodeActivity`instance_id/namespace_generation/maintenance_generation/protocol_version/topology_digest/data_movement_active/dirty usage),拓扑摘要覆盖 pools/sets/drives URL,协议版本不齐拒绝共享缓存锁(`scanner.rs:970-1068`);**数据迁移(rebalance/decommission)期间推迟周期**`scanner_io.rs:2226-2374`);周期结束逐 peer RPC 确认 dirty-usage ack`scanner.rs:2925-2952`)。
### 3.2 遍历模型
- 主遍历是**全量目录 walk**tokio::fs::read_dir 递归,`scanner_folder.rs:1915-2234`),不走 metacachemetacache/`list_path_raw` 仅用于 abandoned children 跨盘核查(:2528-2792)。
- 三级并发:leader → per-set(信号量默认 4)→ per-disk 桶扫描(默认 4)→ 单盘递归;每桶每 set 缓存锁 `.scanner-cycle.lock.pool-N.set-M`(锁丢失取消该桶扫描,锁竞争重排队);每盘单扫描准入(本地盘也走信号量,`scanner_io.rs:3246-3274`)。
- 桶顺序:shuffle 后按 dirty → 未缓存 → 已缓存重排(`scanner_io.rs:2947-2949,462-488`);目录内按名字排序 + resume 提示旋转(`scanner_folder.rs:333-359`)。
- **断点续扫**`DataUsageScanCheckpoint{version,resume_after,reason}` 持久于缓存 info`data_usage_define.rs:68,293-307`);预算耗尽/取消写入,恢复有 Used/Stale/NoHint 指标;续扫单位是目录(无跨周期对象级分页)。
- erasure 语义:发现 `xl.meta` 即对象边界不下钻;UUID data-dir 候选最多探测 64 entry;有数据无元数据 → 记 failed + 高优 healsymlink 目录忽略/环跳过。
- 协作让出:每 N 对象(默认 128)`yield_now`
### 3.3 大桶跳过策略(对标 MinIO compaction
1. 缓存当前性复用:桶与扫描计划未变(name/source/snapshot_complete/plan digest/next_cycle/leader_epoch/cache_key_format 全匹配)整桶跳过(`scanner_io.rs:1062-1109`)。
2. compacted 目录 16 周期轮换窗口:`hash mod (next_cycle, 16)` 命中才重扫,否则从旧缓存拷贝(`scanner_folder.rs:74,2429-2442`)。
3. compaction 阈值:子项 <500 或纯对象叶子压缩为单 entry;子文件夹 ≥2500(根 10000)预压缩;children ≥10000 归约(:75-78,2314-2340,2846-2887)。
4. 失败对象 TTL 跳过:86400s/最多 10000 条(:88-91,1354-1381)。
与 MinIO master 对比:MinIO 的跳过策略同样是 hash-mod-16 周期 + compaction 阈值树(500/10000/2500),**bloom filter 已从 master 删除**。RustFS 的常量与结构与 MinIO 现状同源(MinIO 未采用跨盘 dirty-generation 优先,RustFS 额外多两层跳过——plan digest 与缓存当前性校验)。
### 3.4 data usage 统计
- 维度:每目录 entrysize/objects/versions/delete_markers/大小直方图/版本直方图/复制统计/failed_objects/per-tier stats/children/compacted`data-usage/src/data_usage.rs:661-679`);每对象 SizeSummary(含 per-ARN 复制目标统计、tier 统计,tier 分类:transitioned 完成记入其 tier 否则按 storage classfree version 不计);桶级 `BucketUsageInfo`;集群级 `DataUsageInfo`(含 scanner_cycle/scanner_epoch 围栏 + usage_snapshot_complete)。
- 存储:每桶每 set `{bucket}/.usage-cache.bin`(主 + `.bkp` 备份 + CAS 重试);权威集群快照 `buckets/data-usage/data-usage.json`(每 10 周期同步 `.bkp`,legacy 路径兼容);陈旧快照拒绝写入(epoch/cycle/last_update 三重判定);被竞争 superseded 的观测快照另存 `data-usage-observed.json`
- 消费:`replace_bucket_usage_memory_from_info` 刷新桶用量内存 + 两层缓存失效(`scanner.rs:4142-4152`)→ bucket stats/quota/admin account_info/system;写路径内存实时叠加 overlay;启动读快照判断冷缓存跳过启动延迟。
- 未完成 multipart 不参与统计(与 MinIO 一致,MinIO 也不扫 multipart 桶)。
### 3.5 ILM 集成
- 每对象 `ScannerItem::apply_actions``scanner_folder.rs:747-1032`):`Evaluator::new(lifecycle).with_lock_retention(...).with_replication_config(...).eval()` 批量评估。
- 已实现动作(IlmAction 全集,`scanner-contracts/src/metrics.rs:34-45`):expiry 删除(Delete/DeleteRestored/DeleteRestoredVersion)、全版本删除(DeleteAllVersions/DelMarkerDeleteAllVersions,处理后停止后续版本)、transitionTransition/TransitionVersiontier 列表运行时读取)、noncurrent 批量(DeleteVersionAction → `enqueue_by_newer_noncurrent`)、free-version 清理(`enqueue_free_version`)、object-lock retention 约束。**与 MinIO 的 9 个 ILM 动作一一对应**。
- 执行模型:scanner 是"发现与入队"角色(expiry 队列/transition 队列在 ecstore `bucket_lifecycle_ops.rs`),动作由 worker 池消费——与 MinIO globalExpiryState/globalTransitionState 同型。
- AbortIncompleteMultipartUpload 不在 scanner/ILM 内执行(MinIO 同样不在:`internal/bucket/lifecycle/rule.go` 有 FIXME,实际由 `erasureSets.cleanupStaleUploads` 全局例程承担);RustFS 由 ecstore 独立后台任务 `init_background_stale_multipart_upload_cleanup``bucket_lifecycle_ops.rs:3289-3320`+ 桶删除时 on-demand。
- 集成测试覆盖:transition+restore、free-version、noncurrent、delete-marker、0-day、后台扫描过期(`scanner/tests/lifecycle_integration_test.rs:1071-2095`)。
### 3.6 heal 候选生产(scanner 侧)
- 抽样:`hash mod_alt(next_cycle/prob_div, 1024/prob_div)`,进入 compacted 分支重扫时 prob_div=16 等效概率 ×16(与 MinIO 同款补偿,`scanner_folder.rs:125-127,2117-2122`)。
- deep/normal:周期级 `get_cycle_scan_mode`bitrot_cycle 默认 30d`scanner.rs:1626-1657`)→ 对象级带 `HealScanMode::Deep`;新鲜对象(60s 内修改)降级 Normal:146-155);状态持久 `.background-heal.json``BackgroundHealInfo{bitrot_start_time,bitrot_start_cycle,current_scan_mode}`,与 MinIO 同路径同结构)。
- scanner 只入队不内联执行(内联 heal 已移除,兼容旗标仅告警,`scanner_folder.rs:411-427`);`HealScanMode::Deep` 只是标记,bitrot 校验读发生在 heal 消费端(ecstore Deep 路径)。
- 元数据损坏 → 高优 heal`classify_get_size_failure` → HealMetadata);abandoned children → list_path_raw quorum 核查 + 桶级/对象级高优 heal;healing 盘粘性跳过(`should_heal` :1628-1648)。
- pending-heal 账本:heal 通道满被拒持久化到缓存 info,下轮重试。
- 复制 heal`queue_replication_heal` → replication 队列(走 replication 通道而非 heal channel);per-ARN 复制用量统计。
### 3.7 remote_scanner RPC 协议(RustFS 特有)
请求 ≤16KB msgpackversion/request_id/server_epoch/session_id/session_sequence/bucket/next_cycle/leader_epoch/scan_plan_digest/skip_healing/scan_mode/budget);帧 ≤2MB、HMAC-SHA256 逐帧认证(域 `rustfs-ns-scanner-frame-v3`);进度心跳 1s(预算模式 250ms);阶段播报 Scanning→PersistingRPC 生命周期上限 24h、断连宽限 2min;防重放 session+sequence 缓存(容量 65536);服务端校验 leader fence 与持久化 cycle 一致 + 每 5s fence 复验;结果 Complete/Partial/NamespaceNotFound/CycleAhead;不支持 v4 协议的远端盘回退 leader 本地扫描(`remote_scanner.rs` 全文件;`scanner_io.rs:2750-2812`)。
### 3.8 限速/预算/热更/观测
- DynamicSleeper 比例退避(速度档 fastest/fast/default/slow/slowest,同 MinIO 五档参数);idle_mode 总闸;前台 S3 读流量每请求 10ms 封顶 250ms 额外退避。
- 周期预算 ScannerCycleBudgetmax_duration/max_objects/max_directories(默认 0=不限),partial 周期仍推进 cycle 计数。
- runtime_config 三层来源(env > config > default)逐字段来源标记(Env/Config/ScannerCompatConfig/Default),admin `PUT /v3/config` 热更 → generation+Notify 即时生效;`GET /v3/scanner/status` 返回 enabled/freshness(fresh/stale/unknown)/metrics/cycle_schedule/runtime_config`GET /v3/ilm/expiry/status` 返回 expiry 队列/worker/missed/blocked。
- 指标:leader lock、周期 complete/partial/deferred/superseded、versions scanned、per-sourceUsage/Lifecycle/BucketReplication/SiteReplication/Heal/Bitrot/Alertschecked/executed/queued/missed、checkpoint set/used/stale、当前路径(per-disk+bucket 实时)、缓存 save 系列、并发系列、告警(excess versions/version size/folders)。
---
## 4. 与 MinIO 逐项对标
### 4.1 heal 触发通道对照
| MinIO 通道 | RustFS 对应 | 状态 |
|---|---|---|
| A. 手动 admin healhealSequenceclientToken/forceStart/forceStop | heal channel Start/Query/Cancel + 集群 coordinator + envelope 重放防护 | ✅ 等价且增强(集群路由);序列语义差异见 §6 HS-06 |
| B. 常驻后台 heal 队列(newBgHealSequence + healRoutine worker 池) | HealManager 常驻调度器 + 优先级队列 + bulkhead | ✅ 等价且增强 |
| C. 新盘/换盘自动 resyncmonitorLocalDisksAndHeal 10s + healFreshDisk + healingTracker + waitForFormatErasure 握手) | auto disk scanner10s+ replacement_readiness + durable intent/proof 状态机 + heal_replacement_format | ✅ 等价且增强(identity fence + completion proofMinIO 的 tracker 面向对外可见性更强,见 §6 HS-07) |
| D. MRF(队列 100k + 持久化 list.bin + shutdown 回放 + 读路径 corrupt 投递) | read-repairLow+TTL 去重)+ 写路径 convergence heal 部分承担;`HealType::MRF` 执行体无生产入口 | ⚠️ 部分等价(§6 HS-01) |
| E. Scanner 抽样 heal1/1024 + compacted ×16 补偿)+ abandoned children | 同款抽样 + ×16 补偿 + abandoned children + pending-heal 账本 | ✅ 等价且增强(账本) |
| F. 读路径内联触发 → MRFGetObject part 缺失/损坏、元数据重建 missingBlocks>0 | read repairmissing_shards/decode_error/metadata_read_error 三入口) | ✅ 等价(入 heal 队列而非 MRF 队列) |
### 4.2 对象级 heal 语义对照
| 特性 | MinIO | RustFS | 状态 |
|---|---|---|---|
| mod-time quorum 仲裁 | listOnlineDisks | 同 | ✅ |
| ETag 多数派兜底(时钟漂移) | filterDisksByETag | `filter_by_etag`/`quorum_etag`heal.rs:525-567 | ✅ 已亲验 |
| cannotHeal 的 ETag 豁免 | ETag 全一致豁免重试 | heal.rs:679 | ✅ |
| Normal=CheckPartsstat/ Deep=VerifyFilebitrot | 是 | `disks_with_all_parts` 按 scan_modeops/heal.rs:562-572,978-1024 | ✅ |
| Normal 检出 corrupt 自动升 Deep 重试一次 | erasure-healing.go:1101-1106 | ops/heal.rs:2022-2031 | ✅ |
| dangling 判定(not-found > parity+ 删除审计 | isObjectDangling/deleteIfDangling | `dangling_delete_safety`:1488+ scanner HEAL_DELETE_DANGLING | ✅(审计 tags 细节有差异) |
| 孤儿 data-dir/inline 清理(CleanAbandonedData | CheckAbandonedPartsscanner 抽中 + admin Remove 时显式调用) | heal 路径内 `reclaim_orphan_data_dirs_best_effort`:1428);独立 API 三层 NotImplemented | ⚠️ 部分等价(§6 HS-02 |
| 版本化/delete-marker heal | HealObject versionIDnullVersionID 特判 | 逐版本枚举 + delete-marker latest healB5 回归) | ✅ |
| 对象级 healing 元数据标记(x-minio-healingRenameData 跳过版本清理) | 有 | 无对象级标记;依赖盘级 healing.bin + NSLock + rename 语义 | ⚠️ 评估项(§6 HS-12) |
| Distribution/Index 一致性三处防线 | 有(manual modification 拒绝) | 目标盘格式结果全 ok 校验 + 身份围栏 | ✅(粒度不同) |
| 无 parityEC:0)对象 | bitrot 不可恢复处理 | 判不可恢复(:700-726)+ 写入自校验 | ✅ 增强(写路径自校验) |
| 三层分布不一致拒绝 heal | 有 | heal_walk 归一化 + 页界防御 | ✅(实现方式不同) |
| multipart 孤儿对账 | CheckAbandonedParts 承担 | 显式 NotImplemented(由 lifecycle 清理承担) | ⚠️ §6 HS-02 |
| suspended/decommissioned pool 处理 | IsSuspended 跳过 | deferral 语义(store/heal.rs:192-207PR #5876 | ✅ |
| heal 与并发删除互斥 | NSLock + healing 标记 | NSLock + 写锁 | ✅ |
### 4.3 新盘 resync 对照
| MinIO | RustFS | 状态 |
|---|---|---|
| waitForFormatErasure 四类可恢复错误无限等待握手 | startup 盘解析 + renew_disk 重连路径 | ✅(模型不同:RustFS 不在启动时阻塞等待 format) |
| HealFormat NSLock + errNoHealRequired + refFormat 不一致拒绝 | `heal_format`/`heal_replacement_format` fail-closed + 目标槽位限定(PR #1787 语义) | ✅ 增强 |
| per (pool,set) 分布式锁防并发 resync | set 级队列去重 + bulkheadmanager.rs:2854-2889 | ✅ |
| 全新集群检测(待 heal 盘数==总盘数不触发) | replacement_readiness(独立挂载点/物理设备校验,非 root) | ✅ 增强 |
| healingTracker.healing.binBytes/Items 计数、QueuedBuckets/HealedBuckets、Resume 快照、RetryAttempts ≤4、HealID 联动、diskID 变更重置) | resume/checkpoint schema 化持久层 + durable intent/proofper-task 文件,CAS) | ✅ 等价且增强(崩溃窗口补齐);但**对外快照可见性**弱于 MinIO(§6 HS-07 |
| 跳过 heal 开始后新写入版本(ModTime > Started | 无同款过滤 | ⚠️ §6 HS-13 |
| 跳过 ILM 已过期版本(filterLifecycle | 无同款过滤 | ⚠️ §6 HS-13 |
| worker 数 max(GOMAXPROCS,NR)/4 下限 4heal:drive_workers 覆盖 | 页内并发 8Deep/AutoHeal 强制 1+ per-set bulkhead | ✅(参数模型不同) |
| 每 entry waitForLowHTTPReq 让路 | mainline throttle(前台利用率门控) | ✅ 增强 |
| heal 范围含 `.minio.sys/config``.minio.sys/buckets` 两个伪桶;最新桶优先 | ErasureSet 任务逐 bucket 预处理(含 meta bucket 语义由 heal_bucket 承担) | ✅(顺序无"最新优先" |
| 失败整体重试 ≤4 次(resetHealing + errRetryHealing | schedule_retry 复位双层 + 可恢复重试 ≤3 | ✅ |
### 4.4 scanner 对照
| MinIO | RustFS | 状态 |
|---|---|---|
| 集群单 leaderglobalLeaderLock | leader.lock + 持久化 leader-epoch CAS 围栏 | ✅ 增强(epoch 围栏防脑裂,MinIO 无持久化 epoch |
| `.bloomcycle.bin` 只存 cyclebloom 已删除) | 同路径存 cycle+leader_epochRSCYC001 | ✅ 对齐(v1 误判已修正) |
| folderScanner hash-mod-16 + compaction500/10000/2500 | 同款常量 + plan digest + 缓存当前性校验 + dirty 优先 | ✅ 增强 |
| 每盘扫描并行 ≤GOMAXPROCShealing 盘排除 | per-set/per-disk 信号量 + healing 盘粘性跳过 | ✅ |
| scannerSleeperfactor 2/max 1sspeed 档热更) | DynamicSleeper 同款 + idle_mode + 前台读退避 | ✅ 增强 |
| idle 语义:`scanner:idle_speed=on`(空闲时段才节流,忙时全速) | `RUSTFS_SCANNER_IDLE_MODE=true`(启用限速总闸) | ⚠️ 语义方向相反,§6 HS-14 |
| applyActions 顺序(heal→ILM→复制→告警) | apply_actions 同序(heal 候选→ILM→复制 heal→告警) | ✅ |
| ILM 9 动作 + 批量评估 + DeletePrefixObject 优化 | 同 9 动作 + 批量评估 + expiry 队列 | ✅(DeleteAllVersions 是否单调用优化未逐行核) |
| abandoned childrenlistPathRaw minDisks=N/2 发现漏写盘) | list_path_raw + quorum 核查 + 高优 heal | ✅ |
| incomplete multipart 独立例程(6h 间隔/24h 过期,rename 进 .trash | ecstore 独立后台任务(可配间隔/过期) | ✅(trash 二段清理细节差异,§6 HS-18) |
| usage 维度(size/objects/versions/DM/直方图/复制/tier/bucket 级) | 全覆盖 + 集群快照三重防回退 | ✅ 增强 |
| prefix 级 usageloadPrefixUsageFromBackendconsole 消费) | 缓存内有目录树但仅 flatten 桶级 | ❌ §6 HS-08 |
| 超限事件 s3:ObjectManyVersions/LargeVersions/PrefixManyFolders + 审计 | 仅指标 alert_excess_*(默认 100/1TiB/65538 vs MinIO 100/1TB/50000 | ⚠️ §6 HS-04/HS-17 |
| scanner 指标 v3bucket_scans/directories/objects/versions/last_activity | rustfs_scanner_* 全套 + freshness | ✅(命名体系不同) |
| TraceScanner / realtime metricsmc admin scanner status/trace | 无 trace 通道;/v3/scanner/status 自有结构 | ⚠️ §6 HS-03 |
### 4.5 admin/CLI/API 面对照
| MinIO | RustFS | 状态 |
|---|---|---|
| `POST /minio/admin/v3/heal/...` start/status/cancel | `POST /rustfs/admin/v3/heal/...` 同三态 | ✅(路径前缀不同属预期) |
| `HealStartSuccess`/`HealTaskStatus`/`HealResultItem`/DriveState | 同名字段 JSON 兼容 | ✅ |
| `POST /v3/background-heal/status`BgHealState 聚合) | 同路径 + degraded 语义 + operations 矩阵 | ✅ 增强(MRF per-endpoint 子状态无,因无 MRF |
| `GET /v3/healthinfo` 每 drive `HealInfo *HealingDisk` | 无同款 healthinfo heal 字段(replacement-recovery v4 承担部分) | ⚠️ §6 HS-07 |
| madmin 客户端 HealStart/HealStatus/BackgroundHealStatus/ScannerStatus 方法 | 仅 wire 类型,无客户端方法 | ❌ §6 HS-05 |
| mc admin heal --pool/--set、--scan-mode、--force-start/stop | HealOpts 全字段支持(pool/set/scanMode/forceStart/forceStop | ✅(服务端就绪;缺 mc 侧入口,HS-05) |
| ErrHealAlreadyRunning / ErrHealOverlappingPaths 类型化错误 | 去重合并 + 驱逐语义;无类型化重叠拒绝 | ⚠️ §6 HS-06 |
| 结果 backpressuremaxUnconsumedItems=1000、10s 保活流式、24h 未消费 abort) | 快照式查询(1024 条 + 8MiB 截断 + 10min 保留) | ⚠️ §6 HS-06 |
| `mc support inspect`/healing-bin 离线 dump | 无(inspect.rs 存在但 healing dump 未确认) | ⚠️ P3 |
### 4.6 观测面对照
| 维度 | MinIO | RustFS | 状态 |
|---|---|---|---|
| heal 指标 | minio_heal_objects_total/heal_total/errors_total/time_last_activity + v3 drive_health 2=healing | rustfs_heal_* 全套(admission/queue delay/running/throttle/page concurrency | ✅(RustFS 缺 drive_health=healing 单一 gauge 等价物;DiskInfo.healing 已赋值) |
| scanner 指标 | v3 6 个 + realtime 18 项 | rustfs_scanner_* 全套 + per-source 维度 | ✅ |
| ILM 指标 | v3 5 个(expiry/transition pending/active/missed + versions_scanned | ilm expiry status API + scanner per-source | ✅(指标与 API 形态不同) |
| trace | TraceHealing/TraceScanner 两通道 | 无 | ❌ §6 HS-03 |
| 审计 | HealObject 事件、dangling 删除审计、scanner:manyversions 等 | 结构化日志(event style+ 指标;无 audit log 事件 | ⚠️ §6 HS-04 |
| 进度 | healingTracker Bytes/Items/QueuedBuckets/当前对象 + usage-cache 总量基线 | HealProgress{scanned/healed/failed/bytes/current_object/percentage}bytes_processed 注释为 0、estimated_completion_time 恒 None | ⚠️ §6 HS-07 |
### 4.7 配置面对照(默认值)
| MinIO | RustFS | 备注 |
|---|---|---|
| `heal:bitrotscan`(默认 offon=每轮;Nm=N×30×24h | `heal.bitrot_cycle` / `RUSTFS_SCANNER_BITROT_CYCLE_SECS`(默认 30d=2592000s0/on=每轮 Deepoff=禁用) | ✅ 同语义(RustFS 默认 30dMinIO 默认 off——**默认值不同**RustFS 更激进) |
| `heal:max_io=100`/`max_sleep=250ms`waitForLowIO | mainline throttle 阈值 80%/80%、max_sleep 250ms | ✅ 同型(阈值模型不同) |
| `heal:drive_workers`(默认 -1 自动) | 页内并发 8 + per-set 1 | ✅ 同型 |
| `_MINIO_HEAL_WORKERS`GOMAXPROCS/2 | `RUSTFS_HEAL_MAX_CONCURRENT_HEALS=4` + `_MAX_CONCURRENT_PER_SET=1` | ✅ |
| `_MINIO_AUTO_DRIVE_HEALING`on | `RUSTFS_HEAL_AUTO_HEAL_ENABLE=true` | ✅ |
| `_MINIO_SCANNER`on | `RUSTFS_SCANNER_ENABLED=true` | ✅ |
| `scanner:speed` 五档(default=2x/1s/1m | 同五档同名同参数 | ✅ |
| `scanner:idle_speed`on | `RUSTFS_SCANNER_IDLE_MODE`(true | ⚠️ 语义方向(HS-14 |
| `scanner:alert_excess_versions=100` | 100 | ✅ |
| `scanner:alert_excess_folders=50000` | 65538(兼容 PBS 布局) | ⚠️ HS-17 |
| `ilm:expiration_workers=100`/`transition_workers=100` | ecstore expiry/transition worker 池(键见 ilm 子系统) | ✅(默认值未逐项核对) |
| `api:stale_upload_cleanup_interval=6h`/`expiry=24h` | ecstore 后台任务 env 可配 | ✅(默认值未逐项核对) |
| —(无) | `RUSTFS_HEAL_QUEUE_SIZE=10000``_TASK_TIMEOUT_SECS=300``_INTERVAL_SECS=10``_LOW_PRIORITY_MERGE/DROP``_PAGE_*``_SET_BULKHEAD``_MAINLINE_*``RUSTFS_SCANNER_CYCLE_MAX_*` 预算、`_MAX_CONCURRENT_SET/DISK_SCANS=4``_YIELD_EVERY_N_OBJECTS=128` 等 | RustFS 特有(更细粒度) |
### 4.8 RustFS 超出 MinIO 的部分
1. remote_scanner RPC(扫描执行下放远端 peer 本地,含 HMAC 认证/重放缓存/fence 复验/断连宽限)。
2. 持久化 leader-epoch CAS 围栏 + usage 快照 epoch/cycle 防回退(MinIO 仅锁,无持久 epoch)。
3. 周期预算(max_duration/objects/directories+ partial 周期推进语义。
4. per-set/per-disk 扫描并发闸 + 每桶每 set 缓存锁。
5. pending-heal 账本(heal 通道满不丢候选)。
6. 换盘 durable intent + completion proof 状态机 + 身份围栏(MinIO healingTracker 无 proof)。
7. mainline throttle 前台压力门控(permit 利用率驱动)。
8. 集群 heal control coordinator + envelope 重放防护 + degraded 显式降级。
9. 写路径 shard bitrot 自校验(EC:0 场景)。
10. dirty-usage 快路径唤醒(写路径即时通知 + 脏桶优先)。
11. heal 运行时可观测矩阵(优先级×来源 operations snapshot)。
12. workload admission 联动(heal 调度器读前台压力快照)。
---
## 5. 差距与改进清单
分级定义:P1=行为/运维对齐缺口(影响生产运维或工具链兼容);P2=完善性(功能在但缺一角);P3=清理/低风险。每项含现状证据、MinIO 行为、影响、建议、验收方式。
### P18 项)
**HS-01 MRF/ECDecode/Metadata 三类 heal 任务无生产触发入口,HealEvent 未接线**
- 现状:`HealType::MRF/ECDecode/Metadata` 执行体完整(task.rs:1700-2156)但全仓库无生产触发方;`HealEvent`/`HealEventHandler`event.rs:50-367crate 外零引用(已亲验 grep);channel 转换只产生 Cluster/Object/Bucket/Prefix/ErasureSetchannel.rs:566-601)。
- MinIOmrf.go 独立 MRF 队列(容量 100k,满丢弃计数)、进程退出 msgp 持久化 `.heal/mrf/list.bin` + 启动回放、入队 <1s 延迟 1s(等网络恢复)、healSleeper 限速;读路径 GetObject part 缺失/损坏、元数据重建 missingBlocks>0、Put 部分成功、DeleteObject、multipart、peer client 共 7+ 投递点。
- 影响:RustFS 的 read-repair + 写路径收敛覆盖了主场景,但缺少:① 事件驱动的 Urgent ECDecode 重建入口(ecstore 解码失败时目前仅 Low read-repair);② Metadata-only heal 入口(scanner HealMetadata 分类存在但走普通对象 heal);③ MRF 队列持久化(重启丢未消费修复意图——scanner pending-heal 账本部分缓解)。
- 建议:三选一决策——(a) 接线 HealEvent(在 ecstore 解码失败/metadata 损坏点发事件)+ 实现持久化重试账本;(b) 删除 MRF/ECDecode/Metadata 死代码只保留文档说明;(c) 保留执行体、把 HealEvent 降级为内部 API。推荐 (a) 但需先量化 read-repair 是否已覆盖解码失败场景的响应时间要求。
- 验收:解码失败 → Urgent heal 请求链路 e2e;重启后 pending 修复意图回放;HealEvent 环形缓冲指标。
**HS-02 CheckAbandonedParts 三层 NotImplementedabandoned data 独立对账入口缺失)**
- 现状:`set_disk/ops/heal.rs:2052-2056``core/sets.rs:1144-1148``store/heal.rs:258-266` 三层显式 `Err(NotImplemented)`(已亲验),注释"intentionally retained above the set layer until there is a concrete caller"。
- MinIO`CheckAbandonedParts` → 每盘 `CleanAbandonedData`:读 xl.meta → 列 UUID data-dir + inline entries → 与 getDataDirs 差集 → 删多余 data-dir/inline 并重写 xl.meta;由 scanner 抽中 heal 与 admin heal Remove 时显式调用。
- 影响:RustFS heal 路径内 `reclaim_orphan_data_dirs_best_effort`:1428)覆盖"heal 时回收孤儿目录",但 ① 无独立触发点(MinIO 在对象未到 heal 阈值时也能清 abandoned data);② inline data 孤儿条目清理未确认;③ multipart 孤儿对账明确不做(设计决定,由 lifecycle 承担)。
- 建议:评估把 `reclaim_orphan_data_dirs_best_effort` 提升为 heal_object 固定步骤(若尚非)+ 实现 HealOperations::check_abandoned_parts 真实现(调用同一回收逻辑),或明确文档化"由 lifecycle 承担"并关闭 API 面。
- 验收:构造 data-dir/inline 孤儿 → scanner 抽样/admin heal 后被清理;三层 API 返回成功或显式 NotSupported 文档化。
**HS-03 heal/scanner trace 通道缺失**
- 现状:TraceHealing/TraceScanner 零命中(已亲验 grep 全仓库)。
- MinIO`madmin.TraceHealing`mc admin trace --healingFuncName=heal.Bucket/heal.Object/heal.CheckAbandonedParts,带 dry/remove/mode/version-id/disks/bytes)、`TraceScanner`mc admin scanner trace,支持 --filter-size/--response-duration)。
- 影响:无法实时观测单个 heal/scanner 动作的耗时与参数;排障只能靠指标聚合与日志。
- 建议:在 heal channel 执行与 scanner folder/item 处理埋点,接入现有 admin trace 订阅面(若 rustfs 已有 trace 基建则复用,无则按 madmin TraceType 扩展)。
- 验收:mc 等价工具能订阅 heal/scanner trace 流。
**HS-04 scanner 超限 S3 事件与审计缺失**
- 现状:仅 `rustfs_scanner_excess_*_total` 指标(versions 100/version size 1TiB/folders 65538)。
- MinIO:发 `s3:ObjectManyVersions`>100 版本)、`s3:ObjectLargeVersions`(累计 >1TB)、`s3:PrefixManyFolders`>50000 子目录)事件(UserAgent: Scanner+ scanner:manyversions/largeversions/manyprefixes 审计。
- 影响:依赖事件订阅做容量治理的用户(console/外部审计)收不到告警。
- 建议:scanner_folder 告警点接入 notify 事件发布(复用 lifecycle 事件通道语义)。
- 验收:配置桶通知后超限对象触发事件。
**HS-05 madmin 客户端方法缺失**
- 现状:`crates/madmin/src/heal_commands.rs` 只有 wire 类型(HealDriveInfo/Infos/HealResultItem);无 HealStart/HealStatus/BackgroundHealStatus/ScannerStatus 客户端方法。
- MinIOmadmin-go 提供完整客户端;mc admin heal/scanner/status/trace 都建立在上面。
- 影响:mc 等管理工具无法直接对接 RustFS heal/scanner 管理面;自动化运维只能手写 HTTP。
- 建议:按 madmin-go 接口形状补客户端(服务端已就绪,纯客户端工作)。
- 验收:用 madmin 客户端完成 start→query→cancel 全流程。
**HS-06 admin heal 序列语义与 MinIO 差异**
- 现状:重复/重叠请求被去重合并(返回 canonical task_id)或驱逐;无 ErrHealAlreadyRunning/ErrHealOverlappingPaths 类型化错误(已亲验:manager.rs:1309 的 already_running 是幂等启动保护,非 admin 语义);结果为快照式查询(1024 条/8MiB 截断/10min 保留),非 MinIO 的流式增量(clientToken 拉增量 + maxUnconsumedItems=1000 backpressure + 10s 保活 + 24h 未消费 abort)。
- 影响:mc admin heal 的交互模型(长连接拉增量)对 RustFS 表现为多次快照轮询;自动化脚本难以区分"已合并"与"新启动"。
- 建议:① 增量语义:channel query 支持自上次 clientToken 起的 items 增量(或 cursor);② 重叠请求返回类型化错误码(或 receipt 中显式 merged_into 字段——现有 alias 机制已有基础);③ forceStart 先停旧再启新语义核对。
- 验收:madmin 兼容客户端按 MinIO 模式轮询能取得全量 items。
**HS-07 healing 进度与盘级 healing 状态对外可见性不足**
- 现状:bytes 恢复进度 `progress.bytes_processed = 0 // set to 0 for now`erasure_healer.rs:967);`HealProgress::estimated_completion_time` 恒 None、`HealStatistics::add_healed_objects` 未写入(progress.rs:38,135-139 零调用);healthinfo 无每盘 HealInfo 等价(MinIO HealingDiskBytesDone/Failed/Skipped、ObjectsTotal 基线、QueuedBuckets/HealedBuckets、Resume 快照、当前 object);v3 指标无 drive_health=2(healing) 单一 gauge 等价。
- 影响:换盘重建(可能数小时~天)期间运维无法回答"进行到哪/还剩多少/预计何时完成"。
- 建议:① erasure set heal 统计 bytesheal_object 返回对象大小已可得);② 从 usage-cache 读对象总量基线(MinIO 同款做法);③ admin healthinfo/背景状态暴露每盘 healing 快照(DiskInfo.healing 已有,补聚合暴露);④ ETA 由基线+速率推导。
- 验收:换盘重建中 admin 可见 bytes 进度与 ETAmc info 等价输出 Healing 标志。
**HS-08 prefix 级 usage 未暴露**
- 现状:DataUsageCache 内目录树 entry 存在(hash_path 组织),但 `dui()` 只 flatten 到桶名(data_usage_define.rs:858-915)。
- MinIO`loadPrefixUsageFromBackend`30s cache)从每 set `.usage-cache.bin` 聚合 prefix usageconsole 桶前缀统计消费。
- 影响:console/前端无法展示前缀级用量;大桶定位"哪个前缀占空间"无 API。
- 建议:实现 flatten 前缀查询 API(数据已在缓存内,纯聚合与暴露工作)。
- 验收:ListBuckets/PrefixUsage API 返回与前缀过滤匹配的统计。
### P29 项)
**HS-09 get_disk_status 恒返回 Ok(唯一 TODO**`crates/heal/src/heal/storage.rs:930-943`(已亲验)。当前无生产调用方(低风险)。建议:删除该方法或接 ecstore disk 状态真实现(DiskStatus 枚举已定义)。
**HS-10 HealStorageAPI 约 1/3 方法为死代码**get_object_meta/get_object_data/put_object_data/delete_object/verify_object_integrity/ec_decode_rebuild/get_disk_status/format_disk/heal_bucket_metadata/get_object_size/get_object_checksum/list_objects_for_heal(非分页版,自带 memory_heavy 警告)均 0 调用方。建议:随 HS-01 决策一并清理或接线(死接口误导后续维护者以为存在调用路径)。
**HS-11 bitrot 自检缺失**MinIO 启动时 bitrotSelfTest 对四算法已知向量自检失败即 Fatal(防静默数据损坏)。RustFS 无等价(已亲验 grep)。建议:启动时对 HighwayHash256S 等在用算法做已知向量自检(低成本高价值)。
**HS-12 对象级 healing 元数据标记评估**MinIO heal 期间对象打 `x-minio-healing:true`RenameData 据此跳过版本清理/legacy purge(漏掉会导致 heal 与并发删除互毁)。RustFS 无对象级标记(已亲验 grep object.rs 无 healing 分支),依赖 NSLock + rename 语义。建议:审计 RustFS rename 提交路径是否存在"heal 提交与并发 delete/version 清理竞争"窗口;若无则文档化差异,若有则补标记等价机制。
**HS-13 erasure set heal 无"跳过新写入/ILM 已过期版本"过滤**MinIO resync 跳过 ModTime>tracker.Started 的版本(避免 heal 追新写入尾巴)与 ILM 已过期版本(避免白做)。RustFS erasure_healer 未实现同款过滤(按版本 dedup 有,时间/ILM 过滤无)。影响:重建尾部长尾(持续写入的桶 heal 完成判定被新版本推迟)与无效 heal 工作量。建议:disk-walk 枚举处加 started_at 时间过滤 + evaluator 预检。
**HS-14 scanner idle 语义方向与 MinIO 相反**MinIO `scanner:idle_speed=on`(默认)= 集群空闲时才节流、忙时全速;RustFS `RUSTFS_SCANNER_IDLE_MODE=true`(默认)= 限速总闸(false=完全不休眠)。两者默认行为可能相近(都限速)但参数语义不可互换,迁移文档需显式说明;若追求 mc config 兼容需重命名/重语义。建议:先文档化差异,评估是否对齐语义。
**HS-15 alert_excess_folders 默认值差异**RustFS 65538(兼容 PBS/Proxmox 布局,scanner_folder.rs:79vs MinIO 50000。行为差异默认即触发阈值不同。建议:文档化(保留 65538 有本地理由)。
**HS-16 单机默认周期钩子未启用**:`single_disk_default_cycle_secs(_features) -> None` 恒空(scanner.rs:1428-1430),单机部署无专属默认周期覆盖。建议:决定单机默认周期策略后启用或删除钩子。
**HS-17 DeleteAllVersions 批量优化核对**MinIO 用 DeletePrefix+DeletePrefixObject 单调用代替逐版本 fan-out。RustFS expiry 队列路径是否同款优化未逐行核实(集成测试覆盖行为正确性)。建议:核对 `apply_expiry_rule` 全版本删除路径,若无前缀单调用优化则评估补齐。
### P33 项)
**HS-18 trash/临时目录二段清理细节核对**:MinIO `.minio.sys/tmp/.trash` 清理(delete_cleanup_interval 默认 5m + deleteCleanupSleeper)与 stale uploads rename-into-trash 二段式。RustFS 有 delete_tail_activity.rs 与 stale multipart 任务,二段语义是否完整对齐未逐行核实。建议:对照补齐或文档化。
**HS-19 root heal 直连死路径清理**`should_handle_root_heal_directly` 恒 falseadmin/handlers/heal.rs:1200-1202,测试锁定),store.heal_format 直连分支不可达。建议:删除死分支或恢复直连路径作为集群协调失败的降级。
**HS-20 兼容旗标与死指标清理**:`RUSTFS_SCANNER_INLINE_HEAL_ENABLE`(开启仅告警)+ `rustfs_scanner_inline_heal_total` 死指标 + `rustfs_common::metrics` 中 scanner 域代码分层迁移(backlog #1843 已登记)。建议:随分层迁移一并清理。
### 按设计不追平(7 项,记录以防后续误判为缺口)
1. **bloom filter**MinIO master 已删除;RustFS `.bloomcycle.bin` 复用为 cycle/epoch 围栏与 MinIO 现状一致。
2. **scanner 集群单 leader**:双方一致;RustFS 额外有 epoch 围栏。
3. **heal 不发 S3 bucket notification**:双方一致(heal 结果走 admin status)。
4. **incomplete multipart 不在 scanner/ILM 内执行**:双方一致(独立后台例程)。
5. **内联 heal 移除**RustFS 有意为之(scanner 只入队),MinIO 的 applyHealing 内联路径不做对标。
6. **heal 序列常驻保活(10s 空白回写)**:RustFS 快照式查询模型不同,按 HS-06 处理增量语义即可,不复制流式保活。
7. **`.trash`/`tmp-old` 路径名兼容**:RustFS 布局常量独立,不逐字对齐 MinIO 路径。
---
## 6. 配置默认值总表(RustFS
healenv 前缀 `RUSTFS_HEAL_``crates/config/src/constants/heal.rs`,消费于 `manager.rs:724-800`):
| 配置 | 默认 | 热更新 |
|---|---|---|
| AUTO_HEAL_ENABLE | true | 否 |
| QUEUE_SIZE | 10000 | 否 |
| INTERVAL_SECS | 10 | 否(启动时固定) |
| TASK_TIMEOUT_SECS | 300 | 否 |
| MAX_CONCURRENT_HEALS | 4 | 否 |
| MAX_CONCURRENT_PER_SET | 1(≤min(全局,值) | 否 |
| LOW_PRIORITY_MERGE_ENABLE | true | 否 |
| LOW_PRIORITY_DROP_WHEN_FULL | true | 否 |
| PAGE_OBJECT_CONCURRENCY | 8Deep/AutoHeal 强制 1 | 否 |
| EVENT_DRIVEN_SCHEDULER_ENABLE | true | 否 |
| SET_BULKHEAD_ENABLE | true | 否 |
| PAGE_PARALLEL_ENABLE | true | 否 |
| MAINLINE_THROTTLE_ENABLE | true | 否 |
| MAINLINE_READ/WRITE_UTILIZATION_HIGH_PERCENT | 80/80 | 否 |
| MAINLINE_MAX_SLEEP_MS | 250 | 否 |
| (总开关)RUSTFS_HEAL_ENABLED | true | 否 |
| admin 子系统 heal.bitrot_cycle | 30d | 是(经 scanner runtime config |
scanneradmin 子系统 `scanner``crates/config/src/constants/scanner.rs` + `ecstore/src/config/scanner.rs` + `runtime_config.rs:527-673`):
| 键 | env | 默认 |
|---|---|---|
| speed | RUSTFS_SCANNER_SPEED | default2x/1s/60s |
| delay / max_wait / cycle / start_delay | RUSTFS_SCANNER_* | 派生/空 |
| cycle_max_duration/objects/directories | …_MAX_* | 0(不限) |
| bitrot_cycle | …_BITROT_CYCLE_SECS | 259200030d0/on=每轮,off=禁用) |
| idle_mode | …_IDLE_MODE | true |
| cache_save_timeout | …_CACHE_SAVE_TIMEOUT_SECS | 14s |
| max_concurrent_set_scans / disk_scans | …_MAX_CONCURRENT_* | 4/4 |
| yield_every_n_objects | …_YIELD_EVERY_N_OBJECTS | 128 |
| alert_excess_versions / version_size / folders | …_ALERT_* | 100 / 1TiB / 65538 |
scanner 内部 env`RUSTFS_DATA_USAGE_UPDATE_DIR_CYCLES=16``RUSTFS_HEAL_OBJECT_SELECT_PROB=1024``RUSTFS_SCANNER_DEEP_VERIFY_COOLDOWN_SECS=60``RUSTFS_DATA_USAGE_FAILED_OBJECT_TTL_SECS=86400`/`_MAX=10000``RUSTFS_LOCK_ACQUIRE_TIMEOUT=5s``RUSTFS_SCANNER_ENABLED=true``RUSTFS_SCANNER_INLINE_HEAL_ENABLE=false`(兼容告警)。
全部 17 个 scanner 键支持 env > config 双通道 + admin PUT 热更(generation+Notify 即时生效);heal 运行时参数目前仅 env(无 admin 热更入口,`Arc<RwLock<HealConfig>>` 结构已预留)。
---
## 7. 相关 backlog / 历史索引
- 换盘自动修复系列(已闭环):backlog #1786(冗余假绿算法)、#1787(目标槽位限定)、#1789resume 与 healing marker 绑定 replacement 实例)、#1791(黑白盒验收矩阵)。
- #801 DiskInfo.healing 从未赋值(已修复闭环,现 `set_disk/mod.rs:4988` 有赋值链)。
- #1651 Scanner 指标节点/source/bucket-drive 维度(OPEN,本分析 §3.8/§4.6 相关)。
- #1843 crates/common 83% scanner/heal 域代码分层迁移(OPEN,含 HS-20)。
- 代码注释引用的历史缺陷(现已有防护与回归测试):#856/#799 B7(离线盘误记 healed)、#855/B6/#1033skip 不得标记完成)、#920sub-quorum 并集枚举)、#856 B5(按版本续扫)、#5173bitrot trailing bytes)、#5029(回归节点 stale 版本合并)。
- v1 对标文档:`docs/rustfs-heal-scanner-vs-minio-parity-assessment.md`(本文取代)、落地手册 `docs/rustfs-heal-scanner-vs-minio-improvement-playbook.md`(部分条目已被后续实现超越)。
- 换盘深度分析:`docs/new-disk-replacement-and-healing-deep-analysis-zh.md``docs/node-disk-identity-and-healing-analysis-zh.md`
## 8. 审计方法与局限
- 四路并行审计(heal crate 逐文件、scanner crate 逐文件、ecstore 集成层 wiring、MinIO master 源码研究)+ 主会话对关键"缺失"结论逐条亲验(get_disk_status TODO、HealEvent 零外部引用、.bloomcycle.bin 无 bloom 实现、check_abandoned_parts 三层 NotImplemented、ETag 兜底已实现、trace 通道零命中、already_running 语义)。
- 未逐行核实的点(已在文中标注"未确认/未逐行核"):DeleteAllVersions 前缀单调用优化(HS-17)、trash 二段清理细节(HS-18)、ilm worker 默认值对照、stale multipart 默认值对照、mc CLI flag 逐字拼写(MinIO 侧)。其中 HS-17 与 HS-18 已于 2026-08-19 完成逐行核实,结论见 §9.2/§9.3。
- MinIO 侧引用以其 master `7aac2a2c5b` 为准;RustFS 侧行号以 2026-08-16 工作区为准,后续演进请以符号名检索为准。
## 9. 落地结果(2026-08-19 更新)
本审计衍生的 14 个子 issuebacklog #1865~#1878)已全部闭环。本节为差距清单 HS-01~HS-20 的最终处置记录,也是下一轮对标重审的增量基线。
### 9.1 已落地(PR 均已合并 main)
- HS-01 MRF 接线 + 持久化修复账本(#1865PR #6189):决策选 (a)。common MRF channelbounded 8192、try_send 永不阻塞)+ heal mrf_queue100k 条 / 8MiB 双限环形)+ `buckets/.heal/mrf/journal.bin` CRC 持久化回放(torn tail 截断、回放后删除)+ 三投递点(read decode_error→Urgent ECDecode、scanner 元数据损坏→High Metadata、add_partial→Normal+ `RUSTFS_HEAL_MRF_ENABLE` 一键回退。
- HS-02 abandoned parts/data-dir 对账(#1866PR #6179):接通 abandoned 检查入口,保留 dry-run / reclaim 计数。
- HS-03 heal/scanner trace 通道(#1867PR #6179):进程内 trace bus + `/v3/trace` admin 流式订阅 + heal task / abandoned-parts / scanner folder / ILM / heal-candidate trace producer。
- HS-04 scanner 超限 S3 事件(#1868PR #6176):`s3:Scanner:ManyVersions/LargeVersions/BigPrefix` 三事件 + 24h 边沿冷却;HS-15 阈值差异文档化(`docs/operations/scanner-excess-alerts.md`)。
- HS-05 madmin 客户端一期(#1869PR #6166):SigV4 admin 客户端 heal/scanner 方法;增量消费方法待 follow-up(协议已由 HS-06 并入)。
- HS-06 admin heal 增量语义与类型化重叠(#1870PR #6206):`sinceSeq/nextSeq/minSeq` 增量游标(wire additive、缺省=全量快照)+ `RUSTFS_HEAL_OVERLAP_POLICY`(默认 merge 不变;minio_error 下 AlreadyRunning/OverlappingPaths 类型化拒绝)+ forceStart 先停旧再启新。
- HS-07 healing 进度可见性(#1871PR #6179):data-usage 总量基线 + baseline/current/healed 计数。
- HS-08 prefix usage#1872PR #6171):`GET /v3/usage/{bucket}`
- HS-11 bitrot 启动自检(#1873PR #6165)。
- HS-13 heal 跳过过滤(#1875PR #6179):过滤命中版本不再计为失败。
- HS-16 单机周期钩子(#1878PR #6250):删恒 None 钩子,决策记录见 `docs/operations/heal-scanner-parity-notes-zh.md`
- HS-09/10/19/20 死代码清理批(#1877PR #6256):净 911 行零行为变更;`get_disk_status` TODO(全仓库唯一产品 TODO)清零;HS-01 联动的 `ec_decode_rebuild`/`get_object_meta` 保留并加 Reserved 注释(MRF 当前经 `heal_object` 执行)。
### 9.2 核对后确认"已实现 / 非缺口"(审计期误判修正,累计四例)
- bloom filter(§0 已修正):MinIO master 已删除,双方现状一致。
- ETag 兜底仲裁(§0 已修正):RustFS 已有实现(`set_disk/ops/heal.rs`)。
- HS-17#18762026-08-19 逐行核实后关闭):DeleteAllVersions 前缀单调用优化 RustFS 已完整实现——`apply_expiry_on_non_transitioned_objects``delete_all()` 两 action 设 `delete_prefix + delete_prefix_object` 后单次 `delete_object``bucket_lifecycle_ops.rs:5047-5056`),SetDisks 分支一次写锁 + 一次全版本 quorum 读 + 内联逐版本 object-lock 检查(`set_disk/ops/object.rs:5566-5612`),与 MinIO `expire.go``applyExpiryOnNonTransitionedObjects` 逐行对齐。§8 原列"未逐行核实"的本项已有结论:现状即优化路径,无需实现。
- HS-14#1878PR #6250 附带核对):MinIO"idle=空闲才节流"是 2024-01 minio/minio#18734 之前的行为(`scannerIdleMode` 现为静态配置,`idle_speed=on` 默认即始终按速度档节流,"idle"命名是历史残留);RustFS `RUSTFS_SCANNER_IDLE_MODE` 与 MinIO 当前语义方向一致,且另有 MinIO 没有的前台读退避下限。真实迁移陷阱(变量须 `RUSTFS_` 前缀、`on/off` vs `true/false` 词表、`false` 连前台保护一起关)已文档化于 `docs/operations/heal-scanner-parity-notes-zh.md`
### 9.3 审计型结论(无需改代码)
- HS-12#1874PR #6183):不存在 MinIO 用 `x-minio-healing` 防御的那类竞争——所有同 (bucket, object) 提交面在同一把对象级 ns 写锁互斥,heal 锁 guard 覆盖 rename 提交全程;交付 2 个并发不变量回归测试 + `docs/operations/heal-concurrency-safety-notes-zh.md` 交点矩阵。
- HS-18#18782026-08-19 逐行核实):trash/tmp 三段清理全对齐——stale multipart 隔离-清理等价且更安全(`delete_all_with_quorum` 逐盘递归删即 `move_to_trash` rename 进 `.rustfs.sys/tmp/.trash`,另有锁 + fence)、trash 排空基本等价(无逐条 sleeper 节流,5m 周期天然限频)、tmp 非 trash 24h 回收等价(RustFS 5m 比 MinIO 6h 更及时);周期默认 24h/6h/5m 三项全对齐。§8 原列"未逐行核实"的本项已有结论。
### 9.4 移交 follow-up(汇总于 backlog#1862 评论区)
HS-01 bitrot GET→MRF 全链路 e2e、kill -9 journal 回放 e2e、队列满压测 RSS(≤ 预算+10%);HS-05/06 madmin 增量消费方法 + wire 单一来源化 + embedded e2e + 多轮轮询 soakHS-08 多盘 scanner 周期 e2eHS-04 超限审计条目;HS-18 低于 quorum 的 stale-multipart 崩溃残留窗口(扇出中途崩溃且已清盘数 > parity 时 FileNotFound 不在忽略集导致不自然收敛,修复需专用 quorum 变体)。
下一轮重审建议:跟随 heal/scanner 下一个大特性落地后触发,以本节为增量基线。
@@ -0,0 +1,60 @@
# S3 Tables Durable Backing Cutover Runbook
**Use this when:** moving a table-catalog warehouse from object-backed catalog state to the durable strong snapshot backing (`RUSTFS_TABLE_CATALOG_BACKING=durable-strong`), or rolling the strong snapshot format from version 1 to version 2.
**Source of truth:** the `{warehouse}/catalog/migration` routes registered in `rustfs/src/admin/handlers/table_catalog/routes.rs`; the env constants named below; claims and status labels in [docs/architecture/s3-tables-support-matrix.md](../architecture/s3-tables-support-matrix.md).
## Preconditions
| Requirement | Why |
|---|---|
| A principal with `GetTableCatalogAction` on each table bucket (preflight) and `admin:MigrateTableCatalog` (migration `POST` / `DELETE`). | The mutations are admin-gated. |
| Every catalog writer runs a release that recognizes the durable-backing migration fence. | An older writer does not see the persisted fence and can mutate the object-backed source after the snapshot inventory is captured. |
| An object-backed catalog backup, plus the current metadata pointer and version token for representative tables. | Recovery after a failed cutover is an operator-selected restore, not a restart against the stale pointer. |
| Every mutating object-only operation (maintenance workers, catalog recovery, export, diagnostics, external catalog bridge writes) is inventoried and confirmed supported in durable-strong mode. | Unsupported operations fail closed after cutover rather than continuing against object-backed state. |
## Cutover Procedure
1. Take the object-backed backup and record pointer and version token for representative tables.
2. Run the preflight for each warehouse and treat every `blockers` entry as fail-closed. Repair commit recovery state and backfill the warehouse prefix index before continuing. Requests are SigV4-signed with the catalog's REST signing name; the `/_iceberg/v1` alias accepts the same paths.
```text
GET /iceberg/v1/{warehouse}/catalog/migration
```
3. Drain every catalog writer that predates the migration fence and restart it on a fence-aware release. Keep all writers on that release until cutover completes.
4. Quiesce mutating object-only operations (step 4 of Preconditions).
5. Run the migration `POST` with `admin:MigrateTableCatalog`. It acquires the exclusive migration fence to drain in-flight fence-aware mutations, persists the source fence while exclusivity is held, then copies catalog state and reports `ready_to_enable_durable_strong`.
```text
POST /iceberg/v1/{warehouse}/catalog/migration
```
6. Repeat preflight and materialization for every table bucket. Do not proceed until the preflight reports `SNAPSHOT_MATERIALIZED`, no blockers, and `ready_to_enable_durable_strong: true` for all of them.
7. Restart with `RUSTFS_TABLE_CATALOG_BACKING=durable-strong`, then verify catalog config, table and view loads, commit idempotency, and table data-plane policy resolution before admitting writers.
8. Preserve the object-backed backup until durable strong backing has passed the operator's retention window.
## Cancelling Before Cutover
Before the restart in step 7, `DELETE` on the migration endpoint removes a migration-created target bucket snapshot and releases the source fence. It releases the bucket fence only while the target state has not advanced, and releases the registry fence after the last bucket is cancelled. Retries and `DELETE` may restore a known-absent initial target after an ambiguous first write, but fail closed if a previously existing or materialized global snapshot disappears.
```text
DELETE /iceberg/v1/{warehouse}/catalog/migration
```
After the durable-strong state advances, cancellation fails closed; recovery requires an operator-selected restore or reverse migration.
## Strong Snapshot Version 1 to Version 2
1. Keep snapshot writes on version 1 during a rolling binary upgrade. Current binaries read both versions.
2. After every catalog writer can read version 2, set both `RUSTFS_TABLE_CATALOG_STRONG_SNAPSHOT_V2=true` and `RUSTFS_TABLE_CATALOG_STRONG_SNAPSHOT_V2_FLEET_CONFIRMED=true` and restart the catalog writers. Setting only one gate does not change the write format.
3. Perform a controlled catalog write or migration materialization and confirm the persisted snapshot is version 2 before serving table data-plane traffic. Once v2 is fleet-confirmed, data-plane resolution fails closed until the persisted snapshot is v2.
4. After any version 2 snapshot is persisted, do not roll writers back to a binary that only reads version 1. Current binaries preserve version 2 even when the gates are later disabled.
## Rollback And Collision Repair Rules
- A running process rejects restored version 1 content after observing version 2, but cannot distinguish an older snapshot with the same format version from a deliberate restore. The format high-water mark is process-local: restoring any older snapshot and restarting every writer is a privileged disaster-recovery rollback that must restore a compatible binary and an operator-selected snapshot together.
- Migration preflight rejects an active table/view identifier collision before writing a migration fence. A pre-existing version 1 strong snapshot with such a collision loads in cleanup-only quarantine: ambiguous reads fail closed, each cleanup mutation must reduce the collision set, and unrelated writes stay blocked until all collisions are removed. Drain writers that predate cleanup quarantine before starting the repair, and finish cleanup before the first version 2 write.
## Related
- [S3 Tables support matrix](../architecture/s3-tables-support-matrix.md)
- [Table catalog conformance scripts](../../scripts/table-catalog/README.md) (`failure_coverage.py --print-disaster-recovery-rehearsal` generates the rehearsal for this procedure)
+79 -272
View File
@@ -1,42 +1,18 @@
# Scanner Benchmark Runbook
This runbook describes how to collect reproducible evidence for scanner
pressure, scanner progress, and scanner runtime tuning. It is intended for
maintainers and operators validating scanner changes in an isolated test
deployment.
**Use this when:** you need reproducible before/after evidence that a scanner pacing or cycle change reduces background pressure without stalling lifecycle, replication, heal, or bitrot progress, or you are assembling evidence for a scanner-behavior PR.
Use this runbook together with
[Scanner Runtime Controls](scanner-runtime-controls.md). The runtime controls
document explains each field and configuration key; this document explains how
to run a comparable before/after observation and decide whether the scanner is
healthier after tuning.
**Source of truth:** `scripts/run_scanner_validation_harness.sh` (collection, `scanner-summary.csv` columns), `scripts/run_object_batch_bench.sh` (workload), and [Scanner Runtime Controls](scanner-runtime-controls.md) for the meaning of every status field and configuration key.
## Scope
This runbook verifies that scanner pacing and cycle controls reduce background
pressure while preserving maintenance progress. It is useful for:
This runbook verifies that scanner pacing and cycle controls reduce background pressure while preserving maintenance progress. It covers mostly idle single-node deployments with many small objects, multi-disk or erasure-set nodes, distributed clusters where scanner pressure mixes with lifecycle, replication, heal, or bitrot queues, and backlog investigations for any of those subsystems.
- mostly idle single-node, single-disk deployments with many small objects;
- single-node multi-disk or erasure-set deployments where scanner work is
spread across disks and sets;
- distributed clusters where scanner pressure is mixed with lifecycle,
replication, heal, or bitrot queues;
- lifecycle expiry or transition backlog investigations;
- bucket replication repair backlog investigations;
- scanner-originated heal or bitrot admission investigations;
- pull request evidence when scanner behavior, scanner status, or scanner
controls change.
This runbook does not prove full MinIO parity, site-replication correctness,
or replaced-disk heal correctness. Those flows need dedicated distributed
tests because their failure modes are not limited to scanner pacing.
It does not prove full MinIO parity, site-replication correctness, or replaced-disk heal correctness. Those flows need dedicated distributed tests because their failure modes are not limited to scanner pacing.
## Safety
Run the workload only in a disposable test environment. The commands below can
create many buckets and objects and can overwrite runtime scanner settings.
Record the current scanner and heal configuration before changing anything:
Run the workload only in a disposable test environment. The commands below can create many buckets and objects and overwrite runtime scanner settings. Record the current scanner and heal configuration before changing anything:
```bash
mkdir -p artifacts
@@ -44,75 +20,41 @@ mc admin config get ALIAS scanner > artifacts/scanner-config.before.txt
mc admin config get ALIAS heal > artifacts/heal-config.before.txt
```
Replace `ALIAS`, endpoint, and credentials with values for the test
deployment. Do not paste production credentials into saved artifacts.
The `scanner` and `heal` subsystems are served by `GetConfigKVHandler` (`rustfs/src/admin/handlers/config_admin.rs`, route `/v3/get-config-kv`); this was confirmed by code inspection, not by running `mc` against a live deployment. Replace `ALIAS`, endpoint, and credentials with values for the test deployment. Do not paste production credentials into saved artifacts.
## Required Tools
- `mc` or a compatible admin client for config changes.
- `awscurl` or another SigV4-capable HTTP client for `/v3/scanner/status`.
- `jq` for status extraction.
- `pidstat`, `mpstat`, `iostat`, `top`, or equivalent host telemetry.
- `warp`, `s3bench`, or the repository object benchmark scripts for workload
generation.
| Tool | Purpose |
|---|---|
| `mc` or a compatible admin client | Config snapshots and changes. |
| `awscurl` or another SigV4-capable HTTP client | `/v3/scanner/status` and admin metrics. |
| `jq` | Status extraction. |
| `pidstat`, `mpstat`, `iostat`, `top`, or equivalent | Host telemetry. |
| `warp`, `s3bench`, or `scripts/run_object_batch_bench.sh` | Workload generation. |
## Test Matrix
At minimum, collect two runs on the same RustFS commit and the same workload:
Collect at least two runs on the same RustFS commit and the same workload. Keep hardware, commit, object count, object size, bucket count, scanner-enabled state, and foreground workload constant between runs.
| Run | Purpose | Example scanner settings |
|---|---|---|
| Baseline | Observe current behavior without additional pacing changes. | Existing config. |
| Pacing override | Measure whether cooperative scanner sleeps reduce pressure. | `scanner.delay="30"` and `scanner.max_wait="15"`. |
Add bounded-cycle runs when a single scan cycle is too long:
| Run | Purpose | Example scanner settings |
|---|---|---|
| Duration budget | Bound wall-clock time per cycle. | `scanner.cycle_max_duration="1800"`. |
| Duration budget (when one cycle is too long) | Bound wall-clock time per cycle. | `scanner.cycle_max_duration="1800"`. |
| Object budget | Bound objects processed per cycle. | `scanner.cycle_max_objects="1000000"`. |
| Directory budget | Bound directories entered per cycle. | `scanner.cycle_max_directories="100000"`. |
When comparing runs, keep hardware, RustFS commit, object count, object size,
bucket count, scanner-enabled state, and foreground workload constant.
## Deployment Matrix
Use the smallest deployment that reproduces the symptom, but do not treat
single-node validation as the whole scanner test surface. Scanner changes that
touch queues, admission, or cross-node maintenance should include a distributed
run when practical.
Use the smallest deployment that reproduces the symptom. The single-node, single-disk run is the cheap, repeatable baseline; it is not sufficient for PRs that claim to improve distributed queue behavior, replication repair, or heal/bitrot admission.
| Deployment | What it validates | Minimum evidence |
|---|---|---|
| Single-node, single-disk | Small-object scanner pressure, pacing, cycle interval, and basic progress. | Scanner status time series plus host CPU and disk telemetry. |
| Single-node, multi-disk or erasure set | Set and disk scan concurrency, cycle budgets, checkpoint movement, usage cache persistence, and active path age. | Scanner status time series, per-disk host telemetry, and before/after data usage freshness. |
| Distributed cluster | Lifecycle transition queues, bucket replication repair admission, scanner-originated heal and bitrot admission, and queue/backlog pressure under cross-node work. | Scanner status time series from the cluster, host telemetry from each node, and subsystem-specific queued/skipped/missed counters. |
| Deployment | What it validates | Minimum evidence | Workload shape |
|---|---|---|---|
| Single-node, single-disk | Small-object scanner pressure, pacing, cycle interval, basic progress. | Scanner status time series plus host CPU and disk telemetry. | One node, one data disk, several buckets, at least 100,000 small objects, scanner enabled, no sustained foreground workload during observation. |
| Single-node, multi-disk or erasure set | Set and disk scan concurrency, cycle budgets, checkpoint movement, usage cache persistence, active path age. | Scanner status time series, per-disk host telemetry, before/after data usage freshness. | Same as above across all disks. |
| Distributed cluster | Lifecycle transition queues, bucket replication repair admission, scanner-originated heal and bitrot admission, queue/backlog pressure under cross-node work. | Scanner status time series from the cluster, host telemetry from each node, subsystem-specific queued/skipped/missed counters. | Same structure plus the relevant subsystem condition (lifecycle rules, a replication target, a heal/bitrot scenario); keep status and telemetry cadence identical to the baseline. |
The single-node, single-disk run is the baseline because it is cheap and
repeatable. It is not sufficient for PRs that claim to improve distributed
queue behavior, replication repair, or heal/bitrot admission.
## Workload Shape
For the baseline small-object scanner-pressure run, use a workload that
creates many small objects and then leaves the service mostly idle while the
scanner walks the namespace. A useful minimum shape is:
- one RustFS node;
- one data disk;
- several buckets;
- at least 100,000 small objects total;
- scanner enabled;
- no sustained foreground workload during the observation window.
For a distributed backlog run, use the same structure but add the relevant
subsystem condition, such as lifecycle rules, a bucket replication target, or
a configured heal/bitrot scenario. Keep the status and host telemetry cadence
the same so the result remains comparable with the baseline run.
The repository object benchmark script can generate object traffic if `warp`
or `s3bench` is installed:
Generate object traffic with the repository script if `warp` or `s3bench` is installed; repeat with new buckets or prefixes if one run cannot create enough objects, and record the final object count:
```bash
scripts/run_object_batch_bench.sh \
@@ -129,15 +71,9 @@ scripts/run_object_batch_bench.sh \
--out-dir artifacts/object-load
```
If the object generator cannot create enough objects in one run, repeat the
same command with new buckets or prefixes and record the final object count.
## Status Collection
Capture scanner status before the workload, after the workload finishes, and
throughout the idle observation window.
The repository includes a scanner validation harness for repeatable collection:
Capture scanner status before the workload, after the workload finishes, and throughout the idle observation window. The validation harness does this repeatably and writes scanner/heal config snapshots, scanner status samples, background heal status samples, host telemetry when available, run metadata, `scanner-summary.csv`, and `scanner-validation-report.md`:
```bash
export RUSTFS_ACCESS_KEY="<admin-access-key>"
@@ -153,12 +89,7 @@ scripts/run_scanner_validation_harness.sh \
--out-dir artifacts/scanner-validation
```
The harness writes scanner/heal config snapshots, scanner status samples,
background heal status samples, host telemetry when available, run metadata,
`scanner-summary.csv`, and `scanner-validation-report.md`.
Use `--metrics-endpoints` when the validation needs per-node distributed
evidence. The value is a comma-separated list of RustFS endpoints:
For per-node distributed evidence pass `--metrics-endpoints` (comma-separated). Each sample then stores `/v3/scanner/status`, one `/v3/background-heal/status` response per listed endpoint, and one by-host admin metrics response per listed endpoint; without it, background-heal status is captured only from `--endpoint`:
```bash
scripts/run_scanner_validation_harness.sh \
@@ -172,55 +103,22 @@ scripts/run_scanner_validation_harness.sh \
--out-dir artifacts/scanner-validation-distributed
```
Each sample stores `/v3/scanner/status`, one `/v3/background-heal/status`
response per endpoint listed in `--metrics-endpoints`, and one by-host admin
metrics response per listed endpoint. When `--metrics-endpoints` is omitted,
the harness captures background-heal status only from `--endpoint`.
For ad hoc per-node snapshots outside the harness window, use the by-host `awscurl` loop in [Reading Distributed Metrics](scanner-runtime-controls.md#reading-distributed-metrics); the metrics endpoint reports only the node that handles the request.
For bucket metrics freshness validation, use the same harness around a
post-start bucket creation workload:
### Bucket metrics freshness validation
Use the harness around a post-start bucket creation workload to cover the timing where scanner startup sees no buckets, a bucket is created afterwards, and the first metrics collection must not confuse a cold usage cache with real zero usage:
1. Start RustFS from an empty data path.
2. Start the harness before creating buckets.
3. Create a bucket, upload objects, and keep the harness running until at
least one usage save is observed.
4. Compare `scanner-summary.csv` with
`/rustfs/admin/v3/metrics?types=1&n=1` bucket metrics.
3. Create a bucket, upload objects, and keep the harness running until at least one usage save is observed.
4. Compare `scanner-summary.csv` with `/rustfs/admin/v3/metrics?types=1&n=1` bucket metrics.
This covers the issue 3496 timing where scanner startup sees no buckets, the
bucket is created after startup, and the first metrics collection must not
confuse cold usage cache with real zero usage. The expected evidence is that
dirty usage is marked, `life_time_scan_cycle` or
`life_time_scan_bucket_drive` advances, `life_time_scan_object` advances for
object workloads, and `life_time_save_usage` plus
`usage_last_save_result=success` appear before accepting non-zero bucket usage
metrics as fresh.
Expected evidence: dirty usage is marked, `life_time_scan_cycle` or `life_time_scan_bucket_drive` advances, `life_time_scan_object` advances for object workloads, and `life_time_save_usage` plus `usage_last_save_result=success` appear before non-zero bucket usage metrics are accepted as fresh.
For distributed runs, capture scanner admin metrics from every node with
`by-host=true`. The metrics endpoint reports the node that handles the request;
`by-host=true` preserves that node's host view but does not collect peer nodes.
These per-node artifacts include active path age, checkpoint state, pacing
pressure, source work, and queued/skipped/missed downstream admission counters.
The validation harness can collect these artifacts automatically with
`--metrics-endpoints`; the manual loop below is useful when adding extra nodes
or collecting ad hoc snapshots outside the harness window.
### Manual status sampling
```bash
for endpoint in http://node-a:9000 http://node-b:9000 http://node-c:9000; do
node="${endpoint#http://}"
node="${node%%:*}"
awscurl \
--service s3 \
--region us-east-1 \
--access_key "$RUSTFS_ACCESS_KEY" \
--secret_key "$RUSTFS_SECRET_KEY" \
--request GET \
"${endpoint}/rustfs/admin/v3/metrics?types=1&by-host=true&n=1" \
> "artifacts/scanner-metrics.${node}.$(date -u +%Y%m%dT%H%M%SZ).ndjson"
done
```
Example status request:
Single snapshot:
```bash
awscurl \
@@ -233,7 +131,7 @@ awscurl \
| jq . > "artifacts/scanner-status.$(date -u +%Y%m%dT%H%M%SZ).json"
```
For a time series, sample once per minute:
Time series (stop after the planned observation window):
```bash
mkdir -p artifacts/status
@@ -250,11 +148,9 @@ while sleep 60; do
done
```
Stop the loop after the planned observation window.
## Host Telemetry
Collect host metrics over the same window as scanner status.
Collect host metrics over the same window as scanner status. If `pidstat` is unavailable, use `top`, `ps`, or the platform monitoring system, but record the sampling interval and window in the report.
```bash
pidstat -p "$(pidof rustfs)" 60 > artifacts/pidstat.txt
@@ -262,13 +158,9 @@ iostat -xz 60 > artifacts/iostat.txt
mpstat 60 > artifacts/mpstat.txt
```
If `pidstat` is not available, use `top`, `ps`, or the platform monitoring
system, but keep the sampling interval and observation window in the report.
## Runtime Tuning Examples
Persistent scanner config values use seconds for time fields. Use numeric
strings instead of duration suffixes:
Persistent scanner config values use seconds for time fields; use numeric strings, not duration suffixes. The canonical persistent bitrot cadence belongs to the `heal` subsystem.
```bash
mc admin config set ALIAS scanner delay="30" max_wait="15"
@@ -276,16 +168,10 @@ mc admin config set ALIAS scanner cycle="3600"
mc admin config set ALIAS scanner cycle_max_duration="1800"
mc admin config set ALIAS scanner cycle_max_objects="1000000"
mc admin config set ALIAS scanner cycle_max_directories="100000"
```
The canonical persistent bitrot cadence belongs to the `heal` subsystem:
```bash
mc admin config set ALIAS heal bitrot_cycle="2592000"
```
Environment variables take precedence over persisted config and should be
recorded separately:
Environment variables take precedence over persisted config and should be recorded separately:
```bash
RUSTFS_SCANNER_DELAY=30
@@ -297,8 +183,7 @@ RUSTFS_SCANNER_CYCLE_MAX_DIRECTORIES=100000
RUSTFS_SCANNER_BITROT_CYCLE_SECS=2592000
```
After each config change, read scanner status and confirm the effective value
and `source` under `runtime_config`.
After each config change, read scanner status and confirm the effective value and `source` under `runtime_config`.
## Observation Window
@@ -308,85 +193,45 @@ Use the same window for each run:
2. Wait until foreground workload is idle.
3. Save scanner and heal config.
4. Save one scanner status snapshot.
5. Collect scanner status and host telemetry for at least 30 minutes, or for
one complete scanner cycle when that is practical.
5. Collect scanner status and host telemetry for at least 30 minutes, or for one complete scanner cycle when practical.
6. Save one final scanner status snapshot.
Longer windows are better for cycle interval comparisons. Short windows are
acceptable for quick pressure checks only if the conclusion avoids changing
defaults.
Longer windows are better for cycle interval comparisons. Short windows are acceptable for quick pressure checks only if the conclusion avoids changing defaults.
## Fields To Compare
Compare these fields between baseline and tuned runs:
Field semantics are defined in [Scanner Runtime Controls](scanner-runtime-controls.md); the decision fields for a before/after comparison are:
| Field | Why it matters |
| Field | Decision it supports |
|---|---|
| `runtime_config.*.value` and `runtime_config.*.source` | Confirms the tested settings actually took effect. |
| `metrics.pacing_pressure.primary_pressure` | Shows whether pressure is from queues, budgets, pause activity, active scans, or no scanner pressure. |
| `metrics.pacing_pressure.last_cycle_total_pause_ratio` | Shows how much of the last cycle was cooperative scanner pause time. |
| `metrics.maintenance_control.primary_control` | Shows whether source-level maintenance is blocked, deferred, active, only pacing-limited, or idle. |
| `metrics.maintenance_control.sources` | Shows the source, state, reason, backlog, current or last-cycle missed work, and partial-cycle count for each scanner maintenance source. |
| `metrics.current_cycle_objects_scanned` | Confirms object scan progress during the current cycle. |
| `metrics.current_cycle_directories_scanned` | Confirms directory walk progress during the current cycle. |
| `metrics.last_cycle_result` | Confirms whether the previous cycle completed, stopped partially, or failed. |
| `metrics.last_cycle_partial_reason` | Shows which budget stopped a partial cycle. |
| `metrics.last_cycle_partial_source` | Shows which scanner work source consumed the stopping budget. |
| `metrics.source_work` | Shows cumulative work found, queued, skipped, missed, executed, and failed by source. |
| `metrics.current_cycle_source_work` | Shows which source is consuming the current scan cycle. |
| `metrics.last_cycle_source_work` | Shows which source consumed the previous scan cycle. |
| `metrics.replication_repair` | Splits scanner-discovered replication repair by source, kind, scanner role, and execution owner, including bucket object, delete-marker, version-purge, existing-object repair, and site replication boundary states. |
| `metrics.current_cycle_replication_repair` | Shows which replication repair kind is being discovered or admitted in the current cycle. |
| `metrics.last_cycle_replication_repair` | Shows which replication repair kind consumed the previous cycle. |
| `metrics.lifecycle_expiry.current_queued` | Shows scanner-driven expiry/delete work waiting in the expiry worker queue. |
| `metrics.lifecycle_expiry.current_active` | Shows scanner-driven expiry/delete work currently running in expiry workers. |
| `metrics.lifecycle_expiry.queue_missed` | Shows expiry/delete queue admission failures outside the scanner walk itself. |
| `metrics.lifecycle_expiry.scanner_missed` | Shows scanner-discovered expiry/delete work that could not be queued. |
| `metrics.lifecycle_transition.scanner_missed` | Shows scanner-discovered transition work that could not be queued. |
| `metrics.lifecycle_transition.queue_full` | Shows transition queue pressure outside the scanner walk itself. |
| `metrics.lifecycle_transition.compensation_pending` | Shows transition compensation still pending or running after queue pressure. |
| `metrics.lifecycle_transition.failed` | Shows transition worker failures, which should also surface as lifecycle source failure. |
| `metrics.current_cycle_usage_saves` | Shows usage cache saves produced by the active scan cycle. |
| `metrics.last_cycle_usage_saves` | Shows usage cache saves produced by the previous completed or partial scan cycle. |
| `metrics.usage_freshness.dirty_pending_buckets` | Shows whether bucket/object mutations are still waiting for usage refresh. |
| `metrics.usage_freshness.last_cycle_dirty_buckets` | Shows how many dirty buckets were picked up by the last cycle. |
| `metrics.usage_freshness.last_cycle_cleared_dirty_buckets` | Shows how many dirty bucket marks were cleared by a successful cycle. |
| `metrics.usage_freshness.last_usage_save_result` | Confirms whether the last usage save succeeded, failed, or was skipped. |
| `metrics.life_time_ops.scan_cycle` | Confirms scanner cycles actually started after the workload. |
| `metrics.life_time_ops.scan_bucket_drive` | Confirms bucket-drive scan work reached the storage layer. |
| `metrics.life_time_ops.scan_object` | Confirms object metadata scanning advanced for object workloads. |
| `metrics.life_time_ops.save_usage` | Confirms `DataUsageInfo` save work happened; this is the key freshness signal for bucket metrics. |
| `metrics.scan_checkpoint` | Confirms partial cycles preserve resume context. |
| `metrics.oldest_active_path_age_seconds` | Helps identify scanner paths that may be stuck. |
| `runtime_config.*.value` and `runtime_config.*.source` | The tested settings actually took effect. |
| `metrics.pacing_pressure.primary_pressure`, `last_cycle_total_pause_ratio` | Where pressure comes from and how much of the cycle was cooperative pause. |
| `metrics.maintenance_control.primary_control`, `metrics.maintenance_control.sources` | Whether a maintenance source is blocked, deferred, active, or only pacing-limited. |
| `metrics.current_cycle_objects_scanned`, `metrics.current_cycle_directories_scanned` | Scan progress continues. |
| `metrics.last_cycle_result`, `last_cycle_partial_reason`, `last_cycle_partial_source` | Whether the previous cycle completed, which budget stopped it, and which source consumed it. |
| `metrics.source_work`, `metrics.current_cycle_source_work`, `metrics.last_cycle_source_work` | `missed` growth per source is a downstream admission problem, not pacing. |
| `metrics.replication_repair` (and current/last-cycle variants) | Repair kind, `scanner_role`, and `execution_owner` for replication backlog runs. |
| `metrics.lifecycle_expiry.{current_queued,current_active,queue_missed,scanner_missed}` | Expiry backlog and admission failures. |
| `metrics.lifecycle_transition.{scanner_missed,queue_full,compensation_pending,failed}` | Transition backlog, queue pressure, and worker failures. |
| `metrics.usage_freshness.*`, `metrics.current_cycle_usage_saves`, `metrics.last_cycle_usage_saves` | Bucket metrics freshness; `last_usage_save_result` must be `success`. |
| `metrics.life_time_ops.{scan_cycle,scan_bucket_drive,scan_object,save_usage}` | Cycles, bucket-drive scans, object scans, and `DataUsageInfo` saves actually happened after the workload. |
| `metrics.scan_checkpoint`, `metrics.oldest_active_path_age_seconds` | Partial cycles preserve resume context; stuck paths. |
Do not use a single CPU spike as the conclusion. Compare average and p95 CPU
over the same observation window.
Do not use a single CPU spike as the conclusion; compare average and p95 CPU over the same observation window.
For heal or bitrot pressure investigations, also capture
`/v3/background-heal/status` from every distributed endpoint and compare
`healOperations.queueLength`,
`healOperations.activeTasks`, `healOperations.queuedBySource`,
`healOperations.activeBySource`, `healOperations.queuedByPriority`, and
`healOperations.activeByPriority`. These fields distinguish scanner-submitted
low-priority work from manual admin heal and auto-heal work.
For heal or bitrot pressure investigations, also capture `/v3/background-heal/status` from every distributed endpoint and compare `healOperations.queueLength`, `activeTasks`, `queuedBySource`, `activeBySource`, `queuedByPriority`, and `activeByPriority` (see [Reading Heal Operations](scanner-runtime-controls.md#reading-heal-operations)).
`scanner-summary.csv` includes the heal operation totals needed for quick
before/after comparison. In distributed runs, these fields are aggregated from
the background-heal status snapshots captured across `--metrics-endpoints`.
### `scanner-summary.csv` columns
| Field | Why it matters |
In distributed runs the heal columns are aggregated from the background-heal snapshots captured across `--metrics-endpoints`.
| Column | Meaning |
|---|---|
| `heal_queue_length` | Total queued heal requests at the same timestamp as the scanner status sample. |
| `heal_active_tasks` | Total running heal tasks. |
| `heal_scanner_queued` | Scanner-submitted heal or bitrot work waiting in the queue. |
| `heal_admin_queued` | Manual/admin heal work waiting in the queue. |
| `heal_auto_heal_queued` | Auto-heal work waiting in the queue, typically from disk/set recovery paths. |
`scanner-summary.csv` also includes usage freshness columns for quick
post-start bucket metrics validation:
| Field | Why it matters |
|---|---|
| `current_cycle_usage_saves` | Usage saves during the current cycle. |
| `last_cycle_usage_saves` | Usage saves from the last finished or partial cycle. |
| `usage_dirty_pending_buckets` | Dirty buckets still waiting for scanner refresh. |
@@ -404,74 +249,36 @@ post-start bucket metrics validation:
A useful tuning result has all of these properties:
- average or p95 scanner-related CPU and disk pressure decreases;
- `current_cycle_objects_scanned` or `current_cycle_directories_scanned`
continues to advance;
- `source_work.missed` does not grow unexpectedly for lifecycle, replication,
heal, or bitrot;
- `last_cycle_result` is either `success` or a partial result with a clear
budget reason and checkpoint;
- `current_cycle_objects_scanned` or `current_cycle_directories_scanned` continues to advance;
- `source_work.missed` does not grow unexpectedly for lifecycle, replication, heal, or bitrot;
- `last_cycle_result` is either `success` or a partial result with a clear budget reason and checkpoint;
- data usage freshness remains acceptable for the tested deployment.
Treat these as failure signals:
- CPU drops only because the scanner stops making progress;
- `primary_pressure` stays at `queued_scans` while queues grow;
- `last_cycle_partial_reason` repeats forever with no checkpoint movement;
- lifecycle expiry `queue_missed`, `scanner_missed`, `current_queued`, or
`current_active` grows during a run that was expected to reduce expiry
backlog;
- lifecycle transition `scanner_missed`, `queue_full`,
`compensation_pending`, or `failed` grows during a run that was expected to
reduce backlog;
- bucket metrics show zero usage after post-start uploads while dirty usage
remains pending and `life_time_save_usage` does not advance;
- `bucket_replication` missed work with `scanner_role=repair_admission` grows
while replication worker queues or target failures are also growing; treat
this as downstream replication pressure, not only scanner pacing pressure;
- `site_replication` `active_resync` grows and is interpreted as scanner-owned
repair execution; `scanner_role=boundary_signal` and
`execution_owner=site_replication_runtime` mean active site resync remains
owned by the site replication runtime and admin resync path;
- heal or bitrot work moves from `queued` to `missed` after a scanner pacing
change.
| Signal | Reading |
|---|---|
| CPU drops only because the scanner stops making progress | Not a tuning win. |
| `primary_pressure` stays at `queued_scans` while queues grow | Concurrency, not pacing, is the constraint. |
| `last_cycle_partial_reason` repeats forever with no checkpoint movement | Budget too small or checkpoint not advancing. |
| Lifecycle expiry `queue_missed`, `scanner_missed`, `current_queued`, or `current_active` grows during a run meant to reduce expiry backlog | Downstream expiry pressure. |
| Lifecycle transition `scanner_missed`, `queue_full`, `compensation_pending`, or `failed` grows during a run meant to reduce backlog | Downstream transition pressure. |
| Bucket metrics show zero usage after post-start uploads while dirty usage remains pending and `life_time_save_usage` does not advance | Usage freshness regression. |
| `bucket_replication` missed work with `scanner_role=repair_admission` grows while replication worker queues or target failures also grow | Downstream replication pressure, not only scanner pacing. |
| `site_replication` `active_resync` grows and is read as scanner-owned repair execution | Misreading: `scanner_role=boundary_signal` and `execution_owner=site_replication_runtime` mean active resync remains owned by the site replication runtime. |
| Heal or bitrot work moves from `queued` to `missed` after a scanner pacing change | Heal admission regression. |
## PR Evidence Checklist
For scanner behavior PRs, include this evidence when available:
For scanner behavior PRs, include when available:
- RustFS commit SHA and branch.
- Deployment shape: node count, disk count, disk type, CPU count, memory, and
object count.
- Deployment shape: node count, disk count, disk type, CPU count, memory, object count.
- Workload command or script and benchmark artifact path.
- Scanner and heal config before and after tuning.
- Observation window and sample interval.
- Scanner status snapshots or time series.
- Host CPU and disk telemetry.
- Usage freshness fields from `scanner-summary.csv` when validating bucket
metrics or issue 3496-style timing.
- Short conclusion that separates pressure reduction from scanner progress.
- `scanner-validation-report.md` from the harness when using the scripted
collection path.
## Final Parity Validation Closure
Use the final validation run to prove the scanner control plane is coherent,
not to introduce new runtime behavior. A complete closure package should have
at least these runs:
| Run | Required evidence |
|---|---|
| Single-node, single-disk small-object idle | Scanner status series, host telemetry, `scanner-summary.csv`, and a conclusion that CPU or disk pressure is lower without scan progress stopping. |
| Single-node post-start bucket metrics freshness | Empty data path startup, post-start bucket creation/upload, bucket metrics snapshots, `scanner-summary.csv` usage freshness columns, and evidence that `DataUsageInfo` save work occurred before accepting bucket usage metrics. |
| Single-node erasure or multi-disk | Checkpoint movement, active path age, set/disk scan pressure, data usage freshness, and before/after scanner config. |
| Distributed lifecycle backlog | `maintenance_control`, lifecycle expiry/transition queue fields, source work missed/failed counts, and by-host admin metrics. |
| Distributed replication backlog | Bucket replication repair kind counters, `scanner_role`, `execution_owner`, site replication passive/active boundary counters, source work queued/skipped/missed counts, and by-host admin metrics. |
| Heal or bitrot pressure | Background heal `healOperations` queued/active source and priority counts, scanner source work for heal/bitrot, and by-host admin metrics. |
The expected conclusion is MinIO-style scanner behavior at the operational
contract level: scanner remains enabled, pacing is observable and adjustable,
partial progress is explainable, maintenance work is attributed by source, and
downstream lifecycle, replication, heal, and bitrot backlog can be diagnosed
without guessing from CPU usage alone.
For documentation-only PRs, it is enough to verify links and formatting.
- Usage freshness fields from `scanner-summary.csv` when validating bucket metrics timing.
- A short conclusion that separates pressure reduction from scanner progress.
- `scanner-validation-report.md` from the harness when using the scripted collection path.
+18 -16
View File
@@ -1,37 +1,39 @@
# Scanner Excess Alerts: Metrics, S3 Events, and Thresholds
> 中文版:[scanner-excess-alerts_zh.md](scanner-excess-alerts_zh.md)
**Use this when:** debugging an excess-versions / excess-version-size / excess-folders alert, wiring a notification subscriber for `s3:Scanner:*` events, or explaining why a threshold differs from MinIO.
Date: 2026-08-18 (rustfs/backlog#1868 / HS-04; includes the HS-15 threshold-delta notes)
**Source of truth:** `crates/scanner/src/scanner_folder.rs` (`EVENT_SCANNER_*`, `DEFAULT_SCANNER_ALERT_COOLDOWN_SECS`, `MAX_SCANNER_ALERT_COOLDOWN_KEYS`, `METRIC_SCANNER_EXCESS_*`), `crates/config/src/constants/scanner.rs` (`DEFAULT_SCANNER_ALERT_EXCESS_*`), `crates/s3-types/src/event_name.rs` (`EventName::Scanner*`).
The background scanner detects three classes of "excess" conditions while it walks buckets and surfaces them as alerts. This page documents each alert's trigger condition, the subscribable S3 event, the cooldown semantics, and the threshold differences versus MinIO — for operators debugging alerts and for event consumers wiring up subscriptions.
The background scanner detects three "excess" conditions while it walks buckets and surfaces them as metrics, structured logs, and S3 events. Threshold keys are also listed in the runtime-controls table in [Scanner Runtime Controls](scanner-runtime-controls.md).
## The three alerts
| Alert | Trigger (per scan cycle) | Metric | S3 event (RustFS wire name) | MinIO event name |
|---|---|---|---|---|
| Excess versions | Retained versions of one object `scanner:alert_excess_versions` | `rustfs_scanner_excess_object_versions_total{bucket}` | `s3:Scanner:ManyVersions` | `s3:ObjectManyVersions` |
| Excess version size | Cumulative bytes of all versions of one object `scanner:alert_excess_version_size` | `rustfs_scanner_excess_object_version_size_total{bucket}` | `s3:Scanner:LargeVersions` | `s3:ObjectLargeVersions` |
| Excess folders | Direct subfolders of one directory > `scanner:alert_excess_folders` | `rustfs_scanner_excess_folders_total{root}` | `s3:Scanner:BigPrefix` | `s3:PrefixManyFolders` |
| Excess versions | Retained versions of one object >= `scanner.alert_excess_versions` | `rustfs_scanner_excess_object_versions_total{bucket}` | `s3:Scanner:ManyVersions` | `s3:ObjectManyVersions` |
| Excess version size | Cumulative bytes of all versions of one object >= `scanner.alert_excess_version_size` | `rustfs_scanner_excess_object_version_size_total{bucket}` | `s3:Scanner:LargeVersions` | `s3:ObjectLargeVersions` |
| Excess folders | Direct subfolders of one directory > `scanner.alert_excess_folders` | `rustfs_scanner_excess_folders_total{root}` | `s3:Scanner:BigPrefix` | `s3:PrefixManyFolders` |
Subscribe like any bucket notification: configure a notification on the target bucket with the RustFS wire name above (or the `s3:Scanner:*` wildcard). Events carry `UserAgent: Scanner` as their origin marker, and `req_params` holds the observed value and the threshold (`versions` / `cumulativeSize` / `folders` / `threshold`), so consumers can judge severity directly.
Subscribe like any bucket notification: configure a notification on the target bucket with the RustFS wire name above (or the `s3:Scanner:*` wildcard, `EventName::ObjectScannerAll`). Events carry `UserAgent: Scanner` as their origin marker, and `req_params` holds the observed value and the threshold (`versions` / `cumulativeSize` / `folders` / `threshold`), so consumers can judge severity directly.
## Metrics and events fire on different cadences
- **Metrics and structured logs are level-triggered**: as long as the object stays over the threshold, every scan cycle counts and logs it (default cycle ≈ 60s; see `scanner:speed`).
- **S3 events are edge-triggered with a cooldown**: the same (alert kind, bucket, object) emits at most once per cooldown window — 24 hours by default (`RUSTFS_SCANNER_ALERT_COOLDOWN_SECS`; set it to 0 to emit every cycle). When the window lapses and the object is still over the threshold, the event fires again. The cooldown table lives in process memory with a 4096-entry hard cap; on overflow it is cleared and rebuilt (worst case: one extra emission per still-hot key).
- A process restart resets the cooldown (every still-over-threshold object emits once more after a restart) — deliberately: restarts usually accompany incident response, and the re-emission buys visibility.
| Surface | Cadence |
|---|---|
| Metrics and structured logs | Level-triggered: as long as the object stays over the threshold, every scan cycle counts and logs it (default cycle about 60s; see `scanner.speed`). |
| S3 events | Edge-triggered with a cooldown: the same (alert kind, bucket, object) emits at most once per cooldown window, `RUSTFS_SCANNER_ALERT_COOLDOWN_SECS` (`DEFAULT_SCANNER_ALERT_COOLDOWN_SECS`, 86400; `0` emits every cycle). When the window lapses and the object is still over the threshold, the event fires again. |
| Cooldown table | Process memory, hard cap `MAX_SCANNER_ALERT_COOLDOWN_KEYS` (4096) distinct keys; on overflow it is cleared and rebuilt (worst case one extra emission per still-hot key). A process restart resets it, so every still-over-threshold object emits once more after a restart. |
## Threshold defaults and the MinIO deltas (HS-15)
## Threshold defaults and MinIO deltas
| Config key | ENV | RustFS default | MinIO default | Notes |
|---|---|---|---|---|
| `scanner:alert_excess_versions` | `RUSTFS_SCANNER_ALERT_EXCESS_VERSIONS` | 100 | 100 | Identical |
| `scanner:alert_excess_version_size` | `RUSTFS_SCANNER_ALERT_EXCESS_VERSION_SIZE` | 1 TiB | 1 TB | Same order of magnitude; different unit basis (TiB vs TB) |
| `scanner:alert_excess_folders` | `RUSTFS_SCANNER_ALERT_EXCESS_FOLDERS` | 65538 | 50000 | **Deliberate divergence**: 65538 tolerates the Proxmox Backup Server chunk layout (65536 chunks per directory plus the directory's own entries); MinIO's 50000 would fire continuously for PBS users. Set it to 50000 explicitly to match MinIO behavior |
| `scanner.alert_excess_versions` | `RUSTFS_SCANNER_ALERT_EXCESS_VERSIONS` | 100 (`DEFAULT_SCANNER_ALERT_EXCESS_VERSIONS`) | 100 | Identical. |
| `scanner.alert_excess_version_size` | `RUSTFS_SCANNER_ALERT_EXCESS_VERSION_SIZE` | 1 TiB, 1099511627776 (`DEFAULT_SCANNER_ALERT_EXCESS_VERSION_SIZE`) | 1 TB | Same order of magnitude; different unit basis (TiB vs TB). |
| `scanner.alert_excess_folders` | `RUSTFS_SCANNER_ALERT_EXCESS_FOLDERS` | 65538 (`DEFAULT_SCANNER_ALERT_EXCESS_FOLDERS`) | 50000 | Deliberate divergence: 65538 tolerates the Proxmox Backup Server chunk layout (65536 chunks per directory plus the directory's own entries); MinIO's 50000 would fire continuously for PBS users. Set 50000 explicitly to match MinIO. |
All three keys accept both env and admin config (`PUT /rustfs/admin/v3/config`, `scanner` subsystem); hot updates take effect immediately.
All three keys accept both the environment variable and the `scanner` admin config subsystem (`SCANNER_SUB_SYS`, applied through `apply_scanner_runtime_config`); config updates take effect without a restart.
## Why the event names are mapped
RustFS's event enum (`rustfs_s3_types::EventName::ScannerManyVersions/LargeVersions/BigPrefix`) keeps the repo's established `s3:Scanner:*` wire names (literally different from MinIO's `s3:ObjectManyVersions`; the enum comments preserve the mapping). Subscribers should use the RustFS wire names in this page. If you need MinIO-literal compatibility, map the names on the console/consumer side — do not change the published wire names.
`EventName::ScannerManyVersions` / `ScannerLargeVersions` / `ScannerBigPrefix` keep the repository's established `s3:Scanner:*` wire names, which differ literally from MinIO's `s3:ObjectManyVersions` family; the enum comments preserve the mapping. Subscribers should use the RustFS wire names. If MinIO-literal compatibility is needed, map the names on the console or consumer side rather than changing the published wire names.
@@ -1,37 +0,0 @@
# Scanner 超限告警:指标、S3 事件与阈值
> English version: [scanner-excess-alerts.md](scanner-excess-alerts.md)
日期:2026-08-18rustfs/backlog#1868 / HS-04,含 HS-15 阈值差异说明)
后台 scanner 在扫描过程中检测三类"超限"状态并对外告警。本文说明每类告警的触发条件、可订阅的 S3 事件、冷却语义,以及与 MinIO 的阈值差异,供运维排障与事件消费方对接。
## 三类告警
| 告警 | 触发条件(任一扫描周期) | 指标 | S3 事件(RustFS wire 名) | MinIO 对应事件名 |
|---|---|---|---|---|
| 版本数超限 | 单对象保留版本数 ≥ `scanner:alert_excess_versions` | `rustfs_scanner_excess_object_versions_total{bucket}` | `s3:Scanner:ManyVersions` | `s3:ObjectManyVersions` |
| 版本总大小超限 | 单对象全部版本累计字节 ≥ `scanner:alert_excess_version_size` | `rustfs_scanner_excess_object_version_size_total{bucket}` | `s3:Scanner:LargeVersions` | `s3:ObjectLargeVersions` |
| 子目录数超限 | 单目录直接子目录数 > `scanner:alert_excess_folders` | `rustfs_scanner_excess_folders_total{root}` | `s3:Scanner:BigPrefix` | `s3:PrefixManyFolders` |
订阅方式与普通桶通知一致:对目标桶配置 notification,事件名填上表 RustFS wire 名(或通配 `s3:Scanner:*`)。事件以 `UserAgent: Scanner` 标记来源,`req_params` 携带实际值与阈值(`versions` / `cumulativeSize` / `folders` / `threshold`),便于消费方直接判断严重程度。
## 指标与事件的触发节奏不同
- **指标与结构化日志是电平触发**:只要对象仍在阈值之上,每个扫描周期都会计数/打日志(默认周期约 60s,见 `scanner:speed`)。
- **S3 事件是边沿触发 + 冷却**:同一 (告警类型, 桶, 对象) 在冷却窗口内只发一次,默认 24 小时(`RUSTFS_SCANNER_ALERT_COOLDOWN_SECS`,设 0 表示每周期都发)。窗口过后对象仍超限会再次发出。冷却表在进程内有 4096 条硬顶,超限清空重建(最坏情况是每个仍超限的 key 多发一次)。
- 进程重启会重置冷却(重启后每个仍超限的对象会再发一次)——这是有意为之:重启常伴随排障,重发提供可见性。
## 阈值默认值与 MinIO 差异(HS-15
| 配置键 | ENV | RustFS 默认 | MinIO 默认 | 差异说明 |
|---|---|---|---|---|
| `scanner:alert_excess_versions` | `RUSTFS_SCANNER_ALERT_EXCESS_VERSIONS` | 100 | 100 | 一致 |
| `scanner:alert_excess_version_size` | `RUSTFS_SCANNER_ALERT_EXCESS_VERSION_SIZE` | 1 TiB | 1 TB | 语义同量级,单位口径不同(TiB vs TB) |
| `scanner:alert_excess_folders` | `RUSTFS_SCANNER_ALERT_EXCESS_FOLDERS` | 65538 | 50000 | **有意差异**65538 兼容 Proxmox Backup Server 的 chunk 布局(每目录 65536 个 chunk + 目录自身条目),按 MinIO 的 50000 会对 PBS 用户持续误报。如需与 MinIO 行为一致可显式配置为 50000 |
三个键均支持 env 与 admin config`PUT /rustfs/admin/v3/config``scanner` 子系统)双通道,热更新即时生效。
## 事件名映射的由来
RustFS 的事件枚举(`rustfs_s3_types::EventName::ScannerManyVersions/LargeVersions/BigPrefix`)沿用仓库既有 wire 名 `s3:Scanner:*`(与 MinIO 的 `s3:ObjectManyVersions` 字面不同,枚举注释中保留了映射关系)。订阅方应以本文的 RustFS wire 名为准;如需 MinIO 字面兼容,请在 console/消费侧做名称映射,不要修改已发布的 wire 名。
+210 -333
View File
@@ -1,150 +1,114 @@
# Scanner Runtime Controls
This document describes the runtime controls and status fields for the RustFS
data scanner. It is written for operators who need to reduce scanner pressure,
diagnose slow scan progress, or confirm that background lifecycle, replication,
heal, bitrot, and usage work is still moving.
**Use this when:** tuning scanner pacing or heal runtime knobs, reading `/v3/scanner/status`, deciding whether slow lifecycle/replication/heal progress is a scanner problem or a downstream queue problem, or migrating scanner settings from MinIO.
For reproducible scanner-pressure validation and before/after evidence, see
[Scanner Benchmark Runbook](scanner-benchmark-runbook.md).
**Source of truth:** `crates/config/src/constants/scanner.rs` and `crates/config/src/constants/heal.rs` (keys, env names, `DEFAULT_*` constants), `crates/scanner/src/runtime_config.rs` (resolution order, parsing, hot update), `crates/scanner/src/scanner_folder.rs` (env-only scanner knobs, alerts), `crates/scanner/src/sleeper.rs` (throttling), `crates/heal/src/heal/manager.rs` (`HealConfig::default`), `crates/utils/src/envs.rs` (`EXTERNAL_COMPATIBLE_SUFFIXES`, MinIO env aliases).
For reproducible scanner-pressure validation and before/after evidence, see [Scanner Benchmark Runbook](scanner-benchmark-runbook.md). Alert thresholds and S3 event names are detailed in [Scanner Excess Alerts](scanner-excess-alerts.md).
## What the scanner does
The scanner is the background maintenance loop that walks stored objects and
feeds several subsystems:
- usage accounting and data usage cache updates;
- lifecycle expiry and transition admission;
- bucket replication repair admission;
- scanner-originated heal and bitrot checks;
- namespace alerts for excessive versions, retained version size, and folder
fan-out.
Slowing the scanner can reduce idle CPU and disk pressure, but it also delays
the maintenance work above. Prefer using the status fields below before changing
cycle or pacing values.
The scanner is the background maintenance loop that walks stored objects and feeds usage accounting, lifecycle expiry and transition admission, bucket replication repair admission, scanner-originated heal and bitrot checks, and the namespace excess alerts. Slowing it reduces idle CPU and disk pressure but delays all of that work. Read the status fields below before changing cycle or pacing values.
## Configuration Sources
Scanner runtime config is resolved in this order:
1. Environment variables.
2. Persisted admin config for the `scanner` subsystem.
2. Persisted admin config for the `scanner` subsystem (`SCANNER_SUB_SYS`).
3. Built-in defaults or speed preset-derived values.
Bitrot cycle resolution is slightly different because the canonical persistent
key belongs to the `heal` subsystem:
Bitrot cycle resolution differs because the canonical persistent key belongs to the `heal` subsystem:
1. `RUSTFS_SCANNER_BITROT_CYCLE_SECS`.
2. `heal.bitrot_cycle`.
3. Legacy compatibility key `scanner.bitrot_cycle`.
4. Built-in default.
The `/v3/scanner/status` response reports each effective runtime value with a
`source` of `env`, `config`, `scanner_compat_config`, or `default`.
`/v3/scanner/status` reports each effective runtime value with a `source` of `env`, `config`, `scanner_compat_config`, or `default`. Every key in the table below accepts admin config updates that take effect without a restart (`apply_scanner_runtime_config`).
## Runtime Controls
| Persistent key | Environment variable | Unit | Default | Effect |
|---|---|---:|---:|---|
| `scanner.speed` | `RUSTFS_SCANNER_SPEED` | preset | `default` | Selects the base pacing preset: `fastest`, `fast`, `default`, `slow`, or `slowest`. |
| `scanner.delay` | `RUSTFS_SCANNER_DELAY` | factor | preset-derived | Overrides the sleep multiplier. Valid range is `0` through `10000`. |
| Persistent key | Environment variable | Unit | Default (constant) | Effect |
|---|---|---:|---|---|
| `scanner.speed` | `RUSTFS_SCANNER_SPEED` | preset | `default` (`DEFAULT_SCANNER_SPEED`) | Selects the base pacing preset: `fastest`, `fast`, `default`, `slow`, or `slowest`. |
| `scanner.delay` | `RUSTFS_SCANNER_DELAY` | factor | preset-derived | Overrides the sleep multiplier. Valid range is `0` through `10000` (`MAX_SCANNER_DELAY_FACTOR`). |
| `scanner.max_wait` | `RUSTFS_SCANNER_MAX_WAIT_SECS` | seconds | preset-derived | Caps one scanner sleep. |
| `scanner.cycle` | `RUSTFS_SCANNER_CYCLE` | seconds | preset-derived | Sets the interval between scanner cycles. |
| `scanner.start_delay` | `RUSTFS_SCANNER_START_DELAY_SECS` | seconds | unset | Sets startup delay and, for compatibility, the cycle interval when `scanner.cycle` is unset. |
| `scanner.cycle_max_duration` | `RUSTFS_SCANNER_CYCLE_MAX_DURATION_SECS` | seconds | `1800` | Caps one cycle's runtime. An explicit `0` disables this budget. |
| `scanner.cycle_max_objects` | `RUSTFS_SCANNER_CYCLE_MAX_OBJECTS` | objects | `0` | Caps objects processed by one cycle. `0` disables this budget. |
| `scanner.cycle_max_directories` | `RUSTFS_SCANNER_CYCLE_MAX_DIRECTORIES` | directories | `0` | Caps directories entered by one cycle. `0` disables this budget. |
| `heal.bitrot_cycle` | `RUSTFS_SCANNER_BITROT_CYCLE_SECS` | seconds | `2592000` | Controls periodic deep bitrot scans. `false`, `off`, `no`, or `disabled` disables periodic deep scans; `0`, `true`, `on`, or `yes` runs deep mode every scanner cycle. |
| `scanner.idle_mode` | `RUSTFS_SCANNER_IDLE_MODE` | boolean | `true` | Enables scanner sleeps and cooperative throttling. |
| `scanner.cache_save_timeout` | `RUSTFS_SCANNER_CACHE_SAVE_TIMEOUT_SECS` | seconds | `14` | Timeout for saving scanner cache; runtime enforces a minimum of `1` and keeps the default persistence budget within the distributed publication lease. |
| `scanner.max_concurrent_set_scans` | `RUSTFS_SCANNER_MAX_CONCURRENT_SET_SCANS` | count | `4` | Caps concurrent set-level scanner tasks. `0` keeps topology-derived concurrency. |
| `scanner.max_concurrent_disk_scans` | `RUSTFS_SCANNER_MAX_CONCURRENT_DISK_SCANS` | count | `4` | Caps concurrent disk bucket walks per set. `0` keeps disk-count-derived concurrency. |
| `scanner.yield_every_n_objects` | `RUSTFS_SCANNER_YIELD_EVERY_N_OBJECTS` | objects | `128` | Controls how often object loops yield to the async runtime. `0` disables this extra yield. |
| `scanner.alert_excess_versions` | `RUSTFS_SCANNER_ALERT_EXCESS_VERSIONS` | versions | `100` | Version count threshold for scanner alerts. |
| `scanner.alert_excess_version_size` | `RUSTFS_SCANNER_ALERT_EXCESS_VERSION_SIZE` | bytes | `1099511627776` | Retained version byte threshold for scanner alerts. |
| `scanner.alert_excess_folders` | `RUSTFS_SCANNER_ALERT_EXCESS_FOLDERS` | folders | `65538` | Direct subfolder threshold for scanner alerts. |
| `scanner.start_delay` | `RUSTFS_SCANNER_START_DELAY_SECS` (deprecated alias `RUSTFS_DATA_SCANNER_START_DELAY_SECS`) | seconds | unset | Sets startup delay and, for compatibility, the cycle interval when `scanner.cycle` is unset. |
| `scanner.cycle_max_duration` | `RUSTFS_SCANNER_CYCLE_MAX_DURATION_SECS` | seconds | `1800` (`DEFAULT_SCANNER_CYCLE_MAX_DURATION_SECS`) | Caps one cycle's runtime. An explicit `0` disables this budget. |
| `scanner.cycle_max_objects` | `RUSTFS_SCANNER_CYCLE_MAX_OBJECTS` | objects | `0` (`DEFAULT_SCANNER_CYCLE_MAX_OBJECTS`) | Caps objects processed by one cycle. `0` disables this budget. |
| `scanner.cycle_max_directories` | `RUSTFS_SCANNER_CYCLE_MAX_DIRECTORIES` | directories | `0` (`DEFAULT_SCANNER_CYCLE_MAX_DIRECTORIES`) | Caps directories entered by one cycle. `0` disables this budget. |
| `heal.bitrot_cycle` | `RUSTFS_SCANNER_BITROT_CYCLE_SECS` | seconds | `2592000` (`DEFAULT_HEAL_BITROT_CYCLE_SECS`, 30 days) | Controls periodic deep bitrot scans. `false`, `off`, `no`, or `disabled` disables periodic deep scans; `0`, `true`, `on`, or `yes` runs deep mode every scanner cycle. |
| `scanner.idle_mode` | `RUSTFS_SCANNER_IDLE_MODE` | boolean | `true` (`DEFAULT_SCANNER_IDLE_MODE`) | Master switch for scanner throttling: preset sleeps plus the foreground-read backoff floor. `false` disables both and the scanner runs at full speed. |
| `scanner.cache_save_timeout` | `RUSTFS_SCANNER_CACHE_SAVE_TIMEOUT_SECS` | seconds | `14` (`DEFAULT_SCANNER_CACHE_SAVE_TIMEOUT_SECS`) | Timeout for saving scanner cache; runtime enforces a minimum of `1` and keeps the default persistence budget within the distributed publication lease. |
| `scanner.max_concurrent_set_scans` | `RUSTFS_SCANNER_MAX_CONCURRENT_SET_SCANS` | count | `4` (`DEFAULT_SCANNER_MAX_CONCURRENT_SET_SCANS`) | Caps concurrent set-level scanner tasks. `0` keeps topology-derived concurrency. |
| `scanner.max_concurrent_disk_scans` | `RUSTFS_SCANNER_MAX_CONCURRENT_DISK_SCANS` | count | `4` (`DEFAULT_SCANNER_MAX_CONCURRENT_DISK_SCANS`) | Caps concurrent disk bucket walks per set. `0` keeps disk-count-derived concurrency. |
| `scanner.yield_every_n_objects` | `RUSTFS_SCANNER_YIELD_EVERY_N_OBJECTS` | objects | `128` (`DEFAULT_SCANNER_YIELD_EVERY_N_OBJECTS`) | Controls how often object loops yield to the async runtime. `0` disables this extra yield. |
| `scanner.alert_excess_versions` | `RUSTFS_SCANNER_ALERT_EXCESS_VERSIONS` | versions | `100` (`DEFAULT_SCANNER_ALERT_EXCESS_VERSIONS`) | Version count threshold for scanner alerts. |
| `scanner.alert_excess_version_size` | `RUSTFS_SCANNER_ALERT_EXCESS_VERSION_SIZE` | bytes | `1099511627776` (`DEFAULT_SCANNER_ALERT_EXCESS_VERSION_SIZE`) | Retained version byte threshold for scanner alerts. |
| `scanner.alert_excess_folders` | `RUSTFS_SCANNER_ALERT_EXCESS_FOLDERS` | folders | `65538` (`DEFAULT_SCANNER_ALERT_EXCESS_FOLDERS`) | Direct subfolder threshold for scanner alerts. |
The `fastest`, `fast`, `default`, `slow`, and `slowest` presets set the base
sleep multiplier, maximum wait, and cycle interval. Use `scanner.delay`,
`scanner.max_wait`, and `scanner.cycle` when the preset is close but one axis
needs a precise override.
Speed presets (`crates/config/src/constants/scanner.rs`) set the base sleep multiplier, maximum wait, and cycle interval:
When the cycle duration control is unset, RustFS uses a finite 1800-second
(30-minute) default, matching the scanner benchmark guidance. An explicit `0`
preserves the compatibility behavior of an unbounded cycle; object and
directory budgets likewise remain unbounded when explicitly set to `0`. Invalid
or overflowing duration environment values are configuration errors rather than
silent fallback values.
| Preset | Sleep factor | Max sleep | Cycle interval |
|---|---:|---:|---:|
| `fastest` | 0 | 0 | 1s |
| `fast` | 1x | 100ms | 60s |
| `default` | 2x | 1s | 60s |
| `slow` | 10x | 15s | 60s |
| `slowest` | 100x | 15s | 30m |
When a finite deadline expires, RustFS cancels cooperative scanner work and
waits only for the existing bounded shutdown window. A non-yielding I/O future
is dropped after that window. RustFS then attempts a higher leadership epoch so
late cycle, usage, cache, and remote writes from the old generation fail closed.
If the worker cannot stop cooperatively, the cycle state was not confirmed
durable, or that epoch fence cannot be durably persisted, the scanner reports
`recovery-required`; it does not claim an uncooperative cursor was saved.
Use `scanner.delay`, `scanner.max_wait`, and `scanner.cycle` when the preset is close but one axis needs a precise override. With `idle_mode=true`, directory-level sleep is `1ms x factor` and object-level sleep is `time spent on the object x factor`, both capped at `max_wait`; a foreground-read floor of `FOREGROUND_READ_BACKOFF_PER_REQUEST_MS` (10ms) per concurrent GetObject/streaming read, capped at `FOREGROUND_READ_BACKOFF_MAX_MS` (250ms), is applied on top and can exceed the preset's `max_wait` (`crates/scanner/src/sleeper.rs`).
An explicit `scanner.cycle` or `RUSTFS_SCANNER_CYCLE` is a minimum inter-cycle
cadence: dirty-usage notifications do not bypass that configured interval.
The default adaptive policy continues to use dirty-usage notifications to wake
the scanner between timer-driven cycles.
### Environment-only scanner knobs
These have no persistent key and are read from the environment only.
| Environment variable | Default (constant) | Effect |
|---|---|---|
| `RUSTFS_SCANNER_ENABLED` (deprecated alias `RUSTFS_ENABLE_SCANNER`) | `true` (`scanner_enabled_from_env`, `rustfs/src/module_switches.rs`) | Starts the data scanner at all. The heal manager is initialized whenever heal or scanner is enabled, because scanner-produced heal candidates need a consumer. |
| `RUSTFS_SCANNER_ALERT_COOLDOWN_SECS` | `86400` (`DEFAULT_SCANNER_ALERT_COOLDOWN_SECS`, `scanner_folder.rs`) | Per-(kind, bucket, object) cooldown between S3 excess-alert events; `0` emits every cycle. See [Scanner Excess Alerts](scanner-excess-alerts.md). |
| `RUSTFS_SCANNER_DEEP_VERIFY_COOLDOWN_SECS` | `60` (`DEFAULT_SCANNER_DEEP_VERIFY_COOLDOWN_SECS`, `scanner_folder.rs`) | Objects modified within this window are skipped by deep (bitrot) verification in the current cycle. |
| `RUSTFS_HEAL_OBJECT_SELECT_PROB` | `1024` (`DEFAULT_HEAL_OBJECT_SELECT_PROB`, `scanner_folder.rs`) | Sampling divisor for scanner-originated heal checks: roughly one object in N per cycle is selected for a low-priority heal check. |
| `RUSTFS_DATA_USAGE_UPDATE_DIR_CYCLES` | `16` (`DATA_USAGE_UPDATE_DIR_CYCLES`, `scanner_folder.rs`) | Every N cycles a compacted directory is re-descended instead of reusing its cached usage. `1` forces re-descent every cycle (used by lifecycle e2e lanes). |
| `RUSTFS_DATA_USAGE_FAILED_OBJECT_TTL_SECS` | `86400` (`DEFAULT_FAILED_OBJECT_TTL_SECS`, `scanner_folder.rs`) | Retention of per-bucket failed-object entries in the usage cache. |
| `RUSTFS_DATA_USAGE_FAILED_OBJECTS_MAX` | `10000` (`DEFAULT_FAILED_OBJECTS_MAX`, `scanner_folder.rs`) | Cap on retained failed-object entries per bucket. |
### Cycle budgets and cadence
When the cycle duration control is unset, RustFS uses the finite 1800-second default. An explicit `0` preserves the compatibility behavior of an unbounded cycle; object and directory budgets likewise remain unbounded when explicitly set to `0`. Invalid or overflowing duration environment values are configuration errors rather than silent fallback values.
When a finite deadline expires, RustFS cancels cooperative scanner work and waits only for the existing bounded shutdown window. A non-yielding I/O future is dropped after that window. RustFS then attempts a higher leadership epoch so late cycle, usage, cache, and remote writes from the old generation fail closed. If the worker cannot stop cooperatively, the cycle state was not confirmed durable, or that epoch fence cannot be durably persisted, the scanner reports `recovery-required`; it does not claim an uncooperative cursor was saved.
An explicit `scanner.cycle` or `RUSTFS_SCANNER_CYCLE` is a minimum inter-cycle cadence: dirty-usage notifications do not bypass that configured interval. The default adaptive policy continues to use dirty-usage notifications to wake the scanner between timer-driven cycles.
## Single-disk clean-idle scheduling
An erasure single-disk deployment using the built-in cycle and bitrot defaults
automatically backs off repeated clean idle scans instead of walking the same
unchanged namespace every minute. Each successful timer-driven cycle that
finds no dirty usage or unresolved maintenance work doubles the next interval.
The status endpoint reports the effective interval and multiplier.
An erasure single-disk deployment using the built-in cycle and bitrot defaults automatically backs off repeated clean idle scans instead of walking the same unchanged namespace every minute. Each successful timer-driven cycle that finds no dirty usage or unresolved maintenance work doubles the next interval. The status endpoint reports the effective interval and multiplier.
The backoff is reset to the base interval by object or bucket mutations,
lifecycle or replication configuration changes, partial or failed cycles,
usage persistence failures, and unresolved scanner-originated heal or bitrot
work. Active lifecycle or replication rules keep the base cadence. An explicit
cycle, a non-default persisted speed, any environment speed or start-delay
override, an environment bitrot override, or a non-default persisted active
bitrot cycle also keeps the configured cadence rather than applying the
automatic policy. Persisting `scanner.speed=default` or the default bitrot cycle
is normalized to the built-in default and therefore keeps automatic scheduling
enabled.
The backoff is reset to the base interval by object or bucket mutations, lifecycle or replication configuration changes, partial or failed cycles, usage persistence failures, and unresolved scanner-originated heal or bitrot work. Active lifecycle or replication rules keep the base cadence. An explicit cycle, a non-default persisted speed, any environment speed or start-delay override, an environment bitrot override, or a non-default persisted active bitrot cycle also keeps the configured cadence rather than applying the automatic policy. Persisting `scanner.speed=default` or the default bitrot cycle is normalized to the built-in default and therefore keeps automatic scheduling enabled.
Lifecycle and replication configuration inspection is bounded so a slow
metadata read cannot stall scanner startup or scheduling. A failed or timed-out
inspection keeps the base cadence and is retried after 5 minutes, doubling up
to a maximum of 60 minutes while failures continue. A lifecycle or replication
configuration change wakes the scanner and retries inspection immediately.
Lifecycle and replication configuration inspection is bounded so a slow metadata read cannot stall scanner startup or scheduling. A failed or timed-out inspection keeps the base cadence and is retried after 5 minutes, doubling up to a maximum of 60 minutes while failures continue. A lifecycle or replication configuration change wakes the scanner and retries inspection immediately.
With the default 30-day bitrot cycle, the clean-idle interval is capped at the
bitrot cycle divided by the object selection window. With the default selection
window this is about 42 minutes, which preserves the intended wall-clock bitrot
coverage. If periodic bitrot is disabled, the clean-idle policy cap is 24 hours.
The effective interval is jittered by up to 10 percent to avoid synchronized
scanner starts.
With the default 30-day bitrot cycle, the clean-idle interval is capped at the bitrot cycle divided by the object selection window (about 42 minutes with the default `RUSTFS_HEAL_OBJECT_SELECT_PROB`), which preserves the intended wall-clock bitrot coverage. If periodic bitrot is disabled, the clean-idle cap is 24 hours. The effective interval is jittered by up to 10 percent to avoid synchronized scanner starts.
## Status Endpoint
The scanner status route is:
```text
GET /v3/scanner/status
```
The request must be authenticated with an admin identity that has
`ServerInfoAdminAction`. The JSON response has three scanner-specific top-level
objects:
The request must be authenticated with an admin identity that has `ServerInfoAdminAction`. The JSON response has these scanner-specific top-level objects:
- `runtime_config`: the effective runtime controls and their value sources.
- `cycle_schedule`: the current effective cycle interval and clean-idle
backoff state.
- `metrics`: scanner work, pressure, checkpoint, lifecycle, replication, heal,
bitrot, and alert counters.
- `data_movement_pause`: the global-pause policy, current movement reason,
operation epoch, start time, duration, and estimated movement work items.
- `pause_backlog`: the replicated durable pause ledger, post-pause catch-up
phase, rate window, retry state, thresholds, and active alert reasons.
- `catch_up_estimate`: movement work plus current dirty-usage and already
discovered lifecycle queues.
| Object | Content |
|---|---|
| `runtime_config` | Effective runtime controls and their value sources. |
| `cycle_schedule` | Current effective cycle interval and clean-idle backoff state. |
| `metrics` | Scanner work, pressure, checkpoint, lifecycle, replication, heal, bitrot, and alert counters. |
| `data_movement_pause` | Global-pause policy, current movement reason, operation epoch, start time, duration, and estimated movement work items. |
| `pause_backlog` | Replicated durable pause ledger, post-pause catch-up phase, rate window, retry state, thresholds, and active alert reasons. |
| `catch_up_estimate` | Movement work plus current dirty-usage and already discovered lifecycle queues. |
Example fields to inspect:
@@ -189,28 +153,9 @@ catch_up_estimate.discovered_transition_items
## Usage State Reset
The supported break-glass route for rebuilding scanner usage state is:
The supported break-glass route for rebuilding scanner usage state is `POST /v3/scanner/usage-state/reset` with body `{"mode":"full-rebuild"}`, authenticated as an admin identity holding `ConfigUpdateAdminAction` (route registered in `rustfs/src/admin/route_registration_test.rs`). Use it only after the scanner status shows a usage-floor load failure, a conflicting persisted usage floor, or an operator decision to discard the durable usage baseline and rebuild it from a full scanner pass.
```text
POST /v3/scanner/usage-state/reset
{"mode":"full-rebuild"}
```
The request must be authenticated with an admin identity that has
`ConfigUpdateAdminAction`. Use it only after confirming the scanner status shows
a usage-floor load failure, a conflicting persisted usage floor, or an operator
decision to discard the durable usage baseline and rebuild it from a full
scanner pass.
The reset does not delete metadata files from disk by hand and does not publish
an authoritative zero-usage snapshot. It holds the scanner leader lock, fences
the operation with the storage-owned publication epoch, CAS-publishes a v2
`bootstrap-pending` marker in the primary usage slot, then clears stale backup,
legacy, and observed usage slots by object revision. The next scanner
leadership claim binds that marker to a fresh epoch and the next complete
scanner cycle replaces it with authoritative usage.
The JSON response is machine-readable:
The reset does not delete metadata files by hand and does not publish an authoritative zero-usage snapshot. It holds the scanner leader lock, fences the operation with the storage-owned publication epoch, CAS-publishes a v2 `bootstrap-pending` marker in the primary usage slot, then clears stale backup, legacy, and observed usage slots by object revision. The next scanner leadership claim binds that marker to a fresh epoch, and the next complete scanner cycle replaces it with authoritative usage.
```json
{
@@ -229,92 +174,57 @@ The JSON response is machine-readable:
}
```
If the response is an error mentioning data movement, wait for decommission or
rebalance to leave the scanner metadata path and retry. If it reports that the
scanner cycle state is invalid, run the cycle-state recovery reset first:
```text
POST /v3/scanner/cycle-state/reset
{"mode":"full-rescan"}
```
| Error mentions | Do |
|---|---|
| data movement | wait for decommission or rebalance to leave the scanner metadata path, then retry |
| invalid scanner cycle state | run `POST /v3/scanner/cycle-state/reset` with `{"mode":"full-rescan"}` first |
## Data Movement Pauses
RustFS currently uses a `global_pause` policy while pool decommission or
rebalance can hide scanner metadata. Usage publication, lifecycle discovery,
tier cleanup discovery, scanner-originated heal and bitrot checks, and
replication discovery are deferred together. A failed or canceled
decommission remains a publication barrier until an operator retries or clears
it.
RustFS uses a `global_pause` policy while pool decommission or rebalance can hide scanner metadata: usage publication, lifecycle discovery, tier cleanup discovery, scanner-originated heal and bitrot checks, and replication discovery are deferred together. A failed or canceled decommission remains a publication barrier until an operator retries or clears it. The same pause and estimate objects are included in `GET /v3/ilm/expiry/status`.
`data_movement_pause.reasons` combines the in-process decommission worker state
with the durable pool and rebalance operation metadata. Exhausted operation
epochs or movement generations also fail closed and appear as explicit pause
reasons. Its start time, duration, and movement backlog come from the durable
metadata; a worker-only or exhausted-counter snapshot can therefore report
`paused=true` with zero start time and backlog.
`movement_backlog_work_items` counts remaining movement bucket work units, not
expired objects. `catch_up_estimate` combines that estimate with dirty-usage
buckets and lifecycle items that were already discovered before or during the
pause. The API sets `undiscovered_ilm_items_known=false` because a global pause
cannot count newly expired objects without scanning the namespace. Use
`usage_baseline_unix_secs` to judge the age of that estimate.
| Field | Meaning |
|---|---|
| `data_movement_pause.reasons` | In-process decommission worker state combined with durable pool and rebalance operation metadata. Exhausted operation epochs or movement generations fail closed and appear as explicit reasons. |
| `data_movement_pause.duration_seconds`, start time, `movement_backlog_work_items` | From durable metadata; a worker-only or exhausted-counter snapshot can report `paused=true` with zero start time and backlog. `movement_backlog_work_items` counts remaining movement bucket work units, not expired objects. |
| `catch_up_estimate` | Movement estimate plus dirty-usage buckets and lifecycle items discovered before or during the pause. `undiscovered_ilm_items_known=false` because a global pause cannot count newly expired objects without scanning; use `usage_baseline_unix_secs` to judge the estimate's age. |
| `pause_backlog.persistence_state` | `persistence_unavailable` when the `.scanner-pause-backlog.json` ledger cannot be read or updated; scanner cycles stay gated and persistence is retried every five minutes. |
| `pause_backlog.phase` | `idle`, `paused`, `catching_up`, or `retry_exhausted`. The ledger returns to `idle` only after one successful full namespace scan and zero known dirty-usage, expiry, and transition queues. |
| `membership_repair_pending` | A rejoining decommission source is being re-seeded from the last committed surviving-set ledger before a new full-membership commit is allowed. |
| `pause_backlog.thresholds` | Exact pause-duration, deferred-cycle, backlog-size, rate, and failure limits used by the running binary. |
| `pause_backlog.alert_reasons` | Exceeded thresholds, exhausted counters or retries, replica degradation, and persistence failures. |
The same pause and estimate objects are included in
`GET /v3/ilm/expiry/status`. The gauges
`rustfs_scanner_data_movement_paused`,
`rustfs_scanner_data_movement_pause_duration_seconds`, and
`rustfs_scanner_data_movement_backlog_work_items` expose the local snapshot
without bucket-name labels.
Built-in limits reported under `pause_backlog.thresholds`:
The scanner persists `.scanner-pause-backlog.json` independently on erasure
sets in every surviving pool. A generation becomes authoritative only after
the identical commit record reaches every set named by its membership marker.
When a failed, canceled, or cleared decommission source rejoins, the last
committed surviving-set ledger seeds it before a new full-membership commit is
allowed; a smaller stale source membership cannot override the largest valid
surviving-set proof, and a membership claim is valid only when every declared
member stores the same proof. This repair appears as
`membership_repair_pending`. A partial commit is
rolled back to the previous stable generation after a crash or leader switch.
The ledger never rewrites pool or rebalance movement state. A new scanner
leader recovers the committed writer epoch and generation, counts an
interrupted attempt as a failure, and requires one successful full namespace
scan after movement clears. Known dirty-usage, expiry, and transition queues
must also reach zero before the ledger returns to `idle`. If the ledger cannot
be read or updated, scanner cycles remain gated and persistence is retried
every five minutes; the management status reports `persistence_unavailable`
until recovery.
| Limit | Value |
|---|---|
| Pause-duration alert | 24 hours |
| Movement-deferral alert | 3 deferrals in one unconverged pause episode |
| Backlog alert | 10,000 known pending work items |
| Catch-up rate window | At most 4 attempts per hour, no more than 1 per 5 minutes |
| Retry exhaustion | 5 consecutive failed or interrupted attempts, then a sparse hourly probe |
Catch-up attempts remain subject to the normal cycle duration, object,
directory, sleeper, and foreground-read budgets. The additional durable rate
window admits at most four attempts per hour and no more than one attempt per
five minutes. Five consecutive failed or interrupted attempts move the ledger
to `retry_exhausted`; accelerated retries stop and a sparse hourly probe is
used instead. A successful probe can return to bounded catch-up.
Catch-up attempts remain subject to the normal cycle duration, object, directory, sleeper, and foreground-read budgets.
`pause_backlog.thresholds` reports the exact pause-duration, deferred-cycle,
backlog-size, rate, and failure limits used by the running binary.
`pause_backlog.alert_reasons` identifies exceeded thresholds, exhausted
counters or retries, replica degradation, and persistence failures. The
threshold alerts fire after a 24-hour pause, three movement deferrals in one
unconverged pause episode, or 10,000 known pending work items. The
corresponding unlabeled gauges are:
Unlabeled Prometheus gauges:
- `rustfs_scanner_pause_backlog_phase` (`0` idle, `1` paused, `2` catching up,
`3` retry exhausted);
- `rustfs_scanner_pause_backlog_pause_duration_seconds`;
- `rustfs_scanner_pause_backlog_pending_work_items`;
- `rustfs_scanner_pause_backlog_consecutive_failures`;
- `rustfs_scanner_pause_backlog_rate_limited`;
- `rustfs_scanner_pause_backlog_retry_exhausted`;
- `rustfs_scanner_pause_backlog_alerting`;
- `rustfs_scanner_pause_backlog_replica_degraded`.
| Gauge | Meaning |
|---|---|
| `rustfs_scanner_data_movement_paused` | Local pause snapshot. |
| `rustfs_scanner_data_movement_pause_duration_seconds` | Local pause duration. |
| `rustfs_scanner_data_movement_backlog_work_items` | Local movement backlog. |
| `rustfs_scanner_pause_backlog_phase` | `0` idle, `1` paused, `2` catching up, `3` retry exhausted. |
| `rustfs_scanner_pause_backlog_pause_duration_seconds` | Durable pause duration. |
| `rustfs_scanner_pause_backlog_pending_work_items` | Known pending work items. |
| `rustfs_scanner_pause_backlog_consecutive_failures` | Consecutive failed catch-up attempts. |
| `rustfs_scanner_pause_backlog_rate_limited` | Catch-up currently rate limited. |
| `rustfs_scanner_pause_backlog_retry_exhausted` | Ledger in `retry_exhausted`. |
| `rustfs_scanner_pause_backlog_alerting` | Any alert reason active. |
| `rustfs_scanner_pause_backlog_replica_degraded` | Ledger replica set degraded. |
## Reading Pacing Pressure
`metrics.pacing_pressure.primary_pressure` summarizes the highest-priority
scanner pressure signal:
`metrics.pacing_pressure.primary_pressure` summarizes the highest-priority scanner pressure signal:
| Value | Meaning | Usual response |
|---|---|---|
@@ -324,49 +234,21 @@ scanner pressure signal:
| `active_scans` | Scanner work is active but not currently queued or budget-limited. | Usually healthy; correlate with CPU/disk metrics. |
| `none` | No current scanner pressure was observed. | No scanner pacing action needed. |
The ratio fields are fractions of the last cycle duration:
- `last_cycle_throttle_sleep_ratio`
- `last_cycle_yield_ratio`
- `last_cycle_total_pause_ratio`
If CPU is high but pause ratios are already high, increasing `scanner.delay` or
`scanner.max_wait` may have limited value. Check active paths, source work, and
disk activity before changing the cycle interval.
The ratio fields `last_cycle_throttle_sleep_ratio`, `last_cycle_yield_ratio`, and `last_cycle_total_pause_ratio` are fractions of the last cycle duration. If CPU is high but pause ratios are already high, increasing `scanner.delay` or `scanner.max_wait` may have limited value; check active paths, source work, and disk activity before changing the cycle interval.
## Reading Source Work
`metrics.source_work`, `metrics.current_cycle_source_work`, and
`metrics.last_cycle_source_work` group scanner work by source:
`metrics.source_work`, `metrics.current_cycle_source_work`, and `metrics.last_cycle_source_work` group scanner work by source: `usage`, `lifecycle`, `bucket_replication`, `site_replication`, `heal`, `bitrot`, `alerts`.
- `usage`
- `lifecycle`
- `bucket_replication`
- `site_replication`
- `heal`
- `bitrot`
- `alerts`
Each source has `checked`, `queued`, `executed`, `failed`, `skipped`, and
`missed` counters. `missed` means the scanner found work but could not admit it
to the downstream queue. `skipped` means the work was intentionally merged or
deduplicated.
Use these counters to decide whether scan progress is limited by scanner pacing
or by a downstream subsystem such as lifecycle transition, replication repair,
or heal admission.
Each source has `checked`, `queued`, `executed`, `failed`, `skipped`, and `missed` counters. `missed` means the scanner found work but could not admit it to the downstream queue. `skipped` means the work was intentionally merged or deduplicated. Use these counters to decide whether scan progress is limited by scanner pacing or by a downstream subsystem such as lifecycle transition, replication repair, or heal admission.
## Reading Heal Operations
The background heal status route is:
```text
POST /v3/background-heal/status
```
It reports scanner-driven bitrot state together with heal queue execution
state. `healQueueLength` and `healActiveTasks` keep the legacy totals.
`healOperations` adds the same totals split by request source and priority:
Reports scanner-driven bitrot state together with heal queue execution state. `healQueueLength` and `healActiveTasks` keep the legacy totals; `healOperations` adds the same totals split by request source and priority:
| Field | Meaning |
|---|---|
@@ -377,32 +259,83 @@ state. `healQueueLength` and `healActiveTasks` keep the legacy totals.
| `queuedByPriority` | Queued requests split into `low`, `normal`, `high`, and `urgent`. |
| `activeByPriority` | Running tasks split into `low`, `normal`, `high`, and `urgent`. |
Use this route when `metrics.source_work` shows `heal` or `bitrot` queued or
missed work. Scanner-originated object checks should appear under
`scanner/low` for opportunistic work, while manual admin heal should appear
under `admin/high`. If scanner work grows but admin work remains blocked, treat
that as heal queue pressure rather than scanner pacing pressure.
Use this route when `metrics.source_work` shows `heal` or `bitrot` queued or missed work. Scanner-originated object checks should appear under `scanner/low`, manual admin heal under `admin/high`. If scanner work grows but admin work remains blocked, treat that as heal queue pressure rather than scanner pacing pressure.
## Heal runtime controls
Heal knobs are environment-only and read by `HealConfig::default` (`crates/heal/src/heal/manager.rs`), the MRF queue (`crates/heal/src/heal/mrf_queue.rs`), or the erasure-set healer (`crates/heal/src/heal/erasure_healer.rs`). The admin `heal` config subsystem accepts only `bitrot_cycle` (`HEAL_KEYS`), which is documented in the scanner table above. Constants live in `crates/config/src/constants/heal.rs` unless another file is named.
| Environment variable | Default (constant) | Effect |
|---|---|---|
| `RUSTFS_HEAL_ENABLED` (deprecated alias `RUSTFS_ENABLE_HEAL`) | `true` (`heal_enabled_from_env`, `rustfs/src/module_switches.rs`) | Master switch for the background heal manager. |
| `RUSTFS_HEAL_AUTO_HEAL_ENABLE` | `true` (`DEFAULT_HEAL_AUTO_HEAL_ENABLE`) | Enables automatic healing of detected issues; `false` leaves healing to manual admin requests. |
| `RUSTFS_HEAL_QUEUE_SIZE` | `10000` (`DEFAULT_HEAL_QUEUE_SIZE`) | Heal request queue capacity. |
| `RUSTFS_HEAL_INTERVAL_SECS` | `10` (`DEFAULT_HEAL_INTERVAL_SECS`) | Heal manager polling interval. |
| `RUSTFS_HEAL_TASK_TIMEOUT_SECS` | `300` (`DEFAULT_HEAL_TASK_TIMEOUT_SECS`) | Per-task timeout. |
| `RUSTFS_HEAL_MAX_CONCURRENT_HEALS` | `4` (`DEFAULT_HEAL_MAX_CONCURRENT_HEALS`) | Global concurrent heal task limit. |
| `RUSTFS_HEAL_MAX_CONCURRENT_PER_SET` | `1` (`DEFAULT_HEAL_MAX_CONCURRENT_PER_SET`) | Per-erasure-set limit; effective value is `min(global, per_set)`, each floored at `1`. |
| `RUSTFS_HEAL_LOW_PRIORITY_MERGE_ENABLE` | `true` (`DEFAULT_HEAL_LOW_PRIORITY_MERGE_ENABLE`) | Merge duplicate low-priority requests with the same dedup key. |
| `RUSTFS_HEAL_LOW_PRIORITY_DROP_WHEN_FULL` | `true` (`DEFAULT_HEAL_LOW_PRIORITY_DROP_WHEN_FULL`) | Drop, rather than block on, low-priority requests when the queue is full. |
| `RUSTFS_HEAL_EVENT_DRIVEN_SCHEDULER_ENABLE` | `true` (`DEFAULT_HEAL_EVENT_DRIVEN_SCHEDULER_ENABLE`) | Notify-driven scheduler wakeups. |
| `RUSTFS_HEAL_SET_BULKHEAD_ENABLE` | `true` (`DEFAULT_HEAL_SET_BULKHEAD_ENABLE`) | Per-set bulkhead scheduling. |
| `RUSTFS_HEAL_PAGE_PARALLEL_ENABLE` | `true` (`DEFAULT_HEAL_PAGE_PARALLEL_ENABLE`) | Page-level parallel object healing during erasure-set repair. |
| `RUSTFS_HEAL_PAGE_OBJECT_CONCURRENCY` | `8` (`DEFAULT_HEAL_PAGE_OBJECT_CONCURRENCY`) | Concurrent object heals within one erasure-set page. Forced to `1` when page parallelism is off, for `Deep` scan mode, and for `AutoHeal`-sourced requests (`ErasureSetHealer::effective_heal_page_object_concurrency_for_source`). |
| `RUSTFS_HEAL_MAINLINE_THROTTLE_ENABLE` | `true` (`DEFAULT_HEAL_MAINLINE_THROTTLE_ENABLE`) | Pause best-effort heal task starts while foreground I/O is saturated. |
| `RUSTFS_HEAL_MAINLINE_READ_UTILIZATION_HIGH_PERCENT` | `80` (`DEFAULT_HEAL_MAINLINE_READ_UTILIZATION_HIGH_PERCENT`, capped at 100) | Foreground read-permit utilization at which heal starts pause. |
| `RUSTFS_HEAL_MAINLINE_WRITE_UTILIZATION_HIGH_PERCENT` | `80` (`DEFAULT_HEAL_MAINLINE_WRITE_UTILIZATION_HIGH_PERCENT`, capped at 100) | Foreground write utilization at which heal starts pause. |
| `RUSTFS_HEAL_MAINLINE_MAX_SLEEP_MS` | `250` (`DEFAULT_HEAL_MAINLINE_MAX_SLEEP_MS`) | Recheck delay after deferring heal starts for foreground pressure. |
| `RUSTFS_HEAL_OVERLAP_POLICY` | `merge` (`DEFAULT_HEAL_OVERLAP_POLICY`) | `merge` dedups an admin heal start that overlaps a running or queued heal; `minio_error` returns a typed already-running / overlapping-paths rejection like madmin. |
| `RUSTFS_HEAL_MRF_ENABLE` | `true` (`DEFAULT_HEAL_MRF_ENABLE`) | MRF intent pipeline: error paths deliver repair intents to the heal runtime and unconsumed intents replay from the durable journal after restart. |
| `RUSTFS_HEAL_MRF_QUEUE_SIZE` | `100000` (`DEFAULT_HEAL_MRF_QUEUE_SIZE`) | MRF in-memory queue capacity. |
| `RUSTFS_HEAL_MRF_JOURNAL_MAX_BYTES` | `8388608` (`DEFAULT_HEAL_MRF_JOURNAL_MAX_BYTES`, 8 MiB) | MRF journal size at which compaction runs. |
| `RUSTFS_HEAL_MRF_REPLAY_BATCH` | `256` (`DEFAULT_HEAL_MRF_REPLAY_BATCH`) | Intents per replay push round. |
| `RUSTFS_HEAL_DANGLING_DELETE_GRACE_SECS` | `3600` (`DEFAULT_HEAL_DANGLING_DELETE_GRACE_SECS`, `crates/ecstore/src/set_disk/core/io_primitives.rs`) | A recently modified object is never deleted as dangling inside this window; `0` disables the grace window. |
## Deliberate non-parity with MinIO
These differences from MinIO are design decisions, recorded so they are not re-filed as gaps.
| Area | RustFS behavior | Why it is not a gap |
|---|---|---|
| Bloom filter | `.bloomcycle.bin` (`DATA_USAGE_BLOOM_NAME`, `crates/scanner/src/data_usage_define.rs`) is reused only as the cycle/epoch fence. | MinIO master removed the bloom filter too. |
| Scanner leadership | Single cluster-wide scanner leader plus an epoch fence. | Same model as MinIO; the fence is additive. |
| Heal notifications | Heal emits no S3 bucket notification; results are exposed through admin status. | Same as MinIO. |
| Incomplete multipart cleanup | Runs as an independent background routine, not inside the scanner or ILM. | Same as MinIO. |
| Inline heal | The scanner only enqueues heal candidates; nothing heals inline on the scan path. | MinIO's inline `applyHealing` path is intentionally not a parity target. |
| Heal-sequence keep-alive | Admin heal status is a snapshot query with incremental `sinceSeq`/`nextSeq` semantics (`crates/heal-contracts/src/heal_channel.rs`). | MinIO's 10-second blank keep-alive write-back belongs to its streaming model and is not copied. |
| `.trash` / `tmp-old` paths | Layout constants are RustFS's own (`crates/ecstore/src/disk/local.rs`). | No literal alignment with MinIO path names is intended. |
## Migrating from MinIO scanner settings
Only two MinIO scanner variables are recognized. `apply_external_env_compat` (`crates/utils/src/envs.rs`, called from `rustfs/src/startup_preflight.rs`) copies `MINIO_<suffix>` into `RUSTFS_<suffix>` at startup for suffixes on `EXTERNAL_COMPATIBLE_SUFFIXES`, and only when the `RUSTFS_` key is absent; when both are set with different values the `RUSTFS_` value wins and a `Detected external-prefix compatibility conflicts` warning is logged. The scanner suffixes on that list are `SCANNER_SPEED` and `SCANNER_CYCLE` (tests `scanner_aliases_are_mapped_when_rustfs_missing` in `crates/utils/src/envs.rs`, `test_cycle_interval_supports_minio_speed_alias` and `test_cycle_interval_supports_minio_cycle_alias` in `crates/scanner/src/scanner/tests.rs`). Every other `MINIO_SCANNER_*` or `MINIO_HEAL_*` variable is silently ignored.
| MinIO setting | RustFS setting | Migration |
|---|---|---|
| `MINIO_SCANNER_SPEED` / `scanner speed` | `RUSTFS_SCANNER_SPEED` / `scanner.speed` | Env alias mapped at startup; preset names and the preset table are identical. |
| `MINIO_SCANNER_CYCLE` / `scanner cycle` | `RUSTFS_SCANNER_CYCLE` / `scanner.cycle` | Env alias mapped at startup. |
| `MINIO_SCANNER_IDLE_SPEED` / `scanner idle_speed` (`on` default, `off`) | `RUSTFS_SCANNER_IDLE_MODE` / `scanner.idle_mode` (`true` default, `false`) | Not mapped; must be rewritten. Direction matches (`on` and `true` both mean throttled). Both RustFS channels also accept `on`/`off` as booleans (`parse_config_bool`, `parse_bool_str`). RustFS `false` additionally disables the foreground-read backoff floor that MinIO does not have, so the scanner competes with foreground reads at full speed; use it only for benchmarks or exclusive-I/O windows. |
| `MINIO_HEAL_BITROTSCAN` / `heal bitrotscan` (default `off`) | `RUSTFS_SCANNER_BITROT_CYCLE_SECS` / `heal.bitrot_cycle` (default 30 days) | Not mapped. RustFS deep-scans periodically by default; set `off` or `disabled` to reproduce MinIO's default. |
| `MINIO_API_STALE_UPLOADS_EXPIRY` (24h) | `RUSTFS_API_STALE_UPLOADS_EXPIRY` (`DEFAULT_STALE_UPLOADS_EXPIRY`, 24h, `crates/ecstore/src/bucket/lifecycle/bucket_lifecycle_ops.rs`) | Not mapped; same default. |
| `MINIO_API_STALE_UPLOADS_CLEANUP_INTERVAL` (6h) | `RUSTFS_API_STALE_UPLOADS_CLEANUP_INTERVAL` (`DEFAULT_STALE_UPLOADS_CLEANUP_INTERVAL`, 6h) | Not mapped; same default. |
| `MINIO_API_DELETE_CLEANUP_INTERVAL` (5m) | None; `DELETED_OBJECTS_CLEANUP_INTERVAL` is a 5-minute constant in `crates/ecstore/src/disk/local.rs`. | No knob. Trash draining is not per-entry throttled. |
| `scanner alert_excess_folders` (50000) | `scanner.alert_excess_folders` (65538) | Not mapped; see [Scanner Excess Alerts](scanner-excess-alerts.md). |
| Any other `MINIO_SCANNER_*` | Corresponding `RUSTFS_SCANNER_*` from the tables above | Not mapped; rename explicitly. |
Stale-upload cleanup differs in one crash-recovery detail. RustFS's stale multipart cleanup (`cleanup_stale_multipart_uploads_in_set`) takes a namespace write lock, re-checks the upload under write quorum (`check_multipart_upload_path_exists`), and fans the delete out to every disk, where the local recursive delete renames the directory into `.rustfs.sys/tmp/.trash/<uuid>` (`move_to_trash`). If the process dies mid fan-out after more disks than the parity count have already moved the upload directory, the next cleanup pass's quorum re-check fails (`FileNotFound` is not in `OBJECT_OP_IGNORED_ERRS`) and the candidate is skipped, so the remaining per-disk residue is not reclaimed by that job. The residue is invisible to the S3 API and only consumes disk space; the window is milliseconds wide. MinIO processes each disk independently and converges in the same scenario.
## Replacement Recovery Completion
`POST /v3/background-heal/status` is an execution-queue view. `state=idle`, zero queue and active counts, an online disk, a readable object, or acceptance of an Admin deep-heal request do not independently prove that a replacement disk contains every erasure shard.
`POST /v3/background-heal/status` is an execution-queue view. `state=idle`, zero queue and active counts, an online disk, a readable object, or acceptance of an admin deep-heal request do not independently prove that a replacement disk contains every erasure shard.
Treat replacement recovery as verified only after the repair task has completed for the exact replacement instance and an operator has confirmed the target disk contains the expected `xl.meta` and data parts for every relevant object version. A replacement that is not mounted, is unsafe to format, loses its marker, or returns a partial target outcome must be treated as deferred or incomplete rather than complete.
Treat replacement recovery as verified only after the repair task has completed for the exact replacement instance and an operator has confirmed the target disk contains the expected `xl.meta` and data parts for every relevant object version. A replacement that is not mounted, is unsafe to format, loses its marker, or returns a partial target outcome must be treated as deferred or incomplete. Do not automate destructive replacement actions from an `idle` observation alone. A new node must not infer replacement completion from an old or unavailable peer; regard that information as unknown or degraded until every required peer can report the same replacement instance and verified completion.
The v3 route and its peer status protocol preserve their existing fields for mixed-version clusters. A new node must not infer replacement completion from an old or unavailable peer; regard that information as unknown or degraded until every required peer can report the same replacement instance and verified completion. Do not automate destructive replacement actions from an `idle` observation alone.
`GET /rustfs/admin/v4/heal/replacement-recovery` reports durable automatic replacement records from survivor disks. `local.records[]` entries distinguish `waiting_for_replacement`, `running`, `incomplete`, `unrecoverable`, `cleanup_pending`, `completed`, and `unknown`; `local.definitive=false` or any `unknown` record means the node could not prove a local replacement state. The `cluster` section queries the replacement-recovery peer RPC and sets `cluster.definitive=true` only when the expected peer topology is complete, every peer supports the RPC, every peer snapshot is locally definitive, and all peers report the same replacement records. Old peers, unavailable peers, malformed peer payloads, topology gaps, and generation disagreements are reported as degraded or unknown rather than complete.
`GET /rustfs/admin/v4/heal/replacement-recovery` reports durable automatic replacement records from survivor disks. Its `local.records[]` entries distinguish `waiting_for_replacement`, `running`, `incomplete`, `unrecoverable`, `cleanup_pending`, `completed`, and `unknown`; `local.definitive=false` or any `unknown` record means the node could not prove a local replacement state. Its `cluster` section queries the replacement-recovery peer RPC and sets `cluster.definitive=true` only when the expected peer topology is complete, every peer supports the RPC, every peer snapshot is locally definitive, and all peers report the same replacement records. Old peers, unavailable peers, malformed peer payloads, topology gaps, and generation disagreements are reported as degraded or unknown rather than complete.
Replacement resume and checkpoint files use an independent on-disk schema. A newer reader rejects a future schema rather than continuing with data it cannot interpret, while an older binary cannot safely enforce the new generation fence because it may ignore fields it does not know. Do not roll a cluster back after a replacement generation has started. Complete that recovery with the current-or-newer release; if it cannot complete, keep that version for diagnosis rather than deleting its durable records or continuing with an older binary.
Replacement resume and checkpoint files use an independent on-disk schema. A newer reader rejects a future schema rather than continuing with data it cannot interpret, while an older binary cannot safely enforce the new generation fence. Do not roll a cluster back after a replacement generation has started; complete that recovery with the current-or-newer release, and if it cannot complete, keep that version for diagnosis rather than deleting its durable records or continuing with an older binary.
## Reading Replication Repair
`metrics.replication_repair`, `metrics.current_cycle_replication_repair`, and
`metrics.last_cycle_replication_repair` split scanner-discovered replication
repair work by source and repair kind.
Each entry has the same `checked`, `queued`, `executed`, `failed`, `skipped`,
and `missed` counters used by `source_work`, plus:
`metrics.replication_repair`, `metrics.current_cycle_replication_repair`, and `metrics.last_cycle_replication_repair` split scanner-discovered replication repair work by source and repair kind. Each entry has the same `checked`, `queued`, `executed`, `failed`, `skipped`, and `missed` counters used by `source_work`, plus:
| Field | Meaning |
|---|---|
@@ -411,39 +344,21 @@ and `missed` counters used by `source_work`, plus:
| `scanner_role` | `repair_admission` means scanner found work and attempted to admit it to a worker queue. `boundary_signal` means scanner is reporting state owned by another runtime. |
| `execution_owner` | `bucket_replication_queue` for bucket replication repair execution, or `site_replication_runtime` for site replication resync execution. |
For bucket replication, `queued` means scanner-discovered repair was admitted
to the replication queue, `missed` means the queue or worker path could not
accept it, and `skipped` means the object did not require a new repair task.
The site replication kinds keep passive scanner discovery separate from active
resync. Scanner status may report site replication boundary counters, but the
scanner should not be treated as the active site replication resync controller.
Use this boundary when interpreting replication pressure:
For bucket replication, `queued` means scanner-discovered repair was admitted to the replication queue, `missed` means the queue or worker path could not accept it, and `skipped` means the object did not require a new repair task. The site replication kinds keep passive scanner discovery separate from active resync; the scanner is never the active site replication resync controller.
| Scenario | Scanner source | Repair kind | Scanner role | Execution owner | Operational meaning |
|---|---|---|---|---|---|
| Bucket object, delete-marker, version-purge, or existing-object repair found during a scan | `bucket_replication` | `object`, `delete_marker`, `version_purge`, `existing_object` | `repair_admission` | `bucket_replication_queue` | Scanner found bucket replication repair work and attempted to admit it to the replication queue. |
| Peer-originated or passive site replication work is observed while scanning | `site_replication` | `passive_requeue` | `boundary_signal` | `site_replication_runtime` | Scanner is reporting a passive site-replication boundary signal; it is not taking ownership of active site resync. |
| Admin-triggered or runtime-owned site resync activity is visible in scanner metrics | `site_replication` | `active_resync` | `boundary_signal` | `site_replication_runtime` | Treat this as a boundary/status signal owned by the site replication runtime, not as scanner-controlled repair execution. |
| Admin-triggered or runtime-owned site resync activity is visible in scanner metrics | `site_replication` | `active_resync` | `boundary_signal` | `site_replication_runtime` | A boundary/status signal owned by the site replication runtime, not scanner-controlled repair execution. |
If `site_replication` counters grow while bucket replication counters stay
flat, investigate site replication status and resync state before tuning
scanner pacing. If `bucket_replication` `missed` grows, investigate the bucket
replication worker queue or target health before changing scanner cycle
settings.
If `site_replication` counters grow while bucket replication counters stay flat, investigate site replication status and resync state before tuning scanner pacing. If `bucket_replication` `missed` grows, investigate the bucket replication worker queue or target health before changing scanner cycle settings.
## Reading Maintenance Control
`metrics.maintenance_control` derives a source-level control snapshot from
scanner pacing, partial-cycle state, source work, and lifecycle transition
queue state. It does not change scanner scheduling by itself; it explains why a
source is moving, deferred, or blocked. When no scan cycle is currently active,
source-work controls use the last completed cycle so recently missed work stays
visible between scanner passes.
`metrics.maintenance_control` derives a source-level control snapshot from scanner pacing, partial-cycle state, source work, and lifecycle transition queue state. It does not change scanner scheduling; it explains why a source is moving, deferred, or blocked. When no scan cycle is active, source-work controls use the last completed cycle so recently missed work stays visible between passes.
`metrics.maintenance_control.primary_control` summarizes the highest-priority
source state:
`metrics.maintenance_control.primary_control`:
| Value | Meaning |
|---|---|
@@ -453,49 +368,35 @@ source state:
| `pacing_pressure` | No source-specific state dominated, but scanner pacing pressure is still visible. |
| `none` | No source-level maintenance control pressure was observed. |
Each `metrics.maintenance_control.sources[]` entry has:
Each `metrics.maintenance_control.sources[]` entry:
| Field | Meaning |
|---|---|
| `source` | Scanner source such as `usage`, `lifecycle`, `bucket_replication`, `site_replication`, `heal`, `bitrot`, or `alerts`. |
| `source` | `usage`, `lifecycle`, `bucket_replication`, `site_replication`, `heal`, `bitrot`, or `alerts`. |
| `state` | `idle`, `active`, `deferred`, or `blocked`. |
| `reason` | Derived reason such as `active_work`, `queued_work`, `partial_cycle`, `missed_work`, `expiry_queue_backlog`, `transition_failed`, `transition_compensation_backlog`, `transition_queue_backlog`, or `transition_queue_full`. |
| `reason` | `active_work`, `queued_work`, `partial_cycle`, `missed_work`, `expiry_queue_backlog`, `transition_failed`, `transition_compensation_backlog`, `transition_queue_backlog`, or `transition_queue_full`. |
| `backlog` | Current source-level backlog estimate from queued or missed work. |
| `current_checked` | Current-cycle checked work for this source, or the last completed cycle when no scan cycle is active. |
| `current_queued` | Current-cycle queued work for this source, or the last completed cycle when no scan cycle is active. |
| `current_missed` | Current-cycle work that could not be admitted, or the last completed cycle when no scan cycle is active. |
| `lifetime_missed` | Lifetime missed work counter for context. |
| `current_checked` / `current_queued` / `current_missed` | Current-cycle counters for this source, or the last completed cycle when no scan cycle is active. |
| `lifetime_missed` | Lifetime missed work counter. |
| `partial_cycles` | Partial cycles attributed to this source. |
Use this snapshot before changing scanner controls. For example,
`blocked_source` with `lifecycle/missed_work` points at downstream lifecycle
admission, while `deferred_source` with `usage/partial_cycle` points at scanner
cycle budgets. `lifecycle/expiry_queue_backlog` means scanner-driven expiry or
delete work is still queued or active in the expiry worker pool.
`lifecycle/transition_failed` means transition worker execution failed during
the current or last completed scan cycle, while
`lifecycle/transition_compensation_backlog` means transition compensation is
still pending or running after queue backpressure.
Read this snapshot before changing scanner controls: `blocked_source` with `lifecycle/missed_work` points at downstream lifecycle admission, `deferred_source` with `usage/partial_cycle` points at scanner cycle budgets, `lifecycle/expiry_queue_backlog` means expiry or delete work is still queued or active in the expiry worker pool, `lifecycle/transition_failed` means transition worker execution failed during the current or last completed cycle, and `lifecycle/transition_compensation_backlog` means transition compensation is still pending or running after queue backpressure.
`metrics.lifecycle_expiry` exposes the expiry/delete worker queue observed by
scanner-driven lifecycle work:
`metrics.lifecycle_expiry` exposes the expiry/delete worker queue:
| Field | Meaning |
|---|---|
| `current_queue_capacity` | Effective expiry worker queue capacity for this node. |
| `current_queued` | Expiry/delete tasks currently waiting in the worker queue. |
| `current_active` | Expiry/delete tasks currently running in a worker. |
| `current_queued` | Expiry/delete tasks waiting in the worker queue. |
| `current_active` | Expiry/delete tasks currently running. |
| `current_workers` | Configured expiry worker count. |
| `queue_missed` | Expiry/delete tasks that could not be queued because no worker channel was available or the queue was closed. |
| `queue_missed` | Tasks that could not be queued because no worker channel was available or the queue was closed. |
| `scanner_queued` | Scanner-discovered expiry/delete object versions admitted to the expiry queue. |
| `scanner_missed` | Scanner-discovered expiry/delete object versions that could not be admitted. |
## Reading Distributed Metrics
`/rustfs/admin/v3/scanner/status` and `/rustfs/admin/v3/metrics` report the
node that handles the HTTP request. The metrics endpoint does not fan out to
peer nodes. In distributed deployments, query every node explicitly and keep
`by-host=true` enabled so each response includes that node's host view:
`/rustfs/admin/v3/scanner/status` and `/rustfs/admin/v3/metrics` report the node that handles the HTTP request; the metrics endpoint does not fan out to peers. In distributed deployments, query every node explicitly and keep `by-host=true` so each response includes that node's host view:
```bash
for endpoint in http://node-a:9000 http://node-b:9000 http://node-c:9000; do
@@ -512,19 +413,11 @@ for endpoint in http://node-a:9000 http://node-b:9000 http://node-c:9000; do
done
```
The `aggregated.scanner` payload preserves the same scanner progress,
checkpoint, pacing, source work, maintenance control, lifecycle expiry, and
lifecycle transition fields used by the local scanner status, but only for the
node that returned the response. The `by_host.*.scanner` payload keeps that
node's host view.
Compare the per-node artifacts externally to find old active paths, partial
checkpoints, pacing pressure, source-level control pressure, or downstream
queue admission problems across the deployment.
The `aggregated.scanner` payload preserves the same scanner progress, checkpoint, pacing, source work, maintenance control, lifecycle expiry, and lifecycle transition fields used by the local scanner status, but only for the responding node; `by_host.*.scanner` keeps that node's host view. Compare the per-node artifacts externally to find old active paths, partial checkpoints, pacing pressure, source-level control pressure, or downstream queue admission problems across the deployment.
## Reading Lifecycle Transition Status
`metrics.lifecycle_transition` focuses on scanner-driven lifecycle transition
work:
`metrics.lifecycle_transition`:
| Field | Meaning |
|---|---|
@@ -542,43 +435,28 @@ work:
| `completed` | Transition worker completions. |
| `failed` | Transition worker failures. |
When `scanner_missed` or `queue_full` rises, scanner lifecycle work is finding
transition candidates faster than the transition queue can accept them. That is
a downstream transition pressure signal, not just a scanner walk pressure signal.
When `scanner_missed` or `queue_full` rises, scanner lifecycle work is finding transition candidates faster than the transition queue accepts them: a downstream transition pressure signal, not just a scanner walk pressure signal.
## Tuning Workflow
For symptoms where a mostly idle single-node, single-disk deployment has
sustained CPU usage while the scanner is enabled:
For a mostly idle single-node, single-disk deployment with sustained CPU usage while the scanner is enabled:
1. Read `/v3/scanner/status`.
2. Check `metrics.pacing_pressure.primary_pressure`.
3. Check `metrics.maintenance_control.primary_control` and source entries
before changing runtime controls.
4. Check `runtime_config.delay`, `runtime_config.max_wait_seconds`, and
`runtime_config.cycle_interval_seconds` to confirm the active values and
their sources.
5. Check `metrics.current_cycle_objects_scanned`,
`metrics.current_cycle_directories_scanned`, and active paths to confirm the
scanner is the active work.
6. If `primary_pressure` is `throttle_pause` and pause ratios are low, raise
`scanner.delay` first.
3. Check `metrics.maintenance_control.primary_control` and source entries before changing runtime controls.
4. Check `runtime_config.delay`, `runtime_config.max_wait_seconds`, and `runtime_config.cycle_interval_seconds` to confirm the active values and their sources.
5. Check `metrics.current_cycle_objects_scanned`, `metrics.current_cycle_directories_scanned`, and active paths to confirm the scanner is the active work.
6. If `primary_pressure` is `throttle_pause` and pause ratios are low, raise `scanner.delay` first.
7. If individual sleeps are too short, raise `scanner.max_wait`.
8. If each scan cycle finishes but starts too often, raise `scanner.cycle`.
9. If scans must be broken into bounded chunks, set one of the cycle budgets:
`scanner.cycle_max_duration`, `scanner.cycle_max_objects`, or
`scanner.cycle_max_directories`.
10. Recheck `pacing_pressure`, `maintenance_control`, source work, and
lifecycle transition status after one or more scanner cycles.
9. If scans must be broken into bounded chunks, set one of `scanner.cycle_max_duration`, `scanner.cycle_max_objects`, or `scanner.cycle_max_directories`.
10. Recheck `pacing_pressure`, `maintenance_control`, source work, and lifecycle transition status after one or more scanner cycles.
Do not rely only on a longer cycle interval if lifecycle, replication, heal, or
bitrot work must keep moving. Use source work and transition status to confirm
that background maintenance is still making progress.
Do not rely only on a longer cycle interval if lifecycle, replication, heal, or bitrot work must keep moving; use source work and transition status to confirm that background maintenance still progresses.
## Helm
The Helm chart exposes the scanner environment variables under
`config.rustfs.scanner`. Example:
The Helm chart exposes the scanner environment variables under `config.rustfs.scanner` (`helm/rustfs/values.yaml`):
```yaml
config:
@@ -596,5 +474,4 @@ config:
bitrot_cycle_secs: "2592000"
```
Use `extraEnv` for experimental or unrelated environment variables that are not
represented by chart values.
Use `extraEnv` for environment variables that are not represented by chart values, including every heal knob above.
+91 -216
View File
@@ -1,15 +1,12 @@
# SFTP Operations Guide
The guide covers enabling and operating the SFTP server in RustFS.
It is written for operators who need to expose buckets over SFTP, manage host
keys across platforms, size the server for large transfers, and diagnose
session cleanup behaviour.
**Use this when:** enabling the SFTP front end, managing host keys across platforms, sizing it for large transfers, granting IAM permissions for SFTP users, or diagnosing session cleanup and orphaned multipart uploads.
## Enabling SFTP: Recommended Configuration
**Source of truth:** `crates/config/src/constants/protocols.rs` (`ENV_SFTP_*`, `DEFAULT_SFTP_*`), `crates/protocols/src/sftp/constants.rs` (session, keepalive, handle, and listing limits), `crates/protocols/src/sftp/driver.rs` (abort permit pool), `crates/ecstore/src/bucket/lifecycle/bucket_lifecycle_ops.rs` (`DEFAULT_STALE_UPLOADS_EXPIRY`, `DEFAULT_STALE_UPLOADS_CLEANUP_INTERVAL`).
Two settings are required: the enable flag and the host-key directory. The
listen address has a default of `0.0.0.0:2222` and setting it explicitly is
recommended practice.
## Enabling SFTP
Two settings are required: the enable flag and the host-key directory. Setting the listen address explicitly is recommended even though it has a default.
```bash
RUSTFS_SFTP_ENABLE=true
@@ -17,38 +14,24 @@ RUSTFS_SFTP_ADDRESS=0.0.0.0:2222
RUSTFS_SFTP_HOST_KEY_DIR=/etc/rustfs/sftp-keys
```
- `RUSTFS_SFTP_ENABLE` starts the SFTP listener at server startup. Off by
default.
- `RUSTFS_SFTP_ADDRESS` is the listen address and port. `0.0.0.0` accepts
connections on every interface. Port `2222` avoids the privileged port 22.
- `RUSTFS_SFTP_HOST_KEY_DIR` is the directory the SSH host keys are loaded
from. No default, and startup fails when SFTP is enabled without it.
| Variable | Role |
| --- | --- |
| `RUSTFS_SFTP_ENABLE` | Starts the SFTP listener at server startup. Off by default. |
| `RUSTFS_SFTP_ADDRESS` | Listen address and port. `0.0.0.0` accepts connections on every interface; port `2222` avoids the privileged port 22. |
| `RUSTFS_SFTP_HOST_KEY_DIR` | Directory the SSH host keys are loaded from. No default; startup fails when SFTP is enabled without it. |
Generate a host key before first start. On Unix, `ssh-keygen` also writes a
world-readable `.pub` file into the directory, and the server requires every
file it considers in the host-key directory (regular, non-empty, at most
1 MiB) to be owner-only, so restrict or remove it:
Generate a host key before first start. On Unix, `ssh-keygen` also writes a world-readable `.pub` file into the directory, and the server requires every file it considers in the host-key directory (regular, non-empty, at most 1 MiB) to be owner-only, so restrict or remove it. The server does not read the `.pub` file. Keys must be unencrypted (no passphrase).
```bash
ssh-keygen -t ed25519 -f /etc/rustfs/sftp-keys/ssh_host_ed25519_key -N ""
chmod 600 /etc/rustfs/sftp-keys/ssh_host_ed25519_key.pub
```
The private key itself is already written with owner-only permissions. The
server does not read the `.pub` file, so it can be removed instead. Keys must
be unencrypted (no passphrase).
On success the server logs `SFTP server listening` with the bound address. A
port of `0` in `RUSTFS_SFTP_ADDRESS` is resolved to a free port at startup
and the resolved port appears in that log line.
Every other setting has a tested default and should be left alone unless a
section below gives a concrete reason to change it.
On success the server logs `SFTP server listening` with the bound address. A port of `0` in `RUSTFS_SFTP_ADDRESS` is resolved to a free port at startup and the resolved port appears in that log line. Every other setting has a tested default and should be left alone unless a section below gives a concrete reason to change it.
## Path Model
The SFTP root directory is the account's bucket list. The first path
component names the bucket and the remainder is the object key:
The SFTP root directory is the account's bucket list. The first path component names the bucket and the remainder is the object key:
```text
/reports/2026/q1.pdf
@@ -57,153 +40,78 @@ bucket: reports
object key: 2026/q1.pdf
```
Files cannot be created at the root level. Creating or removing a top-level
directory creates or removes a bucket. The root listing shows at most 10000
buckets and logs a warning when truncated.
Files cannot be created at the root level. Creating or removing a top-level directory creates or removes a bucket. The root listing shows at most 10000 buckets (`ROOT_LISTING_MAX_ENTRIES`) and logs `root READDIR truncated` when cut off.
## Host Keys
`RUSTFS_SFTP_HOST_KEY_DIR` must name an existing directory containing at
least one decodable private key. Startup fails otherwise. There is no
generated fallback key.
`RUSTFS_SFTP_HOST_KEY_DIR` must name an existing directory containing at least one decodable private key. Startup fails otherwise; there is no generated fallback key.
- Any private key that russh can decode is accepted. Ed25519, ECDSA, and RSA
are the expected formats. Passphrase-protected keys cannot be decoded and
do not count. Keys are offered to clients in the order Ed25519, ECDSA,
RSA, then anything else. Multiple keys of one algorithm all load, so during
key rotation clients can be offered the new key as soon as it is added,
not when the old one is removed.
- Empty files and files larger than 1 MiB are skipped entirely.
- On Unix, every other regular file in the directory must have no group or
other permission bits set (owner-only, for example mode `0600`, `0400`, or
`0700`). The check covers non-key files too: a world-readable README or
`.pub` file fails startup with `host key file has insecure permissions`.
Keep only owner-only files in the directory.
- On Windows there is no permission-bit check. The server logs a one-time
warning at startup whose alertable first sentence reads exactly `SFTP host
key file permission enforcement is not active on Windows`. Restrict the
NTFS ACL on the host-key directory to the running rustfs service account,
`NT AUTHORITY\SYSTEM`, and `BUILTIN\Administrators`. Default ProgramData
inheritance grants `BUILTIN\Users` read access. Remove that grant on the
host-key directory.
- Targets that are neither Unix nor Windows do not support SFTP and fail
startup.
Host keys can be hot-reloaded without a restart by setting
`RUSTFS_SFTP_HOST_KEY_RELOAD_ENABLE=true`. The directory is rescanned every
`RUSTFS_SFTP_HOST_KEY_RELOAD_INTERVAL` seconds (default 30, silently raised
to 5 if set lower). A failed rescan, including a permission violation that
would be fatal at startup, keeps the previous keys and logs `SFTP host key
reload failed; keeping previous keys`. Reloaded keys affect new connections
only.
| Rule | Detail |
| --- | --- |
| Accepted keys | Any private key russh can decode: Ed25519, ECDSA, RSA. Passphrase-protected keys cannot be decoded and do not count. |
| Offer order | Ed25519, ECDSA, RSA, then anything else. Multiple keys of one algorithm all load, so during rotation clients see the new key as soon as it is added, not when the old one is removed. |
| Skipped files | Empty files and files larger than 1 MiB. |
| Unix permissions | Every other regular file in the directory must have no group or other permission bits (`0600`, `0400`, `0700`). Non-key files count too: a world-readable README or `.pub` file fails startup with `host key file has insecure permissions`. |
| Windows | No permission-bit check. A one-time startup warning whose first sentence reads exactly `SFTP host key file permission enforcement is not active on Windows` is logged. Restrict the NTFS ACL on the directory to the rustfs service account, `NT AUTHORITY\SYSTEM`, and `BUILTIN\Administrators`; default ProgramData inheritance grants `BUILTIN\Users` read access, so remove that grant. |
| Other targets | Neither Unix nor Windows: SFTP is unsupported and startup fails. |
| Hot reload | `RUSTFS_SFTP_HOST_KEY_RELOAD_ENABLE=true` rescans every `RUSTFS_SFTP_HOST_KEY_RELOAD_INTERVAL` seconds (`DEFAULT_SFTP_HOST_KEY_RELOAD_INTERVAL`, 30; silently raised to 5 if lower). A failed rescan, including a permission violation that would be fatal at startup, keeps the previous keys and logs `SFTP host key reload failed; keeping previous keys`. Reloaded keys affect new connections only. |
## Configuration Reference
Changing values beyond the recommended section is rarely necessary. The
defaults are the configuration the test suites and stress runs exercise.
Tuning values interact (for example part size multiplies against the handle
cap in worst-case memory), so change them deliberately and one at a time.
Changing values beyond the recommended section is rarely necessary. The defaults are the configuration the test suites and stress runs exercise. Tuning values interact (part size multiplies against the handle cap in worst-case memory), so change them deliberately and one at a time.
Invalid values fall into four classes:
1. Fatal at startup. The server refuses to start and names the problem.
2. Warn and use the default. The server starts and logs a warning naming
the variable, the rejected value, and the bounds.
3. Silently clamped. The reload interval is the one variable raised to its
floor without a warning.
4. Silent fallback. Non-numeric text in any numeric variable behaves as if
the variable were unset. Invalid boolean text logs a one-time warning
and uses the default.
| Class | Behavior |
| --- | --- |
| Fatal at startup | The server refuses to start and names the problem. |
| Warn and use the default | The server starts and logs a warning naming the variable, the rejected value, and the bounds. |
| Silently clamped | Only the reload interval is raised to its floor without a warning. |
| Silent fallback | Non-numeric text in any numeric variable behaves as if the variable were unset. Invalid boolean text logs a one-time warning and uses the default. |
| Variable | Default | Valid values | On invalid |
| Variable | Default (constant) | Valid values | On invalid |
| -------- | ------- | ------------ | ---------- |
| `RUSTFS_SFTP_ENABLE` | `false` | boolean | warn, default |
| `RUSTFS_SFTP_ADDRESS` | `0.0.0.0:2222` | see binding note below | fatal |
| `RUSTFS_SFTP_HOST_KEY_DIR` | none | existing directory, required | fatal |
| `RUSTFS_SFTP_IDLE_TIMEOUT` | `600` | seconds, greater than zero | fatal |
| `RUSTFS_SFTP_PART_SIZE` | `16777216` | bytes, 5 MiB to 5 GiB | fatal |
| `RUSTFS_SFTP_READ_ONLY` | `false` | boolean | warn, default |
| `RUSTFS_SFTP_ADDRESS` | `0.0.0.0:2222` (`DEFAULT_SFTP_ADDRESS`) | see binding note below | fatal |
| `RUSTFS_SFTP_HOST_KEY_DIR` | none (`DEFAULT_SFTP_HOST_KEY_DIR`) | existing directory, required | fatal |
| `RUSTFS_SFTP_IDLE_TIMEOUT` | `600` (`DEFAULT_SFTP_IDLE_TIMEOUT`) | seconds, greater than zero | fatal |
| `RUSTFS_SFTP_PART_SIZE` | `16777216` (`DEFAULT_SFTP_PART_SIZE`) | bytes, 5 MiB to 5 GiB | fatal |
| `RUSTFS_SFTP_READ_ONLY` | `false` (`DEFAULT_SFTP_READ_ONLY`) | boolean | warn, default |
| `RUSTFS_SFTP_BANNER` | `SSH-2.0-RustFS` | must start with `SSH-2.0-` | fatal |
| `RUSTFS_SFTP_HANDLES_PER_SESSION` | `64` | `8` to `1024` | warn, default |
| `RUSTFS_SFTP_HANDLES_PER_SESSION` | `64` (`DEFAULT_HANDLES_PER_SESSION`) | `8` to `1024` (`HANDLES_PER_SESSION_MIN`/`_MAX`) | warn, default |
| `RUSTFS_SFTP_BACKEND_OP_TIMEOUT_SECS` | `60` | `5` to `600` | warn, default |
| `RUSTFS_SFTP_READ_CACHE_WINDOW_BYTES` | `4194304` | `0` disables, else 256 KiB to 64 MiB | warn, default |
| `RUSTFS_SFTP_READ_CACHE_TOTAL_MEM_BYTES` | `268435456` | at least 16 MiB | warn, default |
| `RUSTFS_SFTP_HOST_KEY_RELOAD_ENABLE` | `false` | boolean | warn, default |
| `RUSTFS_SFTP_HOST_KEY_RELOAD_INTERVAL` | `30` | seconds, minimum 5 | clamped to 5 |
| `RUSTFS_SFTP_HOST_KEY_RELOAD_ENABLE` | `false` (`DEFAULT_SFTP_HOST_KEY_RELOAD_ENABLE`) | boolean | warn, default |
| `RUSTFS_SFTP_HOST_KEY_RELOAD_INTERVAL` | `30` (`DEFAULT_SFTP_HOST_KEY_RELOAD_INTERVAL`) | seconds, minimum 5 | clamped to 5 |
Notes on individual variables:
- `RUSTFS_SFTP_ADDRESS`: the host part must be a wildcard (`0.0.0.0` or
`[::]`) or an address or hostname assigned to the host, otherwise startup
fails. Use `[::]:2222` to listen on IPv6, which on Linux usually accepts
IPv4 as well via dual-stack. Port `0` auto-assigns a free port.
- `RUSTFS_SFTP_BANNER` is the SSH protocol identification string sent on
connect, not a free-text login banner. Values that do not start with
`SSH-2.0-` fail startup.
- `RUSTFS_SFTP_IDLE_TIMEOUT` cannot be set to `0` to disable idle
disconnects. See the sessions section for what closes idle sessions.
- `RUSTFS_SFTP_PART_SIZE` bounds follow S3 multipart limits. See the large
files section before changing it.
- `RUSTFS_SFTP_READ_CACHE_WINDOW_BYTES` set to exactly `0` disables read
caching, which turns every client read request into one backend call.
- Worst-case buffered write memory per session is the handle cap times the
part size: 64 handles at 16 MiB is 1 GiB per session. The server imposes
no limit on concurrent sessions, so total worst case is that figure
times however many clients connect. Only the read cache has a global
cap. Enforce connection limits externally if that matters.
- `RUSTFS_SFTP_ADDRESS`: the host part must be a wildcard (`0.0.0.0` or `[::]`) or an address or hostname assigned to the host, otherwise startup fails. `[::]:2222` listens on IPv6, which on Linux usually accepts IPv4 as well via dual-stack. Port `0` auto-assigns a free port.
- `RUSTFS_SFTP_BANNER` is the SSH protocol identification string sent on connect, not a free-text login banner.
- `RUSTFS_SFTP_IDLE_TIMEOUT` cannot be `0`; see the sessions section for what closes idle sessions.
- `RUSTFS_SFTP_PART_SIZE` bounds follow S3 multipart limits; see the large files section before changing it.
- `RUSTFS_SFTP_READ_CACHE_WINDOW_BYTES=0` disables read caching, turning every client read request into one backend call.
- Worst-case buffered write memory per session is the handle cap times the part size: 64 handles at 16 MiB is 1 GiB per session. The server imposes no limit on concurrent sessions, so the total worst case is that figure times the number of connected clients. Only the read cache has a global cap. Enforce connection limits externally if that matters.
## Sessions and Cleanup
Several mechanisms close sessions:
| Mechanism | Behavior |
| --- | --- |
| SSH keepalive | Sent every 15 seconds (`KEEPALIVE_INTERVAL_SECS`); the connection is closed after 3 consecutive unanswered keepalives (`KEEPALIVE_MAX`), about 60 seconds after a client stops responding. Not configurable. Cleans up clients that vanish without closing TCP. |
| `RUSTFS_SFTP_IDLE_TIMEOUT` | SSH inactivity timeout (default 600 seconds). Keepalive replies count as SSH traffic and reset it, so a client that answers keepalives is never disconnected by it. |
| Session watchdog (all platforms) | Closes sessions silent at the SFTP request layer for 30 minutes; keepalives do not count. This is what ends a healthy but idle session. Linux logs `wedge watchdog cancelling session` with reason `fallback_silence`; other platforms log `fallback watchdog cancelling session`. |
| Linux TCP-state probe | A socket in `CLOSE_WAIT` across two consecutive 15-second checks while the SFTP layer has been silent for 30 seconds is cancelled, typically 45 to 60 seconds after the client vanished. Two consecutive failed TCP-state probes are treated the same way, so a container that blocks `/proc/net/tcp` can see sessions cancelled on that schedule. The `reason` field names the trigger. |
| Socket dup failure | If duplicating the connection socket fails at accept time (rare, usually file-descriptor exhaustion) the session runs with no watchdog; only keepalive and idle mechanisms apply. Log line: `wedge watchdog: dup_socket failed`. |
| Server shutdown | Cancels every live session immediately; clients are disconnected mid-transfer. The server then waits up to 30 seconds (`SHUTDOWN_DRAIN_TIMEOUT_SECS`) for session cleanup, including aborting in-flight multipart uploads, before remaining tasks are dropped. |
- The server sends an SSH keepalive every 15 seconds and closes the
connection after 3 consecutive unanswered keepalives, about 60 seconds
after a client stops responding. Not configurable. The keepalive check
cleans up clients that vanish without closing TCP.
- `RUSTFS_SFTP_IDLE_TIMEOUT` (default 600 seconds) sets the SSH inactivity
timeout, but keepalive replies count as SSH traffic and reset it, so a
client that answers keepalives is never disconnected by it.
- The session watchdog closes sessions that are silent at the SFTP request
layer for 30 minutes, on every platform. Keepalives do not count as
SFTP-layer activity. The watchdog is the mechanism that ends a healthy
but idle session. On Linux the close logs `wedge watchdog cancelling
session` with reason `fallback_silence`. On other platforms it logs
`fallback watchdog cancelling session`.
- On Linux the watchdog additionally reads kernel TCP state for fast
cleanup of wedged sessions: a socket sitting in `CLOSE_WAIT` across two
consecutive 15-second checks while the SFTP layer has been silent for 30
seconds is cancelled, typically 45 to 60 seconds after the client
vanished. Two consecutive failed TCP-state probes are treated the same
way, so a container that blocks `/proc/net/tcp` can see sessions
cancelled on that schedule. The `reason` field of the log line names the
trigger.
- If duplicating the connection socket fails at accept time (rare, usually
file-descriptor exhaustion) the session runs with no watchdog at all and
only the keepalive and idle mechanisms apply. Log line: `wedge watchdog:
dup_socket failed`.
- Server shutdown cancels every live session immediately. Clients are
disconnected mid-transfer. The server then waits up to 30 seconds for
session cleanup, including aborting in-flight multipart uploads, before
remaining tasks are dropped.
A session ended by any cancellation path logs `SFTP session cancelled
(watchdog or server shutdown)`.
A session ended by any cancellation path logs `SFTP session cancelled (watchdog or server shutdown)`.
## Authentication and Authorization
Authentication is by password only, verified against RustFS IAM users.
Public-key authentication is rejected. Anonymous access is not available.
Failed logins are rejected without delay and there is no lockout, so apply
rate limiting externally when the listener is exposed to untrusted
networks. Accepted logins log `SFTP auth accepted` at info level, rejections
log `SFTP auth rejected` at warn level.
Authentication is by password only, verified against RustFS IAM users. Public-key authentication is rejected. Anonymous access is not available. Failed logins are rejected without delay and there is no lockout, so apply rate limiting externally when the listener is exposed to untrusted networks. Accepted logins log `SFTP auth accepted` at info level, rejections log `SFTP auth rejected` at warn level.
Every SFTP operation is authorized against IAM policy before it reaches
storage. Policy condition keys (for example `aws:SourceIp`) are not
evaluated on the SFTP path. Only unconditional Allow and Deny statements
take effect.
The S3 actions a user needs:
Every SFTP operation is authorized against IAM policy before it reaches storage. Policy condition keys (for example `aws:SourceIp`) are not evaluated on the SFTP path; only unconditional Allow and Deny statements take effect.
| SFTP activity | Required S3 actions |
| ------------- | ------------------- |
@@ -215,66 +123,38 @@ The S3 actions a user needs:
| Stat a bucket | `s3:ListBucket` |
| Delete a file | `s3:DeleteObject` |
| Rename (implemented as copy then delete) | `s3:GetObject` on the source, `s3:PutObject` on the destination, `s3:DeleteObject` on the source |
| Create or remove a top-level directory | `s3:CreateBucket` or `s3:DeleteBucket`, removal also needs `s3:ListBucket` for the emptiness check |
| Create or remove a top-level directory | `s3:CreateBucket` or `s3:DeleteBucket`; removal also needs `s3:ListBucket` for the emptiness check |
| Upload cleanup on disconnect | `s3:AbortMultipartUpload` (see the large files section) |
Upload-only users also need `s3:GetObject`, because SFTP clients stat files
as part of normal transfers.
Upload-only users also need `s3:GetObject`, because SFTP clients stat files as part of normal transfers.
With `RUSTFS_SFTP_READ_ONLY=true` the server rejects all mutating packets
at the protocol layer: opening a file for write, `WRITE`, `REMOVE`,
`MKDIR`, `RMDIR`, `RENAME`, `SETSTAT`, and `FSETSTAT`.
With `RUSTFS_SFTP_READ_ONLY=true` the server rejects all mutating packets at the protocol layer: opening a file for write, `WRITE`, `REMOVE`, `MKDIR`, `RMDIR`, `RENAME`, `SETSTAT`, and `FSETSTAT`.
## Client Compatibility
The server speaks SFTP version 3 and maps onto object storage. Differences
from a filesystem-backed SFTP server:
The server speaks SFTP version 3 (`SFTP_VERSION`) and maps onto object storage. Differences from a filesystem-backed SFTP server:
- Uploads must be a single sequential stream from offset zero. Transfer
resume, append mode, in-place edits (read-write opens), and segmented or
multi-connection uploads of one file are rejected. Configure clients for
whole-file, single-connection transfers.
- In normal (read-write) mode, `SETSTAT` and `FSETSTAT` are accepted and
ignored: chmod, timestamp preservation, and ownership changes silently
have no effect, and listed permissions are fixed server-generated values.
- Symlink operations (`SYMLINK`, `READLINK`) are not supported and return
an unsupported-operation error. Object storage has no symlink equivalent.
- Renaming copies the object server-side and then deletes the source, so
large-file renames are slow and not atomic: a failure after the copy can
leave the file at both paths. Renaming a top-level directory (a bucket)
is not supported.
| Behavior | Detail |
| --- | --- |
| Uploads | Must be a single sequential stream from offset zero. Transfer resume, append mode, in-place edits (read-write opens), and segmented or multi-connection uploads of one file are rejected. Configure clients for whole-file, single-connection transfers. |
| `SETSTAT` / `FSETSTAT` | Accepted and ignored in read-write mode: chmod, timestamp preservation, and ownership changes silently have no effect; listed permissions are fixed server-generated values. |
| Symlinks | `SYMLINK` and `READLINK` return an unsupported-operation error. |
| Rename | Server-side copy then delete: slow for large files and not atomic (a failure after the copy can leave the file at both paths). Renaming a top-level directory (a bucket) is not supported. |
## Large Files and Multipart Uploads
An upload smaller than `RUSTFS_SFTP_PART_SIZE` bytes is buffered in memory
and written with a single `PutObject` when the file is closed. At part size
or larger the handle switches to S3 multipart, flushing a part each time a
full part accumulates.
An upload smaller than `RUSTFS_SFTP_PART_SIZE` bytes is buffered in memory and written with a single `PutObject` when the file is closed. At part size or larger the handle switches to S3 multipart, flushing a part each time a full part accumulates.
The server enforces the S3 limit of 10000 parts per upload, so the largest
single upload is part size times 10000: 156.25 GiB at the default 16 MiB
part size. Raise `RUSTFS_SFTP_PART_SIZE` until that product covers the
largest expected file. A larger part size raises per-session memory. A
write that would exceed the cap is rejected with log line `SFTP write would
exceed the S3 multipart parts limit`. The same cap can reject at file
close, logging `SFTP close rejected: trailing part would exceed S3
multipart parts limit`.
The server enforces the S3 limit of 10000 parts per upload (`S3_MAX_MULTIPART_PARTS`), so the largest single upload is part size times 10000: 156.25 GiB at the default 16 MiB part size. Raise `RUSTFS_SFTP_PART_SIZE` until that product covers the largest expected file; a larger part size raises per-session memory. A write that would exceed the cap logs `SFTP write would exceed the S3 multipart parts limit`; the same cap at file close logs `SFTP close rejected: trailing part would exceed S3 multipart parts limit`.
When a session ends mid-upload the server aborts the in-flight multipart
upload. Four situations leave an orphaned upload behind that the session
itself cannot clean up:
When a session ends mid-upload the server aborts the in-flight multipart upload. Four situations leave an orphaned upload behind that the session itself cannot clean up:
- the user lacks `s3:AbortMultipartUpload`,
- the abort itself fails or times out,
- a burst of simultaneous disconnects exhausts the global abort task pool
(sized at twice the CPU parallelism, between 8 and 128),
- the user lacks `s3:AbortMultipartUpload`;
- the abort itself fails or times out;
- a burst of simultaneous disconnects exhausts the global abort task pool (sized at twice the CPU parallelism, between `ABORT_PERMITS_FLOOR` 8 and `ABORT_PERMITS_CEILING` 128);
- the server process is killed outright.
Orphaned uploads are not permanent. RustFS runs a background cleanup that
aborts stale incomplete multipart uploads (by default, uploads older than
24 hours, checked every 6 hours). An `AbortIncompleteMultipartUpload`
bucket lifecycle rule reclaims them as well and gives per-bucket control of
the window. The log lines when a session-drop abort is skipped or fails:
Orphaned uploads are not permanent. The background stale-upload cleanup aborts incomplete multipart uploads older than `DEFAULT_STALE_UPLOADS_EXPIRY` (24 hours), checked every `DEFAULT_STALE_UPLOADS_CLEANUP_INTERVAL` (6 hours). An `AbortIncompleteMultipartUpload` bucket lifecycle rule reclaims them as well and gives per-bucket control of the window. The log lines when a session-drop abort is skipped or fails:
```text
skipped abort of orphaned multipart upload on session drop, principal lacks s3:AbortMultipartUpload, bucket lifecycle rules must reclaim parts
@@ -283,31 +163,26 @@ failed to abort orphaned multipart upload
Drop abort of orphaned multipart upload timed out; bucket lifecycle rule must reclaim parts
```
Abort failures at file close log with different wording and are retried at
session drop, where the lines above appear.
Abort failures at file close log with different wording and are retried at session drop, where the lines above appear.
## Log Lines Worth Alerting On
Every line below logs at warn level except `SFTP server listening` and
`SFTP auth accepted`, which log at info. The server's default log level is
error, so none of them are visible until the log level is raised to info
(for example `RUSTFS_OBS_LOGGER_LEVEL=info`, or warn to capture alerts
only).
Every line below logs at warn level except `SFTP server listening` and `SFTP auth accepted`, which log at info. The server's default log level is error, so none of them are visible until the log level is raised (for example `RUSTFS_OBS_LOGGER_LEVEL=info`, or `warn` to capture alerts only).
| Message | Meaning |
| ------- | ------- |
| `SFTP server listening` | Startup complete, address bound |
| `SFTP auth rejected` | Failed login, no built-in lockout exists |
| `SFTP host key file permission enforcement is not active on Windows` | Expected once per start on Windows, verify the ACL guidance once |
| `host key file has insecure permissions` | Fatal at startup on Unix, fix file modes. During hot reload the same text appears inside the reload-failed warning and previous keys stay active |
| `SFTP host key reload failed; keeping previous keys` | Hot reload rescan failed, service unaffected, investigate the directory |
| `wedge watchdog cancelling session` | Linux only. `reason` field: `tcp_state_close_wait_confirmed` or `probe_failed_confirmed` mean a wedged or unprobeable socket, `fallback_silence` means routine 30-minute idle cleanup |
| `fallback watchdog cancelling session` | Non-Linux platforms: session silent at the SFTP layer for 30 minutes |
| `wedge watchdog: dup_socket failed` | Rare, that session runs without any watchdog |
| `SFTP session cancelled (watchdog or server shutdown)` | Session ended by cancellation rather than client close |
| `SFTP write would exceed the S3 multipart parts limit` | Client hit the per-upload size cap, raise the part size if legitimate. Also check the close-time variant below |
| `SFTP close rejected: trailing part would exceed S3 multipart parts limit` | Same cap hit at file close |
| `root READDIR truncated` | A principal can see more than 10000 buckets, listing was cut off |
| `skipped abort of orphaned multipart upload` | Orphaned parts on storage until background cleanup or lifecycle rule reclaims them |
| `abort permit pool exhausted on session drop` | Mass-disconnect burst, orphaned parts on storage until reclaimed |
| `RUSTFS_SFTP_` prefix in a warning | A tuning variable was rejected and its default applied |
| `SFTP server listening` | Startup complete, address bound. |
| `SFTP auth rejected` | Failed login; no built-in lockout exists. |
| `SFTP host key file permission enforcement is not active on Windows` | Expected once per start on Windows; verify the ACL guidance once. |
| `host key file has insecure permissions` | Fatal at startup on Unix; fix file modes. During hot reload the same text appears inside the reload-failed warning and previous keys stay active. |
| `SFTP host key reload failed; keeping previous keys` | Hot reload rescan failed; service unaffected; investigate the directory. |
| `wedge watchdog cancelling session` | Linux only. `reason` `tcp_state_close_wait_confirmed` or `probe_failed_confirmed` mean a wedged or unprobeable socket; `fallback_silence` means routine 30-minute idle cleanup. |
| `fallback watchdog cancelling session` | Non-Linux platforms: session silent at the SFTP layer for 30 minutes. |
| `wedge watchdog: dup_socket failed` | Rare; that session runs without any watchdog. |
| `SFTP session cancelled (watchdog or server shutdown)` | Session ended by cancellation rather than client close. |
| `SFTP write would exceed the S3 multipart parts limit` | Client hit the per-upload size cap; raise the part size if legitimate. |
| `SFTP close rejected: trailing part would exceed S3 multipart parts limit` | Same cap hit at file close. |
| `root READDIR truncated` | A principal can see more than 10000 buckets; listing was cut off. |
| `skipped abort of orphaned multipart upload` | Orphaned parts on storage until background cleanup or a lifecycle rule reclaims them. |
| `abort permit pool exhausted on session drop` | Mass-disconnect burst; orphaned parts on storage until reclaimed. |
| `RUSTFS_SFTP_` prefix in a warning | A tuning variable was rejected and its default applied. |
+67 -87
View File
@@ -1,37 +1,36 @@
# Tier / ILM Transition Debugging Guide
How to debug lifecycle tiering (hot → cold transition) issues: inspecting
`xl.meta`, tracing the versionId sent to the remote tier, and known pitfalls.
**Use this when:** a tiered (transitioned) object returns `NoSuchVersion` or fails to restore, a transition run looks stuck, or you need to inspect `xl.meta` and trace the versionId sent to the remote tier.
> Code map (post `#3929` layout):
>
> | Concern | Location |
> |---------|----------|
> | ILM actions (`transition_object`, `expire_transitioned_object`, `get_transitioned_object_reader`, `gen_transition_objname`) | `crates/ecstore/src/bucket/lifecycle/bucket_lifecycle_ops.rs` |
> | Erasure-set transition/restore entry points | `crates/ecstore/src/set_disk/` and `crates/ecstore/src/store/` |
> | `WarmBackend` trait (put/get/remove/in_use) | `crates/ecstore/src/services/tier/warm_backend.rs` |
> | Per-provider tier backends (S3, MinIO, GCS, Azure, …) | `crates/ecstore/src/services/tier/warm_backend_*.rs` |
> | Remote-tier sweep (`delete_object_from_remote_tier`) | `crates/ecstore/src/bucket/lifecycle/tier_sweeper.rs` |
> | `ObjectInfo` / `TransitionedObject` types | `crates/ecstore/src/object_api/types.rs` |
> | `FileMeta` / `FileInfo` / version metadata | `crates/filemeta/src/` |
> | Dual-key internal metadata helpers (`insert_bytes` / `get_bytes`) | `crates/utils/src/http/metadata_compat.rs` |
**Source of truth:** `crates/ecstore/src/bucket/lifecycle/bucket_lifecycle_ops.rs` (transition/expiry actions), `crates/utils/src/http/metadata_compat.rs` (dual-key helpers), `crates/filemeta/src/filemeta/version.rs` (`SUFFIX_TRANSITIONED_VERSION_ID` read pattern), `crates/filemeta/examples/dump_fileinfo.rs`.
## Code map
| Concern | Location |
|---------|----------|
| ILM actions (`transition_object`, `expire_transitioned_object`, `get_transitioned_object_reader`, `gen_transition_objname`) | `crates/ecstore/src/bucket/lifecycle/bucket_lifecycle_ops.rs` |
| Erasure-set transition/restore entry points | `crates/ecstore/src/set_disk/` and `crates/ecstore/src/store/` |
| `WarmBackend` trait (put/get/remove/in_use) | `crates/ecstore/src/services/tier/warm_backend.rs` |
| Per-provider tier backends (S3, MinIO, GCS, Azure, ...) | `crates/ecstore/src/services/tier/warm_backend_*.rs` |
| Remote-tier sweep (`delete_object_from_remote_tier`) | `crates/ecstore/src/bucket/lifecycle/tier_sweeper.rs` |
| Persisted free-version recovery (remote cleanup after local-first expiry) | `crates/ecstore/src/bucket/lifecycle/tier_free_version_recovery.rs` |
| `ObjectInfo` / `TransitionedObject` types | `crates/ecstore/src/object_api/types.rs` |
| `FileMeta` / `FileInfo` / version metadata | `crates/filemeta/src/` |
| Dual-key internal metadata helpers (`insert_bytes` / `get_bytes`) | `crates/utils/src/http/metadata_compat.rs` |
## Metadata key conventions
Internal metadata is stored under **both** `x-rustfs-internal-<suffix>` and
`x-minio-internal-<suffix>` for MinIO interoperability. `get_bytes` prefers
the RustFS key and falls back to the MinIO key.
Internal metadata is stored under both `x-rustfs-internal-<suffix>` and `x-minio-internal-<suffix>` for MinIO interoperability. `get_bytes` prefers the RustFS key and falls back to the MinIO key.
| Suffix | Meaning |
|--------|---------|
| `transition-status` | `"complete"` when tiered |
| `transitioned-object` | tier key path (stored **without** the tier prefix; `get_dest` adds it) |
| `transitioned-object` | tier key path (stored without the tier prefix; `get_dest` adds it) |
| `transitioned-versionID` | S3 version_id returned by tier PUT (16 raw UUID bytes, or absent) |
| `transition-tier` | tier name |
| `tier-free-versionID` | delete-marker version for free-version sweep |
Reading binary values must reject empty/malformed/nil values (regression
covered in `crates/filemeta/src/filemeta/version.rs` tests):
Reading binary values must reject empty, malformed, and nil values (regression covered in `crates/filemeta/src/filemeta/version.rs` tests):
```rust
get_bytes(&self.meta_sys, SUFFIX_TRANSITIONED_VERSION_ID)
@@ -40,9 +39,7 @@ get_bytes(&self.meta_sys, SUFFIX_TRANSITIONED_VERSION_ID)
// None for: absent key, wrong-length bytes, nil UUID
```
`transition_version_id == None` means the tier bucket is unversioned; the
GET/DELETE against the tier must then send **no** `versionId` parameter.
A nil UUID (`00000000-…`) sent as `?versionId=` causes `NoSuchVersion`.
`transition_version_id == None` means the tier bucket is unversioned; the GET/DELETE against the tier must then send no `versionId` parameter. A nil UUID (`00000000-...`) sent as `?versionId=` causes `NoSuchVersion`. Do not use `Uuid::from_slice(..).unwrap_or_default()` here: it converts an empty metadata value into `Uuid::nil()`, which is exactly that failure.
## Inspect xl.meta directly
@@ -52,106 +49,89 @@ cargo build -p rustfs-filemeta --example dump_fileinfo
# Shows: transition_status, transition_tier, transitioned_obj, transition_ver_id
```
- `transition_ver_id: <none>` → no versionId will be sent to the tier
(correct for a non-versioned tier bucket).
- `transition_ver_id: <uuid>` → that UUID will be sent as `?versionId=<uuid>`.
| Output | Meaning |
|---|---|
| `transition_ver_id: <none>` | No versionId will be sent to the tier (correct for a non-versioned tier bucket). |
| `transition_ver_id: <uuid>` | That UUID will be sent as `?versionId=<uuid>`. |
There is one `xl.meta` per erasure shard disk
(`{disk}/{bucket}/{object}/xl.meta`); all shards of a healthy object should
be identical. `dump_versions` (same crate) lists every version in a file.
There is one `xl.meta` per erasure shard disk (`{disk}/{bucket}/{object}/xl.meta`); all shards of a healthy object should be identical. `dump_versions` (same crate) lists every version in a file.
## Trace the versionId at runtime
```bash
RUST_LOG=rustfs_ecstore::bucket::lifecycle=debug rustfs
RUST_LOG=rustfs_ecstore::bucket::lifecycle=debug rustfs ...
```
- `fetching transitioned object from tier` — DEBUG, before the tier request.
- `tier GET failed` — ERROR, includes `tier_version_id`.
| Log line | Level | Meaning |
|---|---|---|
| `fetching transitioned object from tier` | DEBUG | Emitted before the tier request. |
| `tier GET failed` | ERROR | Includes `tier_version_id`. |
If both `x-rustfs-internal-transitioned-versionID` and
`x-minio-internal-transitioned-versionID` are the **empty string**, the object
was transitioned to a non-versioned tier bucket and no versionId must be sent.
If both `x-rustfs-internal-transitioned-versionID` and `x-minio-internal-transitioned-versionID` are the empty string, the object was transitioned to a non-versioned tier bucket and no versionId must be sent.
## Manual transition run
Manual transition run is an operator trigger for the existing lifecycle transition evaluator. It does not force objects that are not due under the bucket lifecycle rule, and it does not bypass versioning, replication, delete-marker, directory-marker, tier, or in-flight transition checks.
The admin endpoint is:
```text
POST /rustfs/admin/v3/ilm/transition/run?bucket=<bucket>&prefix=<prefix>&tier=<tier>&dryRun=true&maxObjects=10000&maxDurationSeconds=30
```
Only `bucket` is required. `prefix`, `tier`, `dryRun`, `maxObjects`, and `maxDurationSeconds` narrow or bound the run. `maxObjects` defaults to `10000` and is capped at `100000`. `maxDurationSeconds` is optional, capped at `3600`, and enforced as a best-effort budget checked between listed object versions and pages; an in-flight listing call is not cancelled.
| Parameter | Contract |
|---|---|
| `bucket` | Required. |
| `prefix`, `tier`, `dryRun` | Narrow the run. |
| `maxObjects` | Defaults to `10000`, capped at `100000`. |
| `maxDurationSeconds` | Optional, capped at `3600`; a best-effort budget checked between listed object versions and pages. An in-flight listing call is not cancelled. |
The current contract is `enqueue_only`: the response reports what this bounded scan evaluated and enqueued into the in-memory transition queue. `state=completed` means the bounded scan reached the end of its current scope without queue pressure or configured budget truncation. It does not mean every remote tier PUT has completed. `state=partial` means the run stopped early because it hit `maxObjects`, `maxDurationSeconds`, or queue pressure (`skipped_queue_full`, `skipped_queue_closed`, or `skipped_queue_timeout`).
The current contract is `enqueue_only`: the response reports what this bounded scan evaluated and enqueued into the in-memory transition queue. `state=completed` means the bounded scan reached the end of its current scope without queue pressure or budget truncation; it does not mean every remote tier PUT has completed. `state=partial` means the run stopped early on `maxObjects`, `maxDurationSeconds`, or queue pressure (`skipped_queue_full`, `skipped_queue_closed`, or `skipped_queue_timeout`).
`job_id` and `status_endpoint` are currently `null`. There is no durable background job, no restart recovery cursor, no cluster-wide single-flight admission, no status polling endpoint, and no cancel endpoint for this trigger yet. Re-running the command is safe only in the normal lifecycle sense: already in-flight object versions are deduplicated in the local process, but separate nodes do not yet share a durable manual-run admission record.
`job_id` and `status_endpoint` are `null`. There is no durable background job, restart recovery cursor, cluster-wide single-flight admission, status polling endpoint, or cancel endpoint for this trigger. Re-running is safe only in the normal lifecycle sense: in-flight object versions are deduplicated in the local process, but separate nodes do not share a durable manual-run admission record.
Recommended operator flow:
Recommended operator flow (external `rc` CLI):
```bash
rc admin ilm transition run local/mybucket --prefix logs/ --tier cold --dry-run --max-objects 1000 --max-duration-seconds 30
rc admin ilm transition run local/mybucket --prefix logs/ --tier cold --max-objects 1000 --max-duration-seconds 30
```
Inspect the aggregate counters before widening scope. Full object-key lists are intentionally not returned by the admin response. If `RUSTFS_RPC_SECRET` or other credentials were pasted into an issue, chat, log, or ticket while debugging tiering, rotate them on every node, restart the cluster with the new value, and redact the exposed copy before sharing more diagnostics.
Inspect the aggregate counters before widening scope. Full object-key lists are intentionally not returned. If `RUSTFS_RPC_SECRET` or other credentials were pasted into an issue, chat, log, or ticket while debugging tiering, rotate them on every node, restart the cluster with the new value, and redact the exposed copy before sharing more diagnostics.
## Reconcile an unknown transition upload
Historical transition transactions in `upload_outcome_unknown` state can use an explicit two-stage operator workflow when the tier probe is ambiguous and the provider supports exact version deletion. The endpoint refuses transactions that are still inside their ownership window or are in any other state.
First inspect the transaction without changing it:
1. Inspect the transaction without changing it:
```text
GET /rustfs/admin/v3/ilm/transition/reconcile/<transaction-id>
```
```text
GET /rustfs/admin/v3/ilm/transition/reconcile/<transaction-id>
```
If independent provider evidence identifies the exact remote version to remove, submit that opaque version identifier with explicit confirmation:
2. If independent provider evidence identifies the exact remote version to remove, submit that opaque version identifier with explicit confirmation. This performs only an exact version delete; the response reports whether the transaction journal was still observed afterwards, since background recovery may have finalized the same transaction concurrently:
```json
POST /rustfs/admin/v3/ilm/transition/reconcile/<transaction-id>
{
"action": "delete_candidate",
"confirm": true,
"remote_version_id": "<exact-provider-version>"
}
```
```json
POST /rustfs/admin/v3/ilm/transition/reconcile/<transaction-id>
{
"action": "delete_candidate",
"confirm": true,
"remote_version_id": "<exact-provider-version>"
}
```
This operation performs only an exact version delete. Its response reports whether the transaction journal was still observed after the delete; background recovery may have finalized the same transaction concurrently. If the journal remains, inspect the transaction again and finalize it only after the live provider probe proves that the candidate is missing:
3. If the journal remains, inspect again and finalize only after the live provider probe proves the candidate is missing:
```json
POST /rustfs/admin/v3/ilm/transition/reconcile/<transaction-id>
{
"action": "finalize_missing",
"confirm": true
}
```
```json
POST /rustfs/admin/v3/ilm/transition/reconcile/<transaction-id>
{
"action": "finalize_missing",
"confirm": true
}
```
`finalize_missing` re-runs the provider probe and fails closed for `unversioned_present`, `versioned_present`, `ambiguous`, `unsupported`, or probe errors. It never accepts an operator assertion in place of a live `missing` result. Providers without an authoritative probe or exact version deletion remain pending; this endpoint does not infer provider capabilities, accept external absence assertions, or select a candidate automatically.
`finalize_missing` re-runs the provider probe and fails closed for `unversioned_present`, `versioned_present`, `ambiguous`, `unsupported`, or probe errors. It never accepts an operator assertion in place of a live `missing` result. Providers without an authoritative probe or exact version deletion remain pending; the endpoint does not infer provider capabilities, accept external absence assertions, or select a candidate automatically.
## Historical fixes (for context, already merged)
## Invariant: local-first expiry ordering
- Expire/GET race (`NoSuchVersion` during expiry of a tiered object):
`expire_transitioned_object` used to delete the remote tier version
**before** local metadata, so a concurrent GET between the two steps read a
stored version_id whose remote version was already gone. Fixed in `#3491`:
local metadata is deleted **first** (making the object unreachable), and
remote-tier cleanup is driven by persisted free-version recovery
(`crates/ecstore/src/bucket/lifecycle/tier_free_version_recovery.rs`).
**Invariant — keep local-first ordering**: never remove a remote tier
version while live local metadata still points at it. Regression test:
`serial_tests::test_expire_transitioned_object_never_races_concurrent_get`
in `crates/scanner/tests/lifecycle_integration_test.rs` (runs in the CI
ILM Integration serial lane) pins both the local-first ordering and the
"concurrent GET never sees `NoSuchVersion`" contract.
- Nil-UUID versionId sent to tier (`NoSuchVersion`): reading code used
`Uuid::from_slice(..).unwrap_or_default()`, converting an empty metadata
value into `Uuid::nil()`. Fixed by the `and_then`/`filter` pattern above.
- `warm_backend_s3sdk` ignored the remote version and range options on
GET/DELETE.
- `copy_object` returned 501 (`NotImplemented`) for tiered objects on
self-copy with `--storage-class`, blocking de-tiering via
`mc cp --storage-class STANDARD obj obj`. Fixed by writing the
tier-fetched `put_object_reader` back through `put_object`.
`expire_transitioned_object` deletes local metadata first (making the object unreachable) and leaves remote-tier cleanup to persisted free-version recovery (`crates/ecstore/src/bucket/lifecycle/tier_free_version_recovery.rs`). Never remove a remote tier version while live local metadata still points at it: doing so lets a concurrent GET read a stored version_id whose remote version is already gone and fail with `NoSuchVersion`.
Regression test: `serial_tests::test_expire_transitioned_object_never_races_concurrent_get` in `crates/scanner/tests/lifecycle_integration_test.rs` (CI ILM Integration serial lane) pins both the local-first ordering and the "concurrent GET never sees `NoSuchVersion`" contract.
+56 -200
View File
@@ -1,23 +1,16 @@
# Two-Factor Authentication
> Scope: the self-service account surface (`/rustfs/admin/v3/account/*`), the
> login gate on `AssumeRole`, and the administrative reset
> (`/rustfs/admin/v3/user/mfa`).
**Use this when:** changing anything under `/rustfs/admin/v3/account/*`, `/v3/mfa/challenge`, `/v3/user/mfa`, or the `AssumeRole` login gate, or before "fixing" a 2FA boundary that looks like a gap. This is the design record of what the second factor protects and why.
This document records what the second factor does and does not protect, and why.
The boundaries are deliberate; several of them look like gaps until the
alternative is spelled out.
**Source of truth:** `crates/iam/src/mfa/` (`record.rs` `MAX_FAILED_ATTEMPTS`, `totp.rs` `TOTP_SKEW_STEPS`, `challenge.rs` `CHALLENGE_TTL_SECONDS`, `recovery.rs` `RECOVERY_CODE_COUNT` / `RECOVERY_CODE_ENTROPY_BITS`), `crates/credentials/src/credentials.rs` (root credential `OnceLock`), `crates/iam/src/root_credentials.rs` (`token_signing_key`), `rustfs/src/admin/route_policy.rs` (`CredentialOnly` routes).
## What is protected
TOTP gates **session minting**: the `AssumeRole` call that turns a long-term
credential into a short-lived STS session. That is the only interactive login
RustFS has — the Console holds nothing but an STS session, and obtains it by
signing an `AssumeRole` request with the access key the user typed.
TOTP gates session minting: the `AssumeRole` call that turns a long-term credential into a short-lived STS session. That is the only interactive login RustFS has; the Console holds nothing but an STS session, obtained by signing an `AssumeRole` request with the access key the user typed.
When an identity has an active enrollment:
```
```text
access key + secret key
@@ -32,268 +25,131 @@ POST / Action=AssumeRole
STS credentials, with the claim x-rustfs-mfa-verified: true
```
Without a `TokenCode`, `AssumeRole` fails with `AccessDenied` and a message
carrying the `MultiFactorAuthRequired` marker. Clients match on that marker to
prompt for a code rather than reporting a failed login — the password *was*
accepted.
`SerialNumber` and `TokenCode` are `AssumeRole`'s own parameters, so an SDK or a
script authenticates the same way the Console does, with no RustFS-specific
protocol.
Without a `TokenCode`, `AssumeRole` fails with `AccessDenied` and a message carrying the `MultiFactorAuthRequired` marker. Clients match on that marker to prompt for a code rather than reporting a failed login; the password was accepted. `SerialNumber` and `TokenCode` are `AssumeRole`'s own parameters, so an SDK or script authenticates the same way the Console does, with no RustFS-specific protocol.
## What is deliberately not protected
**A request signed directly with a long-term access key is not gated.** This is
the most important boundary in the design, and it is intentional:
| Boundary | Rationale |
| --- | --- |
| A request signed directly with a long-term access key is not gated. | Gating it would break every script, SDK client, and `rc` invocation the moment a human enabled 2FA on their own account, and it would add no protection: whoever holds the secret key already has full access and never needs to mint a session. AWS draws the same line: MFA gates `AssumeRole` and is enforced for API calls through the `aws:MultiFactorAuthPresent` policy condition, not by refusing signed requests. |
| OIDC and Keystone sessions are not gated. | Those identities are authenticated by their provider; a RustFS-side enrollment would not be consulted at login and would give a false impression of protection. `CallerIdentity` reports such sessions as `FederatedIdentity` and refuses enrollment. |
| Service-account credentials cannot manage their parent's factor. | A machine credential must not be able to take over the human identity it was minted from. |
- Gating it would break every script, SDK client and `rc` invocation the moment
a human enabled 2FA on their own account. An operator who turned on a security
feature would discover it by way of a production outage.
- It would add no protection. Whoever holds the secret key already has full
access to everything that identity can reach; they never need to present a
code, because they never need to mint a session.
This is the same division AWS draws: MFA gates `AssumeRole` and is enforced for
API calls through the `aws:MultiFactorAuthPresent` policy condition, not by
refusing signed requests.
**Consequence to state plainly:** 2FA raises the cost of a stolen *password*. It
does not contain a stolen *secret key*. Because in RustFS the password **is** the
S3 secret key (see below), those are the same string — so 2FA protects the
console login path against credential reuse and phishing, and nothing more,
until the policy-condition work lands.
The tracked follow-up is an `aws:MultiFactorAuthPresent` condition key populated
from the `x-rustfs-mfa-verified` session claim, which would let an operator write
a policy that denies administrative actions to a session that presented no
second factor. That is the mechanism that makes 2FA meaningful for API access.
**OIDC and Keystone sessions are not gated either.** Those identities are
authenticated by their provider; a RustFS-side TOTP enrollment would not be
consulted at login and would give a false impression of protection. MFA for a
federated identity belongs to its IdP. `CallerIdentity` reports such sessions as
`FederatedIdentity` and refuses enrollment.
**Service-account credentials cannot manage their parent's factor.** A machine
credential must not be able to take over the human identity it was minted from.
Consequence to state plainly: 2FA raises the cost of a stolen password; it does not contain a stolen secret key. Because in RustFS the password is the S3 secret key (below), those are the same string, so 2FA protects the console login path against credential reuse and phishing, and nothing more, until the policy-condition work lands. The tracked follow-up is an `aws:MultiFactorAuthPresent` condition key populated from the `x-rustfs-mfa-verified` session claim, which would let an operator deny administrative actions to a session that presented no second factor. That is the mechanism that makes 2FA meaningful for API access; it is not implemented yet.
## Password reality: there is no password hash
RustFS is an S3 server. SigV4 requires the server to know the secret key itself
in order to recompute a request signature, so secret keys **cannot** be hashed —
not here, and not in any S3-compatible implementation. The "password" a user
types into the Console is their S3 secret key.
What protects it instead:
RustFS is an S3 server. SigV4 requires the server to know the secret key itself to recompute a request signature, so secret keys cannot be hashed, here or in any S3-compatible implementation. The "password" a user types into the Console is their S3 secret key.
| Protection | Mechanism |
| --- | --- |
| At rest | `RUSTFS_IAM_MASTER_KEY` + `encrypt_stream_io` (Argon2id AES-GCM / ChaCha20-Poly1305) |
| At rest | `RUSTFS_IAM_MASTER_KEY` + `encrypt_stream_io` (Argon2id -> AES-GCM / ChaCha20-Poly1305) |
| Length floor | `is_secret_key_valid` (`SECRET_KEY_MIN_LEN`) |
| Maximum length | None, deliberately; capping password length is an anti-pattern. |
| Rotation | `POST /v3/account/password`, requiring the current secret |
| Session cleanup | Every STS session minted from the identity is revoked on rotation |
There is deliberately **no maximum length**. The previous Console capped
passwords at 40 characters, which was a client-side invention with no server
constraint behind it; capping password length is an anti-pattern.
## At-rest protection is mandatory for TOTP secrets
A TOTP secret is credential-equivalent: anyone holding it can mint valid codes
forever. So enrollment is **refused** when `RUSTFS_IAM_MASTER_KEY` is not
configured, rather than writing the secret in plaintext:
A TOTP secret is credential-equivalent: anyone holding it can mint valid codes forever. Enrollment is therefore refused when `RUSTFS_IAM_MASTER_KEY` is not configured, rather than writing the secret in plaintext:
```
```text
POST /v3/account/mfa/enroll → 501 NotImplemented
"two-factor authentication requires RUSTFS_IAM_MASTER_KEY to be configured
so the shared secret can be encrypted at rest"
```
IAM *identities* tolerate a missing master key for backward compatibility with
existing deployments. A new feature has no such history to honour, and a second
factor that can be lifted off a disk is worse than none, because the user
believes they have one.
`GET /v3/account/mfa` reports `enrollment_available: false` with the reason, so
the Console and `rc` explain the remedy instead of offering a control that fails.
IAM identities tolerate a missing master key for backward compatibility with existing deployments. A new feature has no such history to honour, and a second factor that can be lifted off a disk is worse than none because the user believes they have one. `GET /v3/account/mfa` reports `enrollment_available: false` with the reason, so the Console and `rc` explain the remedy instead of offering a control that fails.
## Root credentials cannot be changed at runtime
The root identity comes from `RUSTFS_ACCESS_KEY` / `RUSTFS_SECRET_KEY` and lands
in a process-wide `OnceLock` (`crates/credentials/src/credentials.rs`). It cannot
be rotated while the server runs, and the account surface reports this as
`credentials_source: "env"` with `mutable.password: false`.
The root identity comes from `RUSTFS_ACCESS_KEY` / `RUSTFS_SECRET_KEY` and lands in a process-wide `OnceLock` (`crates/credentials/src/credentials.rs`). It cannot be rotated while the server runs, and the account surface reports this as `credentials_source: "env"` with `mutable.password: false`.
This is not merely a missing feature. The root secret key feeds three things:
The root secret key feeds three things, so rotating it at runtime would invalidate every session cluster-wide and break node-to-node authentication:
1. **STS session token signing** (`root_credentials::token_signing_key`) every
live session in the cluster is HMAC-signed with it.
2. **The internode RPC secret** (`derive_rpc_secret`), unless
`RUSTFS_RPC_SECRET` is set explicitly.
3. **Legacy IAM at-rest decryption** for blobs migrated from MinIO.
1. STS session token signing (`root_credentials::token_signing_key`): every live session in the cluster is HMAC-signed with it.
2. The internode RPC secret (`derive_rpc_secret`), unless `RUSTFS_RPC_SECRET` is set explicitly.
3. Legacy IAM at-rest decryption for blobs migrated from MinIO.
Rotating it at runtime would therefore invalidate every session cluster-wide and
break node-to-node authentication. Making root mutable is a separate piece of
work with those three couplings as prerequisites; it is not a side effect of
adding a profile page.
**Operational recommendation:** treat root as a bootstrap identity. Create a
built-in IAM user with the `consoleAdmin` policy for day-to-day administration.
That identity has a working password change and full 2FA support.
Making root mutable is separate work with those three couplings as prerequisites. Operational recommendation: treat root as a bootstrap identity and create a built-in IAM user with the `consoleAdmin` policy for day-to-day administration; that identity has a working password change and full 2FA support.
## Rate limiting, replay and expiry
| Control | Value | Where |
| Control | Value | Where / why |
| --- | --- | --- |
| Failed attempts before lockout | 5 | `mfa/record.rs` |
| First lockout | 15 minutes, doubling per further run | `mfa/record.rs` |
| Lockout ceiling | 1 hour | so a sustained attack cannot deny the owner indefinitely |
| TOTP clock skew | ±1 step (±30s) | three codes valid at once, no more |
| TOTP replay | Consumed time step is a high-water mark; `step <= last_used` is refused | closes the ~90s window a captured code would otherwise have |
| Failed attempts before lockout | 5 (`MAX_FAILED_ATTEMPTS`) | `crates/iam/src/mfa/record.rs` |
| First lockout | 15 minutes, doubling per further run | `record.rs` |
| Lockout ceiling | 1 hour | A sustained attack cannot deny the owner indefinitely. |
| TOTP clock skew | +/-1 step, +/-30s (`TOTP_SKEW_STEPS`) | Three codes valid at once, no more. |
| TOTP replay | Consumed time step is a high-water mark; `step <= last_used` is refused | Closes the ~90s window a captured code would otherwise have. |
| Recovery code replay | `used_at` stamp, single use | |
| Login challenge TTL | 5 minutes | |
| Pending enrollment TTL | 10 minutes | an abandoned enrollment leaves no usable secret |
| Login challenge TTL | 5 minutes (`CHALLENGE_TTL_SECONDS`) | |
| Pending enrollment TTL | 10 minutes | An abandoned enrollment leaves no usable secret. |
A wrong code, a replayed code and a malformed code are **indistinguishable** on
the wire: all three return `AccessDenied` with the same message. The distinction
survives only in the audit trail, so an operator can tell a guessing attempt from
a replay without an attacker learning that a captured code was genuine.
The lockout is stored in the record and updated under an optimistic
compare-and-set, so it holds across the cluster rather than per node.
A wrong code, a replayed code, and a malformed code are indistinguishable on the wire: all three return `AccessDenied` with the same message. The distinction survives only in the audit trail, so an operator can tell a guessing attempt from a replay without an attacker learning that a captured code was genuine. The lockout is stored in the record and updated under an optimistic compare-and-set, so it holds across the cluster rather than per node.
## Storage
```
```text
.rustfs.sys/config/mfa/<access-key>/totp.json (encrypted with the IAM master key)
```
A sibling of `config/iam/`, not a child: the IAM cache loader walks the whole
`config/iam/` tree on startup and buckets what it finds by first path segment, so
a new prefix under there would be swept into that walk for no benefit.
A sibling of `config/iam/`, not a child: the IAM cache loader walks the whole `config/iam/` tree on startup and buckets what it finds by first path segment, so a new prefix under there would be swept into that walk for no benefit.
Records are **not cached**. Every verification reads from the store, because a
cache would need cluster-wide invalidation to keep the replay mark and the
lockout counter honest, and getting that wrong reopens exactly the holes this
design closes. Verifications are rare enough that the read is not worth
optimising.
Writes are read-modify-write under an `If-Match` precondition with bounded
retries — the same optimistic scheme the IAM lazy-rewrite path uses. It degrades
to a retry rather than to a distributed lock a crashed node would have to time
out.
Records are not cached. Every verification reads from the store, because a cache would need cluster-wide invalidation to keep the replay mark and the lockout counter honest, and getting that wrong reopens exactly the holes this design closes. Verifications are rare enough that the read is not worth optimising. Writes are read-modify-write under an `If-Match` precondition with bounded retries, the same optimistic scheme the IAM lazy-rewrite path uses; it degrades to a retry rather than to a distributed lock a crashed node would have to time out.
## Login challenges are stateless
A challenge is `HMAC-SHA256(root_secret, "rustfs-mfa-challenge:v1" ‖ access_key ‖
issued_at)`, base64url-encoded with its payload.
The obvious alternative is a TTL cache, the way the OIDC flow stores its PKCE
verifiers. That store is node-local, which is fine for OIDC because the whole
authorization round trip returns to the node that started it. A second factor
does not: a cluster behind a load balancer without session affinity would issue
the challenge on one node and receive the code on another, and a node-local
challenge would fail there for reasons no operator could debug.
Statelessness costs nothing, because the challenge is not what makes the exchange
single-use — the consumed TOTP time step is.
A challenge is `HMAC-SHA256(root_secret, "rustfs-mfa-challenge:v1" ‖ access_key ‖ issued_at)`, base64url-encoded with its payload. A node-local TTL cache (the way the OIDC flow stores PKCE verifiers) works for OIDC because the whole round trip returns to the node that started it; a second factor behind a load balancer without session affinity would issue the challenge on one node and receive the code on another. The challenge is not what makes the exchange single-use; the consumed TOTP time step is.
## Authorization model
| Route | Gate |
| --- | --- |
| `GET /v3/account/info` | possession of the credential |
| `POST /v3/account/password` | credential **+ knowledge of the current secret** |
| `POST /v3/account/password` | credential + knowledge of the current secret |
| `GET /v3/account/mfa` | possession of the credential |
| `POST /v3/account/mfa/enroll` | credential, and the credential kind must be mutable |
| `POST /v3/account/mfa/activate` | credential + a valid code from the pending secret |
| `POST /v3/account/mfa/disable` | credential + **a valid code and the account password** |
| `POST /v3/account/mfa/disable` | credential + a valid code and the account password |
| `POST /v3/account/mfa/recovery-codes` | credential + a valid code |
| `GET /v3/mfa/challenge` | possession of the credential |
| `GET /v3/user/mfa` | `admin:GetUser` |
| `DELETE /v3/user/mfa` | `admin:EnableUser` |
| `PUT /v3/set-user-secret-key` | `admin:CreateUser` |
The self-service routes carry **no admin action**. Giving them one would be wrong
in both directions: it would stop an ordinary user from changing their own
password, and it would let any holder of that action change somebody else's.
They are registered as `CredentialOnly` in the route-policy matrix.
The self-service routes carry no admin action and are registered as `CredentialOnly` in the route-policy matrix. Giving them one would be wrong in both directions: it would stop an ordinary user from changing their own password, and it would let any holder of that action change somebody else's.
`POST /v3/account/password` and `POST /v3/account/mfa/disable` require a
proof-of-knowledge step because a signature only proves a credential was *used*.
The Console signs with a short-lived session, so without it a hijacked browser
tab could rewrite the account's credentials or strip its second factor.
`POST /v3/account/password` and `POST /v3/account/mfa/disable` require a proof-of-knowledge step because a signature only proves a credential was used. The Console signs with a short-lived session, so without it a hijacked browser tab could rewrite the account's credentials or strip its second factor; requiring the password makes disabling the factor as hard as the thing the factor protects.
### Why turning the factor off needs the password too
Requiring only a code would mean a single shoulder-surfed number, in a session
someone walked away from, is enough to remove the protection. Requiring the
password makes disabling the factor as hard as the thing the factor protects.
### Break-glass
`DELETE /v3/user/mfa` clears another identity's factor, for a user who lost both
their authenticator and their recovery codes. It is gated on `admin:EnableUser`
rather than a bespoke action, because that is the same capability that can
already re-enable a disabled account — anyone who can do that can already take
the identity over, so a separate action would be a distinction without a security
difference.
The record is deleted outright rather than disabled, so no stale lockout counter
survives to block the user's next enrollment. The acting administrator is
recorded in the audit entry.
Break-glass: `DELETE /v3/user/mfa` clears another identity's factor, for a user who lost both their authenticator and their recovery codes. It is gated on `admin:EnableUser` rather than a bespoke action because that capability can already re-enable a disabled account, and anyone who can do that can already take the identity over. The record is deleted outright rather than disabled, so no stale lockout counter survives to block the user's next enrollment. The acting administrator is recorded in the audit entry.
## Recovery codes
Ten codes, `XXXX-XXXX-XXXX-XXXX-XXXX`, 100 bits of uniform randomness each, in a
Crockford base32 alphabet with `I`, `L`, `O` and `U` removed so a handwritten
code cannot be ambiguous.
Ten codes (`RECOVERY_CODE_COUNT`), `XXXX-XXXX-XXXX-XXXX-XXXX`, 100 bits of uniform randomness each (`RECOVERY_CODE_ENTROPY_BITS`), in a Crockford base32 alphabet with `I`, `L`, `O` and `U` removed so a handwritten code cannot be ambiguous.
Stored as domain-separated SHA-256 digests, **not** a password KDF. With 100 bits
of uniform randomness there is no dictionary to try and no human-chosen pattern
to exploit, so the attacks a slow KDF defends against do not apply — while a
memory-hard KDF would have to run once per stored code on every verification
attempt, turning each guess into an attacker-controlled multiple of that cost.
This is the standard treatment for high-entropy bearer tokens, and the same
reasoning is why there is no per-code salt.
Stored as domain-separated SHA-256 digests, not a password KDF. With 100 bits of uniform randomness there is no dictionary to try and no human-chosen pattern to exploit, so the attacks a slow KDF defends against do not apply, while a memory-hard KDF would have to run once per stored code on every verification attempt, turning each guess into an attacker-controlled multiple of that cost. This is the standard treatment for high-entropy bearer tokens, and the same reasoning is why there is no per-code salt.
Codes are returned in plaintext exactly once. Activation always replaces the set:
reusing a previous one would leave codes valid for a secret they were never
issued against. Disabling clears them, so no live bypass survives a factor the
user believes is gone.
Codes are returned in plaintext exactly once. Activation always replaces the set: reusing a previous one would leave codes valid for a secret they were never issued against. Disabling clears them, so no live bypass survives a factor the user believes is gone.
## Audit
Two `EventName` variants carry the whole surface:
- `iam:Identity:CredentialChanged` — password rotation, enrollment, activation,
disable, recovery-code regeneration, administrative reset.
- `iam:Identity:AuthChallenge` — challenge issuance and second-factor
verification.
| Event | Covers |
| --- | --- |
| `iam:Identity:CredentialChanged` | Password rotation, enrollment, activation, disable, recovery-code regeneration, administrative reset. |
| `iam:Identity:AuthChallenge` | Challenge issuance and second-factor verification. |
The per-operation detail lives in `api.name` and the `iamOperation` tag, which is
what a SIEM filters on. The enum is coarse because `EventName::mask()` gives every
variant its own bit in a `u64` and the budget is nearly spent — 63 of 64 used
after these two. Splitting these per-operation needs `mask()` widened first.
The per-operation detail lives in `api.name` and the `iamOperation` tag, which is what a SIEM filters on. The enum is coarse because `EventName::mask()` gives every variant its own bit in a `u64`; splitting these per operation needs `mask()` widened first.
**Redaction:** no secret key, TOTP secret, provisioning URI, submitted code or
recovery code enters an audit entry — not even hashed, and not on the failure
paths where the submitted value would be the most tempting thing to record.
Failures are described by a closed set of static strings
(`AccountAuditFailure`), so no caller-supplied bytes can reach a log target
through this module.
Redaction: no secret key, TOTP secret, provisioning URI, submitted code, or recovery code enters an audit entry, not even hashed, and not on the failure paths where the submitted value would be the most tempting thing to record. Failures are described by a closed set of static strings (`AccountAuditFailure`), so no caller-supplied bytes can reach a log target through this module.
## Known limitations
1. **2FA does not gate direct SigV4 access.** By design; see above. The fix is
the `aws:MultiFactorAuthPresent` policy condition.
2. **Root cannot rotate its own credentials at runtime.** By design; see above.
3. **GHSA-m77q-r63m-pj89 is unaffected.** STS session tokens are signed with the
root secret key, so anyone holding it can still forge a session token —
including one carrying `x-rustfs-mfa-verified`. 2FA does not close this; a
dedicated STS signing key does, and that advisory is tracked separately.
4. **Username changes are not supported for anyone.** The access key is the
primary key for policy mappings, group membership, service-account parents and
bucket-policy principals. A rename is a migration that orphans service
accounts and silently breaks bucket-policy ARNs, not an edit;
`mutable.username` is `false` for every identity.
| Limitation | Status |
| --- | --- |
| 2FA does not gate direct SigV4 access. | Design boundary (above). The fix is the `aws:MultiFactorAuthPresent` policy condition, not implemented yet. |
| Root cannot rotate its own credentials at runtime. | Design boundary (above); three couplings must be broken first. |
| GHSA-m77q-r63m-pj89 is unaffected. | STS session tokens are signed with the root secret key, so anyone holding it can still forge a session token, including one carrying `x-rustfs-mfa-verified`. A dedicated STS signing key closes this; tracked separately. |
| Username changes are not supported for anyone. | The access key is the primary key for policy mappings, group membership, service-account parents, and bucket-policy principals. A rename is a migration that orphans service accounts and silently breaks bucket-policy ARNs; `mutable.username` is `false` for every identity. |
+26 -18
View File
@@ -1,6 +1,8 @@
# Vault KMS authentication runbook
This runbook covers how the RustFS Vault KMS backends (KV2 and Transit) authenticate to Vault, how to deploy AppRole and Vault Agent token-file authentication, and how the fail-closed credential window behaves in production. For what each backend stores in Vault and the KV2/Transit policy scopes, see [KMS backend security properties](kms-backend-security.md).
**Use this when:** configuring how the Vault KMS backends (KV2 and Transit) authenticate to Vault, rotating AppRole SecretIDs or agent tokens, or diagnosing `KMS credentials unavailable` errors.
**Source of truth:** `crates/kms/src/backends/vault_credentials.rs` (`DEFAULT_TOKEN_FILE_POLL_INTERVAL_SECS`, `refresh_safety_window_secs`, the renewal loop and its log lines), `crates/kms/src/config.rs` (`RUSTFS_KMS_TIMEOUT_SECS` default). For what each backend stores in Vault and the KV2/Transit policy scopes, see [KMS backend security properties](kms-backend-security.md).
## Choosing an authentication method
@@ -11,7 +13,7 @@ This runbook covers how the RustFS Vault KMS backends (KV2 and Transit) authenti
| Kubernetes | `Kubernetes` | Lease-bound token obtained by login; renewed by RustFS | Renew at half TTL, re-login on failure | Production on Kubernetes, with no credential to distribute |
| Agent token file | `TokenFile` | Owned by Vault Agent; RustFS only re-reads the sink file | File re-read once per poll interval | Production with a Vault Agent (or equivalent) managing auth |
Exactly one method must be configured. Setting `RUSTFS_KMS_VAULT_TOKEN_FILE` together with any other method, or `RUSTFS_KMS_VAULT_KUBERNETES_ROLE` together with `RUSTFS_KMS_VAULT_APPROLE_ROLE_ID`, is rejected at startup with a configuration error, because the effective identity would be ambiguous. A leftover `RUSTFS_KMS_VAULT_TOKEN` alongside a configured login method is tolerated and ignored, so a stale variable cannot silently downgrade the identity.
Exactly one method must be configured. Setting `RUSTFS_KMS_VAULT_TOKEN_FILE` together with any other method, or `RUSTFS_KMS_VAULT_KUBERNETES_ROLE` together with `RUSTFS_KMS_VAULT_APPROLE_ROLE_ID`, is rejected at startup with a configuration error because the effective identity would be ambiguous. A leftover `RUSTFS_KMS_VAULT_TOKEN` alongside a configured login method is tolerated and ignored, so a stale variable cannot silently downgrade the identity.
All of these are read the same way whether the service is started with `RUSTFS_KMS_ENABLE=true` or configured later through `POST /rustfs/admin/v3/kms/configure`.
@@ -55,11 +57,15 @@ RustFS logs in at startup, then renews the token at half its TTL in the backgrou
### SecretID delivery and rotation
Deliver the SecretID out of band a secrets-manager-mounted file, an init-container writing `RUSTFS_KMS_VAULT_APPROLE_SECRET_ID_FILE`, or Vault response wrapping unwrapped by your deployment tooling. Treat it like a password: owner-readable file permissions, never in logs or shell history.
Deliver the SecretID out of band (a secrets-manager-mounted file, an init-container writing `RUSTFS_KMS_VAULT_APPROLE_SECRET_ID_FILE`, or Vault response wrapping unwrapped by your deployment tooling). Treat it like a password: owner-readable file permissions, never in logs or shell history.
The secret_id file is re-read on every login attempt, so rotating the SecretID is a two-step operation with no restart: generate a new SecretID (`vault write -f auth/approle/role/rustfs-kms/secret-id`), atomically replace the file, then revoke the old SecretID accessor. The already-issued token keeps renewing; the new SecretID is only needed at the next full re-login.
The secret_id file is re-read on every login attempt, so rotating the SecretID needs no restart:
An empty or missing secret_id file fails the login attempt immediately (no Vault round trip). At startup the error is fatal — provider construction fails and the process exits — so a file missing at boot is recovered by restarting the process, not by an in-process retry. Once RustFS is running, the same failure is retried on the normal refresh cadence, so repairing the file mid-run heals the backend without a restart.
1. Generate a new SecretID: `vault write -f auth/approle/role/rustfs-kms/secret-id`.
2. Atomically replace the file.
3. Revoke the old SecretID accessor. The already-issued token keeps renewing; the new SecretID is only needed at the next full re-login.
An empty or missing secret_id file fails the login attempt immediately (no Vault round trip). At startup the error is fatal: provider construction fails and the process exits, so a file missing at boot is recovered by restarting the process, not by an in-process retry. Once RustFS is running, the same failure is retried on the normal refresh cadence, so repairing the file mid-run heals the backend without a restart.
## Kubernetes authentication
@@ -96,7 +102,7 @@ RUSTFS_KMS_VAULT_KUBERNETES_ROLE=rustfs
RustFS logs in at startup and renews the token at half its TTL, falling back to a fresh login exactly as AppRole does. The ServiceAccount token is re-read from disk on every login rather than cached, so a projected token the kubelet rotates is picked up without a restart.
A missing or empty token file fails the login attempt immediately (no Vault round trip). At startup the error is fatal provider construction fails and the process exits so a token projected late during a slow pod start is recovered by the pod restart loop, not by an in-process retry. Once RustFS is running, a token file that goes missing or turns empty is retried on the normal refresh cadence and heals the backend on its own.
A missing or empty token file fails the login attempt immediately (no Vault round trip). At startup the error is fatal (provider construction fails and the process exits), so a token projected late during a slow pod start is recovered by the pod restart loop, not by an in-process retry. Once RustFS is running, a token file that goes missing or turns empty is retried on the normal refresh cadence and heals the backend on its own.
## Vault Agent token file
@@ -130,22 +136,24 @@ RUSTFS_KMS_VAULT_ADDRESS=https://vault.example.com:8200
RUSTFS_KMS_VAULT_TOKEN_FILE=/run/vault-agent/token
```
The poll interval (`poll_interval_secs` in the `TokenFile` auth configuration, default 30 seconds) controls how often the file is re-read. Each successful read grants the token an observed validity of twice the poll interval and installs a fresh client generation, so an agent-rotated token is picked up within one poll interval of the atomic replace.
The poll interval (`poll_interval_secs` in the `TokenFile` auth configuration, `DEFAULT_TOKEN_FILE_POLL_INTERVAL_SECS`, 30 seconds) controls how often the file is re-read. Each successful read grants the token an observed validity of twice the poll interval and installs a fresh client generation, so an agent-rotated token is picked up within one poll interval of the atomic replace.
Requirements enforced at every read, each failing the refresh without contacting Vault:
- The file must exist and be non-empty after trimming whitespace.
- On Unix, the file must not be readable or writable by group or other (mode `0600` or stricter). Wider permissions are a hard error naming the offending mode, mirroring the SFTP host-key rule. RustFS must run as the file's owner.
If the agent stops refreshing the file that is fine — RustFS re-reads the same token and keeps going as long as the token itself is valid on the Vault side. If the file disappears or turns empty, RustFS keeps serving requests on the last-read token until the fail-closed window trips, and heals automatically once the file is restored.
If the agent stops refreshing the file, RustFS re-reads the same token and keeps going as long as the token itself is valid on the Vault side. If the file disappears or turns empty, RustFS keeps serving requests on the last-read token until the fail-closed window trips, and heals automatically once the file is restored.
## Fail-closed window
For lease-bound credentials (AppRole and Kubernetes tokens, token files), `current()` refuses to hand out a token that is within the safety window of its expiry and has not been refreshed. Requests then fail with `KMS credentials unavailable: ...` instead of being sent with a token that could lapse mid-flight and fail unpredictably on the Vault side.
- Default window: one per-attempt timeout (`RUSTFS_KMS_TIMEOUT_SECS`, default 30s) — a request issued now can legitimately stay in flight that long, so the token must outlive it.
- Override: `refresh_safety_window_secs` on the `AppRole`, `Kubernetes` or `TokenFile` auth configuration.
- Static tokens never trip the window: they carry no lease and are assumed valid until Vault says otherwise.
| Aspect | Value |
| --- | --- |
| Default window | One per-attempt timeout (`RUSTFS_KMS_TIMEOUT_SECS`, default 30s): a request issued now can legitimately stay in flight that long, so the token must outlive it. |
| Override | `refresh_safety_window_secs` on the `AppRole`, `Kubernetes` or `TokenFile` auth configuration. |
| Static tokens | Never trip the window; they carry no lease and are assumed valid until Vault says otherwise. |
The window is a symptom threshold, not the fault itself: by the time it trips, refresh has been failing for roughly half the token TTL (AppRole, Kubernetes) or two poll intervals (token file).
@@ -153,12 +161,12 @@ The window is a symptom threshold, not the fault itself: by the time it trips, r
| Symptom | Log line to look for | Likely cause and fix |
| --- | --- | --- |
| Requests fail with `KMS credentials unavailable` | `Vault credential refresh failed; retrying until the credentials recover` (warn, repeated) | Vault unreachable/sealed, or the credential source is broken; the provider recovers on its own once refresh succeeds — fix the cause, no restart needed |
| Renewal succeeded but re-login later fails | `Vault token renewal failed; falling back to a fresh login` followed by login errors | SecretID expired/revoked or AppRole role changed; rotate the secret_id file |
| Token file mode error at startup or during polls | `has insecure permissions` in the error | Fix the sink `mode` (0600) and the file owner; the next poll heals the provider |
| Token file missing/empty errors | `Failed to read Vault token file` / `token file ... is empty` | Vault Agent down or sink misconfigured; restart the agent, the next poll heals the provider |
| Kubernetes login fails with a permission error | `Vault Kubernetes login failed` | The pod's ServiceAccount is not in the role's `bound_service_account_names`/`_namespaces`, or `auth/kubernetes/config` names the wrong API server |
| Kubernetes ServiceAccount token errors | `Failed to read Kubernetes ServiceAccount token` / `ServiceAccount token ... is empty` | The token is not projected into the pod (check `automountServiceAccountToken` and the volume mount); the next refresh cycle heals the provider |
| Startup fails immediately with a configuration error naming two env vars | | Two auth methods configured at once; keep exactly one of token, AppRole, Kubernetes, token file |
| Requests fail with `KMS credentials unavailable` | `Vault credential refresh failed; retrying until the credentials recover` (warn, repeated) | Vault unreachable/sealed, or the credential source is broken; the provider recovers on its own once refresh succeeds. Fix the cause; no restart needed. |
| Renewal succeeded but re-login later fails | `Vault token renewal failed; falling back to a fresh login` followed by login errors | SecretID expired/revoked or AppRole role changed; rotate the secret_id file. |
| Token file mode error at startup or during polls | `has insecure permissions` in the error | Fix the sink `mode` (0600) and the file owner; the next poll heals the provider. |
| Token file missing/empty errors | `Failed to read Vault token file` / `token file ... is empty` | Vault Agent down or sink misconfigured; restart the agent, the next poll heals the provider. |
| Kubernetes login fails with a permission error | `Vault Kubernetes login failed` | The pod's ServiceAccount is not in the role's `bound_service_account_names`/`_namespaces`, or `auth/kubernetes/config` names the wrong API server. |
| Kubernetes ServiceAccount token errors | `Failed to read Kubernetes ServiceAccount token` / `ServiceAccount token ... is empty` | The token is not projected into the pod (check `automountServiceAccountToken` and the volume mount); the next refresh cycle heals the provider. |
| Startup fails immediately with a configuration error naming two env vars | (none) | Two auth methods configured at once; keep exactly one of token, AppRole, Kubernetes, token file. |
When diagnosing, confirm three clocks/lifetimes in order: the Vault token TTL (`vault token lookup` with the token's accessor), the RustFS refresh cadence (half TTL or the poll interval), and the fail-closed window. The renewal task logs every failed cycle, so a silent gap in warnings combined with `CredentialsUnavailable` errors points at the process clock or a paused runtime rather than Vault.