Commit Graph

5602 Commits

Author SHA1 Message Date
overtrue 1d305aeb5b docs(sse): record SSE-C as measured-readable in the migration matrix
The SSE-C row moves from unverified to supported now that a MinIO-generated SSE-C object reads end to end, with the test that establishes it. Single-part stays unverified for all three schemes — the fixture harness still cannot load objects MinIO inlined into xl.meta.

Also notes the one behavioural difference an operator will observe: a MinIO SSE-C object stores no customer-key MD5, so a wrong key is refused by the decryption rather than by the earlier parameter-mismatch check. Different error, same outcome.

Refs rustfs/backlog#1638.
2026-08-19 09:25:10 +08:00
overtrue fd26567cd7 fix(sse): read MinIO SSE-C objects, and pin the customer-key check
MinIO SSE-C objects were unreadable for the same reason its managed objects were: the dispatch demanded a header MinIO never persists. A MinIO SSE-C object stores exactly three keys — the internal sealed key, IV, and seal algorithm — and keeps the customer algorithm on the request, returning it on the response. Recognizing the internal sealed-key slot as well is the whole fix on that side; the customer key still has to be supplied by the caller.

The stored customer-key MD5 needed narrower handling. MinIO stores none, so comparing against it made every migrated object fail. The comparison is now skipped only when there is nothing to compare — no stored MD5 *and* the object carries MinIO's SSE-C slot — which does not weaken what the check buys: it is an early, friendlier rejection, while the key itself is proven by the object-key unseal, whose AEAD fails on a wrong key. A negative test holds that line by reading the fixture with a well-formed but wrong customer key.

Writing that test surfaced a gap worth closing on its own: disabling the stored-MD5 comparison for *every* object left all 115 tests in this file green, so nothing guarded it for RustFS-written objects either, and a later widening of the skip would have gone unnoticed. ssec_stored_md5_mismatch_is_refused_when_an_md5_is_stored now fails when that happens.

Verified against fixtures from a real MinIO server: the interop suite is 7/7, including SSE-C multipart, and both mutations — widening the MD5 skip, and the earlier slot-vs-key-id inference — turn it red.

Refs rustfs/backlog#1638.
2026-08-19 09:24:30 +08:00
overtrue 035a6f431a docs(sse): state MinIO interop by measured shape, and pin the production key path
The migration warning said RustFS cannot read MinIO-encrypted objects at all. That is no longer true for the shapes this branch fixes, and a blanket 'no' costs a migrating operator a full decrypt-and-recopy they may not need. It is replaced by a table that states support per object shape, with the test that establishes each row.

The table's honesty depends on one caveat worth stating in code rather than prose: every interop reader test injected its master key through __RUSTFS_SSE_SIMPLE_CMK, which is #[cfg(test)]-only and never reaches LocalSseDekProvider::new_from_env. Those tests going green therefore said nothing about a deployment. A new case reads a fixture through RUSTFS_SSE_S3_MASTER_KEY alone — the sole production entry point — so the claim now rests on the path operators actually run.

Single-part and SSE-C objects are listed as unverified rather than unsupported, because the fixture harness cannot yet load them (small objects live inline in xl.meta, sharded across disks) and they have never been read in a test either way. Conflating 'untested' with 'broken' in the other direction would be the same failure the old warning made.

The reverse direction gets an explicit ruling: MinIO cannot read RustFS-written objects, and the MinIO-branded headers RustFS writes are RustFS-internal rather than a compatibility promise — no coexistence plan should assume two-way reads.

StaticConfig's doc comment is corrected on the same basis: it described MinIO's ciphertext as the legacy JSON encoding, and claimed MinIO-written objects cannot be read at all, which now belongs to the object read path rather than to this backend.

Refs rustfs/backlog#1638.
2026-08-19 08:36:18 +08:00
overtrue 0e4953aeea Merge remote-tracking branch 'upstream/overtrue/kms-1638-d2-minio-sse-read' into overtrue/kms-1638-d2-minio-sse-read 2026-08-19 07:11:23 +08:00
overtrue 99ec65247a fix(sse): gate the MinIO data-key trait method behind rio-v2
The method's only call site sits in the rio-v2 branch of the managed read path, so a build without that feature carried a trait method nothing could reach — a warning under default features and, with -D warnings, a hard failure of the sftp lane. The declaration now carries the same gate its implementation and its sibling decrypt_legacy_sse_dek already had.

Verified against the lane that caught it (cargo clippy -p rustfs --features sftp --all-targets -- -D warnings, clean), plus the default build and the rio-v2 interop suite (4 passed).

Refs rustfs/backlog#1638.
2026-08-19 07:11:10 +08:00
Zhengchao An 6feb573f74 Merge branch 'main' into overtrue/kms-1638-d2-minio-sse-read 2026-08-19 07:09:40 +08:00
Zhengchao An 40c15c769b docs(architecture): describe io-core by its surviving surface (#6227)
The Code Map entry still summarised io-core as "buffer pool, storage profiling, admission control", which predates #6201 removing eight zero-consumer modules. The crate now exposes pool, io_profile, config, backpressure, deadlock_detector, lock_optimizer, and progress, so the summary names the policy, lock, and progress helpers as well.

Refs backlog#1824
2026-08-19 07:09:00 +08:00
houseme ff9ac1013a feat(ecstore): add default-off dst dir fsync group commit (#6226)
* feat(ecstore): add dst dir fsync group commit

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(ecstore): tidy dst dir fsync group open

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-19 06:24:45 +08:00
Zhengchao An 850445e957 fix(e2e): add missing serial_test import in replication_extension_test (#6224) 2026-08-19 06:24:15 +08:00
houseme 50c39fec45 perf(get): emit accept-ranges with static header (#6225)
Avoid constructing the typed Accept-Ranges string on the GetObject output path. Inject the static header after CORS wrapping so the final S3 response remains unchanged while the hot path avoids one fixed per-GET allocation/conversion.

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 17:20:16 +00:00
hector c86a94a2dc fix(package): declare /etc/default/rustfs as a deb conffile (#6220)
Co-authored-by: houseme <housemecn@gmail.com>
2026-08-18 15:56:52 +00:00
houseme d6814af2bc Merge branch 'main' into overtrue/kms-1638-d2-minio-sse-read 2026-08-18 23:54:31 +08:00
houseme 905082893f fix(e2e): import serial test attribute (#6222) 2026-08-18 23:12:39 +08:00
houseme eed0ca3612 perf(ecstore): validate local IO paths with openat2 (#6221)
Use Linux openat2 with RESOLVE_BENEATH and RESOLVE_NO_SYMLINKS for LocalDisk I/O path validation while keeping the existing lstat walk as the public-path and unsupported-kernel fallback. Add focused regression coverage for traversal, symlink swaps, missing leaves, recreated parents, high-cardinality prefixes, final symlink leaves, and concurrent validation.

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 22:40:43 +08:00
GatewayJ 91c97f3416 chore(deps): update s3s to upstream main (#6203)
* deps: update s3s to upstream main

* deps: refresh s3s upstream revision

* deps: pin s3s to latest upstream main

---------

Co-authored-by: houseme <housemecn@gmail.com>
Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 21:46:56 +08:00
唐小鸭 1d056d7605 test(e2e): pin bounded physical reads for compressed multipart range GETs (#6167)
Byte-exactness tests stay green if the compressed range seek regresses into
decoding from byte zero: the returned bytes are still correct and only the
read amplification explodes. Assert the cost side as well.

The observation reuses rustfs_io_get_object_shard_read_observed_bytes_total,
already emitted per shard read by the erasure layer, so no production code is
instrumented. The OTLP collector learns to accumulate a second counter, keyed
by its path/role/outcome labels rather than by data-point position, which is
not stable across exports.

Two failure modes the assertions guard against:

- With RUSTFS_OBS_METER_INTERVAL=1, treating one unchanged sample as settled
  measures a delta of zero, because the range read's counter has not been
  exported yet. Settling now requires several consecutive equal samples.
- An upper bound alone passes vacuously on a zero delta, so a lower bound
  turns "measured nothing" into a failure instead of a green run.

Refs rustfs/rustfs#5957, backlog#1848.
2026-08-18 21:46:16 +08:00
Zhengchao An 8315c23d49 test(kms): move the Vault KV2 doc guard into check_fips_wording.sh (#6215)
test(kms): move the Vault KV2 Transit-wrapping doc guard into check_fips_wording.sh

`test_vault_kv2_sources_do_not_claim_transit_wrapping` asserted that four
`include_str!`-pinned files never describe the Vault KV2 backend as wrapping key
material through Vault's Transit engine. The invariant is a documentation-claim
invariant with no behavioral twin by construction, and the test form was weak in
both directions: it saw only four files (the same prose in a fifth file passed
silently) and it stopped compiling — rather than reporting a violation — as soon
as one of them was renamed.

Move the four literals verbatim into `scripts/check_fips_wording.sh`, which
already guards the adjacent cryptographic over-claim class (unsupported FIPS
validation wording) and is anchored to the same policy document. The guard now
greps every file under `crates/kms` for the same four case-sensitive literals and
separately reports a moved pinned source instead of failing to build.

`check_fips_wording.sh` previously ran only in `make pre-commit` / `pre-pr`, so
wire it into the Quick Checks job of both CI workflows to keep the invariant's
failure visibility at least as strong as the deleted test's.
2026-08-18 21:46:00 +08:00
唐小鸭 1cf0f7af15 feat(replication): split oversized hot-path functions, proxy unreplicated reads, and fail SSE-C passthrough closed (#6170)
* refactor(replication): split four oversized hot-path functions into focused helpers

Pure-move decomposition of the four oversized functions flagged by the
replication compatibility review (P1-18), unblocking migration milestone
M2 which requires resyncer moves to stay mechanical:

- resync_bucket (522 lines -> 61-line step sequence): leader lock,
  target resolution, walk/collector/worker spawning, and dispatch loop
  extracted into focused helpers; pure decision helpers (DTO builders,
  HEAD-result classification) separated from IO orchestration.
- replicate_all (411 lines -> 113-line main body): initial target-info
  seeding, read/stat option builders, skip-path notes, target HEAD
  action resolution, and the multipart/single-put payload transport
  extracted as private free functions.
- start_mrf_processor (306 lines -> 46-line spawn body): recovery guard,
  ledger load, per-entry replay (delete/object/metadata), and retained
  entry resolution extracted; retry bookkeeping semantics preserved
  exactly (inner continue-paths push inside helpers, outer Missed push
  stays in the loop).
- apply_iam_item (255 lines -> match dispatch skeleton): one helper per
  IAM item type.

No behavior change: log texts, error paths, event emissions, and metric
counts are byte-identical; existing tests unchanged and green (238
ecstore replication/mrf/resync + 232 rustfs site-replication).

* feat(replication): proxy GET/HEAD/Tagging for unreplicated objects to replication targets (#6172)

* feat(replication): proxy GET/HEAD/Tagging for unreplicated objects to replication targets

Implements the MinIO active-active read-proxy protocol (P1-5 of the
replication compatibility review): when a GET/HEAD/GetObjectTagging/
PutObjectTagging/DeleteObjectTagging request fails locally with
not-found and the bucket has replication targets, the request is proxied
to the targets in rule order, mirroring bucket-replication.go
proxyGetToReplicationTarget/proxyHeadToRepTarget/proxyTaggingToRepTarget.

Protocol surface:
- Anti-loop: inbound {x-rustfs-,x-minio-}source-proxy-request is parsed
  into ObjectOptions (proxy_request + proxy_header_set, matching MinIO
  ProxyRequest/ProxyHeaderSet); a request carrying the marker with ANY
  value is never re-proxied. Outbound client proxy calls send the marker
  as "true"; replication worker convergence HEADs send it as "false" so
  a peer's proxy layer cannot answer a convergence check by proxying
  back to the source (which would fake Completed without a PUT).
- Target selection: new replication_proxy.rs get_proxy_targets — empty
  when the marker is set, versioning is suspended, or no replication
  config; otherwise filter_target_arns -> TargetClient lookup, skipping
  targets with proxying disabled.
- TargetClient gains head_object_for_proxy/get_object (streaming) and
  the three tagging calls. Proxy calls never send the replication-check
  SSE-C exemption header; customer SSE-C keys are forwarded verbatim so
  the target performs real decryption. Conditional (If-*) headers are
  not forwarded (MinIO parity); Range and part_number are, with
  parts_count/tag_count/storage_class/expiration passed through.
- Metrics: proxy counters now count only real client proxy traffic,
  MinIO-aligned (one total per proxied request, one failed when no
  target served it). The previous misattributed counters — replication
  worker HEAD/PUT (#2672) and local tagging operations (#2682) — are
  removed; ReplProxyMetric now maps the tagging counters instead of
  dropping them.

e2e (fake_s3_target extended with tagging + header journaling): proxied
GET body + outbound header contract (marker present, no
replication-check, SSE-C passthrough), HEAD, anti-loop 404 with zero
outbound requests, GetObjectTagging, and metric mapping unit tests.

Rolling note: proxying only activates for buckets with replication
targets; requests carrying the marker keep pre-upgrade behavior.

Refs rustfs/backlog#1675 (P1-5)

* fix(replication): fail SSE-C passthrough closed on targets that drop transport headers (#6178)

SSE-C ciphertext passthrough replicates via X-Rustfs-Replication-* transport
headers. A MinIO/generic-S3 target silently discards them, storing bare
ciphertext with no decryption material — yet the PUT succeeded, so the object
reported COMPLETED with a silently unreadable replica (backlog#1675 N2).

Fail-closed design:
- SsecPassthroughCapability {Unknown, Supported, Unsupported} cached in
  BucketTargetSys per target ARN with a recording timestamp. Entries reset
  whenever the target is rebuilt, edited, or removed (arn_remotes_map
  lifecycle) and expire after SSEC_PASSTHROUGH_CAPABILITY_TTL (10 minutes):
  an expired verdict in either direction is re-earned through the audit, so
  an Unsupported target recovers automatically after an upgrade (at most one
  wasted PUT+HEAD audit per bad target per TTL window) and a Supported
  verdict cannot outlive a backend swapped behind the same endpoint.
- Replication worker (replicate_object and replicate_all): fresh Unsupported
  targets never receive the PUT — the attempt fails immediately into the
  normal MRF retry channel with a "run ?replication-check to re-probe" hint.
  Unknown or expired verdicts are audited: after the PUT the worker HEADs
  the replica back through the replication-check channel (source version id
  mapped through resolve_read_api_version_id, so null-version objects audit
  correctly) and requires SSE-C evidence (the echoed customer-algorithm
  header); missing evidence records Unsupported and fails the attempt.
  Convergence HEADs are audited the same way, so a broken ciphertext replica
  from an earlier attempt can never launder itself into COMPLETED via an
  ETag match. The gate/evidence policy is pure (replication_target_boundary,
  staleness folded in as an input) for the M2 worker migration.
- replication-check grows an SsecPassthrough probe phase: a probe PUT
  carrying the live transport-header shape, HEAD-back for evidence, and a
  machine-readable Code BucketRemoteSsecPassthroughUnsupported on failure.
  The probe verdict is synced into the runtime capability cache. Unlike
  VersionFidelity, a failed SsecPassthrough phase does NOT fail the target
  overall — it is a capability limit, not a broken replication contract,
  and a plaintext-only deployment against such a target must not turn red.
- fake_s3_target: default mode now models a RustFS target (stores the
  transport headers, echoes SSE-C evidence); the new
  drop_unlisted_replication_headers mode models MinIO. The journal records
  whether a request carried transport headers.

Receiver-echo verification: the replication-check HEAD exemption only skips
SSE-C key validation; the response has always built sse-customer-algorithm
from stored metadata (rustfs/src/app/object_usecase.rs), so no receiver
change was needed — pinned end to end by the replication-check e2e against
a real RustFS target.

Rolling-upgrade constraint: RustFS targets older than the replication-check
HEAD exemption (#5898) answer the audit HEAD without SSE-C evidence (or fail
it outright), so SSE-C replication to such targets reports FAILED. This is
deliberate — FAILED-and-retryable beats a silently undecryptable replica —
and self-heals: once the target is upgraded, the next TTL expiry (or a
manual ?replication-check re-probe) re-audits and records Supported.
Plaintext and managed-SSE replication are unaffected. The capability cache
is per-node; each node audits independently.

Known limitations:
- The audit judges evidence from the echoed customer-algorithm header only.
  A hypothetical target that preserves that one header while dropping other
  transport headers (partial-drop) would pass the audit; no known target
  behaves this way — observed targets drop the whole unknown-header family.
- A mixed-version target cluster can flap the verdict between audits routed
  to different target nodes until the rollout completes; the TTL bounds how
  long each stale verdict persists.

New e2e (backlog#1675 C1 + N2, red-first): fail-closed against a
header-dropping fake (FAILED + no second PUT via the capability cache,
journal-asserted; red run showed the old COMPLETED), replication-check
reports the SsecPassthrough phase Code while the target stays OK overall,
SSE-C heal convergence after a real target outage, and SSE-C
existing-object resync landing a REPLICA readable with the customer key.
TTL expiry in both directions is pinned at the cache and gate seams.

* refactor(replication): move resyncer pure decision logic into rustfs-replication (M2) (#6180)

* refactor(replication): move resyncer pure decision logic into rustfs-replication (M2)

Pure-move milestone M2 of the ECStore replication split (backlog#1675
P1-17): relocate the resyncer's IO-free decision helpers, with their unit
tests, into the crates they already belong to by type ownership. No
behavior change.

Moved into crates/replication:
- resync.rs: resync_status_duration
- delete.rs: resync_existing_delete_replication_info,
  replicate_delete_outcome, target_delete_version_id,
  delete_marker_purge_version_id, delete_marker_purge_mrf_entry
- object.rs: version_identity_drifted, is_replication_target_offline_error,
  SsecPassthroughCapability, SsecPassthroughGate, ssec_passthrough_gate,
  ssec_passthrough_evidence_present (param-demoted to the echoed
  customer-algorithm string; ECStore keeps the HeadObjectOutput adapter)
- filemeta.rs: NULL_VERSION_ID wire literal (crate-owned copy per the
  filemeta-independence contract)

ECStore rewiring (Rule #14: imports stay in *_boundary.rs):
- resync/object-decision/target boundaries re-export the moved symbols;
  resyncer call sites are unchanged
- bucket_target_sys keeps only the verdict cache + TTL and re-exports the
  capability enum so existing consumer paths keep compiling

Not moved (signatures carry ECStore or aws-sdk types):
verify_resync_head_result, resync_target_error_detail, the SdkError
classifiers, the replicate_all_* option/info builders, and the env-coupled
bounded_resync_max_jobs admission clamp. README milestone table updated.

* chore(replication): retire the datatypes.rs relay early

README sanctions retiring datatypes.rs ahead of M4. The module was a
pure relay (resync boundary -> datatypes -> mod.rs facade) with no
external consumer importing it directly, so the facade now re-exports
ResyncStatusType from replication_resync_boundary and the relay file is
deleted. Consumers stay behind the ECStore facade, keeping Migration
Rule #15 intact — the original retirement wording ("consumers import
through rustfs-replication directly") conflicted with that rule and is
corrected in the README.

* chore(arch): extend migration guards to the M2-moved decision contracts

The adversarial review of the M2 move found the per-symbol ratchet in
check_architecture_migration_rules.sh was not extended for the moved
symbols, leaving them free to be redefined in ECStore or imported past
their boundary without CI noticing:

- resync definition pin + boundary fences gain resync_status_duration;
- the object-decision boundary fences gain the five delete-family
  helpers (delete_marker_purge_mrf_entry, delete_marker_purge_version_id,
  replicate_delete_outcome, resync_existing_delete_replication_info,
  target_delete_version_id);
- the target-boundary fence gains the SSE-C gate family, the offline
  classifier, and version_identity_drifted;
- a new definition pin rejects ECStore redefinitions of the M2-moved
  fns/enums (ssec_passthrough_evidence_present deliberately excluded:
  ECStore keeps a thin HeadObjectOutput adapter under that name).

Mutation-verified: a probe fn ssec_passthrough_gate under
crates/ecstore/src/bucket/replication trips the new pin.

Also anchors the intentionally-duplicated NULL_VERSION_ID wire literal
from the filemeta side and tightens the M2 README note on
bounded_resync_max_jobs.
2026-08-18 21:45:38 +08:00
Zhengchao An 27d23b6135 test(io-metrics): assert record helper emissions; fix census heuristic (#6217)
* test(io-metrics): assert lib.rs record helper emissions

The 29 assertion-less record_* smoke tests in io-metrics/src/lib.rs called
their helpers and checked nothing; because METRICS_ENABLED defaults to
false they did not even reach the emission bodies. Replace them with six
DebuggingRecorder tests that enable the gate, pin every metric name the
helpers own, and pin the derived values, branch selection and label
mapping (rustfs/backlog#1836).

* test(tooling): anchor assertless-census delegation tokens to name segments
2026-08-18 21:21:34 +08:00
houseme 127b662f3f feat(app): add opt-in small GET body once path (#6216)
Use the merged s3s single-chunk StreamingBlob support for exact-length materialized GET bodies when RUSTFS_GET_SMALL_BODY_ONCE_ENABLE is enabled.

Keep the default path unchanged and fall back to the guarded MemoryTrackedBytesStream on length mismatch.

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 11:53:06 +00:00
Zhengchao An 8a3c66e655 test(admin): replace a source-text guard with a behavior test (#6212) 2026-08-18 18:28:46 +08:00
Zhengchao An 9852e53b4c test(lifecycle,scanner): drop 47 no-op #[serial] markers (backlog#1846 T1) (#6213)
Second batch of the #[serial] sweep started in #6209. nextest is the
repository's authoritative runner and executes every test in its own
process, so serial_test's in-process mutex cannot serialize tests against
each other -- docs/testing/README.md documents this, and the mechanism
that actually serializes across the process boundary is a
.config/nextest.toml [test-groups] entry with max-threads = 1.

Unlike the first batch (e2e, process-isolated by construction), these are
in-crate unit tests that could genuinely share process state under the
`cargo test` fallback runner, where #[serial] IS still effective. Every
marker was therefore reviewed individually and removed only where the
test provably touches neither the process environment nor a process-global.

Removed (47, pure deletions, no test bodies touched):

  crates/lifecycle/src/core.rs   35
  crates/scanner/src/scanner.rs  12

The lifecycle removals are all validate_* / filter_rules_* /
has_active_rules_* / noncurrent_versions_expiration_limit_* tests that
build a local BucketLifecycleConfiguration and call a &self method
walking only that value. The scanner removals are pure duration
arithmetic (randomized_cycle_delay_for, initial_scanner_delay_for with an
explicit Some(secs), the bitrot-disabled early return of
scanner_clean_idle_max_interval) and background_heal_info_for_scan_complete
/ _for_scan_result field comparisons over locally built values.

Retained deliberately -- see the PR body for the full list and reasons:

  crates/ecstore/src/bucket/lifecycle/bucket_lifecycle_ops.rs  101 (all)
  crates/lifecycle/src/core.rs                                  44
  crates/scanner/src/scanner.rs                                 53

bucket_lifecycle_ops.rs keeps every marker: its test module caches a
process-wide `static STALE_MULTIPART_TEST_ENV: OnceLock<(Vec<PathBuf>,
Arc<ECStore>)>`, and its own reregister_env_local_disks helper documents
in-tree that sibling #[serial] tests reset and reshape the shared
local-disk registry for each other. That sharing is real, so the markers
stay.

No test was renamed, added, or deleted; no reserved migration-gate name
substring is affected; no .config/nextest.toml entry references any of
the 47 removed tests.
2026-08-18 18:28:13 +08:00
Zhengchao An a38743caf5 chore(zip): trim the unused extract and create surface (#6210)
rustfs-zip has one workspace consumer, and it uses only
CompressionFormat::{from_extension, extension, get_decoder} and
ArchiveLimits. Remove the tar/zip extract, zip create, and in-memory
compress helpers together with the types and dependencies that only
served them. Trimming public API is semver-major once the stable tag is
cut, so it costs least now.
2026-08-18 18:27:01 +08:00
houseme 382ae9529e feat(ecstore): instrument rename sync tail metrics (#6205)
Add default-off PUT stage attribution for the rename_data sync tail so strict durability probes can split queue wait, fdatasync, directory fsync, rename, per-disk wait, and quorum wait without changing commit ordering or S3-visible behavior.

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 17:20:55 +08:00
Zhengchao An 68547ed7ea test(e2e): drop 187 no-op serial markers from the three densest e2e suites (#6209)
test(e2e): drop no-op serial markers from the three densest e2e suites

serial_test's #[serial] is an in-process mutex. cargo-nextest, this repo's
authoritative runner, executes every test in its own process, so the mutex is
never contended and the attribute is a documented no-op -- see the "Serial
execution & nextest profiles" section of docs/testing/README.md and the header
of .config/nextest.toml. Cross-process serialization is provided only by a
[test-groups] entry with max-threads = 1.

Remove 187 such markers (plus 3 now-unused imports) from the three
marker-densest modules of crates/e2e_test:

  multipart_auth_test.rs             85
  replication_extension_test.rs      68
  object_lock/object_lock_test.rs    34

None of these modules is covered by any [test-groups] entry, so the markers
were carrying no isolation for any lane. Every test in all three files builds
its own server via RustFSTestEnvironment::new(), which gives a UUID temp dir
and a uniquely allocated port -- the .config/nextest.toml comment on
replication_extension_test already states this explicitly ("parallel-safe by
construction"). No test mutates process env, binds a fixed port, or touches
process-global state, so nothing here needed temp_env or a test-group instead.

Pure deletion: 190 lines removed, 0 added, no test renamed, no behaviour
changed. serial_test stays in Cargo.toml -- 338 markers across 90 other files
in the crate still use it.

Refs: backlog#1846 (T1).
2026-08-18 17:05:30 +08:00
Zhengchao An f06a9c9cba fix(deps): record heal's bytes and crc-fast in the lockfile (#6211) 2026-08-18 17:04:55 +08:00
Zhengchao An bd296eff9e chore(io-core): drop eight zero-consumer modules (#6201) 2026-08-18 08:14:05 +00:00
houseme a5800033bd feat(heal): incremental status cursors and typed overlap policy (HS-06) (#6206)
* feat(heal): incremental heal status cursors and typed overlap policy (HS-06)

Incremental results: every retained result item now carries a monotonic
sequence number. The status query accepts a client cursor (sinceSeq on
the admin wire, Option<u64> internally) and returns only newer items,
plus nextSeq (the next cursor) and minSeq (the oldest retained
sequence). A cursor that fell behind the 1024-item retention window is
flagged through the existing truncated signal together with minSeq so
the client can restart from it. Sequencing survives task completion:
the completion archive stores the seq-stamped window. None keeps the
exact legacy full-snapshot behavior, so existing clients see no change.

Typed overlap handling for admin starts: RUSTFS_HEAL_OVERLAP_POLICY
(merge default | minio_error). Under minio_error, an admin start whose
path overlaps an active or queued task rejects with typed
already-running / overlapping-paths admission reasons (surfaced through
reason_label in the admin error body, sharing the existing
OperationAborted site because the s3s footprint ratchet forbids new
s3_error! sites); an exact duplicate start rejects with
already-running instead of silently merging. Scanner/autoheal/
read-repair sources never take the rejection path.

forceStart semantics now match MinIO for admin requests: an admin
forceStart first cancels the overlapping active admin task, then
admits the replacement.

Wire: the heal-control Query command grows an optional sinceSeq
(defaulted and skipped when absent, so older peers stay compatible);
the admin handler accepts the sinceSeq query parameter; the local
channel query gains the same cursor.

Tests: seq monotonicity and incremental slicing, window slide moving
minSeq with lagging-cursor flags, overlap matrix (same/containing/
contained/disjoint x policy x source), forceStart cancel-then-admit,
and the completion-archive window handoff.

Co-Authored-By: heihutu <heihutu@gmail.com>

* style: fmt after main merge

---------

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-08-18 16:09:30 +08:00
cxymds a4ea36b298 fix(ecstore): single-flight remote disk recovery (#6096)
* fix(ecstore): single-flight remote disk recovery

* test(ecstore): cover remote recovery review cases

* test(ecstore): exercise recovery through disk slot

* test(ecstore): match format reads exactly

* test(ecstore): cover recovery teardown races

* style(ecstore): format recovery race tests

* fix(ecstore): remove unused health snapshot helper

* fix(ecstore): group recovery monitor test state

* fix(protos): preserve production source in compatibility checks
2026-08-18 14:49:53 +08:00
cxymds 4f68f117ba fix(lock): reap expired local lease guards (#6094)
* fix(lock): reap expired local lease guards

* fix(lock): reject refresh after guard expiry

* test(lock): isolate expired refresh regression

* style(lock): format expired refresh assertion
2026-08-18 14:49:39 +08:00
cxymds 0f30a75fdb fix(ecstore): resume remote shard reads once (#6091)
* fix(ecstore): preserve CopyObject producer errors

* fix(ecstore): resume remote shard reads once

* fix(app): resume preserved relocation I/O errors

* fix(ecstore): reserve remote read recovery budget
2026-08-18 14:49:19 +08:00
Zhengchao An 355c8d2e22 fix(admin): classify missing kms config by error variant (#6196) 2026-08-18 05:40:10 +00:00
Zhengchao An 60eb139db9 refactor: import x-amz-checksum header names from the shared constants (#6193)
Co-authored-by: houseme <housemecn@gmail.com>
2026-08-18 05:27:38 +00:00
hector 9ef059c908 ci(package): auto-trigger DEB/RPM packaging on releases and upload to GitHub release assets (#6202) 2026-08-18 05:18:13 +00:00
Zhengchao An 51497cb533 fix(ecstore): give peer REST failures op and bucket context (#6200) 2026-08-18 04:57:44 +00:00
Zhengchao An c7a29ec0a7 chore(storage): drop dead backpressure and lock-optimizer wrappers (#6195)
Neither rustfs/src/storage/backpressure.rs nor rustfs/src/storage/lock_optimizer.rs had a production caller: their only non-self references were the pub mod lines in storage/mod.rs and a cfg(test) module, so the object transfer path never applied this backpressure and never took these lock shortcuts.

The six removed tests in concurrent_fix_test.rs duplicated tests that lived inside the deleted files; the shared primitives they shadowed keep their own coverage in rustfs-io-core.
2026-08-18 12:51:44 +08:00
houseme dbef072bfe Merge branch 'main' into overtrue/kms-1638-d2-minio-sse-read 2026-08-18 12:46:12 +08:00
Zhengchao An deb0edb7cc chore: adjudicate 26 bare dead_code allows across five crates (#6187)
Remove every bare `#[allow(dead_code)]` in io-core, object-capacity, targets, rio, and scanner. Each allow was stripped first and clippy was then asked which ones the compiler actually missed, so the verdicts rest on the diagnostic rather than on inspection.

23 were inert: they sat on `pub fn`s inside `pub mod`s, where `dead_code` does not apply, or on scanner integration-test helpers that the tests in the same file do call.

The remaining 3 are in rio's private `compress_index` module and the code behind them is deleted rather than annotated. `remove_index_headers` is dead and also wrong — after skipping the 4-byte chunk header it matches against `S2_INDEX_TRAILER` where `S2_INDEX_HEADER` sits, so it returns `None` for every well-formed index; rio-v2 carries the correct equivalent that is actually in use. `restore_index_headers` is its unreachable counterpart, likewise duplicated live in rio-v2. `Index::reset` is a private method with no caller.

Refs backlog#1823
2026-08-18 12:45:42 +08:00
houseme a08de9229b feat(heal): wire MRF intents with durable repair journal (HS-01) (#6189)
* feat(common): add MRF intent channel and Mrf request source (HS-01)

Introduce the producer-facing half of the mission repair feed: a global
bounded (8192) channel carrying lightweight MrfIntent values from IO
error paths, plus the RUSTFS_HEAL_MRF_ENABLE delivery kill-switch and
config constants for queue/journal sizing. Delivery is strictly
non-blocking (try_send, drop-on-full) so it can sit on decode-failure
and partial-write paths without adding latency. HealRequestSource grows
a 'mrf' variant so admission accounting can attribute replayed intents.

Part of backlog#1865 (option a: wire HealEvent-style intents with a
durable retry ledger).

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(heal): add MRF queue, durable journal, and intent consumer (HS-01)

Consumer half of the mission repair feed: a bounded pending queue
(100k intents / 8 MiB dual ceiling, drop-newest on overflow), a durable
journal at buckets/.heal/mrf/journal.bin holding the unaccepted pending
snapshot, and a consumer task that batches intents off the global
channel, translates them into prioritized heal requests (decode
failure -> Urgent ECDecode, metadata corruption -> High Metadata,
partial write -> Normal object heal), and retries full admissions with
a 5s backoff and a 3-attempt ceiling.

Durability: every journal record carries its own CRC32 and a
format/version header, so a torn tail truncates cleanly at replay; the
journal is deleted after a successful replay and when the pending set
drains (mirroring MinIO's post-replay list.bin unlink). Losing the last
500 ms flush window is acceptable: replayed duplicates merge via the
manager dedup key and read-repair remains the safety net.

Metrics: rustfs_heal_mrf_queue_depth/_queue_bytes, _dropped_total
{reason}, _replayed_total, _journal_bytes, _journal_fsync_total.
The consumer is wired at heal runtime bootstrap right after manager
start, honoring RUSTFS_HEAL_MRF_ENABLE (default on, rollback = off).

Tests: unit tests for the dual ceiling, record roundtrip, torn-tail
truncation, and the priority mapping; integration tests against a real
4-disk ECStore proving channel intents reach the manager queue as
Urgent/mrf-attributed requests and journal replay arms intents, drops
torn tails, and removes the file.

Part of backlog#1865 (option a).

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(ecstore,scanner): deliver MRF intents from error paths (HS-01)

Wire the three production delivery points, each a single non-blocking
try_send next to the existing in-memory heal paths, which stay as the
fast path:

- read.rs decode-error branch: DecodeFailure intent beside the existing
  read-repair submit, so an Urgent ECDecode request survives restarts
  even when the Low-priority read-repair request was dropped or lost.
- add_partial: PartialWrite intent, giving partial-write recovery a
  durable Normal-priority object heal across restarts.
- scanner_folder metadata-corruption classification: MetadataCorruption
  intent beside the existing High-priority scanner heal request.

All three are on error paths only: zero cost on healthy IO.

Part of backlog#1865 (option a).

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix: include mrf heal source counts

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix: keep node heal status wire compatibility

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 12:43:27 +08:00
Zhengchao An abffa5cf1b chore(storage): drop dead io-schedule metrics and helpers (#6199) 2026-08-18 04:35:36 +00:00
Zhengchao An b825c54850 refactor(admin): route kms management auth through shared gate (#6194) 2026-08-18 12:21:49 +08:00
houseme de9145e87a feat(storage): add default-off PUT admission gate (#6197)
Add an experimental fixed-count foreground PutObject admission gate for #1882 Phase 0 validation. The gate is default-off, returns SlowDown before body ingest when saturated, and keeps the admission permit with the spawned store commit owner until store PUT returns.

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 04:15:19 +00:00
houseme 84bd76a3ce chore(deps): refresh cargo dependencies (#6198)
Update workspace Cargo dependency requirements and lockfile after cargo update/upgrade, including rumqttc-next 0.34.0 and MQTT API compatibility adjustments.

Verification:

- cargo update --verbose

- cargo upgrade --verbose

- cargo update -p rumqttc-next --precise 0.34.0 --verbose

- cargo tree --invert rumqttc-next --locked

- cargo metadata --locked --no-deps --format-version 1

- cargo fmt --all --check

- cargo check -p rustfs-targets --all-targets --locked

- cargo test -p rustfs-targets mqtt --locked

- make pre-pr

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 04:11:53 +00:00
houseme 9a2d06b370 test(heal): lock heal vs delete/overwrite race invariants (HS-12) (#6183)
* test(heal): add concurrency invariants for heal vs delete/overwrite races (HS-12)

Audit conclusion for backlog#1874: RustFS does not need a persistent
object-level healing marker (MinIO x-minio-healing) because every path
that can touch the same (bucket, object) commit surface serializes on
the same namespace write lock, and the heal lock guard spans the whole
rename commit including the HEAL_RENAME_INCOMPLETE partial path.

Lock the conclusion in with two race regression tests:

- heal_racing_version_delete_never_resurrects_the_deleted_version:
  shard damage is injected on the doomed version so a Deep heal has real
  reconstruction work while a versioned DELETE runs concurrently; the
  deleted version must stay deleted and the survivor intact.
- heal_racing_unversioned_overwrites_preserves_the_last_commit:
  unversioned overwrites (activating the post-commit tail that deletes
  the replaced data dir without the ns lock) race a Deep heal in a
  loop; the final current version must be exactly the last commit.

Also adds docs/operations/heal-concurrency-safety-notes-zh.md with the
full intersection matrix (17 intersections), lock-coverage argument,
and the residual-window classification (commit tail races are
fail-into-retry safe; bare prefix delete has zero production callers;
admin no_lock is an explicit operator opt-in).

Co-Authored-By: heihutu <heihutu@gmail.com>

* test: remove redundant heal etag clone

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 02:01:04 +00:00
Zhengchao An 00de43528c fix(ecstore): describe peer bucket RPC failures with no details (#6190)
heal_bucket, list_bucket, get_bucket_info and delete_bucket returned Error::other("") when a peer answered success=false without an error payload, so operators saw a bare "io error " after quorum reduction. Route all five bucket RPCs through peer_failure_without_details, which names the operation and bucket while staying identical across the peers of one operation so reduce_errs keeps grouping them into a single dominant error.
2026-08-18 01:55:29 +00:00
overtrue 420bfa859b fix(sse): read objects that MinIO encrypted
RustFS could not read a single MinIO-encrypted object. Two independent blockers, and backlog#1638 could only argue them statically because the fixtures the interop tests consume are generated, not checked in — so those tests had never once run. With the fixture lab working, both are now measured, fixed and covered.

The detection gate required `x-amz-server-side-encryption` to be present. MinIO never persists it: `crypto.S3.CreateMetadata` writes only the `X-Minio-Internal-*` family and the public header is synthesized onto the response by `DecryptObjectInfo`. Every MinIO object therefore fell out of the managed path and failed with "encrypted object metadata is incomplete". The scheme is now inferred from which sealed-key slot is present, which is self-consistent by construction: the slot decides both which header the unseal reads and which domain string the sealing key is derived under, so an inference that disagreed with the slot could not silently derive a wrong key. Inferring from the KMS key id would NOT be safe — MinIO writes `-S3-Kms-Key-Id` on SSE-S3 objects too, which the fixtures show and a mutation test pins.

Past the gate, the data key itself could not be unwrapped. Its wire format is `sealed_bytes || iv[16] || nonce[12]` — the randomness trails the ciphertext rather than leading it — with a per-ciphertext sealing key of `HMAC-SHA256(master, iv)` and the encryption context bound as associated data (`internal/kms/secret-key.go`). Note this is not the `{"aead":...}` JSON that backlog#1638's analysis described: current MinIO writes the raw layout and treats JSON only as a legacy encoding, normalizing it into the same byte order. Both are decoded here, in a decoder of their own — `LocalSseDekEnvelope`'s `deny_unknown_fields` is untouched, since loosening it to admit MinIO's shape would also admit malformed RustFS envelopes that backlog#1567 requires to keep failing closed.

Routing between the two decoders cannot key on metadata: RustFS's own writer fills MinIO's slots while storing a RustFS envelope in them, so neither the slot nor the header name distinguishes writers. It keys on the data key's own shape instead, recognizing the two strict RustFS JSON shapes positively and leaving only the remainder to MinIO — so neither decoder is ever handed the other's format. Three round-trip tests caught an earlier slot-based attempt doing exactly that.

Fail-closed is preserved throughout: a scheme that cannot be established still returns None, and the read plan independently classifies the object as encrypted from its markers and refuses to serve it without material, so no path degrades into returning ciphertext as plaintext.

The interop harness also gets a provider reset. The DEK provider is cached process-wide, so a case that ran earlier kept serving its master key to every later case — which silently made the wrong-key negative test unable to fail. It fails correctly now, and the whole suite is meaningful for the first time.

Refs rustfs/backlog#1638.
2026-08-18 09:32:29 +08:00
houseme 35a30cd614 feat(scanner): emit excess alerts as S3 notification events (HS-04) (#6176)
* feat(scanner): emit excess alerts as S3 notification events

The excess-versions / excess-version-size / excess-folders alerts were
metrics-and-logs only; consoles and external auditors had no way to hear
them (rustfs/backlog#1868, HS-04). MinIO emits s3:ObjectManyVersions /
s3:ObjectLargeVersions / s3:PrefixManyFolders for the same conditions —
RustFS carries those as EventName::Scanner* with s3:Scanner:* wire names
that already existed unpublished.

The three alert sites now also dispatch through the standard event
pipeline (send_event via the storage_api owner facade), carrying the
actual values and thresholds in req_params and UserAgent "Scanner".
Without a cooldown a single over-threshold object would re-emit on every
~60s scan cycle, so emissions are edge-held per (kind, bucket, object)
for 24h (RUSTFS_SCANNER_ALERT_COOLDOWN_SECS, 0 = every cycle), backed by
a process-global map with a 4096-key hard cap that clears rather than
grows. Metrics and structured logs stay level-triggered every cycle;
only the notification events are held back. A restart resets the
cooldown deliberately: one re-emission per still-hot key buys back
visibility after the restarts that accompany incident response.

Tests pin the edge-hold semantics (first fires, immediate re-check held,
independent keys, cooldown expiry re-fires, zero cooldown always emits,
hard bound) in one sequential test for the process-global map, and pin
the emitted wire names against EventName's canonical string forms so a
subscribed bucket notification can never silently stop matching.
docs/operations/scanner-excess-alerts.md documents the three events,
the metric-vs-event cadence difference, and the HS-15 threshold deltas
(alert_excess_folders 65538 vs MinIO 50000 is deliberate: Proxmox
Backup Server chunk layout compatibility).

Closes rustfs/backlog#1868.

Co-Authored-By: heihutu <heihutu@gmail.com>

* docs(operations): split scanner excess alerts into English and Chinese pages

The page shipped Chinese-only; keep it as scanner-excess-alerts_zh.md and
add a faithful English translation at the original path, cross-linked at
the top of both.

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 08:46:32 +08:00
houseme 360bceafce feat(heal): add progress and trace observability (#6179)
* feat(heal): track erasure set progress baseline

Record erasure-set heal byte progress from per-object results and seed progress totals from complete usage-cache snapshots when available.

Keep usage-cache failures observational so heal execution continues without a baseline.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(heal): skip filtered erasure set versions

Skip erasure-set versions written after the durable heal start time, and queue lifecycle-expired versions for expiry before skipping them.

Track new-version and ILM-expired skips separately so progress can explain completed baseline work without treating these skips as retry-blocking failures.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(heal): wire abandoned data-dir cleanup check

Connect check_abandoned_parts through ECStore, pool, and set layers so heal can invoke the existing orphan data-dir reclaim path instead of returning NotImplemented.

Add dry-run support to the reclaim scan and cover dry-run plus scoped set behavior with regression tests.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(obs): add heal scanner trace bus

Introduce an in-process broadcast trace bus with typed heal and scanner events, lazy event construction, and bounded lagged-subscriber behavior.

Cover zero-subscriber publishing, subscription delivery, drop accounting, and lagged receivers with focused common-crate tests.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(obs): stream heal trace events from admin API

Wire the admin trace endpoint to the common trace bus for heal/scanner events, including kind, regex, and threshold filtering.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(obs): emit heal trace events

Publish heal task lifecycle and abandoned-parts cleanup events through the common trace bus so the admin trace stream has live heal diagnostics.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(obs): emit scanner trace events

Publish scanner folder, lifecycle action, and heal-candidate events through the common trace bus for live admin scanner diagnostics.

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(heal): route data usage loader through storage api

Keep ECStore data-usage facade access behind the heal storage_api boundary so architecture migration guards can validate the heal progress path.

Co-Authored-By: heihutu <heihutu@gmail.com>

* perf(heal): avoid lifecycle snapshots on ordinary heal pages

Only request lifecycle object snapshots when the heal pass has lifecycle expiry context. This keeps ordinary listing and disk-walk pages from cloning FileInfo/ObjectInfo payloads while preserving the skip path that queues expired versions.

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(heal): update bug-fix mocks for lifecycle snapshots

Carry the lifecycle snapshot opt-in argument through the remaining heal bug-fix test mocks so all-targets clippy covers the updated storage trait.

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(rustfs): sync heal storage mock signature

Update the rustfs storage RPC test mock for the lifecycle snapshot opt-in argument and cover it with rustfs all-targets clippy.

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(e2e): allocate smoke ports across nextest processes

Serialize E2E port selection with a small /tmp allocator so nextest workers do not reuse the same just-released ephemeral port before RustFS binds it.

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 08:29:29 +08:00
Zhengchao An 7cb91a0190 chore(ecstore): adjudicate 32 bare dead_code allows (#6173)
Replace every bare `#[allow(dead_code)]` in ecstore with either a deletion or a per-item allow carrying a `reason`. Blanket allows at module, struct, and impl level silence the lint for future members too, so each is narrowed to the members that are actually dead.

Delete the dead cluster in `config/heal.rs` (`Config`, its three methods, `RUSTFS_BITROT_CYCLE_IN_MONTHS`, `parse_bitrot_config`) rather than annotate it: it has no callers and is unreachable outside the crate, and `parse_bitrot_config` would panic on its disabled path via `Duration::from_secs_f64(-1.0)`. `DEFAULT_KVS` stays, since the config registry uses it.

Correct two `reason` strings on `Checksum::new` and `PutObjReader::md5_current_hex_string`, which are methods but carried a field-only rationale.

Refs backlog#1823

Co-authored-by: houseme <housemecn@gmail.com>
2026-08-17 23:51:56 +00:00
hector beb6e1383e feat(helm): add TLSRoute passthrough support for gateway api (#6169)
Add an optional TLS passthrough listener to the Gateway API support. When gatewayApi.listeners.tls.enabled is true, the Gateway gets a TLS listener with tls.mode: Passthrough and a TLSRoute is rendered to the RustFS service so TLS terminates at the backend (end-to-end encryption).

Refs rustfs/rustfs#3862.
2026-08-18 01:21:15 +08:00