* test(replication): pin delayed delete-marker purge failure handling (red)
P1-21 (rustfs/backlog#1675 B2): two failing e2e tests that pin the missing
failure handling of the delayed delete-marker purge:
- test_delayed_delete_marker_purge_retries_after_transient_target_failure:
four scripted 503s outlast every existing channel (version-purge
replication + its in-process MRF fast retries + the watcher's single
attempt = 3 target DELETEs, all faulted in the recorded run); the
replicated marker is stranded on the target forever.
- test_delayed_delete_marker_purge_exhaustion_persists_to_mrf_and_replays_on_restart:
exhausted purge intents never reach the durable MRF journal, so a restart
replays nothing (recorded run: 3 faulted attempts, zero post-restart).
Red-light evidence (current main):
- Test A: FAILED, journal shows 3x DeleteObject fault=Status(503), no clean
attempt, target marker still present after 15s.
- Test B: FAILED after 468s, same 3 faulted attempts, no purge DELETE after
restart, marker still present.
Test infra: FakeS3Target::stored_versions() exposes per-key version state so
purge tests assert target state instead of inferring it from the journal;
nextest count comments 36->38 nightly / 56->58 total.
* fix(replication): retry, persist and replay failed delete-marker purges
P1-21 (rustfs/backlog#1675 B2). The delayed delete-marker purge was
fire-and-forget: the target DELETE discarded its result (`let _ =`), a
missing target client was silently skipped, and nothing recorded the intent
— one transient target error stranded the replicated marker on the target
forever. Separately, `replicate_delete_with_outcome` held its outcome
hostage to `!requires_delayed_purge`, pinning every delete-marker MRF entry
to Missed so the durable backlog retained them permanently.
Changes:
- `replicate_delete_marker_purge_to_targets` now reports per-target
results (warn + metrics on failure, including `target_client_missing`),
supports retrying only the failed targets, and treats a target-side
NoSuchKey/NoSuchVersion as purge success (strict-404 targets must not
retain the intent forever).
- The delayed watcher (`watch_and_purge_source_delete_marker`) retries
failed targets across its 5x1s watch window; on exhaustion it persists
the purge intent to the durable MRF journal via the new
`ReplicationPoolTrait::persist_mrf_entry` (journal-only on purpose: live
re-dispatch would loop unboundedly against a down target). Intent entries
are shaped as marker-creation deletes so replay funnels into the stale-
marker branch.
- The stale-marker branch (source marker already gone) now purges the
targets instead of silently returning success — closing a latent leak —
and reports the purge result as the replay outcome. Heal callers retry
for the full window (the startup MRF processor runs before target
clients initialize); live callers attempt once and fall back to a fresh
durable intent, so a down target cannot pin a replication worker.
- The outcome formula (extracted as `replicate_delete_outcome` and pinned
by a unit test) no longer includes the delayed purge, so successfully
replayed delete-marker entries are acknowledged instead of retained
forever.
Verification: red -> green e2e pair (transient-failure retry; exhaustion ->
durable MRF -> restart replay -> second-restart zero-replay ack) plus unit
tests; `make pre-commit`, logging guardrails, clippy (ecstore + e2e_test)
all clean; full ecstore lib suite 3729 passed (3 pre-existing local-DNS
kubernetes endpoint failures reproduce without this change).
Adversarial validation (7 roles): no blocking findings after adding the
outcome-formula guard test. Known residuals recorded in the PR: watcher
shutdown window (intent not yet persisted), rolling-downgrade replay acks
without purging (equals pre-fix behavior), and replay falling back to the
source version id on targets that mint their own version ids (P1-19).
* chore(test): refresh the nextest replication count invariant
The e2e-smoke/e2e-repl-nightly split comment is descriptive metadata
(authority: `cargo nextest list`); refresh it to this branch's
post-rebase total.
* fix(replication): purge the marker version the target actually assigned
Review follow-up (#5864), two real defects:
- The delayed purge watcher was spawned with the pre-merge `dobj`, so the
per-target marker version ids this round recorded were invisible to it.
Against a target that mints its own ids the purge fell back to a
source-derived id, the target answered the versioned DELETE with an
idempotent 204, and that "success" cleared the retry set while the real
marker stayed behind. The watcher now receives the merged replication
state (`drs`), which folds this round's target-assigned ids in.
- A target whose recorded version metadata is inconsistent was skipped
without entering `failed_arns`, so an empty result made both the watcher
and the MRF replay treat a purge that issued no DELETE as successful and
drop the intent. The refusal is now a per-target failure (own metric
label): the leak stays visible and the intent is retained instead of
being acknowledged. The version decision also moved ahead of the client
lookup, so the refusal is decided from metadata alone.
Tests: a new e2e drives a fake target with `assign_own_version_ids`, which
ignores the forwarded source-version header for both objects and delete
markers, and asserts the replicated marker is really gone; a unit test
pins the corrupt-metadata refusal as a failed outcome without any target
client registered. The detached-watcher shutdown window is documented at
the watcher as a known non-durable window with the write-ahead follow-up
spelled out.
Keep GET data-read metadata early-stop and bounded fanout behind explicit environment switches so the default path preserves full fanout read-failure tolerance.
Retain the focused opt-in A/B coverage and the invalid parity full-fanout guard for heterogeneous set layouts.
Co-authored-by: heihutu <heihutu@gmail.com>
The RPM build step fails because fpm's --config-files flag requires
/etc/default/rustfs to exist in the staging area, but unlike the DEB
build (which creates it in its package directory structure), the fpm
command has no prior step creating this file.
Create the config file in a temporary directory and pass it to fpm
via a source=dest mapping, matching the DEB build's behavior.
Track accepted replay cache records by gRPC operation and split Lock/Unlock and ReadVersion methods out of grpc_other so hotpath validation can attribute nonce pressure without changing replay protection semantics.
Co-authored-by: heihutu <heihutu@gmail.com>
Add a replacement recovery peer RPC so Admin v4 can distinguish definitive cluster proofs from unsupported, unavailable, or conflicting peer state without extending the existing background heal v3/v1 status protocol.
Co-authored-by: heihutu <heihutu@gmail.com>
The DEB version substitution used ${VERSION/-/~} which caused bash
to expand ~ to $HOME (e.g. /home/runner), producing an invalid
version string like '1.0.0/home/runnerrc.1'.
Store ~ in a variable first to prevent tilde expansion.
Add a v4 admin status endpoint for local durable automatic replacement recovery records without changing the v3 background heal status or peer v1 payloads.
Co-authored-by: heihutu <heihutu@gmail.com>
Three Kubernetes endpoint-identity tests read the real kernel hostname
and panicked when it is an IP literal (e.g. macOS without a static
HostName, where DHCP/reverse-DNS sets the kernel hostname to an address
like 192.168.1.11).
Add a cfg(test) override seam (force_kernel_hostname_for_test, mirroring
the existing force_local_host_resolution_timeout_for_test pattern) and
route the production read through kernel_hostname_for_endpoint_identity()
so the tests inject deterministic hostnames instead of depending on the
host environment. Production behavior is unchanged.
Complete the encrypted-object replication series (backlog#1783, PR-C of
3, after #5872 and #5885): SSE-C objects replicate as ciphertext
passthrough — the source holds no customer key, so the stored bytes and
their encryption metadata travel verbatim and the replica decrypts only
with the original customer key, single-part and multipart.
- Sender: SSE-C objects read raw (raw_data_movement_read), transfer at
ciphertext size, and range multipart parts over stored part sizes.
- Receiver: authorized replication PUTs restore the stored SSE-C keys
from the transport headers (exact lowercase forms - the read-path
check is case-sensitive), set ObjectOptions.preserve_ciphertext, and
skip compression, bucket-default SSE, and sse_encryption behind one
restore-derived gate. Multipart uses an internal session marker to
store parts verbatim and strips it on complete.
- Convergence: the replication HEAD sends
x-rustfs-source-replication-check; the target authorizes it as
ReplicateObjectAction and skips SSE-C read validation for that
request only, so keyless convergence HEADs see etag/size/mtime
instead of 400 and SSE-C replicas stop re-driving forever.
- e2e: SSE-C contract flips to a key-gated readable replica (no-key and
wrong-key GETs fail - the direct silent-plaintext detector); new
multipart passthrough contract with ETag/marker/stability assertions.
* fix(ecstore): anchor Windows rename publication
* fix(ecstore): complete Windows rename confinement
* test(ecstore): retain Windows retry assertion path
* fix(ecstore): accept configured Windows root paths
* fix(ecstore): size Windows rename buffers correctly
* fix(ecstore): use native relative rename on Windows
* fix(ecstore): preserve Windows rename parent guards
* fix(ecstore): reuse guarded Windows rename trees
* fix(ecstore): compile Windows publication helpers
* fix(ecstore): preserve configured Windows disk roots
* fix(ecstore): flush Windows shards with write access
* fix(ecstore): stage Windows rollback backup replacement
* fix(ecstore): defer Windows staged file cleanup
* fix(ecstore): type Windows staged write result
* fix(ecstore): retry Windows sharing violations
* fix(ecstore): share Windows staged deletes
* fix(ecstore): split Windows staged publication handles
* fix(ecstore): close Windows staged writer before rename
* fix(ecstore): share Windows staged publication deletes
* fix(ecstore): allow guarded Windows child publication
---------
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
Co-authored-by: cxymds <cxymds@gmail.com>
Co-authored-by: houseme <housemecn@gmail.com>
Install tzdata in both published runtime variants and verify IANA timezone resolution during image builds.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Open the managed-SSE replication gate (backlog#1783, PR-B of 3, after
#5872): the replication reader already decrypts through the injected
object-encryption resolver, so the source sends plaintext plus an
encryption intent header (AES256 / aws:kms, never the source key id) and
the target re-encrypts on its normal PUT path with its own KMS. No DEK
crosses sites.
- replication_put_object_options: fail closed only on Unsupported;
insert the SSE intent after the strip loop.
- TargetClient::create_multipart_upload sends the full opts.header()
set, fixing multipart replicas losing content-type/user metadata
(plaintext included).
- Preserve source ETag and mtime on replicas (authorized replication
only): receiver wires x-rustfs-source-etag into preserve_etag for PUT
and CompleteMultipartUpload, resolve_complete_etag consumes it, and
complete options carry source_etag/source_mtime (absent mtime
degrades to epoch, not now_utc). Without this every replication HEAD
comparison re-drives re-encrypted objects forever.
- e2e: managed SSE contracts flip to success on an independent-KMS
dual-process pair (byte-identical plain GET proves target-owned
envelopes; ETag/mtime preserved; version stable across scanner
cycles; resync converges; multipart keeps structure and metadata);
new target-without-KMS fail-closed contract; SSE-C stays FAILED.
Co-authored-by: houseme <housemecn@gmail.com>
Groundwork for encrypted-object replication (backlog#1783, PR-A of 3):
- classify_replication_source_encryption: accept the AES256 marker that
every stored SSE-C object carries; the SseC arm was unreachable.
- Fail closed on sealed material without an SSE marker (MinIO-written
objects) instead of replicating ciphertext as plaintext.
- Replace the dead VALID_SSE_REPLICATION_HEADERS table with a transport
map keyed by the metadata keys the SSE writer actually persists, shared
via the new rustfs_utils::http::object_encryption_keys module.
- Structurally strip all encryption metadata from outbound replication
(x-rustfs-encryption-* envelopes previously passed the filters).
- Skip decrypt_checksums for encrypted objects at the boundary so its
is_multipart=false (a response-path contract) cannot misroute
encrypted multipart objects once managed replication opens.
- Redact X-Rustfs-Replication-* SSE transport values in FileInfo Debug.
A reconciliation test pins that every key encryption_material_to_metadata
produces is either transport-mapped or stripped. All four SSE replication
e2e contracts still assert FAILED unchanged.
* fix(build): support non-Linux Unix targets (illumos/Solaris/*BSD)
Two independent build-infrastructure blockers kept RustFS from building on
non-Linux Unix platforms. Neither touches runtime logic.
1. pulsar regenerates its protobuf bindings in build.rs on every build, which
needs `protoc`. Platforms without a packaged protoc (illumos/Solaris/*BSD)
now enable pulsar's `protobuf-src` feature via a cfg-gated dependency, which
builds a vendored protoc from C++ sources. Mainstream targets keep the lean
dependency and their existing system/CI protoc.
2. clocksource 0.8.3 (pulled in transitively by ratelimit 0.10) used the
Linux-only `CLOCK_MONOTONIC_COARSE`. ratelimit 2.0 dropped the clocksource
dependency entirely, so upgrading removes the portability problem at the
root rather than patching clocksource. The bandwidth throttle's bulk
`consume()` is rewritten onto ratelimit 2.0's `try_wait_n`, preserving the
best-effort partial-consumption semantics.
Verified: cargo check + bandwidth monitor unit tests pass; cargo tree confirms
protobuf-src is enabled only for illumos/Solaris/*BSD and clocksource is gone
from the graph. The final illumos build must be confirmed on-platform.
Closes#3195
* fix(ecstore): guard ratelimit v2 capacity overflow
Co-Authored-By: heihutu <heihutu@gmail.com>
* test(ecstore): avoid slow bandwidth reader timeout
Co-Authored-By: heihutu <heihutu@gmail.com>
* fix(targets): drop vendored pulsar protobuf build
Co-Authored-By: heihutu <heihutu@gmail.com>
---------
Co-authored-by: houseme <housemecn@gmail.com>
Co-authored-by: heihutu <heihutu@gmail.com>