* fix(bucket-repl): persist MRF retry queue to disk and reload on startup
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(bucket-repl): address three blocking MRF issues from review
1. MRF replay loses delete operations — add `MrfOpKind` discriminator to
`MrfReplicateEntry` (Object | Delete, default=Object for backward
compat). `DeletedObjectReplicationInfo::to_mrf_entry` now persists
`op=Delete`, `version_id`, `delete_marker_version_id`, and
`delete_marker`. `start_mrf_processor` branches on `op`: delete
entries skip `get_object_info` and replay via
`schedule_replication_delete` with `ReplicationType::Heal`; object
entries follow the existing heal path.
2. `flush_mrf_to_disk` cleared the in-memory batch even on encode/write
failure — changed return type to `bool` and callers now only
`pending.clear()` on `true`, so a transient storage error retries
on the next tick instead of silently dropping the batch.
3. Add focused tests: encode/decode roundtrips for object, delete-marker,
versioned-delete, and mixed-batch entries; a routing test confirming
op-kind propagates correctly and that the default is Object for
legacy files; a legacy-compat test verifying old entries round-trip
cleanly through the new format.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* style: fix clippy redundant-clone in MRF tests
Replace &[entry.clone()] with std::slice::from_ref(&entry) in two
encode_mrf_file call sites flagged by clippy's redundant_clone lint
under --all-targets --features rio-v2.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(bucket-repl): strengthen legacy MRF compat test with hand-built msgpack
The previous mrf_legacy_file_without_op_field_decoded_as_object test
round-tripped through encode_mrf_file, so it exercised the new format
and never touched a truly-legacy payload.
Replace it with a hand-built msgpack payload that genuinely omits the
"op", "deleteMarker", and "deleteMarkerVersionID" keys — exactly what
the old binary would have written before MrfOpKind existed. The test
now fails if #[serde(default)] is removed from the op field, which
proves real backward compatibility rather than round-trip stability.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: houseme <housemecn@gmail.com>
fix(scanner): skip startup delay when replication is active so failed objects heal promptly after restart
Co-authored-by: houseme <housemecn@gmail.com>
fix(bucket-repl): honor op_type in replicate_object so ExistingObject resync respects DISABLED targets
replicate_object was calling filter_target_arns with hard-coded
op_type: Object and existing_object: false regardless of what was
stored in roi.op_type. This meant that a resync worker setting
roi.op_type = ExistingObject (resync_bucket, line 889) had no effect
on target filtering: all configured targets were included, even ones
whose rule had ExistingObjectReplicationStatus::DISABLED.
Fix: pass op_type: roi.op_type and derive existing_object from it
(true only for ExistingObject, not Heal — Heal intentionally bypasses
the existing-object opt-out to repair past failures).
Also add warn! logs at all four MRF channel-overflow sites that were
previously silently returning Missed with no observability.
Verified with a live two-instance test: after resync, objects reached
the ENABLED target and were correctly blocked from the DISABLED target.
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: houseme <housemecn@gmail.com>
* fix(site-repl): clamp replication_cfg_mismatch to owning deployments and add resync status branch
Fix 3: replication_cfg_mismatch was set on ALL deployments for a bucket when
any replication-config discrepancy existed, even on deployments that simply have
no config. mc computes "in sync" as max_buckets minus per-deployment mismatch
entries; with N deployments all flagged for 1 bucket, the result underflows to
1-N = -1. Now only deployments that own a replication config are flagged.
Fix 4: SiteReplicationResyncOpHandler accepted "start" and "cancel" but returned
an empty body (and later a parse error on the client) for any other operation
string. Add a SITE_REPL_RESYNC_STATUS="status" arm that returns the stored
SRResyncOpStatus or an explicit {"status":"not-found"} object so mc never sees
an unexpected end of JSON input.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(site-repl): derive sync_state from reachability and replication rule completeness
PeerInfo.sync_state was set to SyncStatus::Unknown at every construction site
and never updated from real signals. build_status_info now tracks which peers
were reachable during the metainfo fetch phase and, after all bucket stats are
merged, derives sync_state as:
- Enable : reachable AND no replication_cfg_mismatch for any bucket
- Disable : reachable BUT at least one bucket has incomplete/missing rules
- Unknown : unreachable (fetch failed)
The "Sync" column in mc admin replicate status will now reflect actual state
rather than always showing Unknown.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(site-repl): add SRRotateServiceAccountHandler for split-brain svc-acct repair
When site-replicator-0 gets desynced (different secret on different peers), calls
via the old service account return 403, blocking remove and replicate operations
with no in-band recovery path.
SiteReplicationRemoveHandler already purges local state unconditionally before
sending peer notifications, so force-remove works locally even when peers reject
the 403. The new POST /v3/site-replication/rotate-svc-acct endpoint provides the
missing recovery path: it generates a fresh service-account secret, applies it
locally, and pushes a peer/join to every member. Each peer's SRPeerJoinHandler
accepts the join idempotently (update if exists, create otherwise), repairing the
desynced credential without a full teardown + restart.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(site-repl): back-fill pre-existing buckets and objects on replicate add
SiteReplicationAddHandler and SRPeerJoinHandler both called persist_site_replication_state
and returned success without doing anything for buckets that existed before the
sites were linked. Only buckets created after the link (via site_replication_make_bucket_hook)
ever replicated.
Introduce backfill_existing_buckets_after_add which, after persisting the new
state, iterates every local bucket and:
1. Ensures versioning is enabled (required by replication).
2. Reconciles bucket targets so a target entry exists for each remote peer.
3. Reconciles the replication config so a rule pointing to each peer is present.
4. Broadcasts a make-bucket-hook to peers (idempotent) so they create the bucket.
5. Kicks start_site_bucket_resync toward every remote peer so pre-existing
objects travel across.
Errors per bucket are logged but never abort the overall add — manual resync
remains available as a fallback. Both the initiator and the receiving (join) side
run the backfill so convergence happens regardless of which side held data first.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(site-repl): reconcile replication config to restore bidirectional replication
Root cause: ensure_site_replication_bucket_replication_config bailed with Ok(())
the moment any replication config was found on the bucket. When bucket B was
propagated to site-2 via the make-with-versioning bucket-op, site-2's
configure-replication step loaded the freshly-written config and immediately
returned, never adding the reverse-direction rule pointing back to site-1. Result:
objects uploaded to site-2 failed with "replication head_object fallback failed
... service error" because no rule targeted the originating site.
Fix: drop the early-return. Instead load the existing rules, build the full
desired config via build_site_replication_config, and MERGE — adding only the
rules that are absent (identified by their "site-repl-<deployment_id>" id). Re-
number priorities after each merge to avoid conflicts. Existing non-site-repl
rules are preserved. The write is skipped entirely when all desired rules are
already present, so repeat calls remain cheap.
Together with Fix 1 (backfill), any file written to ANY member now replicates to
ALL members. The offline-and-recover case also benefits: when a peer returns,
the per-bucket rules are complete and the resyncer can catch up.
Also register the new rotate-svc-acct route in route_policy and
route_registration_test so the route inventory assertions stay green.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(replication): remove broken resync worker-signal gate
The resync_bucket task opened a new broadcast receiver via
worker_rx.resubscribe(), which positions the receiver at the current
write-head of the ring buffer — past all 10 bootstrap signals written
in ReplicationResyncer::new(). Every spawned resync task therefore
blocked on recv() forever, making `mc admin replicate resync start`
report "started" while no objects ever moved.
Remove the dead wait entirely. Each resync_bucket call is already
spawned on-demand (tokio::spawn in start_bucket_resync / load_resync),
so no additional gate is needed. The per-object concurrency limit is
already enforced by the inner mpsc worker channels (line ~877). Also
remove the now-dead worker_tx/worker_rx fields, bootstrap loop, and
signal-send in resync_bucket_mark_status.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(site-repl): document rotate-svc-acct idempotency and partial-failure behavior
* style: cargo fmt and remove redundant clones
* fix(site-repl): preserve object-lock state when back-filling existing buckets
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: loverustfs <hello@rustfs.com>
Co-authored-by: houseme <housemecn@gmail.com>
Adds an operator guide for the SFTP server: recommended
configuration, the path model, host keys on Unix and Windows, the
environment variable reference, session cleanup behaviour, IAM
policy requirements per SFTP operation, client compatibility
notes, multipart upload sizing and cleanup, and the log lines
worth alerting on.
Corrects documentation the platform change left stale. The module
overview and the UnsupportedPlatform error text still described
Windows as unsupported. A source comment referenced a document
that does not exist in the repository and now points at the new
guide. The changelog adds the two host-key reload variables
missing from its environment list, describes the banner variable
as the SSH identification string, and corrects the upload size
cap and compliance case count.
The session watchdog now selects its detection method per
platform. It previously probed kernel TCP state through a
Linux-only procfs path, so on macOS every healthy idle session
was killed within a minute, and on Windows the watchdog never
spawned at all, leaving wedged sessions with no cleanup. Linux
keeps its fast-kill watchdog unchanged. Other platforms get a
silence-only backstop that kills a session only at the
documented 30-minute idle ceiling.
The host-key loader now has a Windows arm. It loads OpenSSH
format host keys from the configured directory and logs a
one-time warning to restrict NTFS ACLs on the key directory,
the same operator-managed approach FTPS, WebDAV, KMS, and IAM
already use on Windows. Startup previously aborted with
UnsupportedPlatform because the Unix mode-bit permission check
has no Windows equivalent. tokio's io-uring feature is now
enabled only in Linux builds. io-uring is a Linux kernel
interface and enabling it unconditionally broke the Windows
build.
Co-authored-by: houseme <housemecn@gmail.com>