* test(replication): pin the version-fidelity probe contract (red) P1-19 (rustfs/backlog#1675 B2): the supported replication contract is targets that adopt the source version id — a target that mints its own ids silently breaks every version-addressed operation that follows (version deletes, heal re-drives never match), diverging the two sides with no signal. replication-check already captures the probe PUT's response version id but never compares it. Red evidence (current main): against a FakeS3Target with assign_own_version_ids enabled, ?replication-check returns Status "OK" — the drift is invisible. test_replication_check_flags_version_minting_target expects a VersionFidelity phase that fails with the machine-readable code BucketRemoteTargetVersionMismatch, skips the later mutation phases, and still cleans up the probe via the version id the target actually assigned. Test infra: FakeS3Target gains assign_own_version_ids (models a generic S3 service; validated-but-not-mirrored source version headers) and a prefix+max-keys ListObjectVersions implementation (the probe key allocation requires it); stored_versions accessor duplicated from the P1-21 branch (identical code, resolves clean on merge). * fix(replication): probe the version-identity contract in replication-check P1-19 (rustfs/backlog#1675 B2, plan B). Replication only converges on targets that adopt the source version id: version-addressed deletes and heal re-drives address the source id, so a target that mints its own ids silently diverges — nothing surfaced this. replication-check already captured the probe PUT's response version id but never compared it. - The probe PUT now carries the source version as `?versionId=` (the exact shape live replication uses since P0-5, and the only shape MinIO consumes; the internal source-version-id header alone would let the probe pass against targets the real data path drifts on). Reuses ecstore's append_version_id_query through the api facade. - New VersionFidelity phase: the probe PUT's response version id must equal the sent source id. On mismatch the phase fails with the machine-readable extension key `"Code": "BucketRemoteTargetVersionMismatch"` (new optional Code field on phase statuses; Go decoders ignore unknown keys), the overall target fails, the later version-addressed mutation phases are skipped, and cleanup still removes the probe via the id the target actually assigned (with the existing list-based sweep as backstop when the target returns no version id at all). - Runtime half: TargetClient::put_object now returns the assigned version id (mirroring remove_object), and the replication PUT path audits it — every drifting PUT increments rustfs_replication_version_identity_drift_total and the first drift per target ARN logs a structured warning pointing at ?replication-check. The drift judgment is a pure function with an exemption-matrix test (empty / literal "null" / nil-uuid sources carry no contract). - docs/operations/replication-check.md documents the phase and the code. Red -> green: test_replication_check_flags_version_minting_target (fake target with assign_own_version_ids; on main the check reported Status "OK"). The probe's query shape is pinned by a journal assertion (revert of the query hunk alone fails it), probe-level unit tests cover the mismatch/mirror matrix including cleanup addressing the minted id, and the existing success e2e now asserts VersionFidelity OK against a RustFS target. Adversarial review (seven roles): non-blocking; noted follow-ups are the multipart runtime audit (the probe phase already pins the contract) and per-target re-warning after reconfiguration. * fix(e2e): stop the fake target self-deadlocking on version-id minting The assign_own_version_ids flag was read with a fresh `lock(&self.store)` inside two paths that already hold that guard — delete_object's marker-creation branch and create_multipart_upload — and the store mutex is not reentrant, so both hung forever (CI: the fake target's own multipart and delete-marker tests ran >1560s until the job was cancelled). Read the flag from the live guard instead. The replication e2e paths did not catch this: a version-addressed purge DELETE never mints an id, and the probe PUT reads the flag before taking the guard. * chore(test): refresh the nextest replication count invariant The e2e-smoke/e2e-repl-nightly split comment is descriptive metadata (authority: `cargo nextest list`); refresh it to this branch's post-rebase total.
3.4 KiB
Replication target check
GET /BUCKET?replication-check is a signed S3 extension for validating every
replication target referenced by a bucket replication configuration.
Active mutation warning
Despite using GET, this operation is not read-only. On each target it:
- writes an 8-byte object under
.rustfs.sys/replication-check/<uuid>/<uuid>; - creates a replicated delete marker;
- permanently deletes the probe object version; and
- enumerates that exact probe key and attempts to delete every remaining object version and delete marker.
Callers should obtain operator confirmation before sending the request. Probe
keys use a reserved namespace and two independent random UUIDs. Before writing,
the server verifies that no version or delete marker exists at the exact key,
then uses an atomic If-None-Match: * write so it cannot overwrite a key created
concurrently by an application.
Response contract
The route returns HTTP 200 with JSON after all configured targets have been
checked. Status is FAILED when any target or cleanup phase failed; successful
target results remain present when another target fails.
{
"Status": "FAILED",
"ActiveMutation": true,
"MutationDescription": "Writes a probe object, creates a delete marker, deletes the probe version, and cleans up all probe artifacts on each target.",
"ProbeNamespace": ".rustfs.sys/replication-check/",
"Targets": [
{
"Arn": "arn:minio:replication::target",
"Bucket": "replica",
"Status": "FAILED",
"Error": "probe cleanup failed: target delete object version check failed: AccessDenied",
"Phases": {
"Bucket": { "Status": "OK" },
"Versioning": { "Status": "OK" },
"ObjectLock": { "Status": "OK" },
"Put": { "Status": "OK" },
"VersionFidelity": { "Status": "OK" },
"DeleteMarker": { "Status": "OK" },
"VersionDelete": { "Status": "OK" },
"Cleanup": {
"Status": "FAILED",
"Error": "target delete object version check failed: AccessDenied"
}
}
}
]
}
Phase states are OK, FAILED, or SKIPPED. Errors are single-line, bounded
to 512 bytes, and omit remote messages, endpoints, credentials, signatures, and
authorization material. A cleanup failure is always explicit; it is never
reported as a successful check.
VersionFidelity pins the version-identity contract on both write paths:
the probe PUT carries a source version id (header plus ?versionId= query,
the exact shape live replication uses) and the target must answer with the
same id, and a second probe repeats it through CreateMultipartUpload ->
UploadPart -> CompleteMultipartUpload, where the target fixes the version at
initiate and only reports it on completion. A target can adopt PutObject ids
and still mint its own for multipart, which would leave multipart deletes and
heals addressing a version that never existed; the failure message names the
path that drifted. Targets that
mint their own version ids break every version-addressed operation that
follows (version deletes, heal re-drives), so the phase fails with the
machine-readable extension key "Code": "BucketRemoteTargetVersionMismatch",
the later mutation phases are skipped, and cleanup still removes the probe via
the version id the target actually assigned. Code only appears on failures
that callers are expected to branch on; Go decoders ignore the unknown key.