* fix(odm): declare the remote client retry policy per consumer The SDK retry policy was an inherited default: one logical call could cost three wire requests, so the migration breaker counted logical calls on top of a threefold amplification against a source that was already failing. Make it an explicit RemoteS3EndpointSpec field. Replication targets declare today's standard three attempts and keep their behaviour; the on-demand migration source and its admin probe declare a disabled policy, so one counted failure is exactly one source request and pull.rs owns the only retry budget. * fix(odm): count a stalled inline source as a source timeout The inline tee wraps its source body in the idle guard, but the tee turns a stalled source into an ordinary body read error, so the write-back reported it as a local write failure. Hand commit_inline the guard so the pull is counted under source_timeout instead. The background pump now enforces the idle budget through the same guard rather than a second copy of the timeout loop. * test(odm): cover a stalled source body end to end The fake target can now deliver a GetObject body in slices with a pause between them, so the inline abort can be driven by a stalled source instead of a truncated one. Two fault cases drop the workarounds they carried for the SDK's retries: the scripted fault count and the observed source request count now have to agree. The operations guide records the retry and idle-timeout guarantees.
7.4 KiB
Programmable fake S3 target
This module is the shared failure-injection boundary for replication end-to-end tests and the programmable external source for on-demand-migration (ODM) tests. It runs an in-process, path-style S3 endpoint backed by s3s; no production crate depends on it.
FakeS3Target::start() creates the listener. Add target buckets with create_bucket, point a RustFS remote target at address(), use FAKE_ACCESS_KEY / FAKE_SECRET_KEY, then enqueue per-operation faults with inject. Faults for one operation are consumed in FIFO order and do not consume faults queued for another operation. A fault is consumed only after s3s verifies the full request signature, so anonymous, other-access-key, and bad-signature traffic cannot disturb a script.
Supported data operations are HeadBucket, GetBucketVersioning, ListObjectsV2, PUT/GET/HEAD/DELETE Object, Get/Put/Delete ObjectTagging (tags live per version; Put replaces the whole set, Delete clears it), and create/upload/complete/abort multipart upload. create_bucket models general-purpose buckets in S3's shared global namespace; account-regional namespace buckets and their -an names are intentionally out of scope. Buckets created with create_bucket are versioned: PUT creates a version, DELETE without versionId creates a delete marker, and DELETE with versionId removes exactly that version. Internal source version IDs must be UUIDs and are stored canonically. Source mtime is honored only for source-replication PUT/DELETE requests; absent or invalid values use receipt time, matching RustFS, while multipart completion always uses receipt time. Replicated versions are ordered newest-first by source mtime so late older versions and delete markers do not become current. Equal mtimes prefer objects over delete markers, then canonical UUID order; RustFS's internal FileMeta signature tie-break is intentionally out of scope because it is not part of the target S3 protocol. Multipart part numbers follow S3's 1..=10000 range, and every completed part except the final part must be at least 5 MiB.
create_bucket_with_object_lock(name) creates a versioned bucket whose GetObjectLockConfiguration reports Enabled; every other bucket answers ObjectLockConfigurationNotFoundError, the code RustFS's replication-check classifies as "not enabled". Three switches model remote-target behaviors the fleet has shown, so the outbound target matrix (crates/e2e_test/src/replication_target_matrix_test.rs) can replicate every object shape against each: assign_own_version_ids(true) ignores the source version id and mints its own (AWS S3 / Wasabi); reject_aws_chunked_uploads(true) refuses any PutObject or UploadPart announcing aws-chunked framing (Content-Encoding: aws-chunked, an x-amz-trailer, or a STREAMING-* payload hash) with InvalidRequest before the body is read (SeaweedFS 3.97, rustfs#6853); require_checksum_for_object_lock(true) rejects a PutObject carrying any x-amz-object-lock-* header unless it also carries Content-MD5, an x-amz-checksum-* header, or x-amz-sdk-checksum-algorithm (AWS S3 / MinIO, rustfs#7082). Independently of that switch, a Content-MD5 header is always verified against the body and a mismatch answers BadDigest.
create_bucket_with_mode(name, BucketMode::Unversioned) models a plain migration source: PUT overwrites in place, DELETE removes the key without a delete marker, GetBucketVersioning reports no status, and no x-amz-version-id is returned by PUT, GET, HEAD, tagging, or multipart completion. The only versionId such a bucket accepts is null; any other value is rejected with InvalidArgument. The mode is fixed at creation.
ListObjectsV2 lists current versions only (a key whose newest version is a delete marker is hidden) in byte order and supports prefix, delimiter, max-keys (clamped to 1000), start-after, and continuation-token; common prefixes count toward max-keys, IsTruncated / NextContinuationToken / KeyCount follow S3, and continuation tokens are opaque. encoding-type and fetch-owner are accepted but ignored, and ListObjects (v1) is not implemented. GET and HEAD honor Range in the bytes=first-last, bytes=first-, and bytes=-suffix forms with a 206 status, exact Content-Range, and Accept-Ranges: bytes; unsatisfiable ranges answer 416 InvalidRange with Content-Range: bytes */<length>. PUT and CreateMultipartUpload accept Content-Type, Content-Encoding, Content-Disposition, Content-Language, Cache-Control, Expires, and x-amz-meta-* (names stored lowercased), and HEAD/GET replay them verbatim together with Last-Modified and the ETag (hex MD5 for single PUTs, <md5-of-part-md5s>-<parts> for multipart objects). put_seed_object stores an object directly, bypassing the wire, the fault script, and the journal, so a source can be seeded without polluting the assertions a scenario later makes.
Fault actions cover HTTP 401/403/503 responses (Status), any 4xx/5xx status paired with the matching S3 error code (ResponseStatus), pre-dispatch delay, holding a fully computed successful response before its first byte (Stall), connection abort when a logical request-body threshold is reached, GetObject bodies cut off after N bytes while Content-Length announces the full size (TruncateBodyAt), GetObject bodies delivered in fixed slices with a pause between them (SlowSendBody, a mid-body stall rather than a first-byte one), streaming slow drain, and a deliberately wrong response ETag (including multipart-complete XML). requests() returns the ordered, credential-free request journal for assertions and count_requests(operation, key) counts entries for one exact key. Each record journals the Range and User-Agent request headers, the ListObjectsV2 prefix and continuation-token query values, a TransportSnapshot — whether the body was announced as aws-chunked, the verbatim Content-MD5, the sorted x-amz-checksum-* / x-amz-sdk-checksum-algorithm header names, and whether any x-amz-object-lock-* header was present — and a ProxyHeaderSnapshot — the read-proxy anti-loop marker (x-{rustfs,minio}-source-proxy-request), the replication-check exemption header, and the client SSE-C header family (algorithm and key-MD5 values; for the key itself only its presence) — so proxy tests can pin the exact wire contract.
The listener is loopback-only. It admits at most 64 active connections and two concurrently buffered request bodies; authenticated multipart-complete XML collection and assembly take both body permits. Keep-alive is disabled, request-header reads are bounded to 30 seconds, a parsed request is bounded to 65 seconds, and the complete connection lifetime is bounded to 100 seconds. It retains at most 256 buckets, 4,096 journal entries, 4,096 scripted faults, 4,096 object versions, 256 multipart uploads, and 10,000 multipart parts. Retained identifiers are capped at 1 KiB, user metadata at 2 KiB, and content type and each standard object header at 1 KiB. By default a PUT or uploaded part is capped at 64 MiB and a completed multipart object and all stored object/part data are capped at 128 MiB; FakeS3Target::start_with_options(FakeS3TargetOptions { max_object_bytes }) raises the object cap up to 256 MiB, and the total budget then becomes twice the object cap (never below 128 MiB). Body drain, body-permit waits, delay, stall, and slow-drain execution are bounded to 30 seconds; each slow-drain slice delay must be below that bound.