fix(release): backport main fixes and stabilize tier cleanup tests (#7793)

* fix(ecstore): stop pruning at nonempty directories (#7616)

* fix(ecstore): stop pruning at nonempty directories

* test(ecstore): release pruning fixtures before temp cleanup

(cherry picked from commit 8f150d1d8e)

* fix(heal): preserve retryable batch failures during recovery (#7642)

* fix(heal): preserve retryable batch failures during recovery

* test(heal): pin prebuilt hooks binaries in ci

(cherry picked from commit 5cd58319ed)

* fix(s3): reject oversize single PUT early and map body errors to 4xx (#7635)

* fix(s3): reject oversize single PUT early and map body errors to 4xx

A single PutObject above the 5 GiB single-request ceiling was only
rejected after the client had streamed 5 GiB into s3s's read-time body
budget, and the resulting BodySizeLimitExceeded surfaced from the erasure
writer as 500 InternalError. A body whose connection hit EOF before
Content-Length bytes arrived (hyper's IncompleteBody) was also a 500.
SDKs retry 500s, so one oversize upload was resent from offset 0 five
times.

- PutObject and UploadPart reject a declared length above
  MAX_SINGLE_PUT_OBJECT_SIZE with 400 EntityTooLarge before reading the
  body; the constant moves to rustfs_config so the s3s limit and the
  admission check share one value.
- ApiError maps BodySizeLimitExceeded to EntityTooLarge and a hyper body
  EOF to IncompleteBody across both io::Error conversions.

Fixes #7596.

* test(s3): cover UploadPart admission, aws-chunked length, real s3s limit

- Poll-counting test body proves PutObject and UploadPart reject a
  declared size above the ceiling with zero body polls; exact-cap and
  zero-length parts pass admission.
- A STREAMING-* aws-chunked PUT whose framed Content-Length exceeds the
  cap is admitted when the decoded length is within it and rejected when
  the decoded length is over it.
- The display-based BodySizeLimitExceeded matcher is checked against the
  real error produced by the pinned s3s Body budget.

(cherry picked from commit 50b31bc75b)

* fix(ecstore): make directory mtime fixture portable (#7623)

* fix(ecstore): make directory mtime fixture portable

* style(ecstore): format mtime fixture assertion

---------

Co-authored-by: houseme <housemecn@gmail.com>
Co-authored-by: Zhengchao An <anzhengchao@gmail.com>
(cherry picked from commit b1cc286cac)

* fix(storage): prevent readiness after native migration failures (#7652)

* fix(storage): prevent readiness after native migration failures

* fix(storage): skip unsupported IAM records before reading

* fix(storage): use stable typed migration metadata errors

* fix(storage): include migration record in startup errors

* test(storage): cover native migration startup failures

* test(storage): use array chunks in migration fixture

---------

Co-authored-by: RJ Regenold <214054+rjregenold@users.noreply.github.com>
Co-authored-by: cxymds <cxymds@gmail.com>
(cherry picked from commit 0cbc3ffe61)

* fix(admin): expose OIDC account display fields (#7654)

Expose verified OIDC username and email claims as display-only metadata on self-account responses while preserving the virtual parent as the authorization identity.\n\nKeep rustfs-madmin public response structs unchanged by adding the optional wire fields through private handler response wrappers.

(cherry picked from commit f02bc947cd)

* fix(replication): correct peer joins and remote-state reporting (#7650)

* fix(replication): propagate verified peer deployment identities

* fix(replication): report actual remote peer state

* fix(replication): defer initial sync until all peers join

* test(replication): shut down TLS fixtures cleanly

---------

Co-authored-by: houseme <housemecn@gmail.com>
(cherry picked from commit 853ae63b6a)

* fix(s3): bound stalled UploadPart request bodies (#7659)

* fix(s3): bound stalled UploadPart request bodies

* fix(ci): preserve the S3S footprint ratchet

(cherry picked from commit 666dfd9f9f)

* fix(tables): reject reserved warehouse locations (#7671)

Co-authored-by: cxymds <cxymds@gmail.com>
(cherry picked from commit 414176c47f)

* fix(ci): bind nightly lanes to one resolved source (#7688)

(cherry picked from commit 01d8e4347f)

* fix(replication): close the pre-stable convergence gaps from backlog#2367 (#7626)

* fix(replication): total-order rule sort and honor V1 top-level Prefix

Rule matching had two defects from the pre-GA replication audit
(rustfs/backlog#2367 C-1 and C-2):

- The actionable-rule sort compared same-destination rules by priority but
  answered Equal for any other pair, which is not a total order; the
  standard library sort panics on such comparators once a slice exceeds the
  insertion-sort threshold, so an object matching more than 20 enabled rules
  across two or more targets could panic the PUT or DELETE task. Rules now
  sort by priority descending with destination and id as tie-breakers, and
  filter_target_arns preserves that order instead of draining a HashSet.

- A V1 rule written without a <Filter> carries its prefix at the top level;
  that field was never read, so <Prefix>logs/</Prefix> matched every object.
  ReplicationRuleExt::prefix now falls back to it, with a <Filter> keeping
  precedence. The existing prefix fixtures were built this way and had been
  asserting nothing.

* fix(admin): advertise data-usage and listen capabilities to rc

The rc client gated `rc du` and `rc watch` on a pinned contract that
matched server versions by the string prefix `1.0.0-rc.`; a server that
reports `1.0.0` no longer matches, and the dynamic `advertised` list did
not carry either name, so `rc du` against a GA server fails with an
unsupported-capability error (rustfs/backlog#2367 E-2).

Advertise `admin.data-usage` from the admin route inventory like the IAM
entries, and `listen_notification` for the bucket `?events=` extension
route the admin router dispatches. The client merges advertised entries
ahead of its pinned contract, so no version sniffing is needed.

* fix(site-replication): stop notifying the local site on remove and rotate

The pending-remove and pending-rotation notification loops skipped the
local site by endpoint only, while finalization identifies it by
deployment id or endpoint. The reconcile tick resolves the local peer from
the node's own listen address (and a handler from the request Host), so
`remove --all` dialed the site's registered endpoint, waited out the
request timeout against the lifecycle lock it was holding, and answered
`Partial: failed to notify 1 peer(s)` for a removal that had succeeded
(rustfs/backlog#2367 A-4, backlog#2195 item 3).

Both loops now iterate the peers still awaiting notification through one
helper that applies the finalization identity.

* fix(site-replication): promote and settle IAM retries without a tick of slack

Two retry-queue behaviours kept an IAM change from converging for ten to
twenty minutes after a peer came back (rustfs/backlog#2367 A-1 and A-3,
backlog#2305):

- The lightweight 30-second pass filtered its reachability probe to bucket
  ops, so a backed-off IAM or bucket-metadata snapshot waited for the
  600-second tick to notice the peer. It now probes every backed-off class
  and still replays only bounded bucket ops; promotion is a state flip the
  heavyweight tick acts on.

- Backoffs are multiples of the tick interval, so a failure stamped δ
  seconds after a tick was 600 − δ old at the next tick and slipped a whole
  extra interval. The heavyweight drain now evaluates backoff halfway to its
  next tick.

- An IAM entry first created by a non-deletion failure (the add bootstrap's
  snapshot send, the drain's own replay, an import-iam schedule) was never
  stamped `deletions_recorded`, so a later recorded deletion could not
  settle it and it escalated to the marker only `replicate repair` clears.
  Entries created by this binary now start recorded; a row persisted by an
  older binary keeps the escalation semantics.

* fix(site-replication): reload peer node caches after bucket wiring writes

Every S3 bucket-config write ends by asking the other nodes of the cluster
to reload the bucket's metadata; the site-replication writers never did.
On a multi-node site the node that ran the pairing (or applied a peer's
bucket-meta item) rewrote the bucket targets and the derived replication
rules on disk, while every other node kept serving its cached copy for up
to the 15-minute refresh. A `resync start` routed to such a node reported
every freshly wired bucket as `Config not found` and a bucket whose
operator target the pairing had replaced as `recorded remote target no
longer exists` (rustfs/backlog#2367 A-5, backlog#2195 item 2; functional
SITE-105).

Add one best-effort reload helper in the site-replication hooks and call it
after the bucket setup, versioning, peer bucket-meta apply, removed-peer
cleanup, make-with-versioning, and endpoint-refresh writes; the ensure
helpers now report whether they wrote so unchanged passes stay silent. The
resync manifest and start now read the persisted wiring instead of the
node-local cache, matching the target read the start path already did.

The new four-node e2e pairs two clusters and starts a resync through a
non-coordinator node right after pairing; it also covers an IAM user
created on a non-coordinator node converging to the peer site.

* test(e2e): cover delete-marker replication from a multi-node source

The functional suite reported delete markers created on a 3-node source
never reaching the target (rustfs/backlog#2195 item 4, REP-105). The report
was a probe defect, but the shape had no coverage: the existing
delete-marker e2e runs a single-node source. Pin it against a four-node
source replicating to a four-node peer and to a single-node target, with
the write and the delete issued through different nodes.

* ci(e2e): refresh the distributed selection for the new replication cases

Four distributed cases were added (two site-replication, two delete-marker
replication). The linux digest is derived from the last CI listing of the
lane (34 cases, matching the previous pin) plus the four new names; the
darwin digest is the local listing, which selects the same 38 cases.

(cherry picked from commit aeaba86d73)

* fix(e2e): require a verified server binary for every e2e run (#7687)

* fix(ci): share quick checks and lint workflows

* fix(ci): install actionlint from its verified release

* fix(ci): reject dependencies on required quick checks

* feat(test): verify the E2E server build and source identity

* test(e2e): register verified Darwin test membership

* test(e2e): record verified Linux receipt test membership

* test(e2e): record compiled Darwin receipt test membership

* test(e2e): record compiled Linux receipt test membership

* test(e2e): record Darwin e2e-full membership after merging main

* fix(test): route scanner/heal evidence E2E runs through the verified server binary

The evidence runners built rustfs with plain cargo and then ran e2e_test directly, which now fails without a run receipt. They build through scripts/e2e_binary.py and run the e2e_test invocations under e2e_binary.py run; the obsolete rustfs.features stamp is removed.

* docs(e2e): run server-backed e2e commands through the verified binary wrapper

* test(e2e): record Linux e2e-full membership from the branch CI listing

(cherry picked from commit 2909b1bfe1)

* fix(kms): classify KMS/SSE error contracts and SSE-S3 headers (#7697)

* fix(sse): classify bare SSE-KMS writes when no KMS is available

A `aws:kms` request without a key id, on a bucket without a default key,
returned `500 InternalError` whenever no KMS service was running: the
"no KMS key available" branch exited with an untyped storage error before
the availability classification that the keyed form already received.

Route that branch through the same split: `503 ServiceUnavailable` while
a configured KMS is stopped, `400 InvalidRequest` when KMS was never
configured, and `400 InvalidRequest` naming the missing key id when a
running KMS has no default key. `CreateMultipartUpload` shares the path.

Adds a unit test for the bare form and an e2e module that stops KMS
through the admin API, runs a master-key-only node, and runs a Local KMS
without a default key; refreshes the e2e-full selection digests.

(cherry picked from commit c3259dadc3d603a9185a5b0ad9f83dfb884e61c8)

* fix(sse): keep KMS error classes on the encrypted read path

GetObject, CopyObject and UploadPartCopy on an SSE-KMS object whose key
no longer exists answered `500 InternalError` ("KMS key not found") while
PutObject under the same key already answered `400 KMS.NotFoundException`.
The read path carries its classification through ecstore's
`EncryptionResolutionErrorKind`, which had no kind for a missing key, a
denied KMS grant or a missing backend capability, so all three folded
onto `DecryptionFailed` and the S3 layer reported an internal fault.

Add `KeyNotFound`, `AccessDenied` and `NotImplemented` kinds, map them on
both sides of the boundary, and give an envelope the configured backend
cannot unwrap a diagnosable message while keeping its `500`.

Unit tests cover the kind round trip and the reader wrapping; a new e2e
test deletes a key immediately and checks GET/Copy return 400 with
`KMS.NotFoundException` while HEAD stays 200. The e2e-full selection
digests are refreshed from the current listing (the previous digests
predated the delete-authorization tests) and the e2e `create_default_key`
helper is updated to the accepted `EncryptDecrypt` spelling.

(cherry picked from commit 2523a9814e97caea318d4ff1a51bef3a4d4445b2)

* fix(kms): classify key-management errors on the admin routes

`POST /kms/keys`, the legacy `create-key` alias and `generate-data-key`
reported every backend refusal as `500`: a blank key name (which each
backend failed on differently, the Local backend by writing a key file
with an empty stem), a name already taken, an unknown key, a disabled key
and a capability the backend lacks. `delete` and the lifecycle routes
already classified the same errors.

Refuse a blank or whitespace name in `KmsManager::create_key` before any
backend sees it, and share one `KmsError` to status mapping across
create, delete and generate-data-key (400 for validation and key state,
404 for an unknown key, 409 for a taken name, 501 for a missing
capability, 500 only for damaged material). The XML-error routes carry
the same status explicitly since s3s derives none for a custom code.

The read-only Static backend now reports create, delete and
cancel-deletion as `UnsupportedCapability`, matching its rotate and
enable/disable answers, so the admin API returns 501 for all of them.

(cherry picked from commit e33cac5493c4d9d6662e0d2980b58ba2b24a6d1b)

* fix(sse): stop SSE-S3 responses from naming the wrapping KMS key

`x-amz-server-side-encryption-aws-kms-key-id` is defined for `aws:kms`
objects only, but PutObject, CopyObject, CreateMultipartUpload and
GetObject returned it for `AES256` objects too, carrying the KMS key that
wraps the SSE-S3 data key (the service default, or the literal `default`
on a node without KMS). The write paths copied `kms_key_id` from the
encryption material unconditionally, and the single-decrypt GET
classification did the same after resolving the key for authorization.

Add `EncryptionMaterial::response_kms_key_id`, which yields the id only
for SSE-KMS, use it at the four write-response sites, and gate the GET
classification the same way. CompleteMultipartUpload and HeadObject
already omitted the header.

Unit tests pin both directions; a new e2e test covers Put/Get/Head/Copy
and CreateMultipartUpload for AES256 with an aws:kms control. The
e2e-full selection digests are refreshed from the current listing.

(cherry picked from commit 29d793a63352b0b60fd53c565e80fdbede8964bb)

* fix(s3): validate PutBucketEncryption rules before storing them

A default-encryption rule naming an unknown `SSEAlgorithm` (for example
`AES128`), a rule without `ApplyServerSideEncryptionByDefault`, an empty
rule list, or a `KMSMasterKeyID` on an `AES256` rule was stored as
written: the only algorithm check on the route decided whether to fill
in the default KMS key. `GetBucketEncryption` then advertised that
configuration while the write path encrypted header-less writes under
its `AES256` fallback, so the bucket's declared and actual schemes
disagreed. Two comments claimed the route already refused unknown
algorithms.

Validate the configuration before any of it is applied: `MalformedXML`
for a malformed rule set or unknown algorithm, `InvalidArgument` for a
key id on a non-KMS rule, and nothing stored on refusal. Correct the two
comments to describe when the AES256 fallback is still reachable.

Unit tests cover every refusal and the accepted shapes; an e2e test
checks the refusals leave the previous configuration in place. The
e2e-full selection digests are refreshed from the current listing.

(cherry picked from commit 29e4486dce41197ed93f5253cdbabc57d27a4ddb)

* test(e2e): refresh e2e-full selection for the combined KMS/SSE fixes

* test: align two unit tests with the new KMS and bucket-encryption contracts

`scheduled_deletion_carries_a_deadline_and_can_be_cancelled` still
expects the state error (`InvalidOperation`) for cancelling a key that
is not pending deletion; only the Static backend's mutations moved to
`UnsupportedCapability`. The uninitialized-store PutBucketEncryption
test now sends a well-formed AES256 rule so it reaches the store lookup
instead of the new configuration validation.

(cherry picked from commit e2e6a2535a)

* fix(site-replication): keep an operator's bucket-level target to a peer instead of taking it over (#7709)

* fix(site-replication): keep an operator's bucket-level target to a peer instead of taking it over

Site replication wired each bucket by looking for an existing replication
target "to the same peer" and rewriting the first match in place as its own
same-name target. An operator's bucket-level target that happened to point
at that site (different target bucket, operator credentials) was the first
match whenever it pre-dated the join, and the reconciler repeats the pass
every 600s, so the takeover also depended on target order afterwards. The
operator's rule then named an ARN no target backed and their bucket
replication stopped silently, while the inherited bucket-level reset id
made every site resync report the bucket as owned by another resync
(rustfs/backlog#2479, rustfs/backlog#2489).

Follow MinIO's `getRemoteARN` / `getRemoteARNForPeer` shape instead:

- Wiring updates a target in place only under the same ARN, or when it is
  recognisably the site's own under an older ARN shape (same peer,
  same-name target bucket, site replication service account). Anything
  else gets the site target added next to it.
- The site resync manifest takes the target the derived
  `site-repl-<deployment id>` rule names (same-name shape as fallback), so
  an operator target to the peer neither aborts the bucket as "multiple
  remote targets matched peer" nor gets resynced into.
- Peer removal prunes only targets a pruned derived rule names or the
  same-name target bucket; operator targets stamped with the peer's
  deployment id survive together with their rules.

Unit tests cover the three predicates. e2e
`test_site_replication_keeps_operator_bucket_target_to_peer` runs a
bucket-level replication plus `replication-reset` to the future peer, joins
the sites, and requires the operator target untouched, both paths
delivering, the site resync completing against the site target, and the
operator target and rule surviving `replicate remove --all`; without the
fix it fails at the join with the operator target gone. The repl-nightly
selection digest is refreshed for the new case.

* test(site-replication): drop a redundant clone flagged by clippy

The reconcile unit test cloned the remote peer into the state map although
the binding is not used afterwards; workspace clippy (-D warnings) rejects
that as redundant_clone.

(cherry picked from commit ecdc55fa4b)

* fix: enforce S3 permissions for recursive force deletion (#7661)

* fix: enforce S3 authorization for recursive deletion

* fix: satisfy the s3s footprint guard

* fix: restore list versions policy compatibility (#7686)

(cherry picked from commit 3fd1ce414d)

* fix(ci): repair functional defaults and chain regression checks (#7664)

* fix(ci): default functional suites to nightly packages

* test(ci): follow the fault-tolerance chain handoff

(cherry picked from commit 509a0fa90c)

* fix(ci): align security workflow tests with chain (#7679)

(cherry picked from commit d9e47d2813)

* test(ecstore): keep tier cleanup tests stable after immediate receipt queueing

Release PR #7766 made PUT/CopyObject overwrites queue the tier free-version cleanup receipt immediately, which broke two ecstore tests on release CI. In tier_overwrite_put_and_self_copy_recover_persisted_cleanup_owners the restarted store already runs expiry workers from the second iteration on, so they deleted the remote bytes before the test could assert that the commit leaves them in place; the test now fails the first remote DELETE via set_remove_failure(true) so the cleanup owner stays durable and the later restart still has to rediscover it from xl.meta (the failed remove does not bump remove_count, and the test re-enables removes before the recovery wait). In batch_transitioned_delete_post_commit_failures_roll_back_without_free_version_receipt the convergence loop now treats a transient InsufficientReadQuorum as "not yet converged", because cleanup rewrites xl.meta disk by disk and a racing read can briefly miss quorum (seen in the rio-v2 lane); any other error still panics.

* Revert "fix(ci): align security workflow tests with chain (#7679)"

This reverts commit 4544359f6d.

* Revert "fix(ci): repair functional defaults and chain regression checks (#7664)"

This reverts commit 5ba7ec0291.

* fix(ecstore): pass shard integrity to backported ingest-mode test

The stalled-reader test backported with #7659 used main's four-argument encode_with_ingest_mode, but release's signature takes an optional IntegrityBuilder, so pass None to keep the test focused on ingest-mode cleanup.

---------

Co-authored-by: Henry Guo <marshawcoco@gmail.com>
Co-authored-by: 唐小鸭 <tangtang1251@qq.com>
Co-authored-by: houseme <housemecn@gmail.com>
Co-authored-by: RJ Regenold <rregenold@teamraft.com>
Co-authored-by: RJ Regenold <214054+rjregenold@users.noreply.github.com>
Co-authored-by: cxymds <cxymds@gmail.com>
Co-authored-by: GatewayJ <835269233@qq.com>
Co-authored-by: Jason Kossis <jkossis@gmail.com>
This commit is contained in:
Chris
2026-09-14 07:51:15 +08:00
committed by GitHub
parent 8982b4a3d2
commit 4e16705873
101 changed files with 8084 additions and 1134 deletions
+1 -1
View File
@@ -1 +1 @@
sha256=0e338d305260229e17ccfb2adc48a6212dbdfea36a9ebfb5a4e0d38658e6cc45
sha256=fc653ffc00f85a109be82e494117f0f60930f1c105898a5758af3c1bd998b1e1
+1
View File
@@ -49,6 +49,7 @@ script-tests: ## Run shell script tests
./scripts/test_python_bin.sh
./scripts/check_embedded_secrets.sh --self-test
$(RUSTFS_PYTHON_BIN) ./scripts/check_test_wiring.py --self-test
$(RUSTFS_PYTHON_BIN) ./scripts/test_e2e_binary.py
$(RUSTFS_PYTHON_BIN) ./scripts/ci_gate.py --self-test
$(RUSTFS_PYTHON_BIN) ./scripts/check_security_coverage.py --self-test
$(RUSTFS_PYTHON_BIN) ./scripts/check_scheduled_validation_freshness.py --self-test
+14 -9
View File
@@ -564,7 +564,7 @@ jobs:
digest.update(chunk)
return digest.hexdigest()
argv = ["cargo", "build", "-p", "rustfs", "--bins", "--features", "e2e-test-hooks"]
argv = ["python3", "scripts/e2e_binary.py", "build", "--bins", "--features", "e2e-test-hooks"]
commit, tree = git("rev-parse", "HEAD"), git("rev-parse", "HEAD^{tree}")
clean_before = not git("status", "--porcelain", "--untracked-files=normal")
if not clean_before:
@@ -598,6 +598,7 @@ jobs:
name: rustfs-debug-binary
path: |
target/debug/rustfs
target/debug/rustfs.e2e.json
target/debug/rustfs.e2e-startup-cas-build.json
if-no-files-found: error
retention-days: 1
@@ -630,13 +631,15 @@ jobs:
install-build-packaging-tools: 'false'
- name: Build debug binary with rio-v2
run: cargo build -p rustfs --bins --features rio-v2,e2e-test-hooks
run: python3 scripts/e2e_binary.py build --bins --features rio-v2,e2e-test-hooks
- name: Upload debug binary
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
with:
name: rustfs-debug-binary-rio-v2
path: target/debug/rustfs
path: |
target/debug/rustfs
target/debug/rustfs.e2e.json
if-no-files-found: error
retention-days: 1
@@ -785,7 +788,7 @@ jobs:
NEXTEST_ARCHIVE: ${{ runner.temp }}/rustfs-e2e-smoke.tar.zst
RUSTFS_E2E_LOG_DIR: ${{ runner.temp }}/rustfs-e2e-smoke-logs
run: |
cargo nextest run --profile e2e-smoke --archive-file "${NEXTEST_ARCHIVE}" \
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-smoke --archive-file "${NEXTEST_ARCHIVE}" \
--status-level all --final-status-level all --failure-output final
- name: Upload e2e smoke diagnostics
@@ -821,7 +824,7 @@ jobs:
RUSTFS_TEST_PORT="$(python3 -c 'import socket; s=socket.socket(); s.bind(("127.0.0.1", 0)); print(s.getsockname()[1]); s.close()')"
RUSTFS_TEST_PORT="${RUSTFS_TEST_PORT}" \
RUSTFS_TEST_LOG="${RUN_ROOT}/rustfs.log" \
./scripts/e2e-run.sh ./target/debug/rustfs "${RUN_ROOT}/data"
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- ./scripts/e2e-run.sh ./target/debug/rustfs "${RUN_ROOT}/data"
- name: Upload test logs
if: failure()
@@ -950,6 +953,7 @@ jobs:
if manifest["binary_sha256"] != digest.hexdigest() or manifest["commit"] != commit:
raise SystemExit("downloaded hooks binary identity mismatch")
shutil.copy2(manifest_path, target / manifest_path.name)
shutil.copy2(source.with_name("rustfs.e2e.json"), target / "rustfs.e2e.json")
binary.chmod(0o755)
PYINPUT
@@ -966,10 +970,11 @@ jobs:
# debug binary; each test spawns its own rustfs server on a random port.
- name: Run e2e full suite
env:
CARGO_BIN_EXE_rustfs: ${{ runner.temp }}/rustfs-startup-cas-input/rustfs
RUSTFS_E2E_STARTUP_CAS_BINARY: ${{ runner.temp }}/rustfs-startup-cas-input/rustfs
RUSTFS_E2E_STARTUP_CAS_BUILD_MANIFEST: ${{ runner.temp }}/rustfs-startup-cas-input/rustfs.e2e-startup-cas-build.json
RUSTFS_E2E_STARTUP_CAS_ARTIFACT_DIR: ${{ runner.temp }}/rustfs-startup-cas-evidence
run: cargo nextest run --profile e2e-full -p e2e_test
run: python3 scripts/e2e_binary.py run --binary "$RUSTFS_E2E_STARTUP_CAS_BINARY" --features e2e-test-hooks -- cargo nextest run --profile e2e-full -p e2e_test
- name: Upload junit
if: always()
@@ -1041,7 +1046,7 @@ jobs:
- name: Run end-to-end tests
run: |
s3s-e2e --version
./scripts/e2e-run.sh ./target/debug/rustfs /tmp/rustfs
python3 scripts/e2e_binary.py run --features rio-v2,e2e-test-hooks -- ./scripts/e2e-run.sh ./target/debug/rustfs /tmp/rustfs
- name: Upload test logs
if: failure()
@@ -1084,7 +1089,7 @@ jobs:
S3_PORT="${S3_PORT}" \
DATA_ROOT="${RUN_ROOT}" \
S3TESTS_CONF=artifacts/s3tests-single/s3tests.conf \
./scripts/s3-tests/run.sh
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- ./scripts/s3-tests/run.sh
- name: Upload s3 test artifacts
if: always()
@@ -1166,7 +1171,7 @@ jobs:
S3_PORT="${S3_PORT}" \
DATA_ROOT="${RUN_ROOT}" \
S3TESTS_CONF=artifacts/s3tests-single/s3tests.conf \
./scripts/s3-tests/run.sh
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- ./scripts/s3-tests/run.sh
- name: Upload s3 test artifacts
if: always()
+3 -4
View File
@@ -151,8 +151,7 @@ jobs:
- name: Build rustfs binary
run: |
cargo build -p rustfs --bins
: > target/debug/rustfs.features
python3 scripts/e2e_binary.py build --bins
- name: Verify distributed e2e membership
env:
@@ -168,9 +167,9 @@ jobs:
run: |
set -euo pipefail
if [ -n "${FILTER}" ]; then
cargo nextest run --profile e2e-distributed -p e2e_test -E "${FILTER}"
python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-distributed -p e2e_test -E "${FILTER}"
else
cargo nextest run --profile e2e-distributed -p e2e_test --no-tests=fail
python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-distributed -p e2e_test --no-tests=fail
fi
- name: Upload distributed e2e diagnostics
+10 -11
View File
@@ -89,14 +89,10 @@ jobs:
- name: Verify awscurl
run: test -x "$AWSCURL_PATH"
# Build the rustfs binary once up front. The e2e tests spawn it as a
# child process (crates/e2e_test/src/common.rs) and will build it on
# demand otherwise, but a single explicit build avoids several parallel
# nextest test processes racing to build it at once.
# Build once and carry its source/binary identity into the test invocation.
- name: Build rustfs binary
run: |
cargo build -p rustfs --bins
: > target/debug/rustfs.features
python3 scripts/e2e_binary.py build --bins
- name: Verify replication e2e membership
env:
@@ -108,7 +104,7 @@ jobs:
- name: Run replication e2e nightly suite
env:
RUSTFS_E2E_LOG_DIR: ${{ runner.temp }}/rustfs-e2e-repl-nightly-logs
run: cargo nextest run --profile e2e-repl-nightly -p e2e_test
run: python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-repl-nightly -p e2e_test
- name: Upload nextest junit report
if: always()
@@ -144,8 +140,7 @@ jobs:
- name: Build rustfs binary
run: |
cargo build -p rustfs --bins --features e2e-test-hooks
: > target/debug/rustfs.features
python3 scripts/e2e_binary.py build --bins --features e2e-test-hooks
- name: Verify cluster fault e2e membership
env:
@@ -156,8 +151,9 @@ jobs:
- name: Run cluster fault e2e nightly suite
env:
CARGO_BIN_EXE_rustfs: ${{ github.workspace }}/target/debug/rustfs
RUSTFS_E2E_LOG_DIR: ${{ runner.temp }}/rustfs-e2e-nightly-logs
run: cargo nextest run --profile e2e-nightly -p e2e_test
run: python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-nightly -p e2e_test
- name: Upload cluster fault diagnostics
if: always()
@@ -198,6 +194,9 @@ jobs:
sudo apt-get install -y -qq iproute2
ss -tn state CLOSE-WAIT >/dev/null
- name: Build protocol server
run: python3 scripts/e2e_binary.py build --features "$RUSTFS_BUILD_FEATURES"
# The suite owns fixed protocol ports and serializes its internal cases.
- name: Verify protocol e2e membership
env:
@@ -210,7 +209,7 @@ jobs:
env:
RUSTFS_E2E_LOG_DIR: ${{ runner.temp }}/rustfs-protocol-e2e-logs
run: >-
cargo nextest run -j 1 --profile e2e-protocols -p e2e_test --no-capture
python3 scripts/e2e_binary.py run --features "$RUSTFS_BUILD_FEATURES" -- cargo nextest run -j 1 --profile e2e-protocols -p e2e_test --no-capture
- name: Upload protocol diagnostics
if: always()
+2 -3
View File
@@ -125,14 +125,13 @@ jobs:
- name: Build current RustFS binary
run: |
cargo build --locked -p rustfs --bin rustfs
: > target/debug/rustfs.features
python3 scripts/e2e_binary.py build
- name: Run upgrade compatibility test
env:
RUSTFS_SCANNER_HEAL_G09_EVIDENCE_DIR: ${{ runner.temp }}/rustfs-upgrade-g09-evidence/${{ matrix.artifact }}
run: |
cargo test --locked -p e2e_test \
python3 scripts/e2e_binary.py run -- cargo test --locked -p e2e_test \
"upgrade_compatibility_test::${{ matrix.test }}" \
-- --ignored --exact --nocapture
+28 -4
View File
@@ -42,18 +42,39 @@ env:
NIGHTLY_BUILD_REF: ${{ github.event_name == 'schedule' && (vars.NIGHTLY_BRANCH || 'main') || (inputs.branch || github.ref_name) }}
jobs:
resolve-source:
name: Resolve nightly source
runs-on: ubuntu-latest
timeout-minutes: 10
outputs:
source_sha: ${{ steps.source.outputs.sha }}
source_ref: ${{ steps.source.outputs.ref }}
steps:
- name: Checkout selected source
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
ref: ${{ env.NIGHTLY_BUILD_REF }}
- name: Record immutable source
id: source
run: |
echo "sha=$(git rev-parse HEAD)" >> "$GITHUB_OUTPUT"
echo "ref=${NIGHTLY_BUILD_REF}" >> "$GITHUB_OUTPUT"
build:
needs: resolve-source
name: Build x86_64 GNU
runs-on: sm-standard-4
timeout-minutes: 150
env:
NIGHTLY_BUILD_REF: ${{ needs.resolve-source.outputs.source_ref }}
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
ref: ${{ env.NIGHTLY_BUILD_REF }}
ref: ${{ needs.resolve-source.outputs.source_sha }}
- name: Setup Rust environment
uses: ./.github/actions/setup
@@ -334,9 +355,10 @@ jobs:
CANDIDATE_FILE="${RUNNER_TEMP}/nightly-candidate-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}.json"
jq -n --arg source_sha "${SOURCE_SHA}" \
--arg workflow_sha "${GITHUB_SHA}" --arg source_ref "${NIGHTLY_BUILD_REF}" \
--argjson build_run_id "${GITHUB_RUN_ID}" --argjson build_run_attempt "${GITHUB_RUN_ATTEMPT}" \
--arg package_url "${CANDIDATE_URL}" --arg package_sha256 "${DEB_SHA256}" \
'{schema: 1, source_sha: $source_sha, build_run_id: $build_run_id, build_run_attempt: $build_run_attempt, package_url: $package_url, package_sha256: $package_sha256}' \
'{schema: 2, workflow_sha: $workflow_sha, source_ref: $source_ref, source_sha: $source_sha, build_run_id: $build_run_id, build_run_attempt: $build_run_attempt, package_url: $package_url, package_sha256: $package_sha256}' \
> "${CANDIDATE_FILE}"
echo "candidate_file=${CANDIDATE_FILE}" >> "${GITHUB_OUTPUT}"
@@ -378,6 +400,7 @@ jobs:
# self-hosted fleet is heterogeneous — a docker-dependent workflow has been
# burned by it before (see the banner in e2e-s3tests.yml, rustfs/backlog#1149).
kms-vault-lane:
needs: resolve-source
name: KMS live Vault lane
runs-on: ubuntu-latest
timeout-minutes: 90
@@ -399,7 +422,7 @@ jobs:
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
ref: ${{ env.NIGHTLY_BUILD_REF }}
ref: ${{ needs.resolve-source.outputs.source_sha }}
- name: Setup Rust environment
uses: ./.github/actions/setup
@@ -477,6 +500,7 @@ jobs:
# flake cannot mask the main lane's verdict, and vice versa. The script
# provisions and tears down its own Docker cluster.
kms-vault-ha-failover:
needs: resolve-source
name: KMS Vault HA failover lane
runs-on: ubuntu-latest
timeout-minutes: 60
@@ -488,7 +512,7 @@ jobs:
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
ref: ${{ env.NIGHTLY_BUILD_REF }}
ref: ${{ needs.resolve-source.outputs.source_sha }}
- name: Setup Rust environment
uses: ./.github/actions/setup
@@ -132,7 +132,7 @@ jobs:
s3api create-bucket --bucket "${RUSTFS_ODM_INTEROP_BUCKET}"
- name: Build the RustFS binary under test
run: cargo build --locked -p rustfs --bins
run: python3 scripts/e2e_binary.py build --bins
# The lane selects tests by module, so a rename would quietly shrink it.
# The committed digest in .config/e2e-odm-interop-selection.txt fails
@@ -143,7 +143,7 @@ jobs:
python3 ./scripts/check_test_wiring.py --check-profile e2e-odm-interop "${NEXTEST_LISTING}"
- name: Run the interop cases against MinIO
run: cargo nextest run --profile e2e-odm-interop -p e2e_test --no-tests=fail
run: python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-odm-interop -p e2e_test --no-tests=fail
- name: Build the MinIO interop report
if: always()
@@ -251,7 +251,7 @@ jobs:
- name: Build the RustFS binary under test
if: steps.credentials.outputs.present == 'true'
run: cargo build --locked -p rustfs --bins
run: python3 scripts/e2e_binary.py build --bins
# A filterset that matches nothing is valid, so the count is asserted
# rather than inferred from a green run.
@@ -272,7 +272,7 @@ jobs:
- name: Run the three-case minimum
if: steps.credentials.outputs.present == 'true'
run: |
cargo nextest run --profile e2e-odm-interop -p e2e_test \
python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-odm-interop -p e2e_test \
-E "${CLOUD_CASE_FILTER}" --no-tests=fail
- name: Build the ${{ matrix.provider }} interop report
+5
View File
@@ -24,6 +24,11 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **Fresh multi-pool bootstrap with distinct format creators**: a new deployment whose pools have their first endpoint on different nodes (for example two single-node pools) could never publish its initial `pool.bin`: each node held fresh-bootstrap proof only for the pool it formatted, the deployment-wide proof collapsed to none, and every node died with `pool metadata recovery required: no durable bootstrap identity or pool.bin replica is available` after the startup retry budget. The first pool's creator now mints the pending cluster identity on its own pool, every other creator copies that nonce-bound identity onto the pool it formatted first-hand, and the elected writer publishes `pool.bin` once every pool replica carries the same pending identity. Corrupt or disagreeing replicas, pools that merely have a format, expansion pools joining an initialized deployment, and restarts without first-hand proof still fail closed. Non-elected nodes that start before `pool.bin` exists, and the elected writer while it waits for the other creators, no longer latch their pool-metadata write gate for the life of the process. Refs rustfs/backlog#2338, rustfs/backlog#2375.
- **Lock RPC timeout storms** (#7363): the remote lock client no longer evicts and re-dials the shared internode HTTP/2 channel on every request deadline. A timeout evicts only when the peer has not completed any lock RPC for two deadlines, evictions and transport-failure re-dials are rate limited per peer (`RUSTFS_OBJECT_LOCK_RPC_EVICTION_COOLDOWN_MS`, default 5 s), and a timed-out request is left running instead of being reset (bounded per peer by `RUSTFS_OBJECT_LOCK_RPC_DETACHED_LIMIT`, default 256), so a slow lock endpoint can no longer drive the `RST_STREAM`/`GOAWAY too_many_resets`/reconnect loop. A lock granted after its caller timed out is released immediately, and unlocks that fail the quick retries continue on a deferred 1/2/4/8/16 s schedule before the server lease reclaims them. New `rustfs_remote_lock_*` metrics cover timeouts, evictions, suppressed evictions, detached streams, late completions and late releases per peer. Operator guide at `docs/operations/lock-rpc-storm-protection.md`.
- **KMS failures on the S3 data path carry an actionable status**: only "key not found" and a backend outage were classified; every other KMS failure — a disabled or pending-deletion key, a denied KMS grant, an encryption-context mismatch, an unsupported algorithm, a credential or timeout failure, a capability the backend does not have — collapsed onto `500 InternalError`. SDKs therefore applied exponential backoff to configuration errors that no retry can fix, and monitoring filed every one of them as a server fault. Unusable-key and request-side failures now return `400`, a denied grant `403`, transient backend failures `503` — including a key store the backend could not read, so an outage stays distinguishable from a missing key all the way to the client — and a missing backend capability `501`. Damaged or unreadable key material still returns `500`, which is what it is.
- **Bare SSE-KMS writes on a node without a running KMS**: `x-amz-server-side-encryption: aws:kms` without a key id (and no bucket default key) returned `500 InternalError` while KMS was stopped or never configured, because the "no key available" branch exited before the availability classification that the keyed form already received. `PutObject` and `CreateMultipartUpload` now return `503 ServiceUnavailable` while a configured KMS is stopped, `400 InvalidRequest` when KMS was never configured, and `400 InvalidRequest` naming the missing key id when a running KMS has no default key.
- **Reading an object whose KMS key is gone returned `500`**: `GetObject`, `CopyObject` and `UploadPartCopy` on an SSE-KMS object whose key had been deleted reported `500 InternalError` ("KMS key not found"), while the same condition on `PutObject` already returned `400 KMS.NotFoundException`. The read path carried only four classifications across the storage boundary and folded a missing key, a denied KMS grant and a missing backend capability onto "decryption failed". Those reads now return `400 KMS.NotFoundException`, `403 AccessDenied` and `501 NotImplemented` respectively; `HeadObject` is unaffected because it never unwraps the data key. An envelope the configured backend cannot unwrap (the key was re-created under the same name, or the backend was switched) stays `500` but now says so instead of the generic internal-error text.
- **KMS key-management routes answered `500` for client-side failures**: `POST /rustfs/admin/v3/kms/keys` and the legacy `create-key` alias reported every backend refusal as `500`, including a blank key name (each backend failed differently, the Local backend by writing a key file with an empty stem) and a name that already exists; `POST /rustfs/admin/v3/kms/generate-data-key` did the same for an unknown or disabled key, although `describe` and `delete` already classified those. A blank or whitespace key name is now refused before it reaches any backend (`400`), a taken name is `409`, an unknown key is `404` (`KMS.NotFoundException`), a disabled key `400`, and a capability the backend lacks `501`; damaged key material stays `500`. The read-only Static backend now reports create, delete and cancel-deletion as missing capabilities (`501`), matching its rotate and enable/disable answers, instead of `400`/`500`.
- **SSE-S3 responses named the internal wrapping key**: `PutObject`, `CopyObject`, `CreateMultipartUpload` and `GetObject` for an `AES256` object returned `x-amz-server-side-encryption-aws-kms-key-id` carrying the KMS key that wraps the SSE-S3 data key (the service default key, or the literal `default` on a node without KMS), although the header is defined for `aws:kms` objects only and `CompleteMultipartUpload` and `HeadObject` already omitted it. Those responses now advertise a key id only for `aws:kms` objects.
- **PutBucketEncryption accepted algorithms the server cannot honour**: a default-encryption rule naming an unknown `SSEAlgorithm` (for example `AES128`), a rule without `ApplyServerSideEncryptionByDefault`, an empty rule list, or a `KMSMasterKeyID` on an `AES256` rule was stored as written. `GetBucketEncryption` then reported that configuration while every header-less write was encrypted under the `AES256` fallback, so the bucket's advertised and actual schemes disagreed. Those configurations are now refused with `400` (`MalformedXML` for a malformed rule, `InvalidArgument` for a key id on a non-KMS rule) and nothing is stored.
- **SSE-C on buckets with default encryption**: a `PutObject` carrying a valid SSE-C header triple on a bucket that has default encryption configured no longer fails with `400 InvalidArgument` ("The SSE-C and managed server-side encryption headers cannot be used together"). PUT and the POST-object/extract path resolved the bucket default with a hard-coded "no explicit SSE-C" flag, so the default was layered onto the request and then tripped the request's own mutual-exclusion check; an SSE-C request now suppresses the bucket default on all three write paths, matching COPY and AWS S3. Every bucket with default encryption previously refused SSE-C single PUTs outright, while `CreateMultipartUpload` on the same bucket succeeded.
- **Explicit SSE-S3 on SSE-KMS-default buckets**: `x-amz-server-side-encryption: AES256` against a bucket whose default is `aws:kms` no longer fails with `400 InvalidArgument`. The bucket default's KMS key id was inherited independently of the effective algorithm, producing a self-contradictory `AES256` + key-id pair; the key id is now inherited only when the effective algorithm is `aws:kms`. `PutBucketEncryption` fills in a default key id automatically, so this affected nearly every SSE-KMS-default bucket.
- **Restore of encrypted or compressed multipart objects (silent data corruption)**: restoring a multipart object from a remote tier addressed the tier in *plaintext* coordinates while the copy-back reads the *stored* representation. Every part received a misaligned slice of the remote object whose length still satisfied the range, the hash reader and the completion size check, so the restore reported success and replaced the object's bytes. Restore now accumulates stored part sizes, passes the stored length to the hash reader alongside the plaintext length, and validates against the stored size. Objects restored by an affected release must be re-restored from the tier or recovered from a backup — this release does not detect or repair them retroactively.
Generated
+4 -4
View File
@@ -11210,7 +11210,7 @@ checksum = "9774ba4a74de5f7b1c1451ed6cd5285a32eddb5cccb8cc655a4e50009e06477f"
[[package]]
name = "s3s"
version = "0.15.0"
source = "git+https://github.com/rustfs/s3s.git?rev=bdcb6259339c41369f9f1c60e3a42b5ab8da607b#bdcb6259339c41369f9f1c60e3a42b5ab8da607b"
source = "git+https://github.com/s3s-project/s3s.git?rev=f3e17541f366696bf0cbaf380fcbd8b44c17eba4#f3e17541f366696bf0cbaf380fcbd8b44c17eba4"
dependencies = [
"arc-swap",
"arrayvec",
@@ -11268,7 +11268,7 @@ dependencies = [
[[package]]
name = "s3s-rfc2047"
version = "0.16.0-alpha.1"
source = "git+https://github.com/rustfs/s3s.git?rev=bdcb6259339c41369f9f1c60e3a42b5ab8da607b#bdcb6259339c41369f9f1c60e3a42b5ab8da607b"
source = "git+https://github.com/s3s-project/s3s.git?rev=f3e17541f366696bf0cbaf380fcbd8b44c17eba4#f3e17541f366696bf0cbaf380fcbd8b44c17eba4"
dependencies = [
"base64-simd",
"thiserror 2.0.20",
@@ -11277,7 +11277,7 @@ dependencies = [
[[package]]
name = "s3s-sigv2"
version = "0.16.0-alpha.1"
source = "git+https://github.com/rustfs/s3s.git?rev=bdcb6259339c41369f9f1c60e3a42b5ab8da607b#bdcb6259339c41369f9f1c60e3a42b5ab8da607b"
source = "git+https://github.com/s3s-project/s3s.git?rev=f3e17541f366696bf0cbaf380fcbd8b44c17eba4#f3e17541f366696bf0cbaf380fcbd8b44c17eba4"
dependencies = [
"base64-simd",
"hmac 0.13.0",
@@ -11290,7 +11290,7 @@ dependencies = [
[[package]]
name = "s3s-sigv4"
version = "0.16.0-alpha.1"
source = "git+https://github.com/rustfs/s3s.git?rev=bdcb6259339c41369f9f1c60e3a42b5ab8da607b#bdcb6259339c41369f9f1c60e3a42b5ab8da607b"
source = "git+https://github.com/s3s-project/s3s.git?rev=f3e17541f366696bf0cbaf380fcbd8b44c17eba4#f3e17541f366696bf0cbaf380fcbd8b44c17eba4"
dependencies = [
"arrayvec",
"base64-simd",
+1 -1
View File
@@ -312,7 +312,7 @@ rustify = { version = "0.7", default-features = false }
rustix = { version = "1.1.4" }
rust-embed = { version = "8.12.0" }
rustc-hash = { version = "2.1.3" }
s3s = { git = "https://github.com/rustfs/s3s.git", rev = "bdcb6259339c41369f9f1c60e3a42b5ab8da607b", version = "0.15.0", features = ["minio"] }
s3s = { git = "https://github.com/s3s-project/s3s.git", rev = "f3e17541f366696bf0cbaf380fcbd8b44c17eba4", version = "0.15.0", features = ["minio"] }
serial_test = "4.0.1"
shadow-rs = { default-features = false, version = "2.0.0" }
siphasher = "1.0.3"
@@ -64,6 +64,14 @@ pub const MAX_HEAL_REQUEST_SIZE: usize = 1024 * 1024; // 1 MB
/// memory exhaustion from malicious or misconfigured remote services.
pub const MAX_S3_CLIENT_RESPONSE_SIZE: usize = 10 * 1024 * 1024; // 10 MB
/// Maximum body size accepted by a single `PutObject` or `UploadPart` request (5 GiB).
/// Used for: the s3s streaming-body limit and the request-header admission check.
/// Rationale: matches the AWS S3 single-PUT / single-part ceiling. Larger objects
/// must use multipart upload. The header check rejects an oversize
/// `Content-Length` before any body byte is read so the client gets
/// `EntityTooLarge` immediately instead of streaming 5 GiB into a mid-stream failure.
pub const MAX_SINGLE_PUT_OBJECT_SIZE: u64 = 5 * 1024 * 1024 * 1024; // 5 GiB
/// Maximum size for OIDC provider response bodies (1 MB)
/// Used for: discovery documents, JWKS documents and token endpoint responses
/// Rationale: a hostile or compromised identity provider must not be able to exhaust
+1 -1
View File
@@ -27,4 +27,4 @@ follow.
## Suggested Validation
- `cargo test --package e2e_test`
- `python3 scripts/e2e_binary.py build --features e2e-test-hooks`, then `python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo test --package e2e_test` (running `cargo test -p e2e_test` directly fails with a missing E2E run receipt; see [`README.md`](README.md#how-to-run))
+49 -47
View File
@@ -1,7 +1,7 @@
# e2e_test
End-to-end test suite for RustFS. Each test spawns a **real `rustfs` binary**
(built on demand from the workspace) and drives it over the network with the
(built and identified before the test invocation) and drives it over the network with the
AWS SDK (`aws-sdk-s3`), raw HTTP (`reqwest` / `awscurl`), or a protocol client
(FTPS / WebDAV / SFTP). This is the black-box integration layer: exhaustive
end-to-end behavior lives here, unit behavior stays in the source crates
@@ -34,32 +34,35 @@ The external-tool `storage_metric_ownership_test` validates the OTLP/Collector/P
## How to run
All commands assume repo root. `cargo test` triggers an on-demand build of the
`rustfs` binary from [`src/common.rs`](src/common.rs) (`rustfs_binary_path`) on
first use — the first invocation is slow, later ones reuse the binary.
All commands assume repo root and Python 3.9 or newer on Linux or macOS. Build the server once through the provenance entry point, then run the test command through the same script:
```bash
# Whole crate (default = ignored tests skipped)
cargo nextest run -p e2e_test
python3 scripts/e2e_binary.py build --features e2e-test-hooks
# Whole crate (ignored tests remain skipped)
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run -p e2e_test
# One module
cargo nextest run -p e2e_test -E 'test(list_objects_v2_pagination_test)'
# PR smoke subset (see "CI smoke subset" below)
cargo nextest run --profile e2e-smoke -p e2e_test
# ILM serial lane — ignored lifecycle tests, single-threaded (mirrors CI)
cargo nextest run -j1 --run-ignored ignored-only -p rustfs-scanner -p rustfs \
-E 'binary(lifecycle_integration_test) or (package(rustfs) and test(lifecycle_transition_api_test))'
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run -p e2e_test -E 'test(list_objects_v2_pagination_test)'
# PR smoke subset
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-smoke -p e2e_test
```
The protocols suite has its own contract (fixed bind ports 9022–9301,
single-worker execution, feature-gated scheduling) documented in
[`src/protocols/README.md`](src/protocols/README.md). `RUSTFS_BUILD_FEATURES`
selects which features the spawned binary is built with; leave it unset to run
every protocol entry. Use the exact profile command under
[Troubleshooting](#troubleshooting) for CI-equivalent execution.
Root-heal interruption scenarios use a test-only commit barrier, so build and run them with `e2e-test-hooks`:
```bash
python3 scripts/e2e_binary.py build --features e2e-test-hooks
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run -p e2e_test -E 'test(heal_erasure_disk_rebuild_test)'
```
`build` records the source contents, HEAD, resolved Cargo features, profile, toolchain, and binary SHA-256 beside the executable in `rustfs.e2e.json`. `run` validates that identity before and after the command, preserves command failures, and removes its temporary run receipt on completion. The Rust harness checks that receipt before starting each server; it never compiles a server inside a test process. Source or binary changes during a run invalidate the result, even when the test command succeeds. Use an isolated worktree and keep it unchanged until the command finishes.
The additional `--features` arguments must match between `build` and `run`; Cargo defaults remain enabled. The wrapper supplies `RUSTFS_BUILD_FEATURES` from Cargo's resolved feature list, including features enabled by `full`. Protocol helpers require a subset of that list. `CARGO_TARGET_DIR` and `--profile release` are supported. An in-workspace target directory must be Git-ignored; tracked files are always included in the source identity. `build --bins` preserves CI lanes that compile all RustFS binary targets. For a downloaded artifact, copy both the executable and its sidecar, then use `run`; do not generate a new identity for an arbitrary prebuilt binary. `CARGO_BIN_EXE_rustfs` cannot override the verified executable.
Each build/run holds an exclusive `rustfs.e2e.lock` marker beside the binary; concurrent wrappers fail immediately. Use a private target directory and do not run ordinary Cargo builds against it while tests are active: Cargo does not honor this marker. Interrupted runs fail and terminate their command group. After an uncatchable kill, inspect the PID recorded in a leftover marker and remove it only after confirming its owner has stopped. Embedded file symlinks are hashed through their target; embedded directory symlinks are rejected because their contents cannot be enumerated safely by this entry point.
The protocols suite has its own fixed-port and single-worker contract in [`src/protocols/README.md`](src/protocols/README.md). Use its command under [Troubleshooting](#troubleshooting).
### `#[ignore]` semantics
@@ -125,7 +128,7 @@ via `create_s3_client(idx)` / `create_all_clients()`. See
| `wait_for_server_ready` | Poll readiness before issuing requests |
| `create_s3_client` / `create_test_bucket` / `delete_test_bucket` | aws-sdk-s3 client + bucket lifecycle |
| `find_available_port` | Random free port (isolation primitive) |
| `rustfs_binary_path` / `_with_features` | Locate/build the binary; honors `RUSTFS_BUILD_FEATURES` |
| `rustfs_binary_path` / `_with_features` | Verify this run's binary receipt and required feature subset |
| `requested_rustfs_build_features` / `rustfs_build_feature_enabled` | Feature-gate a test to what the binary was built with |
| `execute_awscurl` / `awscurl_post` / `_get` / `_put` / `_delete` / `awscurl_post_sts_form_urlencoded` | Admin/STS API calls via `awscurl`; missing binaries are test failures |
| `replication_fast_env` | Env vars that shrink replication timers (from repl-4); pass to `start_rustfs_server_with_env` |
@@ -191,35 +194,33 @@ the wiring source of truth. Committed test-ID digests under
**Reproduce a CI failure locally** — run the exact profile/lane:
```bash
# Smoke (e2e-tests job) — includes the 20 fast replication tests
cargo nextest run --profile e2e-smoke -p e2e_test
# Full single-node merge/main lane
cargo nextest run --profile e2e-full -p e2e_test
# Cluster fault nightly lane
cargo nextest run --profile e2e-nightly -p e2e_test
# 4-node 4-disk distributed lane (S3 / lock / versioning / replication / decommission / chaos / upgrade)
# Upgrade cases need RUSTFS_UPGRADE_SOURCE_BINARY; without it they fail closed.
cargo nextest run --profile e2e-distributed -p e2e_test
# Replication nightly lane; awscurl is required for STS paths
cargo nextest run --profile e2e-repl-nightly -p e2e_test
# Fixed-port protocol nightly lane
RUSTFS_BUILD_FEATURES=ftps,webdav,sftp \
cargo nextest run -j 1 --profile e2e-protocols -p e2e_test --no-capture
# ILM serial lane
# Smoke, full, and cluster lanes share a server with fault-test hooks.
python3 scripts/e2e_binary.py build --features e2e-test-hooks
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-smoke -p e2e_test
python3 scripts/e2e_binary.py run --binary "$RUSTFS_E2E_STARTUP_CAS_BINARY" --features e2e-test-hooks -- cargo nextest run --profile e2e-full -p e2e_test
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-nightly -p e2e_test
# Distributed 4-node 4-disk lane uses the default server.
# Upgrade cases require RUSTFS_UPGRADE_SOURCE_BINARY and fail closed without it.
python3 scripts/e2e_binary.py build
python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-distributed -p e2e_test
# Replication nightly uses the default server; awscurl is required for STS.
python3 scripts/e2e_binary.py build
python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-repl-nightly -p e2e_test
# Protocol nightly owns fixed ports.
python3 scripts/e2e_binary.py build --features ftps,webdav,sftp
python3 scripts/e2e_binary.py run --features ftps,webdav,sftp -- cargo nextest run -j 1 --profile e2e-protocols -p e2e_test --no-capture
# The ILM serial lane does not use this server harness.
cargo nextest run -j1 --run-ignored ignored-only -p rustfs-scanner -p rustfs \
-E 'binary(lifecycle_integration_test) or (package(rustfs) and test(lifecycle_transition_api_test))'
# s3s-e2e black box
./scripts/e2e-run.sh ./target/debug/rustfs /tmp/rustfs-e2e-data
```
**Stale binary.** Tests build the `rustfs` binary once and reuse it. To avoid
rebuilding while iterating on tests, `common.rs` reuses an existing binary when
running *inside* the e2e test process even if sources changed
(`can_reuse_inside_e2e`, [`src/common.rs`](src/common.rs) line 98). Downside: if
you changed **server** code, force a rebuild with
`cargo build -p rustfs` (or `touch` a source file outside the reuse window)
before re-running, or CI's freshly built artifact will diverge from your local
one.
The full lane also requires the startup-CAS build manifest generated by the `Build debug binary` step in `.github/workflows/ci.yml`. Preserve that binary and both sidecars as its `Preserve startup CAS binary input` step does, and use the same `RUSTFS_E2E_STARTUP_CAS_*` environment as `Run e2e full suite`. A generic local build alone does not supply that fixture evidence.
**Stale or unverified binary.** Re-run the matching `build` command after changing source or features, then invoke tests through `run`. A missing receipt, copied old executable, or mismatched build identity is a prerequisite failure. Bare Cargo invocations that start a server deliberately fail; unit tests that do not start a server can still run directly.
**Port already in use / orphan processes.** A hard-killed run can leak a
`rustfs` child holding its port. Find and kill it:
@@ -251,7 +252,8 @@ spawn error. Install the pinned CI version before running their profiles.
A subset of this crate runs on every PR via the `e2e-tests` job:
```bash
cargo nextest run --profile e2e-smoke -p e2e_test
python3 scripts/e2e_binary.py build --features e2e-test-hooks
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-smoke -p e2e_test
```
The selection lives in `.config/nextest.toml` under `[profile.e2e-smoke]`
+123 -149
View File
@@ -31,7 +31,6 @@ use rustfs_signer::constants::UNSIGNED_PAYLOAD;
use rustfs_signer::sign_v4;
use s3s::Body;
use serde_json;
use std::ffi::OsStr;
use std::fs as stdfs;
use std::io::ErrorKind;
use std::net::SocketAddr;
@@ -44,7 +43,6 @@ use tokio::net::TcpStream;
use tokio::time::sleep;
use tracing::{error, info, warn};
use uuid::Uuid;
use walkdir::WalkDir;
// Common constants for all E2E tests
pub const DEFAULT_ACCESS_KEY: &str = "rustfsadmin";
@@ -437,59 +435,75 @@ fn resolve_rustfs_binary_path(workspace: &Path, configured_target_dir: Option<&P
path
}
/// Resolve the RustFS binary relative to the workspace, optionally requesting build features.
/// Resolve the server verified by `scripts/e2e_binary.py run` for this test invocation.
/// Requested features are a required subset of the server's resolved Cargo features.
pub fn rustfs_binary_path_with_features(requested_features: Option<&str>) -> PathBuf {
if let Some(path) = std::env::var_os("CARGO_BIN_EXE_rustfs") {
return PathBuf::from(path);
}
let requested_features = requested_features.and_then(normalize_rustfs_build_features);
let workspace = workspace_root();
let configured_target_dir = std::env::var_os("CARGO_TARGET_DIR").map(PathBuf::from);
let binary_path = resolve_rustfs_binary_path(&workspace, configured_target_dir.as_deref());
let binary_path = std::env::var_os("CARGO_BIN_EXE_rustfs")
.map(PathBuf::from)
.unwrap_or_else(|| resolve_rustfs_binary_path(&workspace, configured_target_dir.as_deref()));
let receipt_path = std::env::var_os("RUSTFS_E2E_BINARY_RECEIPT").map(PathBuf::from);
receipt_path
.ok_or_else(|| std::io::Error::new(ErrorKind::NotFound, "missing E2E run receipt"))
.and_then(|receipt| verify_e2e_binary_receipt(&receipt, &workspace, &binary_path, requested_features))
.unwrap_or_else(|error| {
panic!(
"E2E server prerequisite failed: {error}. Build with `python3 scripts/e2e_binary.py build --features <features>` and run tests with `python3 scripts/e2e_binary.py run --features <features> -- cargo nextest run ...`"
)
})
}
let features_match = binary_features_match(&binary_path, requested_features.as_deref());
let source_is_newer = workspace_sources_newer_than_binary(&binary_path);
let can_reuse_inside_e2e = running_inside_e2e_test_binary() && requested_features.is_none() && features_match;
if binary_path.is_file() && features_match && (!source_is_newer || can_reuse_inside_e2e) {
if source_is_newer {
warn!(
"RustFS binary at {:?} appears older than workspace sources; reusing it inside cargo test to avoid nested builds",
binary_path
);
}
info!("Using existing RustFS binary at {:?}", binary_path);
return binary_path;
#[derive(serde::Deserialize)]
#[serde(deny_unknown_fields)]
struct E2eBinaryReceipt {
schema: u32,
workspace: PathBuf,
binary: PathBuf,
size: u64,
modified_ns: u128,
features: Vec<String>,
}
fn verify_e2e_binary_receipt(
receipt_path: &Path,
workspace: &Path,
binary_path: &Path,
requested_features: Option<&str>,
) -> std::io::Result<PathBuf> {
let receipt: E2eBinaryReceipt = serde_json::from_slice(&stdfs::read(receipt_path)?)?;
let binary = binary_path.canonicalize()?;
let metadata = binary.metadata()?;
let modified_ns = metadata
.modified()?
.duration_since(std::time::UNIX_EPOCH)
.map_err(std::io::Error::other)?
.as_nanos();
// The runner hashes source and binary before/after the entire suite. Each
// nextest process checks only this invocation's path, features, and file stat.
if receipt.schema != 1
|| receipt.workspace != workspace.canonicalize()?
|| receipt.binary != binary
|| !metadata.is_file()
|| receipt.size != metadata.len()
|| receipt.modified_ns != modified_ns
{
return Err(std::io::Error::new(
ErrorKind::InvalidData,
"E2E server differs from this run's verified binary",
));
}
info!("Building RustFS binary to ensure it's up to date...");
build_rustfs_binary(requested_features.as_deref(), &binary_path);
info!("Using RustFS binary at {:?}", binary_path);
binary_path
}
fn workspace_sources_newer_than_binary(binary_path: &PathBuf) -> bool {
let Ok(binary_meta) = std::fs::metadata(binary_path) else {
return true;
};
let Ok(binary_modified) = binary_meta.modified() else {
return true;
};
let workspace = workspace_root();
let watch_roots = [
workspace.join("Cargo.toml"),
workspace.join("Cargo.lock"),
workspace.join("rustfs"),
workspace.join("crates"),
];
watch_roots.iter().any(|path| path_is_newer_than(binary_modified, path))
}
fn running_inside_e2e_test_binary() -> bool {
std::env::var("CARGO_PKG_NAME").is_ok_and(|value| value == "e2e_test")
if let Some(requested) = requested_features.and_then(normalize_rustfs_build_features)
&& requested
.split(',')
.any(|feature| !receipt.features.iter().any(|actual| actual == feature))
{
return Err(std::io::Error::new(
ErrorKind::InvalidInput,
"E2E server is missing a requested build feature",
));
}
Ok(binary)
}
pub fn requested_rustfs_build_features() -> Option<String> {
@@ -519,96 +533,6 @@ pub fn rustfs_build_feature_enabled(requested_features: Option<&str>, required_f
.any(|feature| feature.eq_ignore_ascii_case(RUSTFS_FULL_FEATURE) || feature.eq_ignore_ascii_case(required_feature))
}
fn rustfs_binary_features_stamp_path(binary_path: &Path) -> PathBuf {
binary_path.with_extension("features")
}
fn binary_features_match(binary_path: &Path, requested_features: Option<&str>) -> bool {
let stamp_path = rustfs_binary_features_stamp_path(binary_path);
let recorded = stdfs::read_to_string(stamp_path)
.ok()
.and_then(|value| normalize_rustfs_build_features(&value));
let requested = requested_features.and_then(normalize_rustfs_build_features);
match requested.as_deref() {
Some(features) => recorded.as_deref() == Some(features),
None => recorded.is_none(),
}
}
fn path_is_newer_than(binary_modified: std::time::SystemTime, path: &Path) -> bool {
if path.is_file() {
return std::fs::metadata(path)
.and_then(|meta| meta.modified())
.map(|modified| modified > binary_modified)
.unwrap_or(false);
}
if !path.is_dir() {
return false;
}
WalkDir::new(path)
.into_iter()
.filter_entry(|entry| {
let name = entry.file_name();
name != OsStr::new("target") && name != OsStr::new(".git")
})
.filter_map(Result::ok)
.filter(|entry| entry.file_type().is_file())
.any(|entry| {
std::fs::metadata(entry.path())
.and_then(|meta| meta.modified())
.map(|modified| modified > binary_modified)
.unwrap_or(false)
})
}
/// Build the RustFS binary using cargo
fn build_rustfs_binary(requested_features: Option<&str>, binary_path: &Path) {
let workspace = workspace_root();
info!("Building RustFS binary from workspace: {:?}", workspace);
let _profile = if cfg!(debug_assertions) {
info!("Building in debug mode");
"dev"
} else {
info!("Building in release mode");
"release"
};
let mut cmd = Command::new("cargo");
cmd.current_dir(&workspace).args(["build", "--bin", "rustfs"]);
if let Some(features) = requested_features {
cmd.arg("--features").arg(features);
info!("Building with features: {}", features);
}
if !cfg!(debug_assertions) {
cmd.arg("--release");
}
info!(
"Executing: cargo build --bin rustfs {}",
if cfg!(debug_assertions) { "" } else { "--release" }
);
let output = cmd.output().expect("Failed to execute cargo build command");
if !output.status.success() {
let stderr = String::from_utf8_lossy(&output.stderr);
panic!("Failed to build RustFS binary. Error: {stderr}");
}
let stamp_path = rustfs_binary_features_stamp_path(binary_path);
if let Err(err) = stdfs::write(&stamp_path, requested_features.unwrap_or_default()) {
warn!("Failed to write RustFS feature stamp {:?}: {}", stamp_path, err);
}
info!("✅ RustFS binary built successfully");
}
fn awscurl_binary_path() -> PathBuf {
std::env::var_os("AWSCURL_PATH")
.map(PathBuf::from)
@@ -2255,16 +2179,66 @@ mod tests {
}
#[test]
fn binary_feature_stamp_matching_uses_normalized_features() {
let binary_path = std::env::temp_dir().join(format!("rustfs-feature-stamp-test-{}", Uuid::new_v4()));
let stamp_path = rustfs_binary_features_stamp_path(&binary_path);
fn explicit_binary_without_run_receipt_is_rejected() {
const CHILD_ENV: &str = "RUSTFS_E2E_RECEIPT_TEST_CHILD";
if std::env::var_os(CHILD_ENV).is_some() {
rustfs_binary_path_with_features(None);
return;
}
let executable = std::env::current_exe().expect("locate isolated test process");
let output = Command::new(&executable)
.args([
"--exact",
"common::tests::explicit_binary_without_run_receipt_is_rejected",
"--nocapture",
])
.env(CHILD_ENV, "1")
.env("CARGO_BIN_EXE_rustfs", &executable)
.env_remove("RUSTFS_E2E_BINARY_RECEIPT")
.output()
.expect("run the missing-receipt scenario with isolated environment variables");
assert!(!output.status.success(), "an explicit binary must not bypass run verification");
assert!(String::from_utf8_lossy(&output.stderr).contains("missing E2E run receipt"));
}
stdfs::write(&stamp_path, " SFTP, ftps ").expect("write feature stamp");
assert!(binary_features_match(&binary_path, Some("sftp,ftps")));
assert!(binary_features_match(&binary_path, Some(" SFTP, FTPS ")));
assert!(!binary_features_match(&binary_path, Some("sftp")));
stdfs::remove_file(stamp_path).ok();
#[test]
fn e2e_run_receipt_rejects_replaced_binary_and_missing_features() {
let directory = std::env::temp_dir().join(format!("rustfs-e2e-receipt-test-{}", Uuid::new_v4()));
stdfs::create_dir(&directory).expect("create receipt fixture");
let binary = directory.join("rustfs");
let receipt = directory.join("receipt.json");
stdfs::write(&binary, "server").expect("write fixture binary");
let metadata = binary.metadata().expect("stat fixture binary");
let record = serde_json::json!({
"schema": 1,
"workspace": directory.canonicalize().expect("canonical workspace"),
"binary": binary.canonicalize().expect("canonical binary"),
"size": metadata.len(),
"modified_ns": metadata.modified().expect("modified time").duration_since(std::time::UNIX_EPOCH).expect("positive timestamp").as_nanos(),
"features": ["default", "full", "ftps", "webdav", "sftp"]
});
stdfs::write(&receipt, serde_json::to_vec(&record).expect("serialize receipt")).expect("write receipt");
verify_e2e_binary_receipt(&receipt, &directory, &binary, Some("sftp,webdav")).expect("resolved feature subset");
verify_e2e_binary_receipt(&receipt, &directory, &binary, Some("full")).expect("full was actually requested");
assert_eq!(
verify_e2e_binary_receipt(&receipt, &directory, &binary, Some("rio-v2"))
.expect_err("full does not enable rio-v2")
.kind(),
ErrorKind::InvalidInput
);
let other = directory.join("old-server");
stdfs::write(&other, "server").expect("write alternate binary");
assert!(verify_e2e_binary_receipt(&receipt, &directory, &other, None).is_err());
stdfs::write(&binary, "different server").expect("replace fixture binary");
assert!(verify_e2e_binary_receipt(&receipt, &directory, &binary, None).is_err());
stdfs::remove_file(&receipt).expect("remove expired receipt");
assert_eq!(
verify_e2e_binary_receipt(&receipt, &directory, &binary, None)
.expect_err("expired receipt")
.kind(),
ErrorKind::NotFound
);
stdfs::remove_dir_all(directory).expect("remove receipt fixture");
}
/// Build a cluster environment struct in-memory (no ports, no processes) so
@@ -0,0 +1,656 @@
// Copyright 2026 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Standard S3 deletion permissions and the explicit recursive-delete extension.
use crate::common::{
AdminTransport, RustFSTestEnvironment, admin_add_canned_policy_via, admin_attach_user_policy_via, admin_create_user,
init_logging,
};
use aws_sdk_s3::Client;
use aws_sdk_s3::error::{ProvideErrorMetadata, SdkError};
use aws_sdk_s3::operation::delete_object::{DeleteObjectError, DeleteObjectOutput};
use aws_sdk_s3::primitives::ByteStream;
use aws_sdk_s3::types::{BucketVersioningStatus, Delete, ObjectIdentifier, VersioningConfiguration};
use futures::{StreamExt, TryStreamExt, stream};
use serde_json::{Value, json};
use std::collections::BTreeSet;
use std::error::Error;
use uuid::Uuid;
type TestResult<T = ()> = Result<T, Box<dyn Error + Send + Sync>>;
type VersionSnapshot = BTreeSet<(String, String, bool)>;
async fn set_policy(env: &RustFSTestEnvironment, name: &str, policy: &Value) -> TestResult {
admin_add_canned_policy_via(
AdminTransport::Signed,
&env.url,
&env.access_key,
&env.secret_key,
name,
&policy.to_string(),
)
.await
}
async fn policy_user(env: &RustFSTestEnvironment, policy_name: &str, policy: Option<Value>) -> TestResult<Client> {
let username = Uuid::new_v4().simple().to_string();
let secret = Uuid::new_v4().simple().to_string();
admin_create_user(env, &username, &secret).await?;
if let Some(policy) = policy {
set_policy(env, policy_name, &policy).await?;
}
admin_attach_user_policy_via(AdminTransport::Signed, &env.url, &env.access_key, &env.secret_key, policy_name, &username)
.await?;
Ok(env.create_s3_client_with_credentials(&username, &secret))
}
async fn versioning(client: &Client, bucket: &str, status: BucketVersioningStatus) -> TestResult {
client
.put_bucket_versioning()
.bucket(bucket)
.versioning_configuration(VersioningConfiguration::builder().status(status).build())
.send()
.await?;
Ok(())
}
async fn put(client: &Client, bucket: &str, key: &str) -> TestResult<String> {
let result = client
.put_object()
.bucket(bucket)
.key(key)
.body(ByteStream::from_static(b"delete authorization fixture"))
.send()
.await?;
Ok(result.version_id().unwrap_or("null").to_string())
}
async fn versions(client: &Client, bucket: &str, prefix: &str) -> TestResult<VersionSnapshot> {
let mut result = BTreeSet::new();
let mut markers = (None, None);
loop {
let page = client
.list_object_versions()
.bucket(bucket)
.prefix(prefix)
.set_key_marker(markers.0.clone())
.set_version_id_marker(markers.1.clone())
.send()
.await?;
for version in page.versions() {
result.insert((
version.key().ok_or("listed version missing key")?.to_string(),
version.version_id().ok_or("listed version missing ID")?.to_string(),
false,
));
}
for marker in page.delete_markers() {
result.insert((
marker.key().ok_or("listed delete marker missing key")?.to_string(),
marker.version_id().ok_or("listed delete marker missing ID")?.to_string(),
true,
));
}
if page.is_truncated() != Some(true) {
return Ok(result);
}
let next = (
Some(
page.next_key_marker()
.ok_or("truncated versions page missing next key marker")?
.to_string(),
),
page.next_version_id_marker().map(str::to_string),
);
assert_ne!(markers, next, "ListObjectVersions pagination must advance");
markers = next;
}
}
async fn force_delete(
client: &Client,
bucket: &str,
prefix: &str,
) -> Result<DeleteObjectOutput, Box<SdkError<DeleteObjectError>>> {
client
.delete_object()
.bucket(bucket)
.key(prefix)
.customize()
.mutate_request(|request| {
request.headers_mut().insert("x-rustfs-force-delete", "true");
})
.send()
.await
.map_err(Box::new)
}
async fn replica_force_delete(
client: &Client,
bucket: &str,
prefix: &str,
) -> Result<DeleteObjectOutput, Box<SdkError<DeleteObjectError>>> {
client
.delete_object()
.bucket(bucket)
.key(prefix)
.customize()
.mutate_request(|request| {
request.headers_mut().insert("x-rustfs-force-delete", "true");
request.headers_mut().insert("x-amz-replication-status", "REPLICA");
})
.send()
.await
.map_err(Box::new)
}
fn assert_denied<T, E>(result: Result<T, SdkError<E>>)
where
T: std::fmt::Debug,
E: ProvideErrorMetadata + std::fmt::Debug,
{
let error = result.expect_err("request must be denied by its S3 permission");
assert_eq!(
error.as_service_error().and_then(ProvideErrorMetadata::code),
Some("AccessDenied"),
"expected an S3 authorization denial, got {error:?}"
);
}
fn assert_boxed_denied<T, E>(result: Result<T, Box<SdkError<E>>>)
where
T: std::fmt::Debug,
E: ProvideErrorMetadata + std::fmt::Debug,
{
assert_denied(result.map_err(|error| *error));
}
#[tokio::test]
async fn sdk_version_deletion_requires_only_delete_object_version() -> TestResult {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?;
let root = env.create_s3_client();
let bucket = "delete-version-permissions";
root.create_bucket().bucket(bucket).send().await?;
put(&root, bucket, "single-null.txt").await?;
put(&root, bucket, "batch-null.txt").await?;
versioning(&root, bucket, BucketVersioningStatus::Enabled).await?;
let old = put(&root, bucket, "single.txt").await?;
let current = put(&root, bucket, "single.txt").await?;
let batch_version = put(&root, bucket, "batch.txt").await?;
let ordinary_version = put(&root, bucket, "ordinary.txt").await?;
let user = policy_user(
&env,
"version-deleter",
Some(json!({"Version":"2012-10-17","Statement":[
{"Effect":"Allow","Action":"s3:DeleteObjectVersion","Resource":format!("arn:aws:s3:::{bucket}/*")},
{"Effect":"Deny","Action":"s3:DeleteObject","Resource":format!("arn:aws:s3:::{bucket}/*")}
]})),
)
.await?;
user.delete_object()
.bucket(bucket)
.key("single.txt")
.version_id(&old)
.send()
.await?;
user.delete_object()
.bucket(bucket)
.key("single-null.txt")
.version_id("null")
.send()
.await?;
assert_denied(user.delete_object().bucket(bucket).key("ordinary.txt").send().await);
let batch = user
.delete_objects()
.bucket(bucket)
.delete(
Delete::builder()
.objects(
ObjectIdentifier::builder()
.key("batch.txt")
.version_id(&batch_version)
.build()?,
)
.objects(ObjectIdentifier::builder().key("batch-null.txt").version_id("null").build()?)
.objects(ObjectIdentifier::builder().key("ordinary.txt").build()?)
.build()?,
)
.send()
.await?;
assert_eq!(batch.deleted().len(), 2, "both explicit version items must succeed");
assert_eq!(batch.errors().len(), 1, "only the unversioned item must be denied");
assert_eq!(batch.errors()[0].key(), Some("ordinary.txt"));
assert_eq!(batch.errors()[0].code(), Some("AccessDenied"));
assert_eq!(
versions(&root, bucket, "").await?,
BTreeSet::from([
("single.txt".into(), current, false),
("ordinary.txt".into(), ordinary_version, false)
]),
"version-only deletion must preserve the current single-object version and denied object"
);
Ok(())
}
#[tokio::test]
async fn sdk_list_bucket_and_list_bucket_versions_permissions_are_independent() -> TestResult {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?;
let root = env.create_s3_client();
let bucket = "list-version-permissions";
root.create_bucket().bucket(bucket).send().await?;
versioning(&root, bucket, BucketVersioningStatus::Enabled).await?;
put(&root, bucket, "visible.txt").await?;
for (action, name) in [
("s3:ListBucket", "object-lister"),
("s3:ListBucketVersions", "version-lister"),
] {
let user = policy_user(
&env,
name,
Some(json!({"Version":"2012-10-17","Statement":[
{"Effect":"Allow","Action":action,"Resource":format!("arn:aws:s3:::{bucket}")}
]})),
)
.await?;
if action == "s3:ListBucket" {
assert_eq!(user.list_objects_v2().bucket(bucket).send().await?.contents().len(), 1);
assert_denied(user.list_object_versions().bucket(bucket).send().await);
} else {
assert_eq!(user.list_object_versions().bucket(bucket).send().await?.versions().len(), 1);
assert_denied(user.list_objects_v2().bucket(bucket).send().await);
}
}
Ok(())
}
#[tokio::test]
async fn console_admin_force_delete_removes_prefix_versions_and_delete_markers() -> TestResult {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?;
let root = env.create_s3_client();
let bucket = "force-console-admin";
root.create_bucket().bucket(bucket).send().await?;
put(&root, bucket, "folder/null.txt").await?;
versioning(&root, bucket, BucketVersioningStatus::Enabled).await?;
for key in ["folder/a.txt", "folder/deep/b.txt", "single.txt"] {
put(&root, bucket, key).await?;
put(&root, bucket, key).await?;
root.delete_object().bucket(bucket).key(key).send().await?;
}
put(&root, bucket, "keep.txt").await?;
put(&root, bucket, "folder-sibling/keep.txt").await?;
let keep = versions(&root, bucket, "keep.txt").await?;
let sibling = versions(&root, bucket, "folder-sibling/").await?;
let user = policy_user(&env, "consoleAdmin", None).await?;
force_delete(&user, bucket, "folder/").await?;
assert!(
versions(&root, bucket, "folder/").await?.is_empty(),
"force prefix deletion must remove null versions and markers"
);
assert_eq!(versions(&root, bucket, "folder-sibling/").await?, sibling);
assert_eq!(
versions(&root, bucket, "single.txt").await?.len(),
3,
"the separate key must survive folder deletion"
);
force_delete(&user, bucket, "single.txt").await?;
assert!(
versions(&root, bucket, "single.txt").await?.is_empty(),
"explicit force deletion must remove every version of the selected key"
);
assert_eq!(versions(&root, bucket, "keep.txt").await?, keep);
Ok(())
}
#[tokio::test]
async fn force_delete_authorizes_only_its_path_scope_without_list_permissions() -> TestResult {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?;
let root = env.create_s3_client();
let bucket = "force-delete-only";
root.create_bucket().bucket(bucket).send().await?;
versioning(&root, bucket, BucketVersioningStatus::Enabled).await?;
put(&root, bucket, "selected.txt").await?;
put(&root, bucket, "selected.txt/child.txt").await?;
put(&root, bucket, "selected.txt-sibling").await?;
let sibling = versions(&root, bucket, "selected.txt-sibling").await?;
let user = policy_user(
&env,
"delete-only",
Some(json!({"Version":"2012-10-17","Statement":[
{"Effect":"Allow","Action":["s3:DeleteObject","s3:DeleteObjectVersion"],"Resource":[
format!("arn:aws:s3:::{bucket}/selected.txt"), format!("arn:aws:s3:::{bucket}/selected.txt/*")
]},
{"Effect":"Deny","Action":["s3:DeleteObject","s3:DeleteObjectVersion"],"Resource":format!("arn:aws:s3:::{bucket}/selected.txt-sibling")}
]})),
)
.await?;
assert_denied(user.list_objects_v2().bucket(bucket).send().await);
assert_denied(user.list_object_versions().bucket(bucket).send().await);
force_delete(&user, bucket, "selected.txt").await?;
assert_eq!(
versions(&root, bucket, "selected.txt").await?,
sibling,
"force deletion must remove the selected path and descendants without authorizing or deleting its similarly prefixed sibling"
);
Ok(())
}
#[tokio::test]
async fn force_directory_delete_cannot_remove_an_unauthorized_colliding_parent() -> TestResult {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?;
let root = env.create_s3_client();
let bucket = "force-directory-collision";
root.create_bucket().bucket(bucket).send().await?;
versioning(&root, bucket, BucketVersioningStatus::Enabled).await?;
let protected_parent_version = put(&root, bucket, "collision.txt").await?;
for key in ["collision.txt/child", "collision.txt-sibling"] {
put(&root, bucket, key).await?;
}
put(&root, bucket, "collision.txt").await?;
root.delete_object().bucket(bucket).key("collision.txt").send().await?;
let user = policy_user(
&env,
"parent-denier",
Some(json!({"Version":"2012-10-17","Statement":[
{"Effect":"Allow","Action":["s3:DeleteObject","s3:DeleteObjectVersion"],"Resource":format!("arn:aws:s3:::{bucket}/*")},
{"Effect":"Deny","Action":"s3:DeleteObjectVersion","Resource":format!("arn:aws:s3:::{bucket}/collision.txt"),
"Condition":{"StringEquals":{"s3:VersionId":protected_parent_version}}}
]})),
)
.await?;
let mut expected = versions(&root, bucket, "").await?;
expected.retain(|(key, _, _)| key != "collision.txt/child");
force_delete(&user, bucket, "collision.txt/").await?;
assert_eq!(
versions(&root, bucket, "").await?,
expected,
"folder deletion must preserve the denied parent's historical versions and delete marker, plus its sibling"
);
Ok(())
}
#[tokio::test]
async fn force_unversioned_directory_requires_only_delete_object() -> TestResult {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?;
let root = env.create_s3_client();
let bucket = "force-unversioned-permissions";
root.create_bucket().bucket(bucket).send().await?;
for key in ["folder/", "folder/child.txt", "outside.txt"] {
put(&root, bucket, key).await?;
}
let outside = versions(&root, bucket, "outside.txt").await?;
let user = policy_user(
&env,
"unversioned-deleter",
Some(json!({"Version":"2012-10-17","Statement":[
{"Effect":"Allow","Action":"s3:DeleteObject","Resource":format!("arn:aws:s3:::{bucket}/*")},
{"Effect":"Deny","Action":"s3:DeleteObjectVersion","Resource":format!("arn:aws:s3:::{bucket}/*")}
]})),
)
.await?;
force_delete(&user, bucket, "folder/").await?;
assert_eq!(
versions(&root, bucket, "").await?,
outside,
"unversioned force deletion, including a synthetic nil directory marker, must use DeleteObject permission"
);
Ok(())
}
#[tokio::test]
async fn force_delete_denied_child_preserves_every_object_despite_bucket_allow() -> TestResult {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?;
let root = env.create_s3_client();
let bucket = "force-child-denial";
root.create_bucket().bucket(bucket).send().await?;
for key in ["folder/a-allowed.txt", "folder/z-denied.txt"] {
put(&root, bucket, key).await?;
}
let user = policy_user(
&env,
"child-denier",
Some(json!({"Version":"2012-10-17","Statement":[
{"Effect":"Allow","Action":["s3:DeleteObject","s3:DeleteObjectVersion","s3:ReplicateDelete"],"Resource":format!("arn:aws:s3:::{bucket}/*")},
{"Effect":"Deny","Action":["s3:DeleteObject","s3:DeleteObjectVersion","s3:ReplicateDelete"],"Resource":format!("arn:aws:s3:::{bucket}/folder/z-denied.txt")}
]})),
)
.await?;
root.put_bucket_policy()
.bucket(bucket)
.policy(json!({"Version":"2012-10-17","Statement":[
{"Effect":"Allow","Principal":"*","Action":["s3:DeleteObject","s3:DeleteObjectVersion","s3:ReplicateDelete"],"Resource":format!("arn:aws:s3:::{bucket}/*")}
]}).to_string())
.send().await?;
let before = versions(&root, bucket, "folder/").await?;
assert_boxed_denied(force_delete(&user, bucket, "folder/").await);
assert_eq!(
versions(&root, bucket, "folder/").await?,
before,
"a denied descendant must prevent every mutation in the force scope"
);
assert_boxed_denied(replica_force_delete(&user, bucket, "folder/").await);
assert_eq!(
versions(&root, bucket, "folder/").await?,
before,
"the REPLICA header must not bypass a descendant's ReplicateDelete denial"
);
root.delete_bucket_policy().bucket(bucket).send().await?;
let replica_user = policy_user(
&env,
"replica-deleter",
Some(json!({"Version":"2012-10-17","Statement":[
{"Effect":"Allow","Action":"s3:DeleteObject","Resource":format!("arn:aws:s3:::{bucket}/*")},
{"Effect":"Allow","Action":"s3:ReplicateDelete","Resource":format!("arn:aws:s3:::{bucket}/*")}
]})),
)
.await?;
replica_force_delete(&replica_user, bucket, "folder/").await?;
assert!(
versions(&root, bucket, "folder/").await?.is_empty(),
"an authorized replica force request must check ReplicateDelete for its descendants"
);
Ok(())
}
#[tokio::test]
async fn force_delete_denied_historical_version_preserves_versions_and_markers() -> TestResult {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?;
let root = env.create_s3_client();
let bucket = "force-version-denial";
root.create_bucket().bucket(bucket).send().await?;
put(&root, bucket, "folder/null.txt").await?;
versioning(&root, bucket, BucketVersioningStatus::Enabled).await?;
let protected_version = put(&root, bucket, "folder/versioned.txt").await?;
put(&root, bucket, "folder/versioned.txt").await?;
let marker = root.delete_object().bucket(bucket).key("folder/versioned.txt").send().await?;
let marker_version = marker
.version_id()
.ok_or("versioned delete must return a marker version ID")?;
put(&root, bucket, "folder/a-allowed.txt").await?;
let policy = |version: &str| {
json!({"Version":"2012-10-17","Statement":[
{"Effect":"Allow","Action":["s3:DeleteObject","s3:DeleteObjectVersion"],"Resource":format!("arn:aws:s3:::{bucket}/*")},
{"Effect":"Deny","Action":"s3:DeleteObjectVersion","Resource":format!("arn:aws:s3:::{bucket}/folder/*"),
"Condition":{"StringEquals":{"s3:VersionId":version}}}
]})
};
let user = policy_user(&env, "version-denier", Some(policy(&protected_version))).await?;
let before = versions(&root, bucket, "folder/").await?;
for version in [protected_version.as_str(), "null", marker_version] {
set_policy(&env, "version-denier", &policy(version)).await?;
assert_boxed_denied(force_delete(&user, bucket, "folder/").await);
assert_eq!(
versions(&root, bucket, "folder/").await?,
before,
"denial of a historical, null, or delete-marker version must prevent recursive deletion"
);
}
Ok(())
}
#[tokio::test]
async fn sdk_ordinary_deletion_preserves_versions_and_directory_children() -> TestResult {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?;
let root = env.create_s3_client();
let user = policy_user(
&env,
"ordinary-deleter",
Some(json!({"Version":"2012-10-17","Statement":[
{"Effect":"Allow","Action":"s3:DeleteObject","Resource":"arn:aws:s3:::*/*"},
{"Effect":"Deny","Action":"s3:DeleteObjectVersion","Resource":"arn:aws:s3:::*/*"}
]})),
)
.await?;
for state in ["unversioned", "enabled", "suspended"] {
let bucket = format!("ordinary-directory-{state}");
root.create_bucket().bucket(&bucket).send().await?;
if state != "unversioned" {
versioning(&root, &bucket, BucketVersioningStatus::Enabled).await?;
}
let historical = put(&root, &bucket, "object.txt").await?;
if state == "suspended" {
versioning(&root, &bucket, BucketVersioningStatus::Suspended).await?;
put(&root, &bucket, "object.txt").await?;
}
put(&root, &bucket, "folder/").await?;
put(&root, &bucket, "folder/child.txt").await?;
let child_before = versions(&root, &bucket, "folder/child.txt").await?;
user.delete_object().bucket(&bucket).key("folder/").send().await?;
assert_eq!(
versions(&root, &bucket, "folder/").await?,
child_before,
"ordinary {state} directory-key deletion must remove only its synthetic marker and preserve children"
);
let deleted = user.delete_object().bucket(&bucket).key("object.txt").send().await?;
let object_versions = versions(&root, &bucket, "object.txt").await?;
if state == "unversioned" {
assert!(object_versions.is_empty());
} else {
assert_eq!(deleted.delete_marker(), Some(true));
assert_eq!(
object_versions.len(),
2,
"ordinary {state} deletion must retain its historical data version"
);
assert!(object_versions.contains(&("object.txt".into(), historical, false)));
if state == "suspended" {
assert!(object_versions.contains(&("object.txt".into(), "null".into(), true)));
}
}
}
Ok(())
}
#[tokio::test]
async fn sdk_delete_objects_force_header_keeps_explicit_item_scope() -> TestResult {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?;
let root = env.create_s3_client();
let bucket = "batch-force-explicit-scope";
root.create_bucket().bucket(bucket).send().await?;
put(&root, bucket, "folder/").await?;
put(&root, bucket, "folder/child.txt").await?;
let child = versions(&root, bucket, "folder/child.txt").await?;
let user = policy_user(&env, "consoleAdmin", None).await?;
let result = user
.delete_objects()
.bucket(bucket)
.delete(
Delete::builder()
.objects(ObjectIdentifier::builder().key("folder/").build()?)
.build()?,
)
.customize()
.mutate_request(|request| {
request.headers_mut().insert("x-rustfs-force-delete", "true");
})
.send()
.await?;
assert!(result.errors().is_empty());
assert_eq!(result.deleted().len(), 1);
assert_eq!(
versions(&root, bucket, "folder/").await?,
child,
"batch deletion must remove only the explicit directory marker even with the force header"
);
Ok(())
}
#[tokio::test]
async fn force_delete_checks_every_version_page_before_mutation() -> TestResult {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?;
let root = env.create_s3_client();
let bucket = "force-delete-pagination";
root.create_bucket().bucket(bucket).send().await?;
versioning(&root, bucket, BucketVersioningStatus::Enabled).await?;
stream::iter(0..1000)
.map(|index| {
let root = &root;
async move { put(root, bucket, &format!("folder/{index:04}.txt")).await.map(|_| ()) }
})
.buffer_unordered(16)
.try_collect::<Vec<_>>()
.await?;
put(&root, bucket, "folder/z-denied.txt").await?;
let allow = json!({"Effect":"Allow","Action":["s3:DeleteObject","s3:DeleteObjectVersion"],"Resource":format!("arn:aws:s3:::{bucket}/*")});
let user = policy_user(
&env,
"paged-deleter",
Some(json!({"Version":"2012-10-17","Statement":[allow.clone(),
{"Effect":"Deny","Action":"s3:DeleteObjectVersion","Resource":format!("arn:aws:s3:::{bucket}/folder/z-denied.txt")}
]})),
)
.await?;
let before = versions(&root, bucket, "folder/").await?;
assert_eq!(before.len(), 1001, "the denied key must be beyond one default versions page");
assert_boxed_denied(force_delete(&user, bucket, "folder/").await);
assert_eq!(
versions(&root, bucket, "folder/").await?,
before,
"a denial on the second page must preserve the first page too"
);
set_policy(&env, "paged-deleter", &json!({"Version":"2012-10-17","Statement":[allow]})).await?;
force_delete(&user, bucket, "folder/").await?;
assert!(
versions(&root, bucket, "folder/").await?.is_empty(),
"authorized recursive deletion must cover all pages"
);
Ok(())
}
+1
View File
@@ -28,6 +28,7 @@ mod harness;
mod heal_test;
mod object_lock_test;
mod observability_test;
mod replication_delete_marker_test;
mod replication_quota_test;
mod s3_basic_test;
mod s3_during_data_movement_test;
@@ -0,0 +1,137 @@
// Copyright 2026 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Functional REP-105 (rustfs/backlog#2195 item 4): a delete marker created
//! on a multi-node source cluster must replicate to the bucket-replication
//! target. Objects converged in seconds while delete markers did not arrive
//! within 180 s on the shared 3-node functional environment; the single-node
//! e2e never saw it.
use super::harness::{
DistCluster, DistLayout, TestResult, enable_versioning, put_bucket_replication, put_object, set_remote_target, unique_bucket,
wait_for_replicated_bytes, wait_until,
};
use crate::common::{FAST_DATA_USAGE_SCANNER_ENV, RustFSTestEnvironment, init_logging, replication_fast_env, signed_request};
use crate::replication_extension_test::LOOPBACK_REPLICATION_TARGET_ENV;
use aws_sdk_s3::Client;
use http::{Method, StatusCode};
use std::time::Duration;
async fn target_has_delete_marker(client: &Client, bucket: &str, key: &str) -> TestResult<bool> {
let versions = client.list_object_versions().bucket(bucket).prefix(key).send().await?;
Ok(versions.delete_markers().iter().any(|marker| marker.key() == Some(key)))
}
async fn delete_marker_replicates(
source: &DistCluster,
source_bucket: &str,
target_client: &Client,
target_bucket: &str,
) -> TestResult {
let key = "delete-marker/object.bin";
let body = b"delete marker replication payload".to_vec();
// Write through one node, delete through another: behind a load
// balancer consecutive requests land on different nodes.
put_object(&source.client(1)?, source_bucket, key, body.clone()).await?;
wait_for_replicated_bytes(target_client, target_bucket, key, &body, Duration::from_secs(60)).await?;
let delete = source
.client(2)?
.delete_object()
.bucket(source_bucket)
.key(key)
.send()
.await?;
assert_eq!(
delete.delete_marker(),
Some(true),
"a versioned DELETE without versionId must create a marker"
);
wait_until(
Duration::from_secs(90),
|| async { target_has_delete_marker(target_client, target_bucket, key).await },
"delete marker replicated to the target bucket",
)
.await
}
#[tokio::test]
async fn four_node_bucket_replication_replicates_delete_marker_to_peer_cluster() -> TestResult {
init_logging();
let (source, target) = DistCluster::start_replication_pair().await?;
let source_bucket = unique_bucket("dm-src");
let target_bucket = unique_bucket("dm-dst");
source.create_bucket(&source_bucket).await?;
target.create_bucket(&target_bucket).await?;
enable_versioning(&source.client(0)?, &source_bucket).await?;
enable_versioning(&target.client(0)?, &target_bucket).await?;
let arn = set_remote_target(&source.cluster, &source_bucket, &target.cluster, &target_bucket).await?;
put_bucket_replication(&source.cluster, &source_bucket, &arn).await?;
delete_marker_replicates(&source, &source_bucket, &target.client(0)?, &target_bucket).await
}
/// The functional environment replicates from a 3-node site to a single-node
/// target; keep that shape as its own case.
#[tokio::test]
async fn four_node_bucket_replication_replicates_delete_marker_to_single_node_target() -> TestResult {
init_logging();
let mut extra: Vec<(&str, &str)> = replication_fast_env();
extra.extend_from_slice(LOOPBACK_REPLICATION_TARGET_ENV);
extra.extend_from_slice(FAST_DATA_USAGE_SCANNER_ENV);
let source = DistCluster::start_with_env(DistLayout::FourNodeFourDisk, &extra).await?;
let mut target = RustFSTestEnvironment::new().await?;
target.start_rustfs_server_without_cleanup(vec![]).await?;
let source_bucket = unique_bucket("dm-src");
let target_bucket = unique_bucket("dm-dst");
source.create_bucket(&source_bucket).await?;
let target_client = target.create_s3_client();
target_client.create_bucket().bucket(&target_bucket).send().await?;
enable_versioning(&source.client(0)?, &source_bucket).await?;
enable_versioning(&target_client, &target_bucket).await?;
let body = serde_json::json!({
"endpoint": target.address,
"credentials": { "accessKey": target.access_key, "secretKey": target.secret_key },
"targetbucket": target_bucket,
"secure": false,
"type": "replication"
});
let url = format!(
"{}/rustfs/admin/v3/set-remote-target?bucket={}",
source.cluster.nodes[0].url,
urlencoding::encode(&source_bucket)
);
let response = signed_request(
Method::PUT,
&url,
&source.cluster.access_key,
&source.cluster.secret_key,
Some(body.to_string().into_bytes()),
Some("application/json"),
)
.await?;
if response.status() != StatusCode::OK {
let status = response.status();
let body = response.text().await.unwrap_or_default();
return Err(format!("set remote target failed: {status} {body}").into());
}
let arn: String = serde_json::from_slice(&response.bytes().await?)?;
put_bucket_replication(&source.cluster, &source_bucket, &arn).await?;
delete_marker_replicates(&source, &source_bucket, &target_client, &target_bucket).await
}
@@ -126,3 +126,165 @@ async fn four_node_site_replication_replicates_object_to_peer_site() -> TestResu
wait_for_replicated_bytes(&site_a.client(3)?, &bucket, reverse_key, &reverse_body, Duration::from_secs(60)).await?;
Ok(())
}
async fn node_admin(
cluster: &crate::common::RustFSTestClusterEnvironment,
node_idx: usize,
method: Method,
path_and_query: &str,
body: Option<String>,
) -> TestResult<(StatusCode, String)> {
crate::common::admin_request(
&cluster.nodes[node_idx].url,
method,
path_and_query,
body,
&cluster.access_key,
&cluster.secret_key,
)
.await
}
/// Pair two clusters through site A's first node and wait until both report
/// the two-site topology as enabled.
async fn pair_sites(site_a: &DistCluster, site_b: &DistCluster) -> TestResult {
let sites = vec![
PeerSite {
name: "site-a".to_string(),
endpoint: site_a.cluster.nodes[0].url.clone(),
access_key: site_a.cluster.access_key.clone(),
secret_key: site_a.cluster.secret_key.clone(),
..Default::default()
},
PeerSite {
name: "site-b".to_string(),
endpoint: site_b.cluster.nodes[0].url.clone(),
access_key: site_b.cluster.access_key.clone(),
secret_key: site_b.cluster.secret_key.clone(),
..Default::default()
},
];
let add_status = site_replication_add(&site_a.cluster, &sites).await?;
assert!(
add_status.success && add_status.err_detail.is_empty() && add_status.initial_sync_error_message.is_empty(),
"site replication add reported failure: {add_status:?}"
);
wait_for_site_replication_enabled(&site_a.cluster).await?;
wait_for_site_replication_enabled(&site_b.cluster).await?;
Ok(())
}
async fn list_users_contains(
cluster: &crate::common::RustFSTestClusterEnvironment,
node_idx: usize,
access_key: &str,
) -> TestResult<bool> {
let (status, body) = node_admin(cluster, node_idx, Method::GET, "/rustfs/admin/v3/list-users", None).await?;
if !status.is_success() {
return Err(format!("list-users on node {node_idx} failed: {status} {body}").into());
}
let users: serde_json::Value = serde_json::from_str(&body)?;
Ok(users.get(access_key).is_some())
}
/// backlog#2367 A-7 / functional SITE-102: an IAM change handled by a node
/// other than the one that ran `site-replication/add` must still reach the
/// peer site. Behind a load balancer every admin call may land on a
/// different node, so the coordinator node is not special.
#[tokio::test]
async fn four_node_site_replication_converges_iam_user_created_on_a_non_coordinator_node() -> TestResult {
init_logging();
let (site_a, site_b) = DistCluster::start_replication_pair().await?;
pair_sites(&site_a, &site_b).await?;
let user = format!("siteuser-{}", &uuid::Uuid::new_v4().simple().to_string()[..8]);
let body = serde_json::json!({ "secretKey": "siteuser-secret-key-1234", "status": "enabled" }).to_string();
let (status, response) = node_admin(
&site_a.cluster,
1,
Method::PUT,
&format!("/rustfs/admin/v3/add-user?accessKey={user}"),
Some(body),
)
.await?;
assert!(status.is_success(), "add-user on site A node 1 failed: {status} {response}");
let site_b_cluster = &site_b.cluster;
let user_ref = user.as_str();
wait_until(
Duration::from_secs(90),
|| async move { list_users_contains(site_b_cluster, 0, user_ref).await },
"user created on site A node 1 visible on site B",
)
.await?;
assert!(
list_users_contains(&site_a.cluster, 2, &user).await?,
"the user must be visible on every site A node"
);
Ok(())
}
/// backlog#2367 A-5 / functional SITE-105: a resync started right after
/// pairing must not report buckets as failed. The bucket carrying an
/// operator-configured bucket-replication target to the peer (the shape the
/// functional suite leaves behind) and a plain versioned bucket are both
/// wired by the pairing itself.
#[tokio::test]
async fn four_node_site_replication_resync_start_right_after_pairing_reports_no_failed_bucket() -> TestResult {
init_logging();
let (site_a, site_b) = DistCluster::start_replication_pair().await?;
let pre_src = unique_bucket("pre-src");
let pre_dst = unique_bucket("pre-dst");
let plain = unique_bucket("plain");
site_a.create_bucket(&pre_src).await?;
site_b.create_bucket(&pre_dst).await?;
site_a.create_bucket(&plain).await?;
enable_versioning(&site_a.client(0)?, &pre_src).await?;
enable_versioning(&site_b.client(0)?, &pre_dst).await?;
enable_versioning(&site_a.client(0)?, &plain).await?;
let arn = super::harness::set_remote_target(&site_a.cluster, &pre_src, &site_b.cluster, &pre_dst).await?;
super::harness::put_bucket_replication(&site_a.cluster, &pre_src, &arn).await?;
pair_sites(&site_a, &site_b).await?;
let (status, info) = node_admin(&site_a.cluster, 1, Method::GET, "/rustfs/admin/v3/site-replication/info", None).await?;
assert!(status.is_success(), "site-replication/info failed: {status} {info}");
let info: serde_json::Value = serde_json::from_str(&info)?;
let peer = info["sites"]
.as_array()
.and_then(|sites| sites.iter().find(|site| site["name"] == "site-b"))
.cloned()
.ok_or_else(|| format!("site-b peer missing from info: {info}"))?;
// Through a non-coordinator node, like a load-balanced admin call.
let (status, response) = node_admin(
&site_a.cluster,
1,
Method::PUT,
"/rustfs/admin/v3/site-replication/resync/op?operation=start",
Some(peer.to_string()),
)
.await?;
assert!(status.is_success(), "resync start failed: {status} {response}");
let resync: rustfs_madmin::SRResyncOpStatus = serde_json::from_str(&response)?;
let failed: Vec<String> = resync
.buckets
.iter()
.filter(|bucket| bucket.status == "failed")
.map(|bucket| format!("{}: {}", bucket.bucket, bucket.err_detail))
.collect();
assert!(
failed.is_empty(),
"resync right after pairing reported failed buckets: {failed:?} (status={}, detail={})",
resync.status,
resync.err_detail
);
assert!(
resync.buckets.iter().any(|bucket| bucket.bucket == pre_src)
&& resync.buckets.iter().any(|bucket| bucket.bucket == plain),
"both buckets must be part of the resync: {:?}",
resync.buckets
);
Ok(())
}
@@ -1683,6 +1683,16 @@ mod tests {
}
}
// Keep the partial-repair checkpoint stable across readiness and admin
// requests. Endpoint-blackhole tests must prove their own network stall.
let commit_barrier = if scenario != InterruptionScenario::TargetEndpointBlackhole {
let barrier = replaced_disk.join(".rustfs.sys/e2e-heal-commit-barrier");
std::fs::create_dir_all(barrier.parent().ok_or("commit barrier has no parent")?)?;
std::fs::write(&barrier, format!("{bucket}/cluster/online/"))?;
Some(barrier)
} else {
None
};
cluster.start_node_from_binary(1, &server_binary).await?;
for rejected_key in rejected_outage_keys {
let delete_deadline = Instant::now() + Duration::from_secs(60);
@@ -1834,6 +1844,12 @@ mod tests {
sleep(Duration::from_millis(10)).await;
};
if let Some(barrier) = &commit_barrier {
assert!(
barrier.with_extension("admitted").is_file(),
"interruption tests require a server built with e2e-test-hooks"
);
}
let pre_interrupt_status_body = signed_admin_post(&status_url, None, &cluster.access_key, &cluster.secret_key).await?;
let pre_interrupt_status: serde_json::Value = serde_json::from_str(&pre_interrupt_status_body)
.map_err(|err| format!("pre-interrupt background heal status is not JSON ({err}): {pre_interrupt_status_body}"))?;
@@ -2039,6 +2055,9 @@ mod tests {
}
}
}
if let Some(barrier) = &commit_barrier {
std::fs::remove_file(barrier)?;
}
cluster.start_node_from_binary(interruption_node, &server_binary).await?;
if interruption_node == 0 {
let target = cluster.nodes[1]
+17 -22
View File
@@ -57,50 +57,46 @@ Broad integration tests that exercise:
pip install awscurl
```
2. **Build RustFS**
2. **Build RustFS** (from the repository root)
```bash
cargo build
python3 scripts/e2e_binary.py build
```
### Run individual suites
Run every command below from the repository root through `scripts/e2e_binary.py run`; plain `cargo test -p e2e_test` fails with a missing E2E run receipt.
#### Local backend
```bash
cd crates/e2e_test
cargo test test_local_kms_end_to_end -- --nocapture
python3 scripts/e2e_binary.py run -- cargo test -p e2e_test test_local_kms_end_to_end -- --nocapture
```
#### Vault backend
```bash
cd crates/e2e_test
cargo test test_vault_kms_end_to_end -- --nocapture
python3 scripts/e2e_binary.py run -- cargo test -p e2e_test test_vault_kms_end_to_end -- --nocapture
```
#### High availability
```bash
cd crates/e2e_test
cargo test test_vault_kms_high_availability -- --nocapture
python3 scripts/e2e_binary.py run -- cargo test -p e2e_test test_vault_kms_high_availability -- --nocapture
```
#### Comprehensive features (disabled)
```bash
cd crates/e2e_test
# Disabled due to AWS SDK compatibility gaps
# cargo test test_comprehensive_kms_functionality -- --nocapture
# cargo test test_sse_modes_compatibility -- --nocapture
# cargo test test_kms_api_comprehensive -- --nocapture
# python3 scripts/e2e_binary.py run -- cargo test -p e2e_test test_comprehensive_kms_functionality -- --nocapture
# python3 scripts/e2e_binary.py run -- cargo test -p e2e_test test_sse_modes_compatibility -- --nocapture
# python3 scripts/e2e_binary.py run -- cargo test -p e2e_test test_kms_api_comprehensive -- --nocapture
```
### Run all KMS suites
```bash
cd crates/e2e_test
cargo test kms -- --nocapture
python3 scripts/e2e_binary.py run -- cargo test -p e2e_test kms -- --nocapture
```
### Run serially (avoid port conflicts)
```bash
cd crates/e2e_test
cargo test kms -- --nocapture --test-threads=1
python3 scripts/e2e_binary.py run -- cargo test -p e2e_test kms -- --nocapture --test-threads=1
```
## 🔧 Configuration
@@ -120,7 +116,7 @@ export RUST_LOG=debug
### Required binaries
Tests look for:
- `../../target/debug/rustfs` – RustFS server
- RustFS server – the binary verified by `scripts/e2e_binary.py` (default `target/debug/rustfs`)
- `vault` – Vault CLI (must be on PATH)
- `/Users/dandan/Library/Python/3.9/bin/awscurl` – AWS SigV4 helper
@@ -174,14 +170,14 @@ which awscurl # Update the path in tests accordingly
**Q: Tests time out**
```bash
RUST_LOG=debug cargo test test_local_kms_end_to_end -- --nocapture
RUST_LOG=debug python3 scripts/e2e_binary.py run -- cargo test -p e2e_test test_local_kms_end_to_end -- --nocapture
```
### Debug tips
1. **Enable verbose logs**
```bash
RUST_LOG=rustfs_kms=debug,rustfs=info cargo test -- --nocapture
RUST_LOG=rustfs_kms=debug,rustfs=info python3 scripts/e2e_binary.py run -- cargo test -p e2e_test kms -- --nocapture
```
2. **Keep temporary files** – comment out cleanup logic to inspect generated configs
@@ -237,9 +233,8 @@ Designed to run inside CI/CD pipelines:
sudo apt-get install -y vault
pip install awscurl
cargo build
cd crates/e2e_test
cargo test kms -- --nocapture --test-threads=1
python3 scripts/e2e_binary.py build
python3 scripts/e2e_binary.py run -- cargo test -p e2e_test kms -- --nocapture --test-threads=1
```
## 📚 References
@@ -21,6 +21,7 @@
use super::common::LocalKMSTestEnvironment;
use crate::common::{TEST_BUCKET, init_logging};
use aws_sdk_s3::error::ProvideErrorMetadata;
use aws_sdk_s3::primitives::ByteStream;
use aws_sdk_s3::types::{
ChecksumAlgorithm, ChecksumMode, CompletedMultipartUpload, CompletedPart, ServerSideEncryption,
@@ -631,3 +632,92 @@ async fn test_sse_kms_without_key_id_populates_default() -> Result<(), Box<dyn s
info!("Test passed: SSE-KMS without key ID correctly populates default key '{}'", default_key_id);
Ok(())
}
/// A default-encryption configuration the write path cannot honour as written
/// must be refused rather than stored: the write path falls back to AES256 for
/// any algorithm it does not know, so storing `AES128` would make
/// GetBucketEncryption report a scheme no object is encrypted under.
#[tokio::test]
async fn test_put_bucket_encryption_rejects_unknown_algorithm_and_misplaced_key_id()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging();
let mut kms_env = LocalKMSTestEnvironment::new().await?;
let default_key_id = kms_env.start_rustfs_for_local_kms().await?;
kms_env.wait_for_kms_ready().await?;
let s3_client = kms_env.base_env.create_s3_client();
kms_env.base_env.create_test_bucket(TEST_BUCKET).await?;
let config_with = |algorithm: ServerSideEncryption, key_id: Option<&str>| {
let mut by_default = ServerSideEncryptionByDefault::builder().sse_algorithm(algorithm);
if let Some(key_id) = key_id {
by_default = by_default.kms_master_key_id(key_id);
}
ServerSideEncryptionConfiguration::builder()
.rules(
ServerSideEncryptionRule::builder()
.apply_server_side_encryption_by_default(by_default.build().unwrap())
.build(),
)
.build()
.unwrap()
};
// Baseline the bucket on AES256 so a refused update has something to leave untouched.
s3_client
.put_bucket_encryption()
.bucket(TEST_BUCKET)
.server_side_encryption_configuration(config_with(ServerSideEncryption::Aes256, None))
.send()
.await?;
let unknown = s3_client
.put_bucket_encryption()
.bucket(TEST_BUCKET)
.server_side_encryption_configuration(config_with(ServerSideEncryption::from("AES128"), None))
.send()
.await
.expect_err("an unknown SSEAlgorithm must be refused");
assert_eq!(unknown.raw_response().map(|response| response.status().as_u16()), Some(400));
assert_eq!(
unknown.as_service_error().and_then(ProvideErrorMetadata::code),
Some("MalformedXML"),
"unknown algorithm error was {unknown:?}"
);
let misplaced_key = s3_client
.put_bucket_encryption()
.bucket(TEST_BUCKET)
.server_side_encryption_configuration(config_with(ServerSideEncryption::Aes256, Some(&default_key_id)))
.send()
.await
.expect_err("KMSMasterKeyID with AES256 must be refused");
assert_eq!(misplaced_key.raw_response().map(|response| response.status().as_u16()), Some(400));
assert_eq!(
misplaced_key.as_service_error().and_then(ProvideErrorMetadata::code),
Some("InvalidArgument"),
"misplaced key id error was {misplaced_key:?}"
);
// The refused updates left the baseline in place.
let stored = s3_client.get_bucket_encryption().bucket(TEST_BUCKET).send().await?;
let by_default = stored
.server_side_encryption_configuration()
.and_then(|config| config.rules().first())
.and_then(|rule| rule.apply_server_side_encryption_by_default())
.expect("baseline configuration must still be present");
assert_eq!(by_default.sse_algorithm(), &ServerSideEncryption::Aes256);
assert_eq!(by_default.kms_master_key_id(), None);
// aws:kms with a key id stays accepted.
s3_client
.put_bucket_encryption()
.bucket(TEST_BUCKET)
.server_side_encryption_configuration(config_with(ServerSideEncryption::AwsKms, Some(&default_key_id)))
.send()
.await?;
kms_env.base_env.delete_test_bucket(TEST_BUCKET).await?;
Ok(())
}
+25 -1
View File
@@ -176,6 +176,30 @@ pub async fn start_kms(
Ok(())
}
/// Stop the running KMS service via admin API, keeping its configuration so
/// `start_kms` can bring it back.
pub async fn stop_kms(
base_url: &str,
access_key: &str,
secret_key: &str,
) -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
let response = kms_admin_request(
base_url,
http::Method::POST,
"/rustfs/admin/v3/kms/stop",
Some("{}"),
access_key,
secret_key,
)
.await?;
let response: serde_json::Value = serde_json::from_str(&response)?;
if response["success"] != true {
return Err(format!("KMS stop failed: {}", response["message"].as_str().unwrap_or("unknown error")).into());
}
info!("KMS stopped successfully");
Ok(())
}
/// Get KMS status via admin API
pub async fn get_kms_status(
base_url: &str,
@@ -325,7 +349,7 @@ pub async fn create_default_key(
secret_key: &str,
) -> Result<String, Box<dyn std::error::Error + Send + Sync>> {
let create_key_body = serde_json::json!({
"key_usage": "ENCRYPT_DECRYPT",
"key_usage": "EncryptDecrypt",
"description": "Default key for e2e testing"
})
.to_string();
@@ -357,3 +357,97 @@ async fn test_multipart_upload_writes_encrypted_data() -> Result<(), Box<dyn std
Ok(())
}
/// `x-amz-server-side-encryption-aws-kms-key-id` is defined for `aws:kms`
/// objects only. SSE-S3 wraps its data key under the service default key too,
/// but that key is internal: PutObject, CopyObject and CreateMultipartUpload
/// responses for an `AES256` object must not name it.
#[tokio::test]
async fn test_sse_s3_write_responses_carry_no_kms_key_id() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging();
let mut kms_env = LocalKMSTestEnvironment::new().await?;
let default_key = kms_env.start_rustfs_for_local_kms().await?;
kms_env.wait_for_kms_ready().await?;
let s3_client = kms_env.base_env.create_s3_client();
kms_env.base_env.create_test_bucket(TEST_BUCKET).await?;
let put = s3_client
.put_object()
.bucket(TEST_BUCKET)
.key("sse-s3-put")
.body(ByteStream::from_static(b"sse-s3 payload"))
.server_side_encryption(ServerSideEncryption::Aes256)
.send()
.await?;
assert_eq!(put.server_side_encryption(), Some(&ServerSideEncryption::Aes256));
assert_eq!(put.ssekms_key_id(), None, "PutObject AES256 must not advertise the wrapping key");
let get = s3_client.get_object().bucket(TEST_BUCKET).key("sse-s3-put").send().await?;
assert_eq!(get.server_side_encryption(), Some(&ServerSideEncryption::Aes256));
assert_eq!(get.ssekms_key_id(), None, "GetObject AES256 must not advertise the wrapping key");
let head = s3_client.head_object().bucket(TEST_BUCKET).key("sse-s3-put").send().await?;
assert_eq!(head.ssekms_key_id(), None, "HeadObject AES256 must not advertise the wrapping key");
let copy = s3_client
.copy_object()
.bucket(TEST_BUCKET)
.key("sse-s3-copy")
.copy_source(format!("{TEST_BUCKET}/sse-s3-put"))
.server_side_encryption(ServerSideEncryption::Aes256)
.send()
.await?;
assert_eq!(copy.server_side_encryption(), Some(&ServerSideEncryption::Aes256));
assert_eq!(copy.ssekms_key_id(), None, "CopyObject AES256 must not advertise the wrapping key");
let multipart = s3_client
.create_multipart_upload()
.bucket(TEST_BUCKET)
.key("sse-s3-multipart")
.server_side_encryption(ServerSideEncryption::Aes256)
.send()
.await?;
assert_eq!(multipart.server_side_encryption(), Some(&ServerSideEncryption::Aes256));
assert_eq!(
multipart.ssekms_key_id(),
None,
"CreateMultipartUpload AES256 must not advertise the wrapping key"
);
s3_client
.abort_multipart_upload()
.bucket(TEST_BUCKET)
.key("sse-s3-multipart")
.upload_id(multipart.upload_id().expect("upload id"))
.send()
.await?;
// Control: the same responses keep naming the key for an aws:kms object.
let kms_put = s3_client
.put_object()
.bucket(TEST_BUCKET)
.key("sse-kms-put")
.body(ByteStream::from_static(b"sse-kms payload"))
.server_side_encryption(ServerSideEncryption::AwsKms)
.send()
.await?;
assert_eq!(kms_put.ssekms_key_id(), Some(default_key.as_str()));
let kms_multipart = s3_client
.create_multipart_upload()
.bucket(TEST_BUCKET)
.key("sse-kms-multipart")
.server_side_encryption(ServerSideEncryption::AwsKms)
.send()
.await?;
assert_eq!(kms_multipart.ssekms_key_id(), Some(default_key.as_str()));
s3_client
.abort_multipart_upload()
.bucket(TEST_BUCKET)
.key("sse-kms-multipart")
.upload_id(kms_multipart.upload_id().expect("upload id"))
.send()
.await?;
kms_env.base_env.delete_test_bucket(TEST_BUCKET).await?;
Ok(())
}
@@ -21,7 +21,7 @@
//! - Corrupted key files
//! - Recovery from transient failures
use super::common::LocalKMSTestEnvironment;
use super::common::{LocalKMSTestEnvironment, create_default_key, create_key_with_specific_id, kms_admin_request};
use crate::common::{TEST_BUCKET, init_logging};
use aws_sdk_s3::error::ProvideErrorMetadata;
use aws_sdk_s3::types::ServerSideEncryption;
@@ -534,3 +534,129 @@ async fn test_kms_concurrent_encryption_requests() -> Result<(), Box<dyn std::er
kms_env.base_env.delete_test_bucket(TEST_BUCKET).await?;
Ok(())
}
/// Once the key an object was wrapped under is deleted, the object cannot be
/// read until the key is restored. That is the same `400 KMS.NotFoundException`
/// a write under a missing key returns, not a `500`; `HeadObject` never unwraps
/// the data key and keeps answering `200`.
#[tokio::test]
async fn test_reads_under_a_deleted_kms_key_report_key_not_found() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging();
let mut kms_env = LocalKMSTestEnvironment::new().await?;
let default_key_id = "rustfs-e2e-test-default-key";
create_key_with_specific_id(&kms_env.kms_keys_dir, default_key_id).await?;
let key_dir = kms_env.kms_keys_dir.clone();
kms_env
.base_env
.start_rustfs_server_with_env(
vec![
"--kms-enable",
"--kms-backend",
"local",
"--kms-key-dir",
&key_dir,
"--kms-default-key-id",
default_key_id,
],
&[
("RUSTFS_KMS_ALLOW_INSECURE_DEV_DEFAULTS", "true"),
// Immediate deletion is refused on a default server; the test
// needs the key gone now rather than after the waiting window.
("RUSTFS_KMS_ALLOW_IMMEDIATE_DELETION", "true"),
],
)
.await?;
kms_env.wait_for_kms_ready().await?;
let base = &kms_env.base_env;
let s3_client = base.create_s3_client();
base.create_test_bucket(TEST_BUCKET).await?;
let doomed_key_id = create_default_key(&base.url, &base.access_key, &base.secret_key).await?;
let object_key = "wrapped-under-doomed-key";
let payload = b"readable only while the key exists".to_vec();
let put = s3_client
.put_object()
.bucket(TEST_BUCKET)
.key(object_key)
.body(aws_sdk_s3::primitives::ByteStream::from(payload.clone()))
.server_side_encryption(ServerSideEncryption::AwsKms)
.ssekms_key_id(&doomed_key_id)
.send()
.await?;
assert_eq!(put.ssekms_key_id(), Some(doomed_key_id.as_str()));
info!("🗑️ deleting {doomed_key_id} immediately");
kms_admin_request(
&base.url,
http::Method::DELETE,
"/rustfs/admin/v3/kms/keys/delete",
Some(
&serde_json::json!({
"key_id": doomed_key_id,
"force_immediate": true,
"confirm_key_id": doomed_key_id,
})
.to_string(),
),
&base.access_key,
&base.secret_key,
)
.await?;
kms_admin_request(
&base.url,
http::Method::POST,
"/rustfs/admin/v3/kms/clear-cache",
Some("{}"),
&base.access_key,
&base.secret_key,
)
.await?;
let head = s3_client.head_object().bucket(TEST_BUCKET).key(object_key).send().await?;
assert_eq!(head.ssekms_key_id(), Some(doomed_key_id.as_str()));
let get_error = s3_client
.get_object()
.bucket(TEST_BUCKET)
.key(object_key)
.send()
.await
.expect_err("an object wrapped under a deleted key must not be readable");
assert_eq!(get_error.raw_response().map(|response| response.status().as_u16()), Some(400));
assert_eq!(
get_error.as_service_error().and_then(ProvideErrorMetadata::code),
Some("KMS.NotFoundException"),
"GetObject error was {get_error:?}"
);
let copy_error = s3_client
.copy_object()
.bucket(TEST_BUCKET)
.key("copied-from-doomed-source")
.copy_source(format!("{TEST_BUCKET}/{object_key}"))
.server_side_encryption(ServerSideEncryption::AwsKms)
.send()
.await
.expect_err("copying from an object wrapped under a deleted key must fail the same way");
assert_eq!(copy_error.raw_response().map(|response| response.status().as_u16()), Some(400));
assert_eq!(
copy_error.as_service_error().and_then(ProvideErrorMetadata::code),
Some("KMS.NotFoundException"),
"CopyObject error was {copy_error:?}"
);
// The default key is untouched, so the node keeps serving other objects.
let unaffected = s3_client
.put_object()
.bucket(TEST_BUCKET)
.key("wrapped-under-default-key")
.body(aws_sdk_s3::primitives::ByteStream::from_static(b"still fine"))
.server_side_encryption(ServerSideEncryption::AwsKms)
.send()
.await?;
assert_eq!(unaffected.ssekms_key_id(), Some(default_key_id));
base.delete_test_bucket(TEST_BUCKET).await?;
Ok(())
}
@@ -0,0 +1,231 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! SSE-KMS writes against a node whose KMS is stopped, was never configured,
//! or runs without a default key.
//!
//! `docs/operations/kms-backend-security.md` promises `503` for a configured
//! KMS that is not running and `400 InvalidRequest` when KMS was never
//! configured. The bare `aws:kms` form (no key id, no bucket default) used to
//! miss both branches and surface as `500`, so every scenario here is
//! exercised with and without a key id.
use super::common::{
LocalKMSTestEnvironment, assert_s3_error, create_key_with_specific_id, start_kms, stop_kms, wait_for_kms_ready,
};
use crate::common::{RustFSTestEnvironment, TEST_BUCKET, init_logging};
use aws_sdk_s3::Client;
use aws_sdk_s3::error::ProvideErrorMetadata;
use aws_sdk_s3::primitives::ByteStream;
use aws_sdk_s3::types::ServerSideEncryption;
use tracing::info;
const SERVICE_UNAVAILABLE_MESSAGE: &str = "The service is unavailable. Please retry.";
const KMS_NOT_CONFIGURED_MESSAGE: &str = "SSE-KMS requires a configured and running KMS service";
const KMS_NO_DEFAULT_KEY_MESSAGE: &str =
"SSE-KMS requires a KMS key id: the request named none and the KMS service has no default key";
/// Issue an SSE-KMS PutObject and a CreateMultipartUpload, each with and
/// without a key id, and require every one of them to fail with `status`/`code`/`message`.
async fn assert_sse_kms_writes_refused(client: &Client, key_prefix: &str, status: u16, code: &str, message: &str) {
for (label, key_id) in [("bare", None), ("keyed", Some("rustfs-e2e-refused-key"))] {
let object_key = format!("{key_prefix}-{label}");
let mut put = client
.put_object()
.bucket(TEST_BUCKET)
.key(&object_key)
.body(ByteStream::from_static(b"must not be published"))
.server_side_encryption(ServerSideEncryption::AwsKms);
if let Some(key_id) = key_id {
put = put.ssekms_key_id(key_id);
}
assert_s3_error(put.send().await, status, code, message, &format!("{label} SSE-KMS PutObject"));
let mut create = client
.create_multipart_upload()
.bucket(TEST_BUCKET)
.key(&object_key)
.server_side_encryption(ServerSideEncryption::AwsKms);
if let Some(key_id) = key_id {
create = create.ssekms_key_id(key_id);
}
assert_s3_error(
create.send().await,
status,
code,
message,
&format!("{label} SSE-KMS CreateMultipartUpload"),
);
let absence = client
.get_object()
.bucket(TEST_BUCKET)
.key(&object_key)
.send()
.await
.expect_err("a refused SSE-KMS write must not publish an object");
assert_eq!(absence.raw_response().map(|response| response.status().as_u16()), Some(404));
assert_eq!(absence.as_service_error().and_then(ProvideErrorMetadata::code), Some("NoSuchKey"));
}
}
/// A configured KMS that an operator stopped is a transient outage: every
/// SSE-KMS write is refused with `503` until `kms/start`, after which the bare
/// form resolves the service default key again and earlier objects read back.
#[tokio::test]
async fn test_sse_kms_writes_are_refused_with_503_while_kms_is_stopped() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging();
let mut kms_env = LocalKMSTestEnvironment::new().await?;
let default_key_id = kms_env.start_rustfs_for_local_kms().await?;
kms_env.wait_for_kms_ready().await?;
let base = &kms_env.base_env;
let client = base.create_s3_client();
base.create_test_bucket(TEST_BUCKET).await?;
let object_key = "written-before-stop";
let payload = b"encrypted under the service default key".to_vec();
let put = client
.put_object()
.bucket(TEST_BUCKET)
.key(object_key)
.body(ByteStream::from(payload.clone()))
.server_side_encryption(ServerSideEncryption::AwsKms)
.send()
.await?;
assert_eq!(put.server_side_encryption(), Some(&ServerSideEncryption::AwsKms));
assert_eq!(put.ssekms_key_id(), Some(default_key_id.as_str()));
info!("stopping KMS through the admin API");
stop_kms(&base.url, &base.access_key, &base.secret_key).await?;
assert_sse_kms_writes_refused(&client, "while-stopped", 503, "ServiceUnavailable", SERVICE_UNAVAILABLE_MESSAGE).await;
info!("starting KMS again");
start_kms(&base.url, &base.access_key, &base.secret_key).await?;
wait_for_kms_ready(&base.url, &base.access_key, &base.secret_key).await?;
let restored = client
.put_object()
.bucket(TEST_BUCKET)
.key("written-after-start")
.body(ByteStream::from_static(b"service is back"))
.server_side_encryption(ServerSideEncryption::AwsKms)
.send()
.await?;
assert_eq!(restored.ssekms_key_id(), Some(default_key_id.as_str()));
let read_back = client.get_object().bucket(TEST_BUCKET).key(object_key).send().await?;
assert_eq!(read_back.body.collect().await?.into_bytes().as_ref(), payload.as_slice());
base.delete_test_bucket(TEST_BUCKET).await?;
Ok(())
}
/// A node that only carries `RUSTFS_SSE_S3_MASTER_KEY` serves SSE-S3 but has
/// no KMS to name: SSE-KMS is a client configuration error (`400`), and it
/// must never be downgraded onto the local master key.
#[tokio::test]
async fn test_sse_kms_writes_are_refused_with_400_when_kms_was_never_configured()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging();
let mut env = RustFSTestEnvironment::new().await?;
// base64 of 32 zero bytes: a valid master key shape for the SSE-S3 fallback.
env.start_rustfs_server_with_env(vec![], &[("RUSTFS_SSE_S3_MASTER_KEY", "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA=")])
.await?;
let client = env.create_s3_client();
env.create_test_bucket(TEST_BUCKET).await?;
let sse_s3 = client
.put_object()
.bucket(TEST_BUCKET)
.key("sse-s3-fallback")
.body(ByteStream::from_static(b"local master key still serves AES256"))
.server_side_encryption(ServerSideEncryption::Aes256)
.send()
.await?;
assert_eq!(sse_s3.server_side_encryption(), Some(&ServerSideEncryption::Aes256));
assert_sse_kms_writes_refused(&client, "no-kms", 400, "InvalidRequest", KMS_NOT_CONFIGURED_MESSAGE).await;
env.delete_test_bucket(TEST_BUCKET).await?;
Ok(())
}
/// A running KMS without a default key can serve a keyed request but has
/// nothing to resolve a bare `aws:kms` request to; that is the caller's
/// omission, not a server fault.
#[tokio::test]
async fn test_bare_sse_kms_write_is_refused_with_400_when_kms_has_no_default_key()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging();
let mut kms_env = LocalKMSTestEnvironment::new().await?;
let named_key_id = "rustfs-e2e-named-key";
create_key_with_specific_id(&kms_env.kms_keys_dir, named_key_id).await?;
let key_dir = kms_env.kms_keys_dir.clone();
kms_env
.base_env
.start_rustfs_server_with_env(
vec!["--kms-enable", "--kms-backend", "local", "--kms-key-dir", &key_dir],
&[("RUSTFS_KMS_ALLOW_INSECURE_DEV_DEFAULTS", "true")],
)
.await?;
kms_env.wait_for_kms_ready().await?;
let base = &kms_env.base_env;
let client = base.create_s3_client();
base.create_test_bucket(TEST_BUCKET).await?;
let keyed = client
.put_object()
.bucket(TEST_BUCKET)
.key("keyed-without-default")
.body(ByteStream::from_static(b"a named key needs no default"))
.server_side_encryption(ServerSideEncryption::AwsKms)
.ssekms_key_id(named_key_id)
.send()
.await?;
assert_eq!(keyed.ssekms_key_id(), Some(named_key_id));
let bare = client
.put_object()
.bucket(TEST_BUCKET)
.key("bare-without-default")
.body(ByteStream::from_static(b"must not be published"))
.server_side_encryption(ServerSideEncryption::AwsKms)
.send()
.await;
assert_s3_error(
bare,
400,
"InvalidRequest",
KMS_NO_DEFAULT_KEY_MESSAGE,
"bare SSE-KMS PutObject without default key",
);
let bare_multipart = client
.create_multipart_upload()
.bucket(TEST_BUCKET)
.key("bare-without-default")
.server_side_encryption(ServerSideEncryption::AwsKms)
.send()
.await;
assert_s3_error(
bare_multipart,
400,
"InvalidRequest",
KMS_NO_DEFAULT_KEY_MESSAGE,
"bare SSE-KMS CreateMultipartUpload without default key",
);
base.delete_test_bucket(TEST_BUCKET).await?;
Ok(())
}
+3
View File
@@ -39,6 +39,9 @@ mod kms_edge_cases_test;
#[cfg(test)]
mod kms_fault_recovery_test;
#[cfg(test)]
mod kms_service_stop_test;
#[cfg(test)]
mod bucket_default_encryption_test;
+3
View File
@@ -179,6 +179,9 @@ mod compression_test;
#[cfg(test)]
mod delete_objects_versioning_test;
#[cfg(test)]
mod delete_authorization_test;
// Regression test for signed DELETE Object?versionId requests without Content-Length.
#[cfg(test)]
mod delete_object_no_content_length_test;
+2 -1
View File
@@ -24,7 +24,8 @@ The tests cover the following AWS policy variable scenarios:
```bash
# From the project root directory
cargo test -p e2e_test policy:: -- --nocapture
python3 scripts/e2e_binary.py build
python3 scripts/e2e_binary.py run -- cargo test -p e2e_test policy:: -- --nocapture
```
Each test starts an isolated RustFS server on a dynamically allocated local port and cleans it up afterward.
+3 -5
View File
@@ -17,15 +17,13 @@ Use the canonical CI-equivalent protocol command in the parent
For targeted debugging of the core suite only:
```bash
RUSTFS_BUILD_FEATURES=ftps,webdav,sftp cargo test --package e2e_test test_protocol_core_suite -- --test-threads=1 --nocapture
python3 scripts/e2e_binary.py build --features ftps,webdav,sftp
python3 scripts/e2e_binary.py run --features ftps,webdav,sftp -- cargo test --package e2e_test test_protocol_core_suite -- --test-threads=1 --nocapture
```
This targeted command does not cover the full `e2e-protocols` profile.
`RUSTFS_BUILD_FEATURES` controls which features the test rustfs binary is
built with. When this variable is set, the protocol test runner schedules
only entries whose protocol is present in the requested feature list. Leave
it unset to run every protocol entry.
`e2e_binary.py` supplies `RUSTFS_BUILD_FEATURES` from the verified server's resolved Cargo features. The protocol runner schedules only entries present in that feature list; helpers check that their required features are available without rebuilding the server.
`--test-threads=1` is required because every entry spawns a rustfs server
on fixed bind ports.
+230 -14
View File
@@ -6134,6 +6134,215 @@ async fn test_site_replication_allows_private_ca_https_with_ca_cert_pem_real_dua
Ok(())
}
/// rustfs/backlog#2489: an operator's bucket-level replication to a site that
/// later becomes a peer keeps working through the site's add, resync and
/// removal. Site replication wires its own same-name target next to the
/// operator's and never takes the operator's target over; a bucket-level
/// `replication-reset` run before the join must not make the site resync
/// report the bucket as owned by another resync.
#[tokio::test]
async fn test_site_replication_keeps_operator_bucket_target_to_peer() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging();
let process_env = [
("RUSTFS_REPLICATION_ALLOW_LOOPBACK_TARGET", "true"),
("RUSTFS_REPL_RESYNC_POLL_MAX_MS", "100"),
("RUST_LOG", "error"),
];
let mut source_env = RustFSTestEnvironment::new().await?;
source_env.start_rustfs_server_with_env(vec![], &process_env).await?;
let mut target_env = RustFSTestEnvironment::new().await?;
target_env.start_rustfs_server_with_env(vec![], &process_env).await?;
let source_bucket = "site-repl-operator-src";
let operator_target_bucket = "site-repl-operator-dst";
let source_client = source_env.create_s3_client();
let target_client = target_env.create_s3_client();
source_client.create_bucket().bucket(source_bucket).send().await?;
enable_bucket_versioning(&source_env, source_bucket).await?;
target_client.create_bucket().bucket(operator_target_bucket).send().await?;
enable_bucket_versioning(&target_env, operator_target_bucket).await?;
let operator_arn = set_replication_target(&source_env, source_bucket, &target_env, operator_target_bucket).await?;
put_bucket_replication(&source_env, source_bucket, &operator_arn).await?;
let put_and_wait = |key: &'static str, fill: u8, on: Vec<&'static str>| {
let source_client = source_client.clone();
let target_client = target_client.clone();
async move {
source_client
.put_object()
.bucket(source_bucket)
.key(key)
.body(ByteStream::from(vec![fill; 4096]))
.send()
.await?;
for bucket in on {
let body = wait_for_object_on_target(&target_client, bucket, key).await?;
if body != vec![fill; 4096] {
return Err(format!("{key} arrived on {bucket} with the wrong body").into());
}
}
Ok::<(), Box<dyn Error + Send + Sync>>(())
}
};
put_and_wait("before-add.bin", b'a', vec![operator_target_bucket]).await?;
let (reset_arn, reset_id) = start_bucket_replication_reset(&source_env, source_bucket).await?;
assert_eq!(reset_arn, operator_arn, "the reset must target the operator ARN");
wait_for_replication_reset_target(&source_env, source_bucket, &operator_arn, |target| {
target.reset_id == reset_id && matches!(target.status.as_str(), "Completed" | "Failed")
})
.await?;
let add_status = site_replication_add(
&source_env,
&[
PeerSite {
name: "source-site".to_string(),
endpoint: source_env.url.clone(),
access_key: source_env.access_key.clone(),
secret_key: source_env.secret_key.clone(),
..Default::default()
},
PeerSite {
name: "target-site".to_string(),
endpoint: target_env.url.clone(),
access_key: target_env.access_key.clone(),
secret_key: target_env.secret_key.clone(),
..Default::default()
},
],
)
.await?;
assert!(add_status.success, "unexpected site add result: {:?}", add_status);
let source_info = wait_for_site_replication_enabled(&source_env, 2).await?;
wait_for_site_replication_enabled(&target_env, 2).await?;
let remote_peer = source_info
.sites
.into_iter()
.find(|peer| peer.endpoint == target_env.url)
.ok_or("target peer missing from source site replication info")?;
wait_for_bucket_on_target(&target_client, source_bucket).await?;
// Both targets: the operator's (untouched) and the site's same-name one.
let list_targets = || async {
let response = list_replication_targets_request(&source_env, Some(source_bucket)).await?;
let targets: Vec<serde_json::Value> = if response.status() == StatusCode::OK {
response.json().await?
} else {
Vec::new()
};
Ok::<Vec<(String, String)>, Box<dyn Error + Send + Sync>>(
targets
.iter()
.map(|target| {
(
target["arn"].as_str().unwrap_or_default().to_string(),
target["targetbucket"].as_str().unwrap_or_default().to_string(),
)
})
.collect(),
)
};
let deadline = tokio::time::Instant::now() + Duration::from_secs(30);
let site_arn = loop {
let targets = list_targets().await?;
let operator_kept = targets
.iter()
.any(|(arn, target_bucket)| arn == &operator_arn && target_bucket == operator_target_bucket);
let site = targets
.iter()
.find(|(arn, target_bucket)| arn.contains(&remote_peer.deployment_id) && target_bucket == source_bucket)
.map(|(arn, _)| arn.clone());
assert!(operator_kept, "site replication must not take the operator target over: {targets:?}");
if let Some(arn) = site {
break arn;
}
if tokio::time::Instant::now() >= deadline {
return Err(format!("site replication never added its own target next to the operator's: {targets:?}").into());
}
sleep(Duration::from_millis(250)).await;
};
assert_ne!(site_arn, operator_arn);
// The operator's path and the site's path both deliver.
put_and_wait("after-add.bin", b'b', vec![operator_target_bucket, source_bucket]).await?;
let started = site_replication_resync_op(&source_env, "start", &remote_peer).await?;
let entry = started
.buckets
.iter()
.find(|entry| entry.bucket == source_bucket)
.ok_or_else(|| format!("start response lost the bucket: {started:?}"))?;
assert_ne!(
entry.status, "conflict",
"a finished bucket-level resync must not block the site resync: {entry:?}"
);
assert_eq!(
entry.target_arn, site_arn,
"the site resync must drive the site target, not the operator's"
);
assert_eq!(started.status, "success", "unexpected start result: {:?}", started);
let deadline = tokio::time::Instant::now() + Duration::from_secs(60);
let finished = loop {
let status = site_replication_resync_op(&source_env, "status", &remote_peer).await?;
match status.state.as_str() {
"completed" | "failed" => break status,
_ if tokio::time::Instant::now() < deadline => sleep(Duration::from_millis(250)).await,
_ => return Err(format!("site resync did not reach a terminal state in time: {status:?}").into()),
}
};
let entry = finished
.buckets
.iter()
.find(|entry| entry.bucket == source_bucket)
.ok_or_else(|| format!("resync status lost the bucket: {finished:?}"))?;
assert_eq!(
(finished.state.as_str(), entry.status.as_str(), entry.failed_objects),
("completed", "completed", 0),
"{finished:?}"
);
// Leaving site replication removes only the site's own target and rule.
let removed = site_replication_remove(
&source_env,
&SRRemoveReq {
remove_all: true,
..Default::default()
},
)
.await?;
assert!(removed.err_detail.is_empty(), "unexpected remove result: {removed:?}");
let targets = list_targets().await?;
assert!(
targets
.iter()
.any(|(arn, target_bucket)| arn == &operator_arn && target_bucket == operator_target_bucket),
"peer removal must keep the operator target: {targets:?}"
);
assert!(
!targets.iter().any(|(arn, _)| arn == &site_arn),
"peer removal must drop the site target: {targets:?}"
);
let rules = source_client
.get_bucket_replication()
.bucket(source_bucket)
.send()
.await?
.replication_configuration
.map(|config| config.rules)
.unwrap_or_default();
assert!(
rules
.iter()
.any(|rule| rule.destination.as_ref().map(|d| d.bucket.as_str()) == Some(operator_arn.as_str())),
"peer removal must keep the operator rule: {rules:?}"
);
put_and_wait("after-remove.bin", b'c', vec![operator_target_bucket]).await?;
Ok(())
}
#[tokio::test]
async fn test_site_replication_resync_lifecycle_survives_real_server_restart() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging();
@@ -6383,6 +6592,18 @@ async fn test_site_replication_edit_and_status_peer_state_real_three_node() -> R
let relayed_key = "after-edit-from-relay.txt";
let relayed_payload = b"site replication after endpoint edit from relay".to_vec();
// The first joining receiver owns data before the third site has the
// shared account. Initial probes and backfill must wait for every join.
target_client.create_bucket().bucket(bucket).send().await?;
enable_bucket_versioning(&target_env, bucket).await?;
target_client
.put_object()
.bucket(bucket)
.key(baseline_key)
.body(ByteStream::from(baseline_payload.clone()))
.send()
.await?;
let add_status = site_replication_add(
&source_env,
&[
@@ -6410,7 +6631,10 @@ async fn test_site_replication_edit_and_status_peer_state_real_three_node() -> R
],
)
.await?;
assert!(add_status.success, "unexpected site add result: {:?}", add_status);
assert!(
add_status.success && add_status.err_detail.is_empty() && add_status.initial_sync_error_message.is_empty(),
"unexpected site add result: {add_status:?}"
);
let source_info = wait_for_site_replication_enabled(&source_env, 3).await?;
let _target_info = wait_for_site_replication_enabled(&target_env, 3).await?;
@@ -6421,19 +6645,11 @@ async fn test_site_replication_edit_and_status_peer_state_real_three_node() -> R
.find(|peer| peer.endpoint == target_env.url)
.ok_or("target peer missing from source site replication info")?;
source_client.create_bucket().bucket(bucket).send().await?;
enable_bucket_versioning(&source_env, bucket).await?;
wait_for_bucket_on_target(&target_client, bucket).await?;
wait_for_bucket_on_target(&relay_client, bucket).await?;
source_client
.put_object()
.bucket(bucket)
.key(baseline_key)
.body(ByteStream::from(baseline_payload.clone()))
.send()
.await?;
let replicated_baseline = wait_for_object_on_target(&target_client, bucket, baseline_key).await?;
assert_eq!(replicated_baseline, baseline_payload);
for client in [&source_client, &relay_client] {
wait_for_bucket_on_target(client, bucket).await?;
let backfilled = wait_for_object_on_target(client, bucket, baseline_key).await?;
assert_eq!(backfilled, baseline_payload);
}
let old_target_address = target_env.address.clone();
let new_target_port = RustFSTestEnvironment::find_available_port().await?;
+3 -1
View File
@@ -202,7 +202,9 @@ pub mod bucket {
}
pub mod migration {
pub use crate::bucket::migration::{LegacyBlobDecryptFn, try_migrate_bucket_metadata, try_migrate_iam_config};
pub use crate::bucket::migration::{
LegacyBlobDecryptFn, migration_startup_error, try_migrate_bucket_metadata, try_migrate_iam_config,
};
}
pub mod object_lock {
+198 -92
View File
@@ -17,7 +17,7 @@
use crate::bucket::metadata::BUCKET_METADATA_FILE;
use crate::bucket::replication::ReplicationMigrationBridge;
use crate::disk::{BUCKET_META_PREFIX, MIGRATING_META_BUCKET, RUSTFS_META_BUCKET};
use crate::error::Error;
use crate::error::{Error, Result, is_err_strict_not_found, is_err_strict_volume_not_found};
use crate::object_api::{GetObjectReader, ObjectInfo, ObjectOptions, PutObjReader};
use crate::storage_api_contracts::{
bucket::{BucketOperations, BucketOptions},
@@ -33,7 +33,7 @@ use rustfs_utils::path::SLASH_SEPARATOR;
use serde::{Deserialize, Serialize};
use std::sync::Arc;
use time::OffsetDateTime;
use tracing::{debug, info, warn};
use tracing::{debug, info};
/// IAM config prefix under meta bucket (e.g. config/iam/).
const IAM_CONFIG_PREFIX: &str = "config/iam";
@@ -53,6 +53,39 @@ type ListObjectVersionsInfo = StorageListObjectVersionsInfo<ObjectInfo>;
type ObjectInfoOrErr = StorageObjectInfoOrErr<ObjectInfo, Error>;
type WalkOptions = StorageWalkOptions<fn(&FileInfo) -> bool>;
#[derive(Clone, Debug, thiserror::Error)]
enum MigrationMetadataError {
#[error("empty legacy metadata: {0}")]
Empty(String),
#[error("incompatible legacy metadata: {0}")]
Incompatible(String),
}
impl From<MigrationMetadataError> for Error {
fn from(error: MigrationMetadataError) -> Self {
let message = match &error {
MigrationMetadataError::Empty(_) => "empty legacy metadata",
MigrationMetadataError::Incompatible(_) => "incompatible legacy metadata",
};
// Keep the record path in the typed source, not in the quorum grouping key.
Self::other_with_context(message, error)
}
}
/// Converts a migration failure at the startup boundary, rendering the safe
/// record path while leaving storage-layer error grouping stable.
pub fn migration_startup_error(error: Error) -> std::io::Error {
if let Error::Io(io_error) = &error
&& let Some(metadata_error) = io_error
.get_ref()
.and_then(|context| context.source())
.and_then(|source| source.downcast_ref::<MigrationMetadataError>())
{
return std::io::Error::other(metadata_error.clone());
}
std::io::Error::other(error)
}
/// Callback used to decrypt an at-rest config blob during MinIO -> RustFS migration.
///
/// MinIO encrypts IAM identity/service-account files and the server config at rest
@@ -211,7 +244,7 @@ fn normalize_bucket_meta_blob(path: &str, data: &[u8]) -> std::result::Result<Op
/// Uses list_bucket (from disk volumes) to get bucket names, since list_objects_v2 on the legacy
/// meta bucket may not work (legacy format differs from object layer expectations).
/// Skips buckets that already exist in RustFS (idempotent).
pub async fn try_migrate_bucket_metadata<S>(store: Arc<S>)
pub async fn try_migrate_bucket_metadata<S>(store: Arc<S>) -> Result<()>
where
S: BucketOperations<Error = crate::error::Error>
+ ObjectIO<
@@ -231,25 +264,18 @@ where
DeletedObject = DeletedObject,
>,
{
let buckets_list = match store
let buckets_list = store
.list_bucket(&BucketOptions {
no_metadata: true,
..Default::default()
})
.await
{
Ok(b) => b,
Err(e) => {
warn!("list buckets failed (skip migration): {e}");
return;
}
};
.await?;
let buckets: Vec<String> = buckets_list.into_iter().map(|b| b.name).collect();
if buckets.is_empty() {
debug!("No migrating bucket metadata found");
return;
return Ok(());
}
debug!("Found {} migrating bucket metadata, migrating...", buckets.len());
@@ -263,26 +289,40 @@ where
for bucket in buckets {
let meta_path = format!("{BUCKET_META_PREFIX}{SLASH_SEPARATOR}{bucket}{SLASH_SEPARATOR}{BUCKET_METADATA_FILE}");
migrate_one_if_missing(store.clone(), &opts, &h, &meta_path, &format!("bucket metadata: {bucket}")).await;
migrate_one_if_missing(store.clone(), &opts, &h, &meta_path, &format!("bucket metadata: {bucket}")).await?;
let resync_path = format!(
"{BUCKET_META_PREFIX}{SLASH_SEPARATOR}{bucket}{SLASH_SEPARATOR}{REPLICATION_META_DIR}{SLASH_SEPARATOR}{RESYNC_META_FILE}"
);
migrate_one_if_missing(store.clone(), &opts, &h, &resync_path, &format!("bucket replication resync: {bucket}")).await;
migrate_one_if_missing(store.clone(), &opts, &h, &resync_path, &format!("bucket replication resync: {bucket}")).await?;
}
Ok(())
}
async fn migration_target_exists<S: EcstoreObjectOperations>(store: &S, path: &str) -> Result<bool> {
match store
.get_object_info(RUSTFS_META_BUCKET, path, &ObjectOptions::default())
.await
{
Ok(_) => Ok(true),
Err(err) if is_err_strict_not_found(&err) || is_err_strict_volume_not_found(&err) => Ok(false),
Err(err) => Err(err),
}
}
async fn migrate_one_if_missing<S>(store: Arc<S>, opts: &ObjectOptions, headers: &HeaderMap, path: &str, label: &str)
async fn migrate_one_if_missing<S>(
store: Arc<S>,
opts: &ObjectOptions,
headers: &HeaderMap,
path: &str,
label: &str,
) -> Result<()>
where
S: EcstoreObjectIO + EcstoreObjectOperations,
{
if store
.get_object_info(RUSTFS_META_BUCKET, path, &ObjectOptions::default())
.await
.is_ok()
{
if migration_target_exists(store.as_ref(), path).await? {
debug!("{label} already exists in RustFS, skip");
return;
return Ok(());
}
let mut rd = match store
@@ -290,43 +330,31 @@ where
.await
{
Ok(r) => r,
Err(e) => {
debug!("read migrating {label}: {e}");
return;
}
// Ordinary RustFS deployments have no legacy bucket, and optional
// legacy settings (such as replication resync) may not exist.
Err(err) if is_err_strict_not_found(&err) || is_err_strict_volume_not_found(&err) => return Ok(()),
Err(err) => return Err(err),
};
let data = match rd.read_all().await {
Ok(d) if !d.is_empty() => d,
Ok(_) => return,
Err(e) => {
debug!("read migrating {label} body: {e}");
return;
}
};
let data = match normalize_bucket_meta_blob(path, &data) {
Ok(Some(normalized)) => normalized,
Ok(None) => data,
Err(e) => {
warn!("skip {label} migration due to incompatible format: {e}");
return;
}
};
let data = rd.read_all().await?;
if data.is_empty() {
return Err(MigrationMetadataError::Empty(path.to_owned()).into());
}
let data = normalize_bucket_meta_blob(path, &data)
.map_err(|_| MigrationMetadataError::Incompatible(path.to_owned()))?
.unwrap_or(data);
let mut put_data = PutObjReader::from_vec(data);
if let Err(e) = store.put_object(RUSTFS_META_BUCKET, path, &mut put_data, opts).await {
warn!("write {label}: {e}");
} else {
info!("Migrated {label}");
}
store.put_object(RUSTFS_META_BUCKET, path, &mut put_data, opts).await?;
info!("Migrated {label}");
Ok(())
}
/// Migrates IAM config from legacy meta bucket `config/iam/` to RustFS meta bucket.
/// Lists all objects under the IAM prefix in the source, copies each to the target if not present.
/// Skips objects that already exist in RustFS (idempotent).
/// If list_objects_v2 on the legacy bucket fails (e.g. format differs), migration is skipped.
pub async fn try_migrate_iam_config<S>(store: Arc<S>, decrypt_fn: Option<LegacyBlobDecryptFn>)
/// An absent legacy bucket is a no-op; migration errors prevent startup readiness.
pub async fn try_migrate_iam_config<S>(store: Arc<S>, decrypt_fn: Option<LegacyBlobDecryptFn>) -> Result<()>
where
S: ListOperations<
Error = crate::error::Error,
@@ -366,47 +394,36 @@ where
loop {
let list_result = match store
.clone()
.list_objects_v2(MIGRATING_META_BUCKET, &prefix, continuation, None, 500, false, None, false)
.list_objects_v2(MIGRATING_META_BUCKET, &prefix, continuation.clone(), None, 500, false, None, false)
.await
{
Ok(r) => r,
Err(e) => {
debug!("list IAM config from legacy bucket failed (skip migration): {e}");
return;
}
Err(err) if is_err_strict_volume_not_found(&err) => return Ok(()),
Err(err) => return Err(err),
};
for obj in list_result.objects {
let path = &obj.name;
if path.is_empty() || path.ends_with('/') {
// Unsupported records must not trigger target lookups, reads, or decryption.
if path != IAM_FORMAT_FILE_PATH
&& !is_identity_path(path)
&& !is_group_path(path)
&& !is_policy_doc_path(path)
&& !is_policy_mapping_path(path)
{
continue;
}
if store
.get_object_info(RUSTFS_META_BUCKET, path, &ObjectOptions::default())
.await
.is_ok()
{
if migration_target_exists(store.as_ref(), path).await? {
debug!("IAM config already exists in RustFS, skip: {path}");
continue;
}
let mut rd = match store
let mut rd = store
.get_object_reader(MIGRATING_META_BUCKET, path, None, h.clone(), &opts)
.await
{
Ok(r) => r,
Err(e) => {
debug!("read migrating IAM config {path}: {e}");
continue;
}
};
let data = match rd.read_all().await {
Ok(d) if !d.is_empty() => d,
Ok(_) => continue,
Err(e) => {
debug!("read migrating IAM config {path} body: {e}");
continue;
}
};
.await?;
let data = rd.read_all().await?;
if data.is_empty() {
return Err(MigrationMetadataError::Empty(path.to_owned()).into());
}
// MinIO encrypts IAM identity/service-account files at rest. Decrypt
// before normalizing; fall back to the raw bytes when no key applies
// (plaintext blobs, or nothing to decrypt) so existing behavior holds.
@@ -420,22 +437,17 @@ where
debug!("skip unsupported IAM config path during migration: {path}");
continue;
}
Err(e) => {
warn!("skip IAM config migration due to incompatible format, path: {path}, err: {e}");
continue;
}
// Parser errors may contain credential data. Report only the path.
Err(_) => return Err(MigrationMetadataError::Incompatible(path.to_owned()).into()),
};
let mut put_data = PutObjReader::from_vec(data);
if let Err(e) = store.put_object(RUSTFS_META_BUCKET, path, &mut put_data, &opts).await {
warn!("write IAM config {path}: {e}");
} else {
info!("Migrated IAM config: {path}");
total_migrated += 1;
}
store.put_object(RUSTFS_META_BUCKET, path, &mut put_data, &opts).await?;
info!("Migrated IAM config: {path}");
total_migrated += 1;
}
continuation = list_result.next_continuation_token.or(list_result.continuation_token);
if !list_result.is_truncated || continuation.is_none() {
continuation = next_iam_migration_page(list_result.is_truncated, continuation, list_result.next_continuation_token)?;
if continuation.is_none() {
break;
}
}
@@ -443,10 +455,74 @@ where
if total_migrated > 0 {
info!("IAM migration complete: {} object(s) migrated", total_migrated);
}
Ok(())
}
fn next_iam_migration_page(truncated: bool, previous: Option<String>, next: Option<String>) -> Result<Option<String>> {
if !truncated {
return Ok(None);
}
let next = next.filter(|token| !token.is_empty());
if next.is_none() || next == previous {
return Err(Error::other("legacy IAM migration listing did not advance"));
}
Ok(next)
}
#[cfg(test)]
mod tests {
#[test]
fn migration_errors_group_by_cause_and_retain_typed_record_context() {
use super::{Error, MigrationMetadataError};
for (make_error, message) in [
(
MigrationMetadataError::Empty as fn(String) -> MigrationMetadataError,
"empty legacy metadata",
),
(MigrationMetadataError::Incompatible, "incompatible legacy metadata"),
] {
let first: Error = make_error("buckets/first/.metadata.bin".into()).into();
let second: Error = make_error("buckets/second/.metadata.bin".into()).into();
assert_eq!(first, second, "record paths must not fragment error grouping");
assert_eq!(first.clone(), second, "cloning must preserve error grouping");
let io_error = std::io::Error::from(first);
let detail = io_error
.get_ref()
.and_then(|context| context.source())
.expect("record context must remain in the error source");
assert!(detail.downcast_ref::<MigrationMetadataError>().is_some());
assert!(detail.to_string().contains("buckets/first/.metadata.bin"));
let startup_error = super::migration_startup_error(make_error("buckets/startup/.metadata.bin".into()).into());
assert!(
startup_error
.get_ref()
.is_some_and(|source| source.is::<MigrationMetadataError>())
);
assert_eq!(startup_error.to_string(), format!("{message}: buckets/startup/.metadata.bin"));
}
assert_ne!(
Error::from(MigrationMetadataError::Empty("record".into())),
Error::from(MigrationMetadataError::Incompatible("record".into())),
"different migration failures must remain distinguishable"
);
}
#[test]
fn truncated_iam_listing_cannot_report_completed_migration() {
use super::next_iam_migration_page;
assert_eq!(next_iam_migration_page(false, Some("old".into()), None).expect("final page"), None);
assert_eq!(
next_iam_migration_page(true, Some("old".into()), Some("next".into())).expect("advancing page"),
Some("next".into())
);
for next in [None, Some(String::new()), Some("old".into())] {
assert!(next_iam_migration_page(true, Some("old".into()), next).is_err());
}
}
use super::{normalize_bucket_meta_blob, normalize_iam_config_blob};
use crate::bucket::replication::{
BucketReplicationResyncStatus, ReplicationMigrationBridge, ResyncStatusType, TargetReplicationResyncStatus,
@@ -659,6 +735,13 @@ mod tests {
.collect();
crate::bucket::metadata_sys::init_bucket_metadata_sys(ecstore.clone(), existing).await;
super::try_migrate_bucket_metadata(ecstore.clone())
.await
.expect("fresh stores do not require a legacy metadata bucket");
super::try_migrate_iam_config(ecstore.clone(), None)
.await
.expect("fresh stores do not require a legacy IAM bucket");
let meta_path = format!("{BUCKET_META_PREFIX}{SLASH_SEPARATOR}interop{SLASH_SEPARATOR}{BUCKET_METADATA_FILE}");
let put_opts = ObjectOptions::default();
@@ -680,8 +763,31 @@ mod tests {
.await
.expect("seed .minio.sys bucket metadata");
// --- Run the real startup migration. ---
super::try_migrate_bucket_metadata(ecstore.clone()).await;
// A partial import must report failure, even if the main bucket
// metadata copied successfully before an incompatible resync record.
let resync_path = format!("{BUCKET_META_PREFIX}/interop/.replication/resync.bin");
ecstore
.put_object(
MIGRATING_META_BUCKET,
&resync_path,
&mut PutObjReader::from_vec(b"invalid resync metadata".to_vec()),
&put_opts,
)
.await
.expect("seed malformed legacy resync metadata");
assert!(
super::try_migrate_bucket_metadata(ecstore.clone()).await.is_err(),
"incompatible native metadata must not be reported as a completed migration"
);
ecstore
.delete_object(MIGRATING_META_BUCKET, &resync_path, ObjectOptions::default())
.await
.expect("remove invalid optional legacy resync record");
// Retry the real startup migration after repairing the source.
super::try_migrate_bucket_metadata(ecstore.clone())
.await
.expect("native bucket metadata migration completes");
// --- The migrated `.rustfs.sys` blob must carry every MinIO config, ---
// byte-identical to the source (typed XML/JSON parsing of these fields is
+221 -2
View File
@@ -6445,6 +6445,10 @@ impl LocalDisk {
// A missing or still-populated directory is benign here; see
// is_benign_object_rmdir_error (handles the illumos/Solaris EEXIST
// convention, rustfs/rustfs#4978).
if is_dir_not_empty_error(&err) {
// A populated directory keeps its ancestors populated; no further pruning is needed.
return Ok(());
}
if !is_benign_object_rmdir_error(&err) {
warn!(
event = EVENT_DISK_LOCAL_DELETE_FAILED,
@@ -11359,6 +11363,176 @@ mod test {
(disk, dir)
}
#[tokio::test]
async fn delete_pruning_stops_at_live_metadata_below_a_guarded_ancestor() {
// Tuple fields drop in order, releasing the disk's root handle before the temporary directory.
let fixture = new_disk().await;
let (disk, _dir) = &fixture;
let base = disk.get_bucket_path(RUSTFS_META_BUCKET).expect("resolve metadata volume");
let shared = base.join("buckets");
let guard = Arc::new(
os::mkdir_all_below_existing_base_std(&shared, &base, &disk.publication_root)
.expect("retain the shared publication directory"),
);
for owned in [false, true] {
for missing_backup in [false, true] {
let transaction = Uuid::new_v4();
let object = shared.join(".bloomcycle.bin");
let rollback = object.join(transaction.to_string());
let metadata = object.join(STORAGE_FORMAT_FILE);
let backup = rollback.join(STORAGE_FORMAT_FILE_BACKUP);
fs::create_dir_all(&rollback).await.expect("create rollback directory");
fs::write(&metadata, b"committed metadata")
.await
.expect("write live metadata");
if !missing_backup {
fs::write(&backup, b"old metadata").await.expect("write rollback backup");
}
let owner: Option<Arc<dyn Send + Sync>> = if owned { Some(guard.clone()) } else { None };
let result = disk
.delete_with_namespace_owner(
RUSTFS_META_BUCKET,
&format!("buckets/.bloomcycle.bin/{transaction}/{STORAGE_FORMAT_FILE_BACKUP}"),
DeleteOptions::default(),
owner,
)
.await;
assert!(!backup.exists(), "backup must be absent, owned={owned}, missing={missing_backup}");
assert!(!rollback.exists(), "empty rollback directory must be pruned");
assert_eq!(fs::read(&metadata).await.expect("read committed metadata"), b"committed metadata");
result.expect("a nonempty object must stop pruning before the guarded ancestor");
}
}
}
#[tokio::test]
async fn delete_pruning_removes_empty_and_missing_ancestors_but_keeps_the_volume() {
let fixture = new_disk().await;
let (disk, _dir) = &fixture;
ensure_test_volume(disk, "pruning").await;
let base = disk.get_bucket_path("pruning").expect("resolve test volume");
for missing in [false, true] {
let parent = base.join("parent");
let rollback = parent.join("object/transaction");
fs::create_dir_all(&rollback).await.expect("create empty ancestor chain");
let path = if missing {
"parent/object/transaction/missing/xl.meta.bkp"
} else {
fs::write(rollback.join(STORAGE_FORMAT_FILE_BACKUP), b"backup")
.await
.expect("create backup");
"parent/object/transaction/xl.meta.bkp"
};
disk.delete("pruning", path, DeleteOptions::default())
.await
.expect("empty and missing ancestors should be pruned");
assert!(!parent.exists(), "the whole empty chain should be removed");
assert!(base.is_dir(), "pruning must stop at the volume boundary");
}
}
#[tokio::test]
async fn delete_pruning_does_not_remove_the_base_or_an_outside_path() {
let fixture = new_disk().await;
let (disk, dir) = &fixture;
let base = dir.path().join("base");
let outside = dir.path().join("outside");
fs::create_dir(&base).await.expect("create base");
fs::write(&outside, b"outside data").await.expect("create outside file");
disk.delete_file(&base, &base, false, false)
.await
.expect("base path is protected");
disk.delete_file(&base, &outside, false, false)
.await
.expect("outside path is protected");
assert!(base.is_dir(), "the base must not be removed even when empty");
assert_eq!(fs::read(&outside).await.expect("read outside file"), b"outside data");
}
#[cfg(windows)]
#[tokio::test]
async fn delete_pruning_propagates_a_locked_backup_error() {
use std::os::windows::fs::OpenOptionsExt;
use windows_sys::Win32::{Foundation::ERROR_SHARING_VIOLATION, Storage::FileSystem::FILE_SHARE_READ};
let fixture = new_disk().await;
let (disk, _dir) = &fixture;
ensure_test_volume(disk, "pruning").await;
let base = disk.get_bucket_path("pruning").expect("resolve test volume");
let backup = base.join(STORAGE_FORMAT_FILE_BACKUP);
fs::write(&backup, b"backup").await.expect("write backup");
let guard = std::fs::OpenOptions::new()
.read(true)
.share_mode(FILE_SHARE_READ)
.open(&backup)
.expect("hold the backup without delete sharing");
let err = disk
.delete("pruning", STORAGE_FORMAT_FILE_BACKUP, DeleteOptions::default())
.await
.expect_err("a genuine target-file deletion failure must propagate");
let DiskError::Io(err) = err else {
panic!("expected contextual I/O error, got {err:?}");
};
let context = err
.get_ref()
.and_then(|err| err.downcast_ref::<FileAccessDeniedWithContext>())
.expect("preserve the failing path and original OS error");
assert_eq!(context.path, backup);
assert_eq!(
context.source.raw_os_error(),
Some(i32::try_from(ERROR_SHARING_VIOLATION).expect("OS code fits"))
);
assert_eq!(fs::read(&backup).await.expect("backup remains readable"), b"backup");
drop(guard);
}
#[cfg(windows)]
#[tokio::test]
async fn delete_pruning_propagates_a_locked_empty_parent_error() {
use windows_sys::Win32::Foundation::ERROR_SHARING_VIOLATION;
let fixture = new_disk().await;
let (disk, _dir) = &fixture;
ensure_test_volume(disk, "pruning").await;
let base = disk.get_bucket_path("pruning").expect("resolve test volume");
let parent = base.join("parent");
let guard = os::mkdir_all_below_existing_base_std(&parent, &base, &disk.publication_root)
.expect("retain an empty parent without delete sharing");
let backup = parent.join(STORAGE_FORMAT_FILE_BACKUP);
fs::write(&backup, b"backup").await.expect("write backup");
let err = disk
.delete("pruning", "parent/xl.meta.bkp", DeleteOptions::default())
.await
.expect_err("a real parent failure without a nonempty boundary must still propagate");
let DiskError::Io(err) = err else {
panic!("expected contextual I/O error, got {err:?}");
};
let context = err
.get_ref()
.and_then(|err| err.downcast_ref::<FileAccessDeniedWithContext>())
.expect("preserve parent failure context");
assert_eq!(context.path, parent);
assert_eq!(
context.source.raw_os_error(),
Some(i32::try_from(ERROR_SHARING_VIOLATION).expect("OS code fits"))
);
assert!(!backup.exists(), "the target was removed before the parent error");
assert!(parent.is_dir(), "the guarded parent remains");
drop(guard);
disk.delete("pruning", "parent/xl.meta.bkp", DeleteOptions::default())
.await
.expect("pruning should succeed once the actual guard is released");
assert!(!parent.exists());
assert!(base.is_dir());
}
// #948: a genuinely missing source is benign and must still return Ok.
#[tokio::test]
async fn windows_and_unix_move_to_trash_missing_source_is_ok() {
@@ -11842,14 +12016,59 @@ mod test {
/// stale deterministically, instead of sleeping and hoping the filesystem
/// timestamp granularity (or a backward wall-clock step) cooperates.
fn backdate_mtime(path: &Path, age: Duration) {
use std::fs::{File, FileTimes};
use std::fs::{FileTimes, OpenOptions};
let mtime = std::time::SystemTime::now() - age;
File::open(path)
let mut options = OpenOptions::new();
options.read(true);
#[cfg(windows)]
{
use std::os::windows::fs::OpenOptionsExt;
use windows_sys::Win32::Storage::FileSystem::{FILE_FLAG_BACKUP_SEMANTICS, FILE_WRITE_ATTRIBUTES};
// Directories need backup semantics, and changing mtime needs attribute-write access.
options
.access_mode(FILE_WRITE_ATTRIBUTES)
.custom_flags(FILE_FLAG_BACKUP_SEMANTICS);
}
options
.open(path)
.expect("path should open to backdate its mtime")
.set_times(FileTimes::new().set_modified(mtime))
.expect("mtime should rewind into the past");
}
#[test]
fn cleanup_tmp_on_startup_backdate_mtime_preserves_files_and_directory_contents() {
use std::time::SystemTime;
let root = tempfile::tempdir().expect("create timestamp fixture root");
let directory = root.path().join("directory");
let file = directory.join("payload");
std::fs::create_dir(&directory).expect("create timestamp fixture directory");
std::fs::write(&file, b"unchanged payload").expect("write timestamp fixture payload");
let age = Duration::from_secs(60);
// Filesystems may round stored timestamps; do not require subsecond precision or sleep.
let rounding = Duration::from_secs(2);
for path in [&file, &directory] {
let earliest = SystemTime::now() - age - rounding;
backdate_mtime(path, age);
let latest = SystemTime::now() - age + rounding;
let modified = std::fs::metadata(path)
.expect("read backdated path metadata")
.modified()
.expect("read backdated modification time");
assert!(modified >= earliest && modified <= latest, "mtime must be backdated for {path:?}");
}
let moved = root.path().join("moved");
std::fs::rename(&directory, &moved).expect("mtime helper must release its handles before cleanup");
assert_eq!(
std::fs::read(moved.join("payload")).expect("read preserved payload"),
b"unchanged payload"
);
}
#[tokio::test]
async fn startup_cleanup_barrier_and_tmp_trash_cleanup_cover_noop_and_delete_paths() {
use tempfile::tempdir;
+41
View File
@@ -47,6 +47,43 @@ use tokio::fs;
use tracing::{info, warn};
use uuid::Uuid;
/// Hold later repair publications after admitting one baseline object. The
/// fixture arms this on one replacement disk before rejoining the cluster.
#[cfg(feature = "e2e-test-hooks")]
async fn wait_for_heal_commit_test_barrier(root: &Path, bucket: &str, object: &str) -> Result<()> {
use tokio::io::AsyncWriteExt;
let barrier = root.join(".rustfs.sys/e2e-heal-commit-barrier");
let prefix = match fs::read_to_string(&barrier).await {
Ok(prefix) => prefix,
Err(error) if error.kind() == ErrorKind::NotFound => return Ok(()),
Err(error) => return Err(error.into()),
};
let key = format!("{bucket}/{object}");
if prefix.is_empty() || !key.starts_with(&prefix) {
return Ok(());
}
let admitted = barrier.with_extension("admitted");
match fs::OpenOptions::new().write(true).create_new(true).open(&admitted).await {
Ok(mut file) => {
file.write_all(key.as_bytes()).await?;
return Ok(());
}
Err(error) if error.kind() == ErrorKind::AlreadyExists => {}
Err(error) => return Err(error.into()),
}
let deadline = tokio::time::Instant::now() + std::time::Duration::from_secs(120);
loop {
if !fs::try_exists(&barrier).await? || fs::read_to_string(&admitted).await? == key {
return Ok(());
}
if tokio::time::Instant::now() >= deadline {
return Err(std::io::Error::new(ErrorKind::TimedOut, "heal commit test barrier was not released").into());
}
tokio::time::sleep(std::time::Duration::from_millis(10)).await;
}
}
fn rollback_committed_rename_std(
dst_file_path: &Path,
new_data_path: Option<&Path>,
@@ -253,6 +290,10 @@ impl LocalDisk {
state: &mut RenameDataState,
) -> Result<RenameDataResp> {
crate::hp_guard!("LocalDisk::rename_data");
#[cfg(feature = "e2e-test-hooks")]
if fi.is_healing() {
wait_for_heal_commit_test_barrier(&self.root, dst_volume, dst_path).await?;
}
let mut fi = fi;
// A non-force DeleteBucket must not remove a directory while a local
// object commit is publishing into it. The peer's empty scan remains
@@ -1737,6 +1737,72 @@ mod tests {
aborting_full_queue_settles_pending_send().await;
}
#[tokio::test(start_paused = true)]
async fn delayed_reader_error_keeps_source_and_drops_every_encode_path() {
#[derive(Debug, thiserror::Error)]
#[error("injected request body inactivity")]
struct BodyInactivity;
#[derive(Debug)]
struct StalledReader {
data: Cursor<Vec<u8>>,
timer: Option<Pin<Box<tokio::time::Sleep>>>,
dropped: Arc<std::sync::atomic::AtomicBool>,
}
impl AsyncRead for StalledReader {
fn poll_read(mut self: Pin<&mut Self>, cx: &mut Context<'_>, buf: &mut ReadBuf<'_>) -> Poll<std::io::Result<()>> {
if self.data.position() < self.data.get_ref().len() as u64 {
return Pin::new(&mut self.data).poll_read(cx, buf);
}
let timer = self
.timer
.get_or_insert_with(|| Box::pin(tokio::time::sleep(Duration::from_secs(300))));
std::task::ready!(timer.as_mut().poll(cx));
Poll::Ready(Err(std::io::Error::other(BodyInactivity)))
}
}
impl Drop for StalledReader {
fn drop(&mut self) {
self.dropped.store(true, std::sync::atomic::Ordering::Release);
}
}
// Explicit entry points select the paths; environment caches and input
// size heuristics cannot silently turn this into repeated Vec coverage.
for path in ["direct", "vec", "bytesmut", "batched"] {
let dropped = Arc::new(std::sync::atomic::AtomicBool::new(false));
let reader = StalledReader {
data: Cursor::new(vec![7; 64]),
timer: None,
dropped: Arc::clone(&dropped),
};
let committed = Arc::new(Mutex::new(Vec::new()));
let mut writers = (0..4)
.map(|_| Some(bitrot_writer(DeferredCommitWriter::new(Arc::clone(&committed)), 32)))
.collect::<Vec<_>>();
let erasure = Arc::new(Erasure::new(2, 2, 64));
let result = match path {
"direct" => erasure.encode_single_block_non_inline(reader, &mut writers, 2).await,
"vec" => erasure.encode_with_ingest_mode(reader, &mut writers, 2, false, None).await,
"bytesmut" => erasure.encode_with_ingest_mode(reader, &mut writers, 2, true, None).await,
"batched" => erasure.encode_batched(reader, &mut writers, 2).await,
_ => unreachable!(),
};
let error = result.expect_err("stalled input must fail before shard commit");
assert!(error.get_ref().is_some_and(|source| source.is::<BodyInactivity>()), "{path}: {error:?}");
assert!(
dropped.load(std::sync::atomic::Ordering::Acquire),
"{path} must release its reader/producer before returning"
);
assert!(
committed.lock().expect("committed bytes").is_empty(),
"{path} must not commit partial shards"
);
}
}
#[tokio::test]
async fn helper_writers_cover_flush_and_shutdown_paths() {
let mut failing_write = FailingWriteWriter;
@@ -34,6 +34,14 @@ pub enum EncryptionResolutionErrorKind {
InvalidRequest,
InvalidMetadata,
ServiceUnavailable,
/// The key named by the object's metadata no longer exists in the KMS.
/// A client error on the read (the object is unreadable until the key is
/// restored), distinct from a damaged envelope.
KeyNotFound,
/// The KMS refused the unwrap for the caller's principal.
AccessDenied,
/// The configured KMS backend lacks the capability the unwrap needs.
NotImplemented,
DecryptionFailed,
}
@@ -5821,6 +5821,61 @@ mod tests {
);
}
#[tokio::test]
async fn capped_staging_queue_does_not_poll_the_part_reader() {
use futures::StreamExt;
use std::sync::atomic::{AtomicUsize, Ordering};
let (_temp_dirs, disks, set_disks) = hermetic_set_disks(4).await;
let bucket = "multipart-staging-body-demand";
let object = "object";
make_bucket_on_all(&disks, bucket).await;
let mut options = ObjectOptions::default();
insert_str(&mut options.user_defined, "max-total-object-size", "1024".to_owned());
let upload = set_disks
.new_multipart_upload(bucket, object, &options)
.await
.expect("capped upload");
let upload_path = SetDisks::get_upload_id_dir(bucket, object, &upload.upload_id);
let semaphore = capped_multipart_staging_semaphore(&upload_path);
let held = Arc::clone(&semaphore).acquire_owned().await.expect("hold staging permit");
let owners = Arc::strong_count(&semaphore);
let polls = Arc::new(AtomicUsize::new(0));
let body_polls = Arc::clone(&polls);
let stream = futures::stream::iter([Ok::<Bytes, std::io::Error>(Bytes::from(vec![7; 512]))]).inspect(move |_| {
body_polls.fetch_add(1, Ordering::Relaxed);
});
let input = tokio_util::io::StreamReader::new(stream);
let mut reader = PutObjReader::new(HashReader::from_stream(input, 512, 512, None, None, false).expect("part reader"));
let task = tokio::spawn(async move {
set_disks
.put_object_part(bucket, object, &upload.upload_id, 1, &mut reader, &ObjectOptions::default())
.await
});
tokio::time::timeout(Duration::from_secs(10), async {
while Arc::strong_count(&semaphore) == owners {
tokio::task::yield_now().await;
}
})
.await
.expect("part must reach the actual staging semaphore");
tokio::time::pause();
tokio::time::advance(Duration::from_secs(600)).await;
tokio::time::resume();
assert_eq!(polls.load(Ordering::Relaxed), 0, "staging admission must not create read demand");
assert!(!task.is_finished());
drop(held);
let part = tokio::time::timeout(Duration::from_secs(10), task)
.await
.expect("staging permit released")
.expect("part task")
.expect("queued part");
assert_eq!(part.size, 512);
assert_eq!(polls.load(Ordering::Relaxed), 1);
drop(semaphore);
remove_capped_multipart_staging_semaphore(&upload_path);
}
#[tokio::test]
async fn put_object_part_recovers_transaction_with_one_faulty_disk_at_write_quorum() {
use tokio::io::AsyncReadExt as _;
+13
View File
@@ -364,6 +364,19 @@ impl ECStore {
Ok(pieces.into_guard(bucket, registration.token))
}
/// Hold this guard through recursive-delete authorization and mutation so
/// writers cannot introduce an unchecked object into the deletion scope.
pub async fn lock_bucket_for_recursive_delete(&self, bucket: &str) -> Result<rustfs_lock::NamespaceLockGuard> {
if self.ctx.lock_manager().is_disabled() {
return Err(StorageError::InvalidArgument(
bucket.to_owned(),
String::new(),
"Recursive deletion requires namespace locking".to_owned(),
));
}
self.acquire_bucket_lifecycle_write_lock(bucket).await
}
pub(crate) async fn acquire_bucket_lifecycle_write_lock(&self, bucket: &str) -> Result<rustfs_lock::NamespaceLockGuard> {
let lock = self.new_ns_lock(bucket, BUCKET_LIFECYCLE_LOCK_OBJECT).await?;
lock.get_write_lock(get_lock_acquire_timeout())
+13 -5
View File
@@ -2921,6 +2921,12 @@ mod tests {
)
.await
.expect("seed transitioned source with locally restored bytes");
// The overwrite queues its committed cleanup owner right
// away, and from the second iteration on the restarted
// store already runs expiry workers. Fail that first remote
// DELETE so the owner stays durable and the restart below
// still has to rediscover it from xl.meta.
backend.set_remove_failure(true);
let expected = if self_copy {
payload.clone()
} else {
@@ -16426,11 +16432,13 @@ mod tests {
);
tokio::time::timeout(Duration::from_secs(30), async {
loop {
let metadata_absent = set
.load_file_info_versions_exact(bucket, object)
.await
.expect("retry cleanup metadata should remain readable")
.is_none();
// Cleanup rewrites xl.meta disk by disk, so a read racing it
// can briefly miss quorum; any other error is a real failure.
let metadata_absent = match set.load_file_info_versions_exact(bucket, object).await {
Ok(versions) => versions.is_none(),
Err(StorageError::InsufficientReadQuorum(..)) => false,
Err(err) => panic!("retry cleanup metadata should remain readable: {err:?}"),
};
if metadata_absent && backend.remove_count().await == 1 {
return;
}
+77 -7
View File
@@ -859,9 +859,78 @@ async fn delete_recursive_prefix_with_tier_delete_journal(
}
}
}
// A trailing slash selects a directory, not the object at its parent key.
// Raw filesystem recursion would also remove that object's metadata and
// data. Preserve it by purging the selected keys individually when they
// share this physical directory. The bucket write lock covers both scans.
if object.ends_with('/') && !is_meta_bucketname(bucket) {
let parent = object.strip_suffix('/').unwrap_or(object);
for pool in &store.pools {
for set in &pool.disk_set {
let page = set
.clone()
.inner_list_object_versions_for_recursive_delete(bucket, parent, None, None, 1)
.await?;
if page.objects.iter().any(|info| info.name == parent) {
return delete_directory_keys_with_tier_delete_journal(store, bucket, object, opts, tier_journal_api).await;
}
}
}
}
delete_prefix_with_tier_delete_journal(store, bucket, object, opts, tier_journal_api).await
}
async fn delete_directory_keys_with_tier_delete_journal(
store: &ECStore,
bucket: &str,
prefix: &str,
opts: &ObjectOptions,
tier_journal_api: Option<&Arc<ECStore>>,
) -> Result<()> {
for pool in &store.pools {
for set in &pool.disk_set {
let mut previous_keys = std::collections::BTreeSet::new();
loop {
// Restart after each bounded batch: its version markers have
// been deleted, and the bucket write lock excludes new keys.
let page = set
.clone()
.inner_list_object_versions_for_recursive_delete(
bucket,
prefix,
None,
None,
RECURSIVE_DELETE_VERSION_SCAN_PAGE_SIZE,
)
.await?;
let keys = page
.objects
.into_iter()
.map(|info| info.name)
.filter(|key| key.starts_with(prefix))
.collect::<std::collections::BTreeSet<_>>();
if keys.is_empty() {
break;
}
if keys == previous_keys {
return Err(Error::other("directory deletion did not advance"));
}
for key in &keys {
let encoded_key = encode_dir_object(key);
let mut exact_opts = opts.clone();
exact_opts.delete_prefix_object = true;
let _guard = store
.acquire_object_write_lock_if_needed("delete_object", bucket, &encoded_key, &mut exact_opts)
.await?;
delete_prefix_with_tier_delete_journal(store, bucket, &encoded_key, &exact_opts, tier_journal_api).await?;
}
previous_keys = keys;
}
}
}
Ok(())
}
/// A GET whose object identity has been resolved while its namespace read lock
/// remains held, but whose body reader has not been constructed yet.
///
@@ -4701,13 +4770,14 @@ impl ECStore {
return Err(Error::other("lifecycle delete-all requires namespace locking"));
}
let _bucket_lifecycle_guard = if is_meta_bucketname(bucket) {
None
} else if opts.delete_prefix {
Some(self.acquire_bucket_lifecycle_write_lock(bucket).await?)
} else {
Some(self.acquire_bucket_lifecycle_read_lock(bucket).await?)
};
let _bucket_lifecycle_guard =
if is_meta_bucketname(bucket) || (opts.delete_prefix && opts.bucket_lifecycle_lock_fence.is_some()) {
None
} else if opts.delete_prefix {
Some(self.acquire_bucket_lifecycle_write_lock(bucket).await?)
} else {
Some(self.acquire_bucket_lifecycle_read_lock(bucket).await?)
};
let object = if opts.delete_prefix && !opts.delete_prefix_object {
object.to_owned()
} else {
+4 -1
View File
@@ -563,7 +563,10 @@ fn retry_budget_for_result(task: &HealTask, result: &Result<()>, retryable_batch
}
let error = err.to_string();
if !err.is_recoverable_heal() {
// Batch aggregation preserves the typed classification in its counters,
// while the returned task error retains only the first error's display text.
let retryable_batch_result = retryable_batch_failure && matches!(err, Error::TaskExecutionFailed { .. });
if !retryable_batch_result && !err.is_recoverable_heal() {
return None;
}
+54 -19
View File
@@ -2675,27 +2675,62 @@ fn test_retry_request_for_recoverable_error_stops_at_limit() {
#[tokio::test]
async fn test_retry_request_rescans_batch_when_all_exhausted_objects_are_retryable() {
let storage: Arc<dyn HealStorageAPI> = Arc::new(MockStorage);
let task = HealTask::from_request(HealRequest::bucket("bucket".to_string()), storage);
let result = Err(task
.record_batch_failure(BatchHealFailure {
scope: "bucket:bucket".to_string(),
failed: 1,
retryable: 1,
permanent: 0,
first_object: "object".to_string(),
first_error: "Lock acquisition timeout".to_string(),
})
.await);
for source_error in [
Error::Disk(DiskError::FaultyDisk),
Error::Disk(DiskError::FaultyRemoteDisk),
Error::Storage(EcstoreError::SlowDown),
Error::TaskExecutionFailed {
message: "Lock acquisition timeout".to_string(),
},
] {
assert!(source_error.is_recoverable_heal());
let first_error = source_error.to_string();
let storage: Arc<dyn HealStorageAPI> = Arc::new(MockStorage);
let task = HealTask::from_request(HealRequest::bucket("bucket".to_string()), storage);
let result = Err(task
.record_batch_failure(BatchHealFailure {
scope: "bucket:bucket".to_string(),
failed: 1,
retryable: 1,
permanent: 0,
first_object: "object".to_string(),
first_error: first_error.clone(),
})
.await);
let (retry_request, retry_delay, error) = retry_request_for_result_with_budget(&task, &result)
.await
.expect("all-retryable batch failure should rescan within the manager retry budget");
let (retry_request, retry_delay, error) = retry_request_for_result_with_budget(&task, &result)
.await
.expect("all-retryable batch failure should rescan within the manager retry budget");
assert_eq!(retry_request.id, task.id);
assert_eq!(retry_request.retry_attempts, 1);
assert!(retry_delay > Duration::ZERO);
assert!(error.contains("Lock acquisition timeout"));
assert_eq!(retry_request.id, task.id);
assert_eq!(retry_request.retry_attempts, 1);
assert!(retry_delay > Duration::ZERO);
assert!(error.contains(&first_error));
}
}
#[tokio::test]
async fn test_retry_request_does_not_rescan_cancelled_or_timed_out_retryable_batch() {
for terminal_error in [Error::TaskCancelled, Error::TaskTimeout] {
let storage: Arc<dyn HealStorageAPI> = Arc::new(MockStorage);
let task = HealTask::from_request(HealRequest::bucket("bucket".to_string()), storage);
let _ = task
.record_batch_failure(BatchHealFailure {
scope: "bucket:bucket".to_string(),
failed: 1,
retryable: 1,
permanent: 0,
first_object: "object".to_string(),
first_error: Error::Disk(DiskError::FaultyDisk).to_string(),
})
.await;
assert!(
retry_request_for_result_with_budget(&task, &Err(terminal_error))
.await
.is_none()
);
}
}
#[tokio::test]
+34 -1
View File
@@ -100,6 +100,34 @@ async fn minio_permanent_identities_survive_migration_and_repeated_iam_loads() {
.await;
env.make_bucket(LEGACY_META_BUCKET, false).await;
for (path, body) in [
("config/iam/empty.json", Vec::new()),
("config/iam/users/ignored/extra.json", b"not JSON".to_vec()),
] {
env.put_object_bytes(LEGACY_META_BUCKET, path, body).await;
}
try_migrate_iam_config(
env.ecstore.clone(),
Some(std::sync::Arc::new(|_| panic!("unsupported IAM records must not be decrypted"))),
)
.await
.expect("unsupported IAM records, including empty objects, must be skipped");
let format_path = "config/iam/format.json";
for body in [Vec::new(), b"invalid IAM format".to_vec()] {
env.put_object_bytes(LEGACY_META_BUCKET, format_path, body).await;
let error = try_migrate_iam_config(env.ecstore.clone(), None)
.await
.expect_err("empty or incompatible supported IAM metadata must prevent startup readiness");
let io_error = std::io::Error::from(error);
let detail = io_error
.get_ref()
.and_then(|context| context.source())
.expect("failure must retain the supported record in its source");
assert!(detail.to_string().contains(format_path), "failure must identify the supported record");
}
seed_legacy_iam_object(&env, format_path, &json!({"version": 1})).await;
let regular_source = json!({
"version": 1,
"credentials": {
@@ -155,7 +183,12 @@ async fn minio_permanent_identities_survive_migration_and_repeated_iam_loads() {
)
.await;
try_migrate_iam_config(env.ecstore.clone(), None).await;
try_migrate_iam_config(env.ecstore.clone(), None)
.await
.expect("legacy IAM migration completes after source repair");
try_migrate_iam_config(env.ecstore.clone(), None)
.await
.expect("completed legacy IAM migration is idempotent");
let store = ObjectStore::new(env.ecstore);
assert_identity_survives(
+2 -1
View File
@@ -22,8 +22,9 @@ When changing key-management behavior, verify compatibility with:
For local KMS end-to-end tests, keep proxy bypass settings:
```bash
python3 scripts/e2e_binary.py build
NO_PROXY=127.0.0.1,localhost HTTP_PROXY= HTTPS_PROXY= http_proxy= https_proxy= \
cargo test --package e2e_test test_local_kms_end_to_end -- --nocapture --test-threads=1
python3 scripts/e2e_binary.py run -- cargo test --package e2e_test test_local_kms_end_to_end -- --nocapture --test-threads=1
```
### Black-box behavior suite and the Vault lane
+10 -8
View File
@@ -299,14 +299,16 @@ async fn static_backend_stateless_contract() {
assert_eq!(decrypted.plaintext, data_key.plaintext_key);
assert_key_state(backend, key_id, KeyState::Enabled).await;
expect_invalid_key_state(backend.create_key(create_request("another-key".to_string())).await, "read-only");
expect_invalid_key_state(backend.delete_key(schedule_request(key_id)).await, "read-only");
expect_invalid_key_state(backend.cancel_key_deletion(cancel_request(key_id)).await, "read-only");
// Enable/disable, rotation and rewrap are capability gaps at the product
// surface, not state-machine rejections. A single fixed key has no second
// version to rewrap onto, so reporting the gap is the only honest answer —
// re-wrapping with the same material would look like progress while
// changing nothing.
// Every mutation of the key set is a capability gap at the product surface,
// not a state-machine rejection: the backend has exactly one externally
// supplied key and no way to add, remove or alter it, so the admin API
// reports 501 for all of them rather than 400 for some.
expect_unsupported(backend.create_key(create_request("another-key".to_string())).await);
expect_unsupported(backend.delete_key(schedule_request(key_id)).await);
expect_unsupported(backend.cancel_key_deletion(cancel_request(key_id)).await);
// A single fixed key has no second version to rewrap onto, so reporting the
// gap is the only honest answer: re-wrapping with the same material would
// look like progress while changing nothing.
expect_unsupported(backend.enable_key(key_id).await);
expect_unsupported(backend.disable_key(key_id).await);
expect_unsupported(backend.rotate_key(key_id).await);
+8 -8
View File
@@ -335,7 +335,7 @@ impl KmsBackend for StaticKmsBackend {
if key_name == self.key_id {
return Err(KmsError::key_already_exists(&self.key_id));
}
Err(KmsError::invalid_operation("Static KMS is read-only: cannot create new keys"))
Err(KmsError::unsupported_capability("static", "create_key"))
}
async fn encrypt(&self, request: EncryptRequest) -> Result<EncryptResponse> {
@@ -405,14 +405,14 @@ impl KmsBackend for StaticKmsBackend {
if request.key_id != self.key_id {
return Err(KmsError::key_not_found(&request.key_id));
}
Err(KmsError::invalid_operation("Static KMS is read-only: cannot delete keys"))
Err(KmsError::unsupported_capability("static", "delete_key"))
}
async fn cancel_key_deletion(&self, request: CancelKeyDeletionRequest) -> Result<CancelKeyDeletionResponse> {
if request.key_id != self.key_id {
return Err(KmsError::key_not_found(&request.key_id));
}
Err(KmsError::invalid_operation("Static KMS is read-only: cannot cancel key deletion"))
Err(KmsError::unsupported_capability("static", "cancel_key_deletion"))
}
async fn health_check(&self) -> Result<bool> {
@@ -654,7 +654,7 @@ mod tests {
async fn test_create_key_returns_error_for_other_keys() {
let (backend, _key_id, _key) = create_test_backend().await;
// Creating any other key should return invalid operation (read-only)
// Creating any other key is a capability the read-only backend lacks.
let result = KmsBackendTrait::create_key(
&backend,
CreateKeyRequest {
@@ -663,9 +663,8 @@ mod tests {
},
)
.await;
assert!(result.is_err());
let err_msg = result.expect_err("should be Err").to_string();
assert!(err_msg.contains("read-only") || err_msg.contains("cannot create"));
let error = result.expect_err("should be Err");
assert!(matches!(error, KmsError::UnsupportedCapability { .. }), "got {error:?}");
}
#[tokio::test]
@@ -778,7 +777,8 @@ mod tests {
)
.await;
assert!(result.is_err());
assert!(result.expect_err("should be Err").to_string().contains("read-only"));
let error = result.expect_err("should be Err");
assert!(matches!(error, KmsError::UnsupportedCapability { .. }), "got {error:?}");
}
#[tokio::test]
+7
View File
@@ -214,6 +214,13 @@ impl KmsManager {
}
async fn create_key_inner(&self, request: CreateKeyRequest) -> Result<CreateKeyResponse> {
// A blank name is neither "generate one" (that is `None`) nor a usable
// id: the Local backend would write a key file with an empty stem and
// the Vault backends would address their mount root, each failing with
// a different backend-specific error.
if request.key_name.as_deref().is_some_and(|name| name.trim().is_empty()) {
return Err(KmsError::validation_error("key name must not be empty or whitespace"));
}
let response = self.backend.create_key(request).await?;
// Cache the key metadata if enabled
+30 -7
View File
@@ -43,7 +43,7 @@ use common::{
};
use rustfs_kms::{
CancelKeyDeletionRequest, CreateKeyRequest, DecryptRequest, DeleteKeyRequest, DescribeKeyRequest, EncryptRequest,
GenerateDataKeyRequest, KeySpec, KeyState, KeyStatus, KeyUsage, KmsManager, ListKeysRequest,
GenerateDataKeyRequest, KeySpec, KeyState, KeyStatus, KeyUsage, KmsError, KmsManager, ListKeysRequest,
};
async fn describe_state(kms: &KmsManager, key_id: &str) -> KeyState {
@@ -113,6 +113,29 @@ async fn created_key_is_enabled_and_fully_described() {
);
}
/// A blank name is refused by the manager before any backend sees it, so every
/// backend answers the same `ValidationError` instead of its own failure mode
/// (an empty Local key file stem, a Vault mount root, a Transit 405).
#[tokio::test]
async fn create_key_refuses_a_blank_name_before_reaching_the_backend() {
for kms in [TestKms::local().await, TestKms::static_backend().await] {
let manager = kms.kms().await;
for name in ["", " ", "\t\n"] {
let result = manager
.create_key(CreateKeyRequest {
key_name: Some(name.to_string()),
..Default::default()
})
.await;
assert!(
matches!(result, Err(KmsError::ValidationError { .. })),
"{:?} with key_name {name:?} must be a ValidationError, got {result:?}",
kms.config().backend
);
}
}
}
#[tokio::test]
async fn auto_generated_key_ids_are_unique() {
let kms = TestKms::local().await;
@@ -607,14 +630,14 @@ async fn static_backend_refuses_every_lifecycle_mutation() {
"static must advertise no lifecycle capability: {caps:?}"
);
assert_invalid_operation(
assert_unsupported_capability(
manager
.create_key(CreateKeyRequest {
key_name: Some("another-key".to_string()),
..Default::default()
})
.await,
"read-only",
"create_key",
);
// Re-creating the configured key is a conflict, not a generic refusal.
assert_key_already_exists(
@@ -628,7 +651,7 @@ async fn static_backend_refuses_every_lifecycle_mutation() {
);
let key_id = kms.config().static_config().expect("static config").key_id.clone();
assert_invalid_operation(
assert_unsupported_capability(
manager
.delete_key(DeleteKeyRequest {
key_id: key_id.clone(),
@@ -637,13 +660,13 @@ async fn static_backend_refuses_every_lifecycle_mutation() {
confirm_key_id: None,
})
.await,
"read-only",
"delete_key",
);
assert_invalid_operation(
assert_unsupported_capability(
manager
.cancel_key_deletion(CancelKeyDeletionRequest { key_id: key_id.clone() })
.await,
"read-only",
"cancel_key_deletion",
);
assert_unsupported_capability(manager.enable_key(&key_id).await, "enable_key");
assert_unsupported_capability(manager.disable_key(&key_id).await, "disable_key");
+93 -16
View File
@@ -730,12 +730,16 @@ impl ReplicationConfigurationExt for ReplicationConfiguration {
}
}
// Highest priority first, like MinIO's `FilterActionableRules`. The
// tie-breakers make this a total order: a comparator that only
// orders same-destination pairs is not transitive, and the standard
// library sort panics on such inputs past its insertion-sort
// threshold (backlog#2367 C-1).
rules.sort_by(|a, b| {
if a.destination == b.destination {
b.priority.cmp(&a.priority)
} else {
std::cmp::Ordering::Equal
}
b.priority
.cmp(&a.priority)
.then_with(|| a.destination.bucket.cmp(&b.destination.bucket))
.then_with(|| a.id.cmp(&b.id))
});
rules
@@ -813,24 +817,19 @@ impl ReplicationConfigurationExt for ReplicationConfiguration {
return vec![role.to_string()];
}
let mut arns = Vec::new();
let mut targets_map: HashSet<String> = HashSet::new();
let rules = self.filter_actionable_rules(obj);
for rule in rules {
// Rule order (priority descending) is the ARN order: callers that
// iterate targets see the highest-priority destination first.
let mut arns: Vec<String> = Vec::new();
for rule in self.filter_actionable_rules(obj) {
if rule.status == ReplicationRuleStatus::from_static(ReplicationRuleStatus::DISABLED) {
continue;
}
let arn = rule.destination.bucket.trim();
if !arn.is_empty() && !targets_map.contains(arn) {
targets_map.insert(arn.to_string());
if !arn.is_empty() && !arns.iter().any(|seen| seen == arn) {
arns.push(arn.to_string());
}
}
for arn in targets_map {
arns.push(arn);
}
arns
}
@@ -1908,6 +1907,84 @@ mod tests {
assert_eq!(decisions, vec![(target_a.to_string(), false), (target_b.to_string(), true)]);
}
// backlog#2367 C-1: the actionable-rule sort must be a total order. A
// comparator that answers `Equal` for different destinations but orders
// same-destination rules by priority is not transitive, and the standard
// library sort panics on such inputs once the slice is past the
// insertion-sort threshold (> 20 rules).
#[test]
fn actionable_rule_sort_is_a_total_order_across_destinations() {
let targets = ["arn:target:a", "arn:target:b", "arn:target:c"];
let mut seed: u64 = 0x2367;
for _ in 0..200 {
let rule_count = 21 + (seed % 200) as usize;
let rules = (0..rule_count)
.map(|index| {
seed = seed.wrapping_mul(6364136223846793005).wrapping_add(1442695040888963407);
let target = targets[(seed >> 33) as usize % targets.len()];
delete_marker_rule(&format!("r{index}"), target, "", index as i32, true)
})
.collect();
let config = ReplicationConfiguration {
role: String::new(),
rules,
};
let ordered = config.filter_actionable_rules(&ObjectOpts {
name: "logs/app.log".to_string(),
op_type: ReplicationType::Object,
..Default::default()
});
assert_eq!(ordered.len(), rule_count);
assert!(
ordered.windows(2).all(|pair| pair[0].priority >= pair[1].priority),
"actionable rules must be ordered by descending priority"
);
}
}
// backlog#2367 C-2: a V1 rule carries its prefix at the top level (no
// <Filter>). Ignoring it made `<Prefix>logs/</Prefix>` match every object.
#[test]
fn top_level_rule_prefix_scopes_matching_without_a_filter() {
let arn = "arn:target:a";
let config = ReplicationConfiguration {
role: String::new(),
rules: vec![delete_marker_rule("v1-prefix", arn, "logs/", 1, true)],
};
assert_eq!(config.rules[0].prefix(), "logs/");
let matching = config.filter_actionable_rules(&ObjectOpts {
name: "logs/app.log".to_string(),
op_type: ReplicationType::Object,
..Default::default()
});
assert_eq!(matching.len(), 1);
let outside = config.filter_actionable_rules(&ObjectOpts {
name: "data/app.log".to_string(),
op_type: ReplicationType::Object,
..Default::default()
});
assert!(outside.is_empty(), "an object outside the V1 prefix must not match: {outside:?}");
assert!(
config
.filter_target_arns(&ObjectOpts {
name: "data/app.log".to_string(),
op_type: ReplicationType::Object,
..Default::default()
})
.is_empty()
);
// A <Filter> still wins over the deprecated top-level element.
let mut filtered = delete_marker_rule("filtered", arn, "logs/", 1, true);
filtered.filter = Some(s3s::dto::ReplicationRuleFilter {
prefix: Some("photos/".to_string()),
..Default::default()
});
assert_eq!(filtered.prefix(), "photos/");
}
#[test]
fn force_delete_targets_use_overlapping_rules_and_highest_priority_switch() {
let target_a = "arn:target:a";
+5 -1
View File
@@ -22,6 +22,10 @@ pub trait ReplicationRuleExt {
}
impl ReplicationRuleExt for ReplicationRule {
/// The rule's key prefix: `Filter.Prefix`, else `Filter.And.Prefix`, else
/// the deprecated top-level `Prefix` of a V1 rule written without a
/// `<Filter>` (backlog#2367 C-2). A rule that carries both keeps AWS's
/// precedence: the `<Filter>` is authoritative.
fn prefix(&self) -> &str {
if let Some(filter) = &self.filter {
if let Some(prefix) = &filter.prefix {
@@ -32,7 +36,7 @@ impl ReplicationRuleExt for ReplicationRule {
""
}
} else {
""
self.prefix.as_deref().unwrap_or("")
}
}
+1 -1
View File
@@ -13,7 +13,7 @@ ODM configuration is stored in two additional keys in the existing bucket metada
Upgrade every node before enabling ODM. Before any rollback to rc.5, stop new migration work, retain a secure copy of the original full configuration and credentials, and disable ODM on every bucket and node. The redacted configuration GET and metadata export are not credential backups. Objects still present only at the source cannot be read through RustFS while ODM is disabled or rc.5 is running; finish migration first, redirect those reads to the source, or plan a maintenance window. After all nodes return to a compatible release, reapply and validate the saved configuration; already stored local objects remain local. Turning the global module switch off alone does not make an old metadata writer preserve these keys.
The ignored `upgrade_compatibility_test::rc5_rollback_requires_restoring_odm_configuration` test pins release commit `40a2470feb567201165a5b809b7598bb4b1f68f5`, restarts against the same data directory, writes bucket tags through rc.5, and verifies configuration recovery after returning to the current binary. Set `RUSTFS_UPGRADE_SOURCE_BINARY` to that release's executable and run `cargo test -p e2e_test rc5_rollback_requires_restoring_odm_configuration -- --ignored --test-threads=1`. The test records a known old-writer limitation; it does not certify mixed-version ODM operation.
The ignored `upgrade_compatibility_test::rc5_rollback_requires_restoring_odm_configuration` test pins release commit `40a2470feb567201165a5b809b7598bb4b1f68f5`, restarts against the same data directory, writes bucket tags through rc.5, and verifies configuration recovery after returning to the current binary. Set `RUSTFS_UPGRADE_SOURCE_BINARY` to that release's executable, build the current binary with `python3 scripts/e2e_binary.py build`, and run `python3 scripts/e2e_binary.py run -- cargo test -p e2e_test rc5_rollback_requires_restoring_odm_configuration -- --ignored --test-threads=1`. The test records a known old-writer limitation; it does not certify mixed-version ODM operation.
## Optional Google dependencies
+7 -1
View File
@@ -102,10 +102,16 @@ Cloudflare's proxy may buffer the entire request body before forwarding and can
1. Bypass the proxy. Send the failing request to `http://<host>:9000` directly. Success confirms the fault is in the proxy/CDN path.
2. Bypass the CDN, keep the proxy. Point the proxy straight at the origin (Cloudflare grey cloud / direct DNS). If it now works, the CDN was buffering or re-chunking the body.
3. Check idle reuse. Intermittent failures that correlate with upload size are almost always the keep-alive mismatch. Lower the proxy keepalive (or disable it) and retry.
4. Check for a truncated body. If the upload hangs indefinitely rather than resetting, the proxy is forwarding a partial body and then going silent without closing the connection. RustFS bounds this wait with `RUSTFS_HTTP_REQUEST_BODY_READ_TIMEOUT` (`DEFAULT_HTTP_REQUEST_BODY_READ_TIMEOUT`, 300; `0` disables) and on timeout logs `put_object_body_read_stalled` with the received/expected byte counts.
4. Check for a stalled body. A client or intermediary can stop forwarding data without closing the connection. `RUSTFS_HTTP_REQUEST_BODY_READ_TIMEOUT` defaults to 300 seconds. `PutObject` logs `put_object_body_read_stalled` when its body-read guard expires. `UploadPart` logs `upload_part_body_read_stalled` and returns `RequestTimeout` (HTTP 400); its log records `raw_bytes_received`, `expected_decoded_bytes`, `timeout_secs`, bucket, key, and request ID. The event identifies missing input progress, without attributing the cause to a particular proxy.
5. Compare bytes. Confirm the proxy forwards exactly `Content-Length` body bytes with no compression or transformation.
6. Confirm signed headers survive. `Host` and `x-amz-*` must reach RustFS unchanged; a `SignatureDoesNotMatch` (rather than a hang) points here.
For HTTP `UploadPart`, the inactivity budget counts time waiting for raw request-body bytes while storage is requesting input. Positive raw bytes reset the budget, including fragments of a signed AWS chunk that has not yet finished decoding. Foreground admission, capped-session staging, and storage backpressure do not consume the budget. Finishing the declared payload does not bypass the signed terminator, required trailers, or final body validation. This is an inactivity limit, so an upload making progress can take longer than 300 seconds overall.
Setting the timeout to `0` disables it for ordinary uploads. Multipart sessions with an explicit total-object-size cap retain a minimum 300-second timeout, including when the configured value is `0`. After a body-stall timeout, HTTP/1 uses the existing raw-body drain and closes the connection; HTTP/2 releases the affected stream and keeps the connection usable.
`UploadPart` requires a known logical byte length. RustFS uses the length normalized by S3S after authentication and decoding, with an exact logical stream length as a fallback. A bare `x-amz-decoded-content-length` or `Content-Encoding: aws-chunked` declaration cannot supply this length by itself. Requests reaching an ordinary upload session without a known length return `MissingContentLength` (HTTP 411) before body ingestion; capped sessions retain their `UnexpectedContent` rejection. Preserve the client's framing and signed headers through the proxy. The 5 GiB limit applies to each part request, not the combined size of an ordinary multipart upload.
## Known failure signatures
| Symptom | Forwarding fault | Issue |
+5 -3
View File
@@ -126,13 +126,15 @@ promtool test rules storage-rules.test.yml
The native pipeline test in
`crates/e2e_test/src/storage_metric_ownership_test.rs` requires pinned Collector,
Prometheus, previous-release RustFS, and current RustFS executables. Set
`RUSTFS_OTELCOL_BINARY`, `RUSTFS_PROMETHEUS_BINARY`,
`RUSTFS_METRICS_BASELINE_BINARY`, and `CARGO_BIN_EXE_rustfs` to those files.
`RUSTFS_OTELCOL_BINARY`, `RUSTFS_PROMETHEUS_BINARY`, and
`RUSTFS_METRICS_BASELINE_BINARY` to those files; `scripts/e2e_binary.py run`
supplies the current RustFS as `CARGO_BIN_EXE_rustfs`.
Optionally set `RUSTFS_METRICS_E2E_ARTIFACTS` to retain logs and Prometheus data.
Run only this external-tool test:
```bash
cargo test --locked -p e2e_test storage_metric_ownership_pipeline -- --ignored --nocapture
python3 scripts/e2e_binary.py build
python3 scripts/e2e_binary.py run -- cargo test --locked -p e2e_test storage_metric_ownership_pipeline -- --ignored --nocapture
```
The test first reproduces duplicated global details with the previous release,
+2 -2
View File
@@ -11,8 +11,8 @@ Pick the lowest layer that can prove the change; add a higher-layer test only wh
|---|---|---|---|
| Unit & crate integration | Per-crate logic and in-process integration tests | `cargo nextest run --all --exclude e2e_test` (or `-p <crate>`); `make test` wraps it | Every PR, required (`Test and Lint`, `ci` profile) |
| ecstore black-box | Erasure-coded read/write/recovery validation; profiles `quick` / `full` / `destructive` / `fuzz` | `scripts/run_ecstore_validation_suite.sh --profile quick` | Local and release validation only; not wired into any workflow. Contract: [ecstore-validation-suite-design.md](ecstore-validation-suite-design.md) |
| e2e (`e2e_test` crate) | A real `rustfs` binary per test, driven over the S3, admin, and protocol APIs | `cargo nextest run --profile e2e-smoke -p e2e_test` | PR: `e2e-smoke` (report-only); merge queue / main push: `e2e-full`; nightly: `e2e-repl-nightly`, `e2e-nightly`, `e2e-protocols`, `e2e-distributed`. Guide: [`crates/e2e_test/README.md`](../../crates/e2e_test/README.md); 4-node 4-disk map: [distributed-e2e.md](distributed-e2e.md) |
| Outbound target matrix | Replication of every object shape (empty, plain, retention, legal hold, multipart) against every remote-target failure mode the fake target models; an explicit expectation table pins known-red cells to an open issue | `cargo nextest run -p e2e_test -E 'test(/^replication_target_matrix_test::/)'` (build `target/debug/rustfs` first) | With `e2e-repl-nightly`; required locally for any change to outbound client defaults (SOP: [`docs/postmortems/2026-09-03-replication-checksum-default-regression.md`](../postmortems/2026-09-03-replication-checksum-default-regression.md)) |
| e2e (`e2e_test` crate) | A real `rustfs` binary per test, driven over the S3, admin, and protocol APIs | `python3 scripts/e2e_binary.py build --features e2e-test-hooks`, then `python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-smoke -p e2e_test` | PR: `e2e-smoke` (report-only); merge queue / main push: `e2e-full`; nightly: `e2e-repl-nightly`, `e2e-nightly`, `e2e-protocols`, `e2e-distributed`. Guide: [`crates/e2e_test/README.md`](../../crates/e2e_test/README.md); 4-node 4-disk map: [distributed-e2e.md](distributed-e2e.md) |
| Outbound target matrix | Replication of every object shape (empty, plain, retention, legal hold, multipart) against every remote-target failure mode the fake target models; an explicit expectation table pins known-red cells to an open issue | `python3 scripts/e2e_binary.py build`, then `python3 scripts/e2e_binary.py run -- cargo nextest run -p e2e_test -E 'test(/^replication_target_matrix_test::/)'` | With `e2e-repl-nightly`; required locally for any change to outbound client defaults (SOP: [`docs/postmortems/2026-09-03-replication-checksum-default-regression.md`](../postmortems/2026-09-03-replication-checksum-default-regression.md)) |
| s3s-e2e conformance | External S3 conformance tool against a live server | `./scripts/e2e-run.sh ./target/debug/rustfs <data-dir>` | PR, report-only (second half of the `End-to-End Tests` job) |
| S3 compatibility | `ceph/s3-tests` (boto3; allow-list `scripts/s3-tests/implemented_tests.txt`) and MinIO `mint` | `scripts/s3-tests/run.sh`; mint via `.github/workflows/mint.yml` | s3-tests: PR report-only plus a weekly full sweep; mint: weekly, report-only |
| Chaos / fault-injection | Single-node disk fault injection (`crates/e2e_test/src/chaos.rs`, `crates/e2e_test/src/fault_proxy.rs`) plus the 4-node kill/fresh-drive/blackhole cases in `crates/e2e_test/src/distributed/chaos_test.rs` | Part of the e2e crate (`e2e-reliability` and `e2e-distributed`) | Reliability cases with `e2e-full`; 4-node chaos on storage-sensitive PRs and nightly via `e2e-distributed` |
+7 -7
View File
@@ -40,9 +40,9 @@ The aggregate requires the validation lanes already selected by `ci.yml`; this c
| PR, non-doc change | `ILM Integration (serial)` | `ci.yml` `test-ilm-integration-serial` | Via aggregate | exact command in the job |
| PR, non-doc change | `Test and Lint (rio-v2)`, `Test and Lint (swift)`, `Test and Lint (sftp)` | `ci.yml` `test-and-lint-rio-v2`, `test-and-lint-protocols` | Via aggregate | `cargo nextest run` with the job's `--features` |
| PR, non-doc change | `Connect Short Credential Boundary` | `ci.yml` `connect-short-credential-boundary` | Via aggregate | `cargo test -p rustfs --test connect_registration --features connect-e2e-short-credentials`; `cargo check -p rustfs --release --features connect-e2e-short-credentials` must fail |
| PR, non-doc change | `Build RustFS Debug Binary` | `ci.yml` `build-rustfs-debug-binary` | Via aggregate; prerequisite for black-box jobs | `cargo build -p rustfs --bins --features e2e-test-hooks` |
| PR, non-doc change | `Build RustFS Debug Binary` | `ci.yml` `build-rustfs-debug-binary` | Via aggregate; prerequisite for black-box jobs | `python3 scripts/e2e_binary.py build --bins --features e2e-test-hooks` (binary plus its `rustfs.e2e.json` sidecar) |
| PR, non-doc change | `io_uring Integration (real)` | `ci.yml` `uring-integration` | Via aggregate | `cargo test -p rustfs-ecstore --lib uring_ -- --test-threads=1 --nocapture` |
| PR, non-doc change | `End-to-End Tests` | `ci.yml` `e2e-tests` | Via aggregate | `cargo nextest run --profile e2e-smoke -p e2e_test`, then `./scripts/e2e-run.sh ./target/debug/rustfs <data-dir>`; membership guards `scripts/check_test_wiring.py --check-profile e2e-smoke <listing.json>` and `scripts/check_security_smoke_count.sh check <listing.json>` |
| PR, non-doc change | `End-to-End Tests` | `ci.yml` `e2e-tests` | Via aggregate | `python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-smoke -p e2e_test`, then `./scripts/e2e-run.sh ./target/debug/rustfs <data-dir>` under the same wrapper; membership guards `scripts/check_test_wiring.py --check-profile e2e-smoke <listing.json>` and `scripts/check_security_smoke_count.sh check <listing.json>` |
| PR, non-doc change | `S3 Implemented Tests` | `ci.yml` `s3-implemented-tests` | Via aggregate | build `rustfs`, then `scripts/s3-tests/run.sh` with the job's `DEPLOY_MODE` / `TEST_MODE` / `MAXFAIL` env |
| PR, non-doc change | `S3 Lifecycle Behavior Tests` | `ci.yml` `s3-lifecycle-behavior-tests` | Via aggregate | `scripts/s3-tests/run.sh` with the job's accelerated-scanner env |
| PR touching `paths` in `audit.yml` | `Cargo Deny`, `Workflow Pin Report`, `Dependency Review` | `audit.yml` `cargo-deny`, `workflow-pin-report`, `dependency-review` | Report-only | `cargo deny check`; `scripts/security/check_workflow_pins.sh` |
@@ -51,11 +51,11 @@ The aggregate requires the validation lanes already selected by `ci.yml`; this c
| PR touching `paths` in `fuzz.yml` | `Build Fuzz Harness`, `Smoke / <target>` | `fuzz.yml` `fuzz-build`, `pr-fuzz-smoke` | Report-only | `MAX_TOTAL_TIME=60 ./scripts/fuzz/run.sh` |
| PR touching `paths` in `windows-filesystem.yml` | `Rename Safety` | `windows-filesystem.yml` `rename-safety` | Report-only | the `cargo test -p rustfs-ecstore --lib <filter>` commands in the job, on Windows |
| PR touching `paths` in `coverage.yml` | `Workspace line coverage` | `coverage.yml` `coverage` | Report-only | `make coverage`; `python3 scripts/check_security_coverage.py target/llvm-cov/coverage.json` |
| PR touching `paths` in `e2e-upgrade.yml` | `Direct upgrade from the previous release`, `Mixed-version rolling upgrade from the previous release`, `Bucket configuration survives the upgrade`, `Rollback reads current bucket metadata` | `e2e-upgrade.yml` `upgrade` matrix | Report-only | the `cargo test --locked -p e2e_test` command in the job with `RUSTFS_UPGRADE_SOURCE_BINARY` pointing at the pinned previous release (`UPGRADE_SOURCE_VERSION`) |
| PR touching `paths` in `e2e-upgrade.yml` | `Direct upgrade from the previous release`, `Mixed-version rolling upgrade from the previous release`, `Bucket configuration survives the upgrade`, `Rollback reads current bucket metadata` | `e2e-upgrade.yml` `upgrade` matrix | Report-only | `python3 scripts/e2e_binary.py build`, then the job's `python3 scripts/e2e_binary.py run -- cargo test --locked -p e2e_test` command with `RUSTFS_UPGRADE_SOURCE_BINARY` pointing at the pinned previous release (`UPGRADE_SOURCE_VERSION`) |
| PR touching `paths` in `oidc-keycloak.yml` | `OIDC Keycloak live gate` | `oidc-keycloak.yml` `oidc-keycloak-live` | Report-only | `cargo build --locked -p rustfs --bin rustfs`, then `bash scripts/test/oidc_keycloak_live.sh ./target/debug/rustfs` |
| PR touching `paths` in `targets-integration.yml` | `PostgreSQL, MySQL, AMQP, and NATS` | `targets-integration.yml` `targets-live` | Report-only | start the containers as in the job, export the `RUSTFS_TEST_*` DSNs, then the job's `cargo test --locked -p rustfs-targets --test <name> -- --ignored --test-threads=1` commands |
| PR, documentation-only selection | `Quick Checks`, `Typos`, `Test and Lint` | `ci.yml` `quick-checks`, `typos`, `required-checks` | Required directly or via aggregate | Quick Checks commands; `python3 scripts/ci_gate.py --self-test` |
| `merge_group`; push to `main` | `End-to-End Tests (full merge gate)` | `ci.yml` `e2e-full` | Via aggregate on these events | `cargo nextest run --profile e2e-full -p e2e_test` |
| `merge_group`; push to `main` | `End-to-End Tests (full merge gate)` | `ci.yml` `e2e-full` | Via aggregate on these events | `python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-full -p e2e_test` (CI adds `--binary "$RUSTFS_E2E_STARTUP_CAS_BINARY"` for the downloaded binary and sidecar) |
e2e filters live in `.config/nextest.toml`; extend a profile instead of adding a second selector. Before a profile runs, `scripts/check_test_wiring.py` compares its listing to the committed digest in `.config/e2e-<profile>-selection.txt`, so a silent test drop fails closed.
@@ -75,12 +75,12 @@ Scheduled lanes never block a PR. Their workflow-local gate fails the run, sched
|---|---|---|---|---|
| `ci.yml` (weekly) | full matrix, including the schedule/dispatch-only rio-v2 jobs `build-rustfs-debug-binary-rio-v2` and `e2e-tests-rio-v2` | strict aggregate; the full E2E lane runs on dispatch, merge groups, and main pushes | yes | dispatch `ci.yml` |
| `build.yml` (weekly) | `build-rustfs` over the six-target platform matrix in `prepare-platform-matrix` (four Linux, macOS aarch64, Windows x86_64) | build/package integrity | yes | dispatch `build.yml` with an exact platform set |
| `e2e-replication-nightly.yml` (nightly) | `repl-nightly`, `cluster-nightly`, `protocols-nightly` | three independent gates; JUnit, membership listing, server logs | yes | `cargo nextest run --profile e2e-repl-nightly -p e2e_test`; `--profile e2e-nightly`; `-j 1 --profile e2e-protocols` |
| `e2e-distributed.yml` (storage-sensitive PRs + nightly) | `distributed` | fail-closed 4-node 4-disk S3, durability, replication, movement, fault, and direct/rolling upgrade gate; JUnit, membership listing, per-node server logs | yes, with `never_ran_grace_until` | download the pinned previous release as in the workflow, export `RUSTFS_UPGRADE_SOURCE_BINARY`, then `cargo nextest run --profile e2e-distributed -p e2e_test` |
| `e2e-replication-nightly.yml` (nightly) | `repl-nightly`, `cluster-nightly`, `protocols-nightly` | three independent gates; JUnit, membership listing, server logs | yes | `python3 scripts/e2e_binary.py build --bins`, then `python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-repl-nightly -p e2e_test`; `--features e2e-test-hooks` on both steps with `--profile e2e-nightly`; `--features ftps,webdav,sftp` on both steps with `-j 1 --profile e2e-protocols` |
| `e2e-distributed.yml` (storage-sensitive PRs + nightly) | `distributed` | fail-closed 4-node 4-disk S3, durability, replication, movement, fault, and direct/rolling upgrade gate; JUnit, membership listing, per-node server logs | yes, with `never_ran_grace_until` | download the pinned previous release as in the workflow, export `RUSTFS_UPGRADE_SOURCE_BINARY`, run `python3 scripts/e2e_binary.py build --bins`, then `python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-distributed -p e2e_test` |
| `e2e-s3tests.yml` (weekly) | `s3tests` (single and distributed, four shards each), `upstream-head-canary` | compatibility gate; report, JUnit, node IDs, server logs | yes | `scripts/s3-tests/run.sh` against an existing single or distributed target |
| `fuzz.yml` (nightly) | `nightly-fuzz-corpus` per target | gate; corpus and crash artifacts | yes | `MAX_TOTAL_TIME=<seconds> ./scripts/fuzz/run.sh` |
| `minio-interop.yml` (nightly) | `minio-interop` | EC + SSE read-parity gate | yes, with `never_ran_grace_until` | pinned Docker fixture steps in the workflow |
| `on-demand-migration-interop.yml` (nightly) | `minio-source`, `cloud-source` (`aws`, `r2`, `gcs`) | report-only provider interop; one JSON report per provider naming cases, timings and source request counts, plus JUnit and MinIO logs. A cloud provider whose `ODM_INTEROP_*` secrets are absent is skipped with a summary note, not failed | no | start the pinned MinIO container as in the job, export the `RUSTFS_ODM_INTEROP_*` variables, then `cargo nextest run --profile e2e-odm-interop -p e2e_test` |
| `on-demand-migration-interop.yml` (nightly) | `minio-source`, `cloud-source` (`aws`, `r2`, `gcs`) | report-only provider interop; one JSON report per provider naming cases, timings and source request counts, plus JUnit and MinIO logs. A cloud provider whose `ODM_INTEROP_*` secrets are absent is skipped with a summary note, not failed | no | start the pinned MinIO container as in the job, export the `RUSTFS_ODM_INTEROP_*` variables, run `python3 scripts/e2e_binary.py build --bins`, then `python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-odm-interop -p e2e_test` |
| `performance-ab.yml` (nightly) | `warp-ab` | regression-budget gate; A/B summaries and server logs | yes | `bash scripts/run_hotpath_warp_abba.sh --help` |
| `nightly-gnu.yml` (nightly) | `build`, `kms-vault-lane`, `kms-vault-ha-failover` | build, live Vault, and HA failover gates | yes | commands and pinned Vault images in the workflow |
| `audit.yml` (nightly) | `cargo-deny`, `workflow-pin-report` | dependency and workflow-pin gates | yes | `cargo deny check`; `scripts/security/check_workflow_pins.sh` |
+4 -4
View File
@@ -23,7 +23,7 @@ The expansion fixture is an all-current-binary fleet, so it initializes pool met
## What this lane covers
`cargo nextest run --profile e2e-distributed -p e2e_test` selects `distributed::*`:
`python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-distributed -p e2e_test` selects `distributed::*`:
- S3 put / get / head / list / copy / rename / delete / presign, range and conditional reads, special keys, metadata, tags, pagination, empty objects, multipart complete and abort
- Object Lock COMPLIANCE, GOVERNANCE and bypass, legal hold, bucket default retention, and non-lock bucket rejection
@@ -56,7 +56,7 @@ Hardware power-loss, physical NIC pull, authenticated inter-node partition, firm
## Run
```bash
cargo build -p rustfs --bins
python3 scripts/e2e_binary.py build --bins
# Expansion/decommission/rebalance cases require four paths on distinct filesystems.
# If you do not already have four disks, sized tmpfs is enough:
# for p in 0 1 2 3; do
@@ -66,13 +66,13 @@ cargo build -p rustfs --bins
export RUSTFS_E2E_POOL_ROOTS=/mnt/rustfs-pool-0:/mnt/rustfs-pool-1:/mnt/rustfs-pool-2:/mnt/rustfs-pool-3
# Upgrade cases require the pinned previous binary (CI downloads it).
export RUSTFS_UPGRADE_SOURCE_BINARY=/path/to/rustfs-1.0.0-rc.2
cargo nextest run --profile e2e-distributed -p e2e_test
python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-distributed -p e2e_test
```
Without `RUSTFS_UPGRADE_SOURCE_BINARY` the two `distributed::upgrade_test::*` cases fail closed. Without four distinct `RUSTFS_E2E_POOL_ROOTS`, the expansion and data-movement cases fail closed. Filter upgrades out for a local run that is not checking upgrade:
```bash
cargo nextest run --profile e2e-distributed -p e2e_test -E 'not test(/^distributed::upgrade_test::/)'
python3 scripts/e2e_binary.py run -- cargo nextest run --profile e2e-distributed -p e2e_test -E 'not test(/^distributed::upgrade_test::/)'
```
The upgrade topology is `ClusterTopology::single_pool(4)` (4 nodes × 1 drive). That matches the proven mixed-version fixture in `upgrade_compatibility_test`; 4×4 localhost drives are rejected by the previous release's same-device disk check.
+3 -3
View File
@@ -25,9 +25,9 @@ Every fixed RustFS GitHub Security Advisory maps to at least one named regressio
| Layer | Command | Lane | Guard |
| --- | --- | --- | --- |
| Unit and crate tests (`ghsa_r5qv_*`, the m77q pins, `ghsa_5354_*`, `ghsa_3ppv_*`, `ghsa_6r96_*`, `ghsa_v9cp_*`, `ghsa_g3vq_*`, `ghsa_g8w9_*`) | `cargo nextest run --profile ci --all --exclude e2e_test` | every PR, `Test and Lint` (required) | none needed; the workspace pass runs every unit and crate test |
| S3-API negative-auth e2e (`negative_sigv4_test`, `presigned_negative_test`, `admin_auth_test`) | `cargo nextest run --profile e2e-smoke -p e2e_test` | every PR, `End-to-End Tests` (report-only) | `scripts/check_security_smoke_count.sh` with the floor in `.config/security-smoke-floor.txt`, run in the `e2e-tests` job; fails when a rename drops one of these modules out of the smoke filter |
| Other S3 e2e guards (`anonymous_access_test`) | `cargo nextest run --profile e2e-smoke -p e2e_test` | every PR, `End-to-End Tests` (report-only) | `scripts/check_test_wiring.py --check-profile e2e-smoke` digest |
| Protocol e2e (`protocols::test_protocol_core_suite`, GHSA-3p3x) | `RUSTFS_BUILD_FEATURES=ftps,webdav,sftp cargo nextest run -j 1 --profile e2e-protocols -p e2e_test` | nightly, `e2e-replication-nightly.yml` job `protocols-nightly`; not PR-gated | `scripts/check_test_wiring.py --check-profile e2e-protocols` digest |
| S3-API negative-auth e2e (`negative_sigv4_test`, `presigned_negative_test`, `admin_auth_test`) | `python3 scripts/e2e_binary.py build --features e2e-test-hooks`, then `python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-smoke -p e2e_test` | every PR, `End-to-End Tests` (report-only) | `scripts/check_security_smoke_count.sh` with the floor in `.config/security-smoke-floor.txt`, run in the `e2e-tests` job; fails when a rename drops one of these modules out of the smoke filter |
| Other S3 e2e guards (`anonymous_access_test`) | same build and run as the row above | every PR, `End-to-End Tests` (report-only) | `scripts/check_test_wiring.py --check-profile e2e-smoke` digest |
| Protocol e2e (`protocols::test_protocol_core_suite`, GHSA-3p3x) | `python3 scripts/e2e_binary.py build --features ftps,webdav,sftp`, then `python3 scripts/e2e_binary.py run --features ftps,webdav,sftp -- cargo nextest run -j 1 --profile e2e-protocols -p e2e_test` | nightly, `e2e-replication-nightly.yml` job `protocols-nightly`; not PR-gated | `scripts/check_test_wiring.py --check-profile e2e-protocols` digest |
Notes:
+59 -12
View File
@@ -38,7 +38,7 @@ use super::supervise_admin_mutation;
use crate::admin::auth::validate_admin_request;
use crate::admin::router::{AdminOperation, Operation, S3Router};
use crate::admin::runtime_sources::{current_action_credentials, current_ready_iam_handle, object_store_from_req};
use crate::admin::service::caller_identity::CallerIdentity;
use crate::admin::service::caller_identity::{CallerIdentity, oidc_profile_fields};
use crate::admin::storage_api::s3::{self, Body, S3ErrorCode, S3Request, S3Response, S3Result};
use crate::admin::utils::read_compatible_admin_body;
use crate::auth::constant_time_eq;
@@ -73,6 +73,16 @@ pub fn register_account_route(r: &mut S3Router<AdminOperation>) -> std::io::Resu
/// `GET /rustfs/admin/v3/account/info`
pub struct SelfAccountInfoHandler {}
#[derive(Debug, serde::Serialize)]
struct SelfAccountInfoResponse {
#[serde(flatten)]
account: SelfAccountInfo,
#[serde(skip_serializing_if = "Option::is_none")]
username: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
email: Option<String>,
}
#[async_trait::async_trait]
impl Operation for SelfAccountInfoHandler {
async fn call(&self, req: S3Request<Body>, _params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> {
@@ -124,17 +134,22 @@ impl Operation for SelfAccountInfoHandler {
None => return Err(s3::error(S3ErrorCode::ServiceUnavailable, "the object store is not ready")),
};
let info = SelfAccountInfo {
access_key: caller.access_key.clone(),
identity_type: caller.identity_type,
session_access_key: caller.session_access_key.clone(),
is_admin: caller.is_owner,
status,
member_of,
policies,
credentials_source: caller.credentials_source,
mutable: caller.mutability(),
mfa,
let (username, email) = oidc_profile_fields(&caller.credentials);
let info = SelfAccountInfoResponse {
account: SelfAccountInfo {
access_key: caller.access_key.clone(),
identity_type: caller.identity_type,
session_access_key: caller.session_access_key.clone(),
is_admin: caller.is_owner,
status,
member_of,
policies,
credentials_source: caller.credentials_source,
mutable: caller.mutability(),
mfa,
},
username,
email,
};
admin_json_response(req.uri.path(), &caller.credentials.secret_key, StatusCode::OK, &info)
@@ -548,6 +563,38 @@ fn validate_new_secret_key(request: &ChangePasswordRequest) -> S3Result<()> {
mod tests {
use super::*;
use crate::server::ADMIN_PREFIX;
use rustfs_madmin::account::{AccountMutability, CredentialsSource};
#[test]
fn self_account_info_response_adds_oidc_display_fields_without_changing_base_type() {
let mut response = SelfAccountInfoResponse {
account: SelfAccountInfo {
access_key: "virtual-parent".to_string(),
identity_type: IdentityType::Sts,
session_access_key: Some("temporary-key".to_string()),
is_admin: false,
status: "enabled".to_string(),
member_of: Vec::new(),
policies: Vec::new(),
credentials_source: CredentialsSource::Iam,
mutable: AccountMutability::default(),
mfa: AccountMfaSummary::default(),
},
username: Some("oidc-user".to_string()),
email: Some("oidc-user@example.test".to_string()),
};
let value = serde_json::to_value(&response).expect("serialize account response");
assert_eq!(value["access_key"], "virtual-parent");
assert_eq!(value["username"], "oidc-user");
assert_eq!(value["email"], "oidc-user@example.test");
response.username = None;
response.email = None;
let legacy_shape = serde_json::to_value(&response).expect("serialize account response without OIDC fields");
assert!(!legacy_shape.as_object().unwrap().contains_key("username"));
assert!(!legacy_shape.as_object().unwrap().contains_key("email"));
}
fn change_request(current: &str, new: &str) -> ChangePasswordRequest {
ChangePasswordRequest {
+37 -2
View File
@@ -15,6 +15,7 @@
use crate::admin::auth::authenticate_request;
use crate::admin::router::{AdminOperation, Operation, S3Router};
use crate::admin::runtime_sources::{current_action_credentials, object_store_from_req};
use crate::admin::service::caller_identity::oidc_profile_fields;
use crate::admin::storage_api::bucket::versioning_sys::BucketVersioningSys;
use crate::admin::storage_api::contract::admin::StorageAdminApi;
use crate::admin::storage_api::contract::bucket::{BucketOperations, BucketOptions};
@@ -52,6 +53,16 @@ pub struct AccountInfo {
pub struct AccountInfoHandler {}
#[derive(Debug, Serialize)]
struct AccountInfoResponse {
#[serde(flatten)]
account: rustfs_madmin::AccountInfo,
#[serde(skip_serializing_if = "Option::is_none")]
username: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
email: Option<String>,
}
pub fn register_account_info_route(r: &mut S3Router<AdminOperation>) -> std::io::Result<()> {
r.insert(
Method::GET,
@@ -242,6 +253,7 @@ impl Operation for AccountInfoHandler {
let policy_str = serde_json::to_string(&effective_policy)
.map_err(|_e| S3Error::with_message(S3ErrorCode::InternalError, "parse policy failed"))?;
let (username, email) = oidc_profile_fields(&cred);
let mut account_info = rustfs_madmin::AccountInfo {
account_name,
server: StorageAdminApi::backend_info(store.as_ref()).await,
@@ -288,8 +300,12 @@ impl Operation for AccountInfoHandler {
}
}
let data = serde_json::to_vec(&account_info)
.map_err(|_e| S3Error::with_message(S3ErrorCode::InternalError, "parse accountInfo failed"))?;
let data = serde_json::to_vec(&AccountInfoResponse {
account: account_info,
username,
email,
})
.map_err(|_e| S3Error::with_message(S3ErrorCode::InternalError, "parse accountInfo failed"))?;
let mut header = HeaderMap::new();
header.insert(CONTENT_TYPE, HeaderValue::from_static("application/json"));
@@ -305,6 +321,25 @@ mod tests {
use rustfs_policy::policy::BucketPolicy;
use s3s::dto::{Destination, ReplicationRule};
#[test]
fn accountinfo_response_adds_optional_oidc_display_fields() {
let mut response = AccountInfoResponse {
account: rustfs_madmin::AccountInfo::default(),
username: Some("oidc-user".to_string()),
email: Some("oidc-user@example.test".to_string()),
};
let value = serde_json::to_value(&response).expect("serialize accountinfo response");
assert_eq!(value["username"], "oidc-user");
assert_eq!(value["email"], "oidc-user@example.test");
response.username = None;
response.email = None;
let legacy_shape = serde_json::to_value(&response).expect("serialize accountinfo response without OIDC fields");
assert!(!legacy_shape.as_object().unwrap().contains_key("username"));
assert!(!legacy_shape.as_object().unwrap().contains_key("email"));
}
#[test]
fn test_account_info_structure() {
// Test AccountInfo struct creation and serialization
+79 -13
View File
@@ -389,7 +389,7 @@ impl Operation for CreateKeyHandler {
error = %e,
"admin kms keys state"
);
Err(s3_error!(InternalError, "failed to create key: {}", e))
Err(key_admin_s3_error("create key", &e))
}
}
}
@@ -461,7 +461,7 @@ impl Operation for DescribeKeyHandler {
error = %e,
"admin kms keys state"
);
Err(s3_error!(InternalError, "failed to describe key: {}", e))
Err(key_admin_s3_error("describe key", &e))
}
}
}
@@ -496,8 +496,8 @@ mod tests {
CancelKmsKeyDeletionRequest, CancelKmsKeyDeletionResponse, CreateKeyApiRequest, CreateKeyApiResponse,
CreateKmsKeyRequest, CreateKmsKeyResponse, DeleteKmsKeyRequest, DeleteKmsKeyResponse, DescribeKeyApiResponse,
DescribeKmsKeyResponse, GenerateDataKeyApiRequest, GenerateDataKeyApiResponse, ListKeysApiResponse, ListKmsKeysResponse,
delete_key_error_status, delete_request_from_query, extract_key_id, extract_query_params, key_impact_if_requested,
key_list_filters, kms_create_key_actions, kms_delete_key_actions, kms_describe_key_actions,
delete_request_from_query, extract_key_id, extract_query_params, key_admin_error_status, key_admin_s3_error,
key_impact_if_requested, key_list_filters, kms_create_key_actions, kms_delete_key_actions, kms_describe_key_actions,
kms_generate_data_key_actions, kms_list_keys_actions, legacy_create_key_name, parse_list_limit, scoped_key_id,
stable_json_value, wants_key_impact,
};
@@ -717,23 +717,59 @@ mod tests {
KmsError::invalid_operation("immediate deletion of key key-a is not allowed"),
KmsError::validation_error("bad input"),
] {
assert_eq!(delete_key_error_status(&error), StatusCode::BAD_REQUEST, "{error} must be a 400");
assert_eq!(key_admin_error_status(&error), StatusCode::BAD_REQUEST, "{error} must be a 400");
}
assert_eq!(delete_key_error_status(&KmsError::key_not_found("key-a")), StatusCode::NOT_FOUND);
assert_eq!(key_admin_error_status(&KmsError::key_not_found("key-a")), StatusCode::NOT_FOUND);
assert_eq!(
delete_key_error_status(&KmsError::backend_error("vault is down")),
key_admin_error_status(&KmsError::backend_error("vault is down")),
StatusCode::INTERNAL_SERVER_ERROR
);
}
/// One mapping serves create, delete and generate-data-key: a backend
/// without the capability answers 501, a taken name 409, a blank or
/// malformed name 400, and only damaged material stays a server fault.
#[test]
fn key_admin_error_status_contract() {
assert_eq!(
key_admin_error_status(&KmsError::unsupported_capability("static", "create_key")),
StatusCode::NOT_IMPLEMENTED
);
assert_eq!(key_admin_error_status(&KmsError::key_already_exists("key-a")), StatusCode::CONFLICT);
assert_eq!(
key_admin_error_status(&KmsError::validation_error("key name must not be empty or whitespace")),
StatusCode::BAD_REQUEST
);
assert_eq!(key_admin_error_status(&KmsError::invalid_key("bad name")), StatusCode::BAD_REQUEST);
assert_eq!(
key_admin_error_status(&KmsError::material_corrupt("key-a", "truncated")),
StatusCode::INTERNAL_SERVER_ERROR
);
// The XML-error routes carry the same status explicitly, since s3s
// derives none for a custom code.
let missing = key_admin_s3_error("generate data key", &KmsError::key_not_found("key-a"));
assert_eq!(missing.status_code(), Some(StatusCode::NOT_FOUND));
assert_eq!(
*missing.code(),
super::s3::S3ErrorCode::Custom(crate::error::KMS_KEY_NOT_FOUND_ERROR_CODE.into())
);
let unsupported = key_admin_s3_error("create key", &KmsError::unsupported_capability("static", "create_key"));
assert_eq!(unsupported.status_code(), Some(StatusCode::NOT_IMPLEMENTED));
assert_eq!(*unsupported.code(), super::s3::S3ErrorCode::NotImplemented);
let blank = key_admin_s3_error("create key", &KmsError::validation_error("key name must not be empty"));
assert_eq!(blank.status_code(), Some(StatusCode::BAD_REQUEST));
assert_eq!(*blank.code(), super::s3::S3ErrorCode::InvalidRequest);
}
/// A key the deployment still points at is refused with 409, not 400: the
/// request is well formed and the key exists, and what has to change to
/// make it succeed is the configuration, not the request.
#[test]
fn a_still_referenced_key_reports_a_conflict() {
let error = KmsError::key_still_referenced("key-a", vec!["bucket:sse-bucket".to_string()]);
assert_eq!(delete_key_error_status(&error), StatusCode::CONFLICT);
assert_eq!(key_admin_error_status(&error), StatusCode::CONFLICT);
}
#[test]
@@ -1600,7 +1636,7 @@ impl Operation for GenerateDataKeyHandler {
error = %e,
"admin kms keys state"
);
Err(s3_error!(InternalError, "failed to generate data key: {}", e))
Err(key_admin_s3_error("generate data key", &e))
}
}
}
@@ -1728,6 +1764,7 @@ impl Operation for CreateKmsKeyHandler {
error = %e,
"admin kms keys state"
);
let status = key_admin_error_status(&e);
let response = CreateKmsKeyResponse {
success: false,
message: format!("failed to create key: {e}"),
@@ -1741,7 +1778,7 @@ impl Operation for CreateKmsKeyHandler {
let mut headers = HeaderMap::new();
headers.insert(CONTENT_TYPE, "application/json".parse().expect("operation should succeed"));
Ok(S3Response::with_headers((StatusCode::INTERNAL_SERVER_ERROR, Body::from(data)), headers))
Ok(S3Response::with_headers((status, Body::from(data)), headers))
}
}
}
@@ -1831,10 +1868,22 @@ fn delete_request_from_query(uri: &hyper::Uri) -> Result<DeleteKmsKeyRequest, Bo
/// A rejected waiting window and a refused immediate deletion both arrive as
/// [`KmsError::InvalidOperation`], and both are the caller's input to fix, so
/// they must surface as 400 rather than as a server fault.
fn delete_key_error_status(error: &KmsError) -> StatusCode {
/// HTTP status for a KMS error on a key-management route, where the key id is
/// the resource being addressed (so a missing key is `404`, unlike the S3 data
/// path where it is a request error). Shared by create, delete and
/// generate-data-key so the same backend error does not read as a client
/// error on one route and a server fault on another.
fn key_admin_error_status(error: &KmsError) -> StatusCode {
match error {
KmsError::KeyNotFound { .. } => StatusCode::NOT_FOUND,
KmsError::InvalidOperation { .. } | KmsError::ValidationError { .. } => StatusCode::BAD_REQUEST,
KmsError::InvalidOperation { .. } | KmsError::ValidationError { .. } | KmsError::InvalidKey { .. } => {
StatusCode::BAD_REQUEST
}
// The request is well formed; the name is simply taken.
KmsError::KeyAlreadyExists { .. } => StatusCode::CONFLICT,
// A permanent gap in the configured backend (for example the read-only
// Static backend), never a missing resource and never retryable.
KmsError::UnsupportedCapability { .. } => StatusCode::NOT_IMPLEMENTED,
// Damaged or missing key material is an integrity fault of an existing
// key: it must surface as a server error, never as NOT_FOUND (the key
// exists) and never as a retryable backend outage.
@@ -1850,6 +1899,23 @@ fn delete_key_error_status(error: &KmsError) -> StatusCode {
}
}
/// The same classification for the routes that answer with an S3 error
/// document instead of a JSON body. s3s derives no status for a custom code,
/// so the status is set explicitly from `key_admin_error_status`.
fn key_admin_s3_error(action: &str, error: &KmsError) -> s3::S3Error {
let status = key_admin_error_status(error);
let code = match status {
StatusCode::NOT_FOUND => s3::S3ErrorCode::Custom(crate::error::KMS_KEY_NOT_FOUND_ERROR_CODE.into()),
StatusCode::BAD_REQUEST => s3::S3ErrorCode::InvalidRequest,
StatusCode::CONFLICT => s3::S3ErrorCode::Custom("KMS.AlreadyExistsException".into()),
StatusCode::NOT_IMPLEMENTED => s3::S3ErrorCode::NotImplemented,
_ => s3::S3ErrorCode::InternalError,
};
let mut s3_error = s3::error(code, format!("failed to {action}: {error}"));
s3_error.set_status_code(status);
s3_error
}
/// Delete a KMS key
pub struct DeleteKmsKeyHandler;
@@ -1984,7 +2050,7 @@ impl Operation for DeleteKmsKeyHandler {
error = %e,
"admin kms keys state"
);
let status = delete_key_error_status(&e);
let status = key_admin_error_status(&e);
let response = DeleteKmsKeyResponse {
success: false,
message: format!("Failed to delete key: {e}"),
File diff suppressed because it is too large Load Diff
+17
View File
@@ -73,6 +73,7 @@ const SITE_REPLICATION_RESYNC_ROUTE: &str = "/rustfs/admin/v3/site-replication/r
const SITE_REPLICATION_REPAIR_ROUTE: &str = "/rustfs/admin/v3/site-replication/repair";
const SITE_REPLICATION_REPAIR_STATUS_ROUTE: &str = "/rustfs/admin/v3/site-replication/repair/status";
const IAM_POLICY_ATTACH_ROUTE: &str = "/rustfs/admin/v3/idp/builtin/policy/attach";
const DATA_USAGE_INFO_ROUTE: &str = "/rustfs/admin/v3/datausageinfo";
const IAM_POLICY_DETACH_ROUTE: &str = "/rustfs/admin/v3/idp/builtin/policy/detach";
const IAM_POLICY_ENTITIES_ROUTE: &str = "/rustfs/admin/v3/idp/builtin/policy-entities";
const IAM_ACCESS_KEYS_BULK_ROUTE: &str = "/rustfs/admin/v3/list-access-keys-bulk";
@@ -1077,12 +1078,25 @@ fn advertised_admin_capabilities() -> Vec<AdvertisedAdminCapability> {
("admin.account.mfa", HttpMethod::Get, ACCOUNT_MFA_ROUTE),
("admin.mfa.challenge", HttpMethod::Get, MFA_CHALLENGE_ROUTE),
("admin.user.mfa", HttpMethod::Get, USER_MFA_ROUTE),
// `rc du` is gated on this name. Before it was advertised the client
// inferred it from a `1.0.0-rc.` version prefix, which no longer
// matches once the server reports `1.0.0` (backlog#2367 E-2).
("admin.data-usage", HttpMethod::Get, DATA_USAGE_INFO_ROUTE),
]
.into_iter()
.map(|(name, method, route)| AdvertisedAdminCapability {
name,
status: admin_route_capability(method, route),
})
.chain(std::iter::once(AdvertisedAdminCapability {
// `rc watch` streams `GET /{bucket}?events=`, a misc extension route
// dispatched by `admin::router` rather than an admin policy route,
// so its status is not an inventory lookup (same version-prefix
// inference on the client as `admin.data-usage`).
name: "listen_notification",
status: CapabilityStatus::supported()
.with_reason("bucket listen notification (?events=) is dispatched by the admin router"),
}))
.collect()
}
@@ -1258,6 +1272,9 @@ mod tests {
"admin.iam.access-keys-bulk",
"admin.iam.access-keys-bulk.ldap",
"admin.iam.access-keys-bulk.openid",
// rc pinned these two by version prefix until 1.0.0 (backlog#2367 E-2).
"admin.data-usage",
"listen_notification",
];
for name in expected_supported {
let entry = response
@@ -2515,6 +2515,11 @@ fn normalize_table_credential_object_prefix(object_prefix: &str) -> S3Result<Str
if object_prefix.is_empty() {
return Err(s3_error!(InvalidRequest, "table credential scope prefix is empty"));
}
if crate::table_catalog::is_reserved_table_object_key(object_prefix) {
return Err(S3Error::from(ApiError::invalid_request(
"table credential scope overlaps the reserved table catalog prefix",
)));
}
if object_prefix.contains('\\') {
return Err(s3_error!(
InvalidRequest,
@@ -10618,6 +10618,10 @@ fn table_credential_scope_rejects_cross_bucket_or_unsafe_prefix() {
entry.warehouse_location = "s3://warehouse/tables/../table-id".to_string();
assert!(table_credential_scope(&entry).is_err());
let mut entry = table_entry_for_credentials();
entry.warehouse_location = "s3://warehouse/.rustfs-table".to_string();
assert!(table_credential_scope(&entry).is_err());
let mut entry = table_entry_for_credentials();
entry.metadata_location = "s3://other/.rustfs-table/metadata/00001.metadata.json".to_string();
assert!(table_credential_scope(&entry).is_err());
@@ -33,6 +33,7 @@ use rustfs_credentials::Credentials;
use rustfs_iam::federation::OIDC_VIRTUAL_PARENT_CLAIM;
use rustfs_iam::sys::is_rustfs_oidc_claims;
use rustfs_madmin::account::{AccountMutability, CredentialsSource, IdentityType};
use serde_json::Value;
/// Claim written by the Keystone middleware onto its synthesized credentials.
const KEYSTONE_ROLES_CLAIM: &str = "keystone_roles";
@@ -54,6 +55,23 @@ pub(crate) fn session_parent_identity(credentials: &Credentials) -> Option<&str>
.and_then(|value| value.as_str())
}
/// Human-readable OIDC identity metadata for self-service responses. These
/// values never replace the issuer-scoped virtual parent used for authorization.
pub(crate) fn oidc_profile_fields(credentials: &Credentials) -> (Option<String>, Option<String>) {
let Some(claims) = credentials.claims.as_ref().filter(|claims| is_rustfs_oidc_claims(claims)) else {
return (None, None);
};
let string_claim = |name| {
claims
.get(name)
.and_then(Value::as_str)
.filter(|value| !value.trim().is_empty())
.map(ToOwned::to_owned)
};
(string_claim("preferred_username"), string_claim("email"))
}
/// Why a credential may not change its own authentication material.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub(crate) enum CredentialMutationDenial {
@@ -398,6 +416,48 @@ mod tests {
assert!(!caller.mutability().password);
}
#[test]
fn oidc_profile_fields_return_normalized_display_claims() {
let mut credentials = sts_session("TEMPKEY", "oidc-parent");
credentials.claims = Some(HashMap::from([
("iss".to_string(), Value::String("rustfs-oidc".to_string())),
("oidc_provider".to_string(), Value::String("entraid".to_string())),
("sub".to_string(), Value::String("subject-123".to_string())),
("preferred_username".to_string(), Value::String("j.bruijns@pay.nl".to_string())),
("email".to_string(), Value::String("fallback@pay.nl".to_string())),
]));
assert_eq!(
oidc_profile_fields(&credentials),
(Some("j.bruijns@pay.nl".to_string()), Some("fallback@pay.nl".to_string()))
);
}
#[test]
fn oidc_profile_fields_omit_missing_blank_and_non_string_values() {
let mut credentials = sts_session("TEMPKEY", "oidc-parent");
credentials.claims = Some(HashMap::from([
("iss".to_string(), Value::String("rustfs-oidc".to_string())),
("oidc_provider".to_string(), Value::String("keycloak".to_string())),
("sub".to_string(), Value::String("subject-123".to_string())),
("preferred_username".to_string(), Value::String(" ".to_string())),
("email".to_string(), Value::Array(vec![Value::String("user@example.test".to_string())])),
]));
assert_eq!(oidc_profile_fields(&credentials), (None, None));
}
#[test]
fn oidc_profile_fields_ignore_non_oidc_claim_shapes() {
let mut credentials = sts_session("TEMPKEY", "ordinary-parent");
credentials.claims = Some(HashMap::from([
("preferred_username".to_string(), Value::String("attacker".to_string())),
("email".to_string(), Value::String("attacker@example.test".to_string())),
]));
assert_eq!(oidc_profile_fields(&credentials), (None, None));
}
#[test]
fn keystone_session_is_reported_as_federated() {
let mut credentials = sts_session("TEMPKEY", "keystone-parent");
+1 -1
View File
@@ -535,7 +535,7 @@ pub(crate) mod replication {
OperatorRuleContract, REMOTE_TARGET_CAPABILITY_CONTRACT_VERSION, REMOTE_TARGET_READ_ONLY_HISTORICAL_FIELDS,
REMOTE_TARGET_UNSUPPORTED_FIELDS, REMOTE_TARGET_WRITABLE_FIELDS, REPLICATION_CAPABILITY_CONTRACT_VERSION,
REPLICATION_READ_ONLY_HISTORICAL_FIELDS, REPLICATION_WRITABLE_FIELDS, assign_site_replication_rule_priorities,
merge_incoming_replication_config, replication_target_arn_deployment_id,
merge_incoming_replication_config, replication_target_arn_deployment_id, site_replication_rule_deployment_id,
};
pub(crate) type BucketReplicationResyncStatus = super::ecstore_bucket::replication::BucketReplicationResyncStatus;
pub(crate) type BucketStats = super::ecstore_bucket::replication::BucketStats;
+102 -3
View File
@@ -117,7 +117,8 @@ use s3s::dto::{
PutBucketNotificationConfigurationInput, PutBucketNotificationConfigurationOutput, PutBucketPolicyInput,
PutBucketPolicyOutput, PutBucketReplicationInput, PutBucketReplicationOutput, PutBucketTaggingInput, PutBucketTaggingOutput,
PutBucketVersioningInput, PutBucketVersioningOutput, PutPublicAccessBlockInput, PutPublicAccessBlockOutput,
ReplicationConfiguration, ServerSideEncryption, Tagging, Timestamp, UserMetadata, VersioningConfiguration,
ReplicationConfiguration, ServerSideEncryption, ServerSideEncryptionConfiguration, Tagging, Timestamp, UserMetadata,
VersioningConfiguration,
};
use s3s::region::Region;
use s3s::xml;
@@ -2233,6 +2234,8 @@ impl DefaultBucketUsecase {
..
} = req.input;
validate_bucket_encryption_configuration(&server_side_encryption_configuration)?;
// When SSE-KMS is set without a specific key ID, populate the default
// KMS key so that GetBucketEncryption responses include it. Clients like
// mc rely on the presence of KMSMasterKeyID to distinguish SSE-KMS from
@@ -2997,6 +3000,46 @@ impl DefaultBucketUsecase {
}
}
/// Refuse a default-encryption configuration the write path could not honour
/// as written. `bucket_default_write_sse` falls back to AES256 for any
/// algorithm it does not know, so storing one would make GetBucketEncryption
/// advertise a scheme no object is encrypted under. s3s parses `SSEAlgorithm`
/// as an open string, so the schema check has to happen here.
fn validate_bucket_encryption_configuration(config: &ServerSideEncryptionConfiguration) -> S3Result<()> {
if config.rules.is_empty() {
return Err(S3Error::with_message(
S3ErrorCode::MalformedXML,
"ServerSideEncryptionConfiguration must contain at least one Rule".to_string(),
));
}
for rule in &config.rules {
let Some(by_default) = rule.apply_server_side_encryption_by_default.as_ref() else {
return Err(S3Error::with_message(
S3ErrorCode::MalformedXML,
"Rule must contain ApplyServerSideEncryptionByDefault".to_string(),
));
};
let names_kms_key = by_default.kms_master_key_id.as_deref().is_some_and(|id| !id.is_empty());
match by_default.sse_algorithm.as_str() {
ServerSideEncryption::AWS_KMS => {}
ServerSideEncryption::AES256 if names_kms_key => {
return Err(S3Error::with_message(
S3ErrorCode::InvalidArgument,
"KMSMasterKeyID can only be specified when SSEAlgorithm is aws:kms".to_string(),
));
}
ServerSideEncryption::AES256 => {}
other => {
return Err(S3Error::with_message(
S3ErrorCode::MalformedXML,
format!("SSEAlgorithm {other} is not supported; expected AES256 or aws:kms"),
));
}
}
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
@@ -3005,7 +3048,8 @@ mod tests {
use s3s::dto::{
BucketVersioningStatus, CORSConfiguration, Destination, ExcludedPrefix, FilterRule, FilterRuleName, LifecycleExpiration,
NoncurrentVersionTransition, PublicAccessBlockConfiguration, QueueConfiguration, ReplicationRule, S3KeyFilter,
ServerSideEncryptionConfiguration, Tag, Transition, TransitionStorageClass,
ServerSideEncryptionByDefault, ServerSideEncryptionConfiguration, ServerSideEncryptionRule, Tag, Transition,
TransitionStorageClass,
};
use std::sync::Arc;
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
@@ -3020,6 +3064,59 @@ mod tests {
.unwrap_or_default()
}
fn sse_config(rules: Vec<ServerSideEncryptionRule>) -> ServerSideEncryptionConfiguration {
ServerSideEncryptionConfiguration { rules }
}
fn sse_rule(algorithm: &str, kms_key_id: Option<&str>) -> ServerSideEncryptionRule {
ServerSideEncryptionRule {
apply_server_side_encryption_by_default: Some(ServerSideEncryptionByDefault {
sse_algorithm: ServerSideEncryption::from(algorithm.to_string()),
kms_master_key_id: kms_key_id.map(|id| id.to_string()),
}),
blocked_encryption_types: None,
bucket_key_enabled: None,
}
}
/// The stored configuration must be one the write path honours as written:
/// only AES256 and aws:kms exist, and a key id belongs to aws:kms alone.
#[test]
fn put_bucket_encryption_refuses_configurations_the_write_path_cannot_honour() {
validate_bucket_encryption_configuration(&sse_config(vec![sse_rule("AES256", None)])).expect("AES256 is valid");
validate_bucket_encryption_configuration(&sse_config(vec![sse_rule("aws:kms", Some("bucket-key"))]))
.expect("aws:kms with a key is valid");
validate_bucket_encryption_configuration(&sse_config(vec![sse_rule("aws:kms", None)]))
.expect("aws:kms without a key is valid (the default key is filled in)");
validate_bucket_encryption_configuration(&sse_config(vec![sse_rule("AES256", Some(""))]))
.expect("an empty key id on AES256 is how some clients spell 'none'");
let unknown = validate_bucket_encryption_configuration(&sse_config(vec![sse_rule("AES128", None)]))
.expect_err("AES128 is not an algorithm this server encrypts with");
assert_eq!(*unknown.code(), S3ErrorCode::MalformedXML);
let misplaced = validate_bucket_encryption_configuration(&sse_config(vec![sse_rule("AES256", Some("bucket-key"))]))
.expect_err("a key id only makes sense for aws:kms");
assert_eq!(*misplaced.code(), S3ErrorCode::InvalidArgument);
let empty = validate_bucket_encryption_configuration(&sse_config(Vec::new())).expect_err("no rule, no default");
assert_eq!(*empty.code(), S3ErrorCode::MalformedXML);
let bare_rule = validate_bucket_encryption_configuration(&sse_config(vec![ServerSideEncryptionRule {
apply_server_side_encryption_by_default: None,
blocked_encryption_types: None,
bucket_key_enabled: None,
}]))
.expect_err("a rule without ApplyServerSideEncryptionByDefault configures nothing");
assert_eq!(*bare_rule.code(), S3ErrorCode::MalformedXML);
// A second rule is checked too, so a malformed one cannot hide behind a valid first rule.
let second_bad =
validate_bucket_encryption_configuration(&sse_config(vec![sse_rule("AES256", None), sse_rule("garbage", None)]))
.expect_err("every rule is validated");
assert_eq!(*second_bad.code(), S3ErrorCode::MalformedXML);
}
#[tokio::test]
async fn bucket_usecase_task_finishes_post_commit_hooks_after_parent_cancellation() {
let admission = Arc::new(Semaphore::new(1));
@@ -5033,9 +5130,11 @@ mod tests {
#[tokio::test]
async fn execute_put_bucket_encryption_returns_internal_error_when_store_uninitialized() {
// A well-formed rule, so the request reaches the store lookup instead
// of being refused by configuration validation first.
let input = PutBucketEncryptionInput::builder()
.bucket("test-bucket".to_string())
.server_side_encryption_configuration(ServerSideEncryptionConfiguration::default())
.server_side_encryption_configuration(sse_config(vec![sse_rule("AES256", None)]))
.build()
.unwrap();
+203 -88
View File
@@ -69,7 +69,7 @@ use super::storage_api::multipart_usecase::{
};
use crate::app::object::{
ConcurrencyManager, ForegroundWriteAdmission, get_concurrency_manager, guard_put_object_body_read_timeout,
put_object_body_read_timeout,
put_object_body_read_timeout, reject_oversize_single_upload,
};
use crate::app::object_data_cache::{
ObjectDataCacheAdapter, invalidate_object_data_cache_after_complete_multipart_success,
@@ -90,6 +90,7 @@ use crate::auth::{
use crate::capacity::record_capacity_write;
use crate::error::ApiError;
use crate::table_catalog;
#[cfg(test)]
use bytes::Bytes;
use futures::StreamExt;
use http::{HeaderMap, HeaderValue, Uri};
@@ -104,7 +105,7 @@ use rustfs_utils::http::{
SUFFIX_MAX_TOTAL_OBJECT_SIZE, SUFFIX_PLAINTEXT_CHECKSUM, SUFFIX_REPLICATION_GENERATION,
SUFFIX_REPLICATION_PRESERVE_CIPHERTEXT, SUFFIX_REPLICATION_STATUS, SUFFIX_REPLICATION_TIMESTAMP,
SUFFIX_SOURCE_REPLICATION_REQUEST, contains_key_str, get_consistent_str, get_header, get_source_scheme,
headers::{AMZ_CHECKSUM_TYPE, AMZ_DECODED_CONTENT_LENGTH, AMZ_OBJECT_TAGGING, AMZ_STORAGE_CLASS},
headers::{AMZ_CHECKSUM_TYPE, AMZ_OBJECT_TAGGING, AMZ_STORAGE_CLASS},
insert_str,
};
use s3s::dto::{
@@ -114,6 +115,7 @@ use s3s::dto::{
ServerSideEncryption, StreamingBlob, Timestamp, UploadPartCopyInput, UploadPartCopyOutput, UploadPartInput, UploadPartOutput,
};
use s3s::header::{X_AMZ_OBJECT_LOCK_LEGAL_HOLD, X_AMZ_OBJECT_LOCK_MODE, X_AMZ_OBJECT_LOCK_RETAIN_UNTIL_DATE};
use s3s::stream::ByteStream;
use s3s::{S3Error, S3ErrorCode, S3Request, S3Response, S3Result, s3_error};
use std::collections::{HashMap, HashSet};
use std::str::FromStr;
@@ -377,42 +379,31 @@ fn extract_request_host(headers: &HeaderMap, uri: &Uri) -> Option<String> {
.or_else(|| uri.authority().map(|authority| authority.as_str().to_string()))
}
fn decoded_content_length_from_headers(headers: &HeaderMap) -> S3Result<Option<i64>> {
let Some(val) = headers.get(AMZ_DECODED_CONTENT_LENGTH) else {
return Ok(None);
};
fn resolve_upload_part_size(content_length: Option<i64>, body: Option<&StreamingBlob>) -> S3Result<Option<i64>> {
if let Some(length) = content_length {
return Ok(Some(length));
}
body.and_then(|body| body.remaining_length().exact())
.map(i64::try_from)
.transpose()
.map_err(|_| s3_error!(UnexpectedContent))
}
match atoi::atoi::<i64>(val.as_bytes()) {
Some(x) => Ok(Some(x)),
None => Err(s3_error!(UnexpectedContent)),
fn require_upload_part_size(size: Option<i64>, capped: bool) -> S3Result<i64> {
match size {
Some(size) if size >= 0 => Ok(size),
Some(_) => Err(s3_error!(UnexpectedContent)),
None if capped => Err(s3_error!(UnexpectedContent)),
None => Err(S3Error::new(S3ErrorCode::MissingContentLength)),
}
}
fn request_uses_aws_chunked(headers: &HeaderMap) -> bool {
let has_aws_chunked = |header_name: &str| {
headers
.get(header_name)
.and_then(|value| value.to_str().ok())
.is_some_and(|value| value.split(',').any(|part| part.trim().eq_ignore_ascii_case("aws-chunked")))
};
has_aws_chunked("content-encoding") || has_aws_chunked("transfer-encoding")
}
fn resolve_upload_part_size(headers: &HeaderMap, content_length: Option<i64>) -> S3Result<Option<i64>> {
let decoded_content_length = decoded_content_length_from_headers(headers)?;
let size = match (request_uses_aws_chunked(headers), decoded_content_length, content_length) {
(true, Some(decoded), _) => Some(decoded),
(_, _, Some(length)) => Some(length),
(_, Some(decoded), None) => Some(decoded),
_ => None,
};
if size == Some(-1) {
return Err(s3_error!(UnexpectedContent));
fn upload_part_body_read_timeout(configured: Duration, capped: bool) -> Duration {
if capped {
configured.max(Duration::from_secs(rustfs_config::DEFAULT_HTTP_REQUEST_BODY_READ_TIMEOUT))
} else {
configured
}
Ok(size)
}
fn build_complete_multipart_location(headers: &HeaderMap, uri: &Uri, bucket: &str, key: &str) -> String {
@@ -1035,7 +1026,7 @@ impl DefaultMultipartUsecase {
let (effective_sse, effective_kms_key_id) = match prepared_material {
Some(material) => {
let server_side_encryption = Some(material.server_side_encryption.clone());
let ssekms_key_id = material.kms_key_id.clone();
let ssekms_key_id = material.response_kms_key_id();
let mut encryption_metadata = encryption_material_to_metadata(&material)?;
if material.key_kind == EncryptionKeyKind::Object {
@@ -1168,7 +1159,10 @@ impl DefaultMultipartUsecase {
validate_table_catalog_object_mutation(&bucket, &key).await?;
let mut size = resolve_upload_part_size(&req.headers, content_length)?;
let size = resolve_upload_part_size(content_length, body.as_ref())?;
if let Some(size) = size {
reject_oversize_single_upload(size)?;
}
let mut body_stream = body.ok_or_else(|| s3_error!(IncompleteBody))?;
let Some(store) = self.object_store() else {
return Err(S3Error::with_message(S3ErrorCode::InternalError, "Not init".to_string()));
@@ -1178,20 +1172,15 @@ impl DefaultMultipartUsecase {
.await
.map_err(ApiError::from)?;
let max_total_object_size = multipart_max_total_object_size(&fi.user_defined)?;
if max_total_object_size.is_some() && size.is_some_and(|size| size < 0) {
return Err(S3Error::new(S3ErrorCode::UnexpectedContent));
}
if max_total_object_size.is_some() && size.is_none() {
return Err(S3Error::new(S3ErrorCode::UnexpectedContent));
}
if let (Some(limit), Some(size)) = (max_total_object_size, size)
let mut size = require_upload_part_size(size, max_total_object_size.is_some())?;
if let Some(limit) = max_total_object_size
&& u64::try_from(size).is_ok_and(|size| size > limit)
{
return Err(S3Error::new(S3ErrorCode::EntityTooLarge));
}
let upload_part_admission = match self
.concurrency_manager()
.admit_multipart_part(size.unwrap_or(-1))
.admit_multipart_part(size)
.await
.map_err(|_| S3Error::with_message(S3ErrorCode::InternalError, "foreground write admission closed"))?
{
@@ -1208,42 +1197,30 @@ impl DefaultMultipartUsecase {
));
}
};
if max_total_object_size.is_some() {
let request_id = req
.extensions
.get::<super::storage_api::multipart_usecase::request_context::RequestContext>()
.map(|ctx| ctx.request_id.clone())
.unwrap_or_default();
body_stream = guard_put_object_body_read_timeout(
body_stream,
let request_id = req
.extensions
.get::<super::storage_api::multipart_usecase::request_context::RequestContext>()
.map(|ctx| ctx.request_id.as_str())
.unwrap_or_default();
let timeout = upload_part_body_read_timeout(put_object_body_read_timeout(), max_total_object_size.is_some());
let raw_control = req.extensions.get::<super::object::request_body::BodyReadControl>().cloned();
let observe_read_demand = if let Some(control) = &raw_control {
control.activate(
timeout,
&bucket,
&key,
&request_id,
content_length,
put_object_body_read_timeout().max(Duration::from_secs(rustfs_config::DEFAULT_HTTP_REQUEST_BODY_READ_TIMEOUT)),
);
}
if size.is_none() {
let mut total = 0i64;
let mut buffer = bytes::BytesMut::new();
while let Some(chunk) = body_stream.next().await {
let chunk = chunk.map_err(|e| ApiError::from(s3s_body_error_to_io(e)))?;
total += chunk.len() as i64;
buffer.extend_from_slice(&chunk);
request_id,
u64::try_from(size).map_err(|_| s3_error!(UnexpectedContent))?,
)
} else {
// Direct protocol callers have no raw HTTP body. Retain their
// existing capped-session guard without inventing a client cause.
if max_total_object_size.is_some() {
body_stream = guard_put_object_body_read_timeout(body_stream, &bucket, &key, request_id, Some(size), timeout);
}
false
};
if total <= 0 {
return Err(s3_error!(UnexpectedContent));
}
size = Some(total);
let combined = buffer.freeze();
let stream = futures::stream::once(async move { Ok::<Bytes, std::io::Error>(combined) });
body_stream = StreamingBlob::wrap(stream);
}
let mut size = size.ok_or_else(|| s3_error!(UnexpectedContent))?;
let ingress_stage_start = rustfs_io_metrics::put_stage_metrics_enabled().then(std::time::Instant::now);
// Apply adaptive buffer sizing based on part size for optimal streaming performance.
@@ -1388,6 +1365,11 @@ impl DefaultMultipartUsecase {
};
reader = write_plan.apply(reader, actual_size).map_err(ApiError::from)?;
if observe_read_demand && let Some(control) = raw_control {
use rustfs_rio::HashReaderMut;
let inner = reader.take_inner();
reader.inner = rustfs_rio::boxed_reader(super::object::request_body::DemandReader::new(inner, control));
}
let mut reader = PutObjReader::new(reader);
@@ -1932,6 +1914,8 @@ fn passthrough_part_actual_size(headers: &HeaderMap) -> Option<i64> {
#[cfg(test)]
mod tests {
use super::*;
mod body_read_tests;
use http::{Extensions, HeaderMap, Method, Uri, header::HeaderValue};
use rustfs_filemeta::ObjectPartInfo;
use rustfs_utils::http::{
@@ -2136,23 +2120,45 @@ mod tests {
}
#[test]
fn resolve_upload_part_size_uses_decoded_length_for_aws_chunked() {
let mut headers = HeaderMap::new();
headers.insert("content-encoding", HeaderValue::from_static("aws-chunked"));
headers.insert(AMZ_DECODED_CONTENT_LENGTH, HeaderValue::from_static("5242880"));
let size = resolve_upload_part_size(&headers, Some(5242962)).expect("decoded size should parse");
assert_eq!(size, Some(5242880));
fn resolve_upload_part_size_uses_normalized_logical_length() {
assert_eq!(resolve_upload_part_size(Some(5242880), None).expect("DTO length"), Some(5242880));
assert_eq!(resolve_upload_part_size(None, None).expect("unknown length"), None);
let body = StreamingBlob::from(Bytes::from_static(b"abc"));
assert_eq!(resolve_upload_part_size(None, Some(&body)).expect("exact bytes"), Some(3));
assert_eq!(resolve_upload_part_size(Some(0), Some(&body)).expect("explicit length wins"), Some(0));
let body = StreamingBlob::wrap(futures::stream::iter([Ok::<_, std::io::Error>(Bytes::from_static(b"abc"))]));
assert_eq!(futures::Stream::size_hint(&body), (1, Some(1)), "the stream knows its item count");
assert_eq!(resolve_upload_part_size(None, Some(&body)).expect("unknown byte length"), None);
}
#[test]
fn resolve_upload_part_size_preserves_regular_content_length() {
let headers = HeaderMap::new();
fn upload_part_length_contract_rejects_unknown_and_all_negative_lengths() {
assert_eq!(
require_upload_part_size(None, false).expect_err("ordinary unknown").code(),
&S3ErrorCode::MissingContentLength
);
assert_eq!(
require_upload_part_size(None, true).expect_err("capped unknown").code(),
&S3ErrorCode::UnexpectedContent
);
for capped in [false, true] {
for size in [-1, -2, i64::MIN] {
assert_eq!(
require_upload_part_size(Some(size), capped).expect_err("negative").code(),
&S3ErrorCode::UnexpectedContent
);
}
assert_eq!(require_upload_part_size(Some(0), capped).expect("zero is valid"), 0);
}
}
let size = resolve_upload_part_size(&headers, Some(5242880)).expect("regular size should parse");
assert_eq!(size, Some(5242880));
#[test]
fn upload_part_timeout_policy_preserves_disabled_and_capped_floor() {
for seconds in [0, 1, 299, 300, 601] {
let configured = Duration::from_secs(seconds);
assert_eq!(upload_part_body_read_timeout(configured, false), configured);
assert_eq!(upload_part_body_read_timeout(configured, true), Duration::from_secs(seconds.max(300)));
}
}
#[test]
@@ -3209,6 +3215,36 @@ mod tests {
assert_eq!(err.code(), &S3ErrorCode::IncompleteBody);
}
/// issue #7596: a part whose declared length exceeds the 5 GiB
/// single-request ceiling is rejected before the body is polled or the
/// store is consulted. Exact-cap and zero-length parts pass admission.
#[tokio::test]
async fn execute_upload_part_rejects_oversize_declared_part_before_reading_the_body() {
let ceiling = i64::try_from(rustfs_config::MAX_SINGLE_PUT_OBJECT_SIZE).expect("ceiling fits i64");
for (declared, expect_too_large) in [(ceiling + 1, true), (ceiling, false), (0, false)] {
let (body, polls) = crate::app::object::PollCountingBody::streaming_blob();
let input = UploadPartInput::builder()
.bucket("bucket".to_string())
.key("object".to_string())
.upload_id("upload-id".to_string())
.part_number(1)
.body(Some(body))
.content_length(Some(declared))
.build()
.unwrap();
let req = build_request(input, Method::PUT);
let err = make_usecase().execute_upload_part(req).await.unwrap_err();
if expect_too_large {
assert_eq!(err.code(), &S3ErrorCode::EntityTooLarge, "declared {declared}");
assert_eq!(polls.load(std::sync::atomic::Ordering::SeqCst), 0, "body must not be polled");
} else {
assert_ne!(err.code(), &S3ErrorCode::EntityTooLarge, "declared {declared} must pass admission");
}
}
}
#[tokio::test]
async fn execute_upload_part_rejects_invalid_part_number_before_body_lookup() {
for part_number in [-1, 0, 10001] {
@@ -3227,6 +3263,85 @@ mod tests {
}
}
#[tokio::test]
#[serial_test::serial]
async fn execute_upload_part_rejects_unknown_length_before_admission_or_body_polling() {
use crate::app::storage_api::test::contract::bucket::{BucketOperations, MakeBucketOptions};
let store = crate::app::gating_test_env::shared_gating_ecstore().await;
let ambient = crate::app::gating_test_env::shared_gating_ambient().await;
let context = Arc::new(AppContext::new(Arc::clone(&store), ambient.iam(), ambient.kms()));
let bucket = format!("upload-part-length-{}", Uuid::new_v4().simple());
store
.make_bucket(&bucket, &MakeBucketOptions::default())
.await
.expect("bucket");
let concurrency_manager = Arc::new(ConcurrencyManager::with_large_put_admission_for_test(
true,
1,
rustfs_config::DEFAULT_PUT_LARGE_FOREGROUND_ADMISSION_MIN_SIZE_BYTES,
Duration::ZERO,
));
let held = concurrency_manager
.admit_multipart_part(1024)
.await
.expect("hold the only permit");
let usecase = DefaultMultipartUsecase::with_context_and_concurrency_manager(Some(context), concurrency_manager);
for capped in [false, true] {
let mut options = ObjectOptions::default();
if capped {
insert_str(&mut options.user_defined, SUFFIX_MAX_TOTAL_OBJECT_SIZE, "1024".to_owned());
}
let upload = store
.new_multipart_upload(&bucket, "object", &options)
.await
.expect("upload session");
for content_length in [None, Some(-1)] {
for declared_chunk_encoding in [false, true] {
let (body, polls) = crate::app::object::PollCountingBody::streaming_blob();
let input = UploadPartInput::builder()
.bucket(bucket.clone())
.key("object".to_owned())
.upload_id(upload.upload_id.clone())
.part_number(1)
.content_length(content_length)
.body(Some(body))
.build()
.expect("part request");
let mut request = build_request(input, Method::PUT);
request
.headers
.insert("x-amz-decoded-content-length", HeaderValue::from_static("1024"));
if declared_chunk_encoding {
request
.headers
.insert("content-encoding", HeaderValue::from_static("aws-chunked"));
}
let error = usecase
.execute_upload_part(request)
.await
.expect_err("unknown or negative logical size");
assert_eq!(
error.code(),
&if capped || content_length.is_some() {
S3ErrorCode::UnexpectedContent
} else {
S3ErrorCode::MissingContentLength
}
);
assert_eq!(polls.load(std::sync::atomic::Ordering::Relaxed), 0);
}
}
let parts = store
.list_object_parts(&bucket, "object", &upload.upload_id, None, 1000, &ObjectOptions::default())
.await
.expect("list rejected session");
assert!(parts.parts.is_empty());
}
drop(held);
}
#[tokio::test]
#[serial_test::serial]
async fn execute_upload_part_rejects_when_foreground_write_admission_is_full() {
@@ -0,0 +1,310 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
use super::*;
use crate::app::object::request_body::{BodyReadControl, ObservedBody};
use crate::app::storage_api::s3::Body as S3Body;
use crate::app::storage_api::test::contract::bucket::{BucketOperations, MakeBucketOptions};
use crate::app::storage_api::test::contract::object::ObjectIO;
use http_body::Frame;
use std::sync::atomic::{AtomicUsize, Ordering};
use tokio::sync::{mpsc, oneshot};
fn part_request(bucket: &str, upload: &str, body: StreamingBlob, size: i64) -> S3Request<UploadPartInput> {
build_request(
UploadPartInput::builder()
.bucket(bucket.to_owned())
.key("object".to_owned())
.upload_id(upload.to_owned())
.part_number(1)
.content_length(Some(size))
.body(Some(body))
.build()
.expect("part input"),
Method::PUT,
)
}
type BodySender = mpsc::UnboundedSender<Result<Frame<Bytes>, std::io::Error>>;
fn observed_request(
bucket: &str,
upload: &str,
size: usize,
) -> (S3Request<UploadPartInput>, BodySender, oneshot::Receiver<()>, Arc<AtomicUsize>) {
let (sender, mut receiver) = mpsc::unbounded_channel();
let (started, waiting) = oneshot::channel();
let mut started = Some(started);
let polls = Arc::new(AtomicUsize::new(0));
let body_polls = Arc::clone(&polls);
let stream = futures::stream::poll_fn(move |cx| {
body_polls.fetch_add(1, Ordering::Relaxed);
let result = receiver.poll_recv(cx);
if result.is_pending()
&& let Some(started) = started.take()
{
let _ = started.send(());
}
result
});
let control = BodyReadControl::default();
let body = ObservedBody::new(http_body_util::StreamBody::new(stream), control.clone());
let mut request = part_request(bucket, upload, StreamingBlob::from(S3Body::http_body_unsync(body)), size as i64);
request.extensions.insert(control);
(request, sender, waiting, polls)
}
async fn temporary_entries(disks: &[std::path::PathBuf]) -> std::collections::BTreeSet<std::path::PathBuf> {
let mut entries = std::collections::BTreeSet::new();
for disk in disks {
let path = disk.join(".rustfs.sys/tmp");
let mut directory = match tokio::fs::read_dir(path).await {
Ok(directory) => directory,
Err(error) if error.kind() == std::io::ErrorKind::NotFound => continue,
Err(error) => panic!("temporary directory: {error}"),
};
while let Some(entry) = directory.next_entry().await.expect("temporary entry") {
if entry.file_name() != ".trash" {
entries.insert(entry.path());
}
}
}
entries
}
#[test]
#[serial_test::serial]
fn upload_part_body_timeout_cleans_storage_preserves_old_part_and_releases_permits() {
crate::app::gating_test_env::run_large_stack_test("upload-part-timeout-direct", || async {
assert_body_timeout_storage_lifecycle(4096, 512).await;
});
}
#[test]
#[serial_test::serial]
fn upload_part_body_timeout_pipeline_storage_lifecycle() {
// Fresh processes with the existing ingest and batching settings cover
// Vec, BytesMut, and batched pipelines independently.
crate::app::gating_test_env::run_large_stack_test("upload-part-timeout-pipeline", || async {
assert_body_timeout_storage_lifecycle(3 * 1024 * 1024, 2 * 1024 * 1024 + 512).await;
});
}
async fn assert_body_timeout_storage_lifecycle(part_size: usize, partial_size: usize) {
use metrics_util::debugging::{DebugValue, DebuggingRecorder};
struct RestoreMetrics(bool);
impl Drop for RestoreMetrics {
fn drop(&mut self) {
rustfs_io_metrics::set_put_stage_metrics_enabled(self.0);
}
}
let _restore = RestoreMetrics(rustfs_io_metrics::put_stage_metrics_enabled());
let recorder = DebuggingRecorder::new();
let snapshotter = recorder.snapshotter();
let _recorder = metrics::set_default_local_recorder(&recorder);
rustfs_io_metrics::set_put_stage_metrics_enabled(true);
let path = if part_size == 4096 {
"multipart_write_single_block_non_inline"
} else if part_size >= rustfs_utils::get_env_usize("RUSTFS_MULTIPART_PUT_LARGE_BATCH_MIN_SIZE_BYTES", 128 * 1024 * 1024) {
"multipart_write_pipeline_batched_large"
} else {
"multipart_write_pipeline"
};
let path_count = || {
snapshotter
.snapshot()
.into_vec()
.into_iter()
.filter_map(|(key, _, _, value)| {
if key.key().name() == "rustfs_s3_put_object_path_total"
&& key.key().labels().any(|label| label.key() == "path" && label.value() == path)
&& let DebugValue::Counter(count) = value
{
Some(count)
} else {
None
}
})
.sum::<u64>()
};
let (disks, store) = crate::app::gating_test_env::shared_gating_ecstore_and_disk_paths().await;
let ambient = crate::app::gating_test_env::shared_gating_ambient().await;
let context = Arc::new(AppContext::new(Arc::clone(&store), ambient.iam(), ambient.kms()));
let manager = Arc::new(ConcurrencyManager::with_large_put_admission_for_test(
true,
1,
rustfs_config::DEFAULT_PUT_LARGE_FOREGROUND_ADMISSION_MIN_SIZE_BYTES,
Duration::ZERO,
));
let usecase = Arc::new(DefaultMultipartUsecase::with_context_and_concurrency_manager(
Some(context),
Arc::clone(&manager),
));
let bucket = format!("body-stall-{}", Uuid::new_v4().simple());
store
.make_bucket(&bucket, &MakeBucketOptions::default())
.await
.expect("bucket");
for capped in [false, true] {
let mut options = ObjectOptions::default();
if capped {
insert_str(&mut options.user_defined, SUFFIX_MAX_TOTAL_OBJECT_SIZE, (2 * part_size).to_string());
}
let upload = store
.new_multipart_upload(&bucket, "object", &options)
.await
.expect("upload session");
let mut old_etag = None;
for (replacement, retry_after_failure) in [(false, true), (true, true), (true, false)] {
let baseline = temporary_entries(&disks).await;
let _ = path_count();
let (request, sender, waiting, _) = observed_request(&bucket, &upload.upload_id, part_size);
sender
.send(Ok(Frame::data(Bytes::from(vec![7; partial_size]))))
.expect("partial body");
let mut upload_future = Box::pin(usecase.execute_upload_part(request));
tokio::time::timeout(Duration::from_secs(10), async {
tokio::select! {
result = &mut upload_future => panic!("upload finished before storage requested raw input: {:?}", result.err().map(|error| error.code().clone())),
result = waiting => result.expect("raw reader"),
}
})
.await
.expect("storage must request raw input");
// Advance only after actual storage demand; filesystem setup and
// cleanup run on a real clock and cannot race auto-advance.
tokio::time::pause();
tokio::time::advance(Duration::from_secs(300)).await;
tokio::time::resume();
let error = tokio::time::timeout(Duration::from_secs(10), upload_future)
.await
.expect("inline cleanup must complete")
.expect_err("stalled part");
assert_eq!(error.code(), &S3ErrorCode::RequestTimeout);
assert_eq!(path_count(), 1, "failure must exercise {path}");
assert!(sender.is_closed(), "producer must release the failed raw body");
assert_eq!(
temporary_entries(&disks).await,
baseline,
"failed part must clean temporary shards inline"
);
let parts = store
.list_object_parts(&bucket, "object", &upload.upload_id, None, 1000, &ObjectOptions::default())
.await
.expect("list parts after error");
assert_eq!(parts.parts.len(), usize::from(replacement));
if replacement {
assert_eq!(parts.parts[0].etag, old_etag, "failed overwrite must preserve the committed part");
}
if !retry_after_failure {
continue;
}
let payload = vec![if replacement { 9 } else { 8 }; part_size];
let request = part_request(&bucket, &upload.upload_id, StreamingBlob::from(Bytes::from(payload)), part_size as i64);
let retry = tokio::time::timeout(Duration::from_secs(10), usecase.execute_upload_part(request))
.await
.expect("foreground and capped staging permits must be released")
.expect("same-number retry");
old_etag = retry.output.e_tag.map(|etag| etag.value().to_owned());
}
let parts = store
.list_object_parts(&bucket, "object", &upload.upload_id, None, 1000, &ObjectOptions::default())
.await
.expect("successful retry");
assert_eq!(parts.parts.len(), 1);
assert_eq!(parts.parts[0].etag, old_etag);
assert_eq!(parts.parts[0].size, part_size);
store
.clone()
.complete_multipart_upload(
&bucket,
"object",
&upload.upload_id,
vec![CompletePart {
part_num: 1,
etag: old_etag,
..CompletePart::default()
}],
&ObjectOptions::default(),
)
.await
.expect("failed replacement must leave the prior part completable");
let mut object = store
.get_object_reader(&bucket, "object", None, HeaderMap::new(), &ObjectOptions::default())
.await
.expect("completed old part remains readable");
let mut restored = Vec::new();
object.stream.read_to_end(&mut restored).await.expect("read all old bytes");
assert_eq!(restored, vec![9; part_size]);
}
eprintln!(
"verified storage lifecycle: path={path}, bytesmut={}",
rustfs_utils::get_env_bool("RUSTFS_ERASURE_ENCODE_BYTESMUT_INGEST", true)
);
}
#[test]
#[serial_test::serial]
fn upload_part_foreground_queue_does_not_consume_body_timeout() {
crate::app::gating_test_env::run_large_stack_test("upload-part-timeout-queue", assert_foreground_queue);
}
async fn assert_foreground_queue() {
let store = crate::app::gating_test_env::shared_gating_ecstore().await;
let ambient = crate::app::gating_test_env::shared_gating_ambient().await;
let context = Arc::new(AppContext::new(Arc::clone(&store), ambient.iam(), ambient.kms()));
let manager = Arc::new(ConcurrencyManager::with_multipart_admission_queue_for_test(
1,
Duration::from_secs(1200),
1,
));
let held = manager.admit_multipart_part(4096).await.expect("hold foreground permit");
let usecase = DefaultMultipartUsecase::with_context_and_concurrency_manager(Some(context), Arc::clone(&manager));
let bucket = format!("body-queue-{}", Uuid::new_v4().simple());
store
.make_bucket(&bucket, &MakeBucketOptions::default())
.await
.expect("bucket");
let upload = store
.new_multipart_upload(&bucket, "object", &ObjectOptions::default())
.await
.expect("upload session");
let (request, sender, waiting, polls) = observed_request(&bucket, &upload.upload_id, 4096);
let task = tokio::spawn(async move { usecase.execute_upload_part(request).await });
tokio::time::timeout(Duration::from_secs(10), async {
while manager.put_object_admission_snapshot().queued != Some(1) {
tokio::task::yield_now().await;
}
})
.await
.expect("request must enter the actual foreground queue");
tokio::time::pause();
tokio::time::advance(Duration::from_secs(600)).await;
tokio::time::resume();
assert_eq!(polls.load(Ordering::Relaxed), 0, "queued requests must not poll the raw body");
assert!(!task.is_finished());
drop(held);
waiting.await.expect("read begins after admission");
sender
.send(Ok(Frame::data(Bytes::from(vec![7; 4096]))))
.expect("body after admission");
drop(sender);
tokio::time::timeout(Duration::from_secs(10), task)
.await
.expect("queued upload completes")
.expect("upload task")
.expect("queue time is not client inactivity");
}
+1 -1
View File
@@ -693,7 +693,7 @@ impl DefaultObjectUsecase {
if let Some(material) = sse_encryption(encryption_request).await? {
effective_sse = Some(material.server_side_encryption.clone());
effective_kms_key_id = material.kms_key_id.clone();
effective_kms_key_id = material.response_kms_key_id();
write_plan = write_plan.with_encryption(material.write_encryption(None));
+204 -42
View File
@@ -16,6 +16,70 @@
use super::*;
async fn authorize_recursive_delete<T>(
req: &mut S3Request<T>,
store: &Arc<ECStore>,
bucket: &str,
prefix: &str,
versioned: bool,
replica: bool,
) -> S3Result<()> {
let original_info = req_info_ref(req)?.clone();
let descendant_prefix = if prefix.ends_with('/') {
prefix.to_owned()
} else {
format!("{prefix}/")
};
let result = async {
let mut marker = None;
let mut version_marker = None;
loop {
let page = store
.clone()
.list_object_versions(bucket, prefix, marker.clone(), version_marker.clone(), None, 1000)
.await
.map_err(ApiError::from)?;
for object in page.objects {
// Disk prefix deletion follows path boundaries, while S3's
// string-prefix listing can also return unrelated siblings.
if object.name != prefix && !object.name.starts_with(&descendant_prefix) {
continue;
}
let object_version = object.version_id.filter(|version| !version.is_nil());
let version_id = (versioned || object_version.is_some())
.then(|| object_version.map_or_else(|| "null".to_owned(), |id| id.to_string()));
let info = req_info_mut(req)?;
info.object = Some(object.name.clone());
info.version_id = version_id.clone();
let action = if replica {
Action::S3Action(S3Action::ReplicateDeleteAction)
} else {
delete_object_authorize_action(version_id.as_deref())
};
authorize_request(req, action).await?;
if has_bypass_governance_header(&req.headers) {
authorize_request(req, Action::S3Action(S3Action::BypassGovernanceRetentionAction)).await?;
}
validate_table_catalog_object_mutation(bucket, &object.name).await?;
}
if !page.is_truncated {
return Ok(());
}
if page.next_marker.is_none() || (marker == page.next_marker && version_marker == page.next_version_idmarker) {
return Err(S3Error::with_message(
S3ErrorCode::InternalError,
"Recursive delete listing did not advance",
));
}
marker = page.next_marker;
version_marker = page.next_version_idmarker;
}
}
.await;
req.extensions.insert(original_info);
result
}
fn successful_delete_audit_objects(
delete: &s3s::dto::Delete,
successful_results: impl IntoIterator<Item = bool>,
@@ -411,14 +475,6 @@ impl DefaultObjectUsecase {
));
}
let is_owner = req_info_ref(&req).map(|info| info.is_owner).unwrap_or(false);
if !recursive_force_delete_is_authorized(&req.headers, is_owner, false) {
return Err(S3Error::with_message(
S3ErrorCode::AccessDenied,
"Recursive force-delete is restricted to administrative requests",
));
}
let Some(store) = self.object_store() else {
return Err(S3Error::with_message(S3ErrorCode::InternalError, "Not init".to_string()));
};
@@ -474,7 +530,7 @@ impl DefaultObjectUsecase {
req_info.version_id = version_id.clone();
}
let auth_res = authorize_request(&mut req, Action::S3Action(S3Action::DeleteObjectAction)).await;
let auth_res = authorize_request(&mut req, delete_object_authorize_action(version_id.as_deref())).await;
if auth_res.is_err() {
if !bulk_denial_logged {
bulk_denial_logged = true;
@@ -882,11 +938,11 @@ impl DefaultObjectUsecase {
authorize_request(&mut req, Action::S3Action(S3Action::ReplicateDeleteAction)).await?;
}
let is_owner = req_info_ref(&req).map(|info| info.is_owner).unwrap_or(false);
if !recursive_force_delete_is_authorized(&req.headers, is_owner, replica) {
let authenticated = req_info_ref(&req).is_ok_and(|info| info.is_owner || info.cred.is_some());
if !recursive_force_delete_has_authenticated_caller(&req.headers, authenticated, replica) {
return Err(S3Error::with_message(
S3ErrorCode::AccessDenied,
"Recursive force-delete is restricted to internal or administrative requests",
"Recursive force-delete requires an authenticated caller",
));
}
validate_table_catalog_object_mutation(&bucket, &key).await?;
@@ -901,6 +957,22 @@ impl DefaultObjectUsecase {
};
validate_bucket_exists(&store, &bucket).await?;
// Lock order is bucket lifecycle, then object/commit locks in storage.
// Keep this guard alive through the physical delete: a preflight without
// writer exclusion could authorize one subtree and delete a newer one.
let recursive_delete_guard = if rustfs_utils::http::get_header(&req.headers, rustfs_utils::http::SUFFIX_FORCE_DELETE)
.is_some_and(|value| value == "true")
{
Some(
store
.lock_bucket_for_recursive_delete(&bucket)
.await
.map_err(ApiError::from)?,
)
} else {
None
};
let metadata = extract_metadata(&req.headers);
// Clone version_id before it's moved
let version_id_clone = version_id.clone();
@@ -924,6 +996,12 @@ impl DefaultObjectUsecase {
apply_bucket_generation_guard(&req, &bucket, &mut opts)?;
let force_delete = opts.delete_prefix;
if let Some(guard) = &recursive_delete_guard {
opts.add_bucket_lifecycle_lock_guard(guard);
authorize_recursive_delete(&mut req, &store, &bucket, &key, opts.versioned || opts.version_suspended, replica)
.await?;
}
// let mut vid = opts.version_id.clone();
if replica {
@@ -1026,6 +1104,7 @@ impl DefaultObjectUsecase {
}
}
};
drop(recursive_delete_guard);
if force_delete {
let _ = invalidate_object_data_cache_prefix_after_delete(&cache_adapter, &bucket, &key).await;
@@ -2212,14 +2291,122 @@ mod tests {
}
#[test]
fn recursive_force_delete_requires_administrative_or_replica_context() {
#[serial_test::serial]
fn recursive_delete_holds_writer_exclusion_after_authorization() {
crate::app::gating_test_env::run_large_stack_test("recursive-delete-writer-exclusion", || async {
use crate::app::storage_api::test::contract::bucket::{
BucketOperations as _, DeleteBucketOptions, MakeBucketOptions,
};
use std::time::Duration;
let store = crate::app::gating_test_env::shared_gating_ecstore().await;
if current_app_context().is_none() {
crate::app::runtime_sources::install_test_app_context(Arc::clone(&store)).await;
}
let context = current_app_context().expect("recursive delete test requires an AppContext");
let bucket = format!("recursive-delete-writer-{}", Uuid::new_v4().simple());
store
.make_bucket(&bucket, &MakeBucketOptions::default())
.await
.expect("create test bucket");
let mut reader = PutObjReader::from_vec(b"old".to_vec());
store
.put_object(&bucket, "folder/old", &mut reader, &ObjectOptions::default())
.await
.expect("seed old object");
let policy_json = format!(
r#"{{"Version":"2012-10-17","Statement":[{{"Effect":"Allow","Principal":{{"AWS":"*"}},"Action":["s3:DeleteObject","s3:DeleteObjectVersion"],"Resource":["arn:aws:s3:::{bucket}/*"]}}]}}"#
);
let mut metadata = (*crate::storage::get_bucket_metadata(&bucket)
.await
.expect("load test metadata"))
.clone();
metadata.policy_config = Some(serde_json::from_str(&policy_json).expect("parse test policy"));
metadata.policy_config_json = policy_json.into_bytes();
crate::storage::storage_api::set_bucket_metadata(bucket.clone(), metadata)
.await
.expect("publish test policy");
let input = DeleteObjectInput::builder()
.bucket(bucket.clone())
.key("folder/".to_owned())
.build()
.expect("build force delete");
let mut req = build_request(input, Method::DELETE);
req.headers.insert("x-rustfs-force-delete", HeaderValue::from_static("true"));
req.extensions.insert(crate::storage::access::ReqInfo {
is_owner: true,
bucket: Some(bucket.clone()),
object: Some("folder/".to_owned()),
..Default::default()
});
let loaded = Arc::new(tokio::sync::Barrier::new(2));
let resume = Arc::new(tokio::sync::Barrier::new(2));
install_delete_source_test_hook(bucket.clone(), Arc::clone(&loaded), Arc::clone(&resume));
let usecase = DefaultObjectUsecase::with_context(Some(context));
let delete = tokio::spawn(async move { usecase.execute_delete_object(req).await });
tokio::time::timeout(Duration::from_secs(30), loaded.wait())
.await
.expect("force delete reaches authorized pre-commit pause");
let writer_store = Arc::clone(&store);
let writer_bucket = bucket.clone();
let mut writer = tokio::spawn(async move {
let mut reader = PutObjReader::from_vec(b"new".to_vec());
writer_store
.put_object(&writer_bucket, "folder/new", &mut reader, &ObjectOptions::default())
.await
});
let before_delete = tokio::time::timeout(Duration::from_secs(1), &mut writer).await;
resume.wait().await;
tokio::time::timeout(Duration::from_secs(30), delete)
.await
.expect("force delete completes without lock recursion")
.expect("delete task joins")
.expect("force delete succeeds");
assert!(
before_delete.is_err(),
"a writer must not enter the authorized subtree before deletion commits"
);
tokio::time::timeout(Duration::from_secs(30), writer)
.await
.expect("writer resumes after deletion")
.expect("writer task joins")
.expect("writer succeeds");
store
.get_object_info(&bucket, "folder/new", &ObjectOptions::default())
.await
.expect("post-delete writer's object survives");
assert!(
store
.get_object_info(&bucket, "folder/old", &ObjectOptions::default())
.await
.is_err(),
"authorized old object is removed"
);
store
.delete_bucket(
&bucket,
&DeleteBucketOptions {
force: true,
..Default::default()
},
)
.await
.expect("remove test bucket");
});
}
#[test]
fn recursive_force_delete_requires_authenticated_or_replica_context() {
let mut headers = HeaderMap::new();
headers.insert("x-rustfs-force-delete", HeaderValue::from_static("true"));
assert!(!recursive_force_delete_is_authorized(&headers, false, false));
assert!(recursive_force_delete_is_authorized(&headers, true, false));
assert!(recursive_force_delete_is_authorized(&headers, false, true));
assert!(recursive_force_delete_is_authorized(&HeaderMap::new(), false, false));
assert!(!recursive_force_delete_has_authenticated_caller(&headers, false, false));
assert!(recursive_force_delete_has_authenticated_caller(&headers, true, false));
assert!(recursive_force_delete_has_authenticated_caller(&headers, false, true));
assert!(recursive_force_delete_has_authenticated_caller(&HeaderMap::new(), false, false));
}
#[tokio::test]
@@ -2240,31 +2427,6 @@ mod tests {
assert_eq!(err.code(), &S3ErrorCode::AccessDenied);
}
#[tokio::test]
async fn execute_delete_objects_rejects_untrusted_force_delete_before_store_access() {
let input = DeleteObjectsInput::builder()
.bucket("test-bucket".to_string())
.delete(Delete {
objects: vec![ObjectIdentifier {
key: "prefix/object".to_string(),
version_id: None,
..Default::default()
}],
quiet: None,
})
.build()
.unwrap();
let mut req = build_request(input, Method::POST);
req.headers.insert("x-rustfs-force-delete", HeaderValue::from_static("true"));
req.extensions.insert(crate::storage::access::ReqInfo::default());
let err = DefaultObjectUsecase::without_context()
.execute_delete_objects(req)
.await
.expect_err("untrusted force-delete must be rejected before storage lookup");
assert_eq!(err.code(), &S3ErrorCode::AccessDenied);
}
// backlog#929 (HP-8): the pre-delete stat may only be skipped when every
// consumer of its result is provably idle. Each guard flips one condition
// to prove the skip is fenced on all four data dependencies.
+1 -1
View File
@@ -2503,7 +2503,7 @@ impl DefaultObjectUsecase {
.await
) {
effective_sse = Some(material.server_side_encryption.clone());
effective_kms_key_id = material.kms_key_id.clone();
effective_kms_key_id = material.response_kms_key_id();
write_plan = write_plan.with_encryption(material.write_encryption(None));
let encryption_metadata = extract_try!(encryption_material_to_metadata(&material));
metadata.extend(encryption_metadata.clone());
+8 -4
View File
@@ -21,8 +21,9 @@ use crate::storage_api::table::get_bucket_metadata;
use super::storage_api::object_usecase::access::{
PostObjectRequestMarker, apply_bucket_generation_guard, apply_copy_source_bucket_generation_guard, authorize_request,
has_bypass_governance_header, load_bucket_generation_from_store, odm_read_generation, prepare_odm_read_generation,
recursive_force_delete_is_authorized, replication_request_authorized, req_info_mut, req_info_ref,
delete_object_authorize_action, has_bypass_governance_header, load_bucket_generation_from_store, odm_read_generation,
prepare_odm_read_generation, recursive_force_delete_has_authenticated_caller, replication_request_authorized, req_info_mut,
req_info_ref,
};
#[cfg(test)]
use super::storage_api::object_usecase::bucket::quota::BucketQuota;
@@ -64,7 +65,7 @@ pub(crate) use super::storage_api::object_usecase::concurrency::{
#[cfg(test)]
use super::storage_api::object_usecase::contract::http::HTTPPreconditions;
use super::storage_api::object_usecase::contract::namespace::NamespaceLocking;
use super::storage_api::object_usecase::contract::object::{ObjectIO as _, ObjectOperations as _};
use super::storage_api::object_usecase::contract::object::{ListOperations as _, ObjectIO as _, ObjectOperations as _};
use super::storage_api::object_usecase::contract::range::HTTPRangeSpec;
use super::storage_api::object_usecase::data_usage::{
quota_object_size, record_bucket_delete_marker_memory, record_bucket_object_delete_memory,
@@ -211,6 +212,7 @@ mod head;
mod internal_put;
mod on_demand_migration_put;
mod put;
pub(crate) mod request_body;
mod restore;
pub(crate) mod shared;
#[cfg(test)]
@@ -223,8 +225,10 @@ pub(crate) use self::extract::*;
pub(crate) use self::get::*;
pub(crate) use self::internal_put::*;
pub(crate) use self::on_demand_migration_put::*;
#[cfg(test)]
pub(crate) use self::put::PollCountingBody;
use self::put::*;
pub(crate) use self::put::{guard_put_object_body_read_timeout, put_object_body_read_timeout};
pub(crate) use self::put::{guard_put_object_body_read_timeout, put_object_body_read_timeout, reject_oversize_single_upload};
#[cfg(test)]
pub(crate) use self::restore::RestoreStatusCommitBarrier;
pub(crate) use self::shared::*;
+145 -1
View File
@@ -109,6 +109,21 @@ fn resolve_put_object_authoritative_size(headers: &HeaderMap, content_length: Op
Ok(size)
}
/// Reject a declared upload length above the single-request ceiling
/// ([`rustfs_config::MAX_SINGLE_PUT_OBJECT_SIZE`]) with `EntityTooLarge`.
///
/// Applies to `PutObject` and `UploadPart`. A negative or unknown length is
/// left to the caller's existing validation.
pub(crate) fn reject_oversize_single_upload(size: i64) -> S3Result<()> {
if u64::try_from(size).is_ok_and(|size| size > rustfs_config::MAX_SINGLE_PUT_OBJECT_SIZE) {
return Err(S3Error::with_message(
S3ErrorCode::EntityTooLarge,
ApiError::error_code_to_message(&S3ErrorCode::EntityTooLarge),
));
}
Ok(())
}
/// Resolve the S3 request-body inter-chunk read timeout from the environment.
///
/// Returns `Duration::ZERO` when disabled (`RUSTFS_HTTP_REQUEST_BODY_READ_TIMEOUT=0`),
@@ -1287,6 +1302,12 @@ impl DefaultObjectUsecase {
// Resolve the authoritative decoded/plain object length (rejecting negative/unknown) before anything else consumes it.
let size = resolve_put_object_authoritative_size(&req.headers, content_length)?;
// The streaming-body limit (s3s `put_object_max_size`) only fires once the
// client has already streamed 5 GiB. The declared length is authoritative,
// so reject an oversize single PUT here, before any body byte is read
// (issue #7596).
reject_oversize_single_upload(size)?;
if let Some(limit) = max_content_length
&& u64::try_from(size).is_ok_and(|size| size > limit)
{
@@ -1846,7 +1867,7 @@ impl DefaultObjectUsecase {
if let Some(material) = encryption_material {
effective_sse = Some(material.server_side_encryption.clone());
effective_kms_key_id = material.kms_key_id.clone();
effective_kms_key_id = material.response_kms_key_id();
write_plan = write_plan.with_encryption(material.write_encryption(None));
@@ -3318,6 +3339,77 @@ mod tests {
assert_eq!(err.code(), &S3ErrorCode::InvalidStorageClass);
}
/// issue #7596: a single PUT whose declared length exceeds the 5 GiB
/// ceiling must be rejected from the headers, before any body byte is
/// requested.
#[tokio::test]
async fn execute_put_object_rejects_oversize_content_length_before_reading_the_body() {
let ceiling = i64::try_from(rustfs_config::MAX_SINGLE_PUT_OBJECT_SIZE).expect("ceiling fits i64");
let (body, polls) = PollCountingBody::streaming_blob();
let input = PutObjectInput::builder()
.bucket("test-bucket".to_string())
.key("huge.bin".to_string())
.body(Some(body))
.content_length(Some(ceiling + 1))
.build()
.unwrap();
let req = build_request(input, Method::PUT);
let usecase = DefaultObjectUsecase::without_context();
let fs = FS::new();
let err = Box::pin(usecase.execute_put_object(&fs, req)).await.unwrap_err();
assert_eq!(err.code(), &S3ErrorCode::EntityTooLarge);
assert_eq!(polls.load(std::sync::atomic::Ordering::SeqCst), 0, "body must not be polled");
}
/// Admission uses the logical object size, not the wire length: a signed
/// aws-chunked request whose framed `Content-Length` exceeds the cap but
/// whose decoded length is within it must not be rejected as oversize,
/// while a decoded length above the cap must be.
#[tokio::test]
async fn execute_put_object_oversize_admission_uses_decoded_length_for_aws_chunked() {
let ceiling = i64::try_from(rustfs_config::MAX_SINGLE_PUT_OBJECT_SIZE).expect("ceiling fits i64");
let framing_overhead = 1_000_000;
for (decoded, expect_too_large) in [(ceiling, false), (ceiling + 1, true)] {
let (body, polls) = PollCountingBody::streaming_blob();
let input = PutObjectInput::builder()
.bucket("test-bucket".to_string())
.key("huge.bin".to_string())
.body(Some(body))
.content_length(Some(decoded + framing_overhead))
.build()
.unwrap();
let mut req = build_request(input, Method::PUT);
req.headers
.insert(http::header::CONTENT_ENCODING, HeaderValue::from_static("aws-chunked"));
req.headers.insert(
HeaderName::from_static("x-amz-content-sha256"),
HeaderValue::from_static("STREAMING-AWS4-HMAC-SHA256-PAYLOAD"),
);
req.headers.insert(
HeaderName::from_static("x-amz-decoded-content-length"),
HeaderValue::from_str(&decoded.to_string()).unwrap(),
);
let usecase = DefaultObjectUsecase::without_context();
let fs = FS::new();
let err = Box::pin(usecase.execute_put_object(&fs, req)).await.unwrap_err();
if expect_too_large {
assert_eq!(err.code(), &S3ErrorCode::EntityTooLarge, "decoded {decoded}");
assert_eq!(polls.load(std::sync::atomic::Ordering::SeqCst), 0, "body must not be polled");
} else {
assert_ne!(
err.code(),
&S3ErrorCode::EntityTooLarge,
"framed wire length above the cap must not reject a decoded length at the cap"
);
}
}
}
#[tokio::test]
async fn execute_put_object_rejects_post_object_sse_kms_from_headers() {
let input = PutObjectInput::builder()
@@ -4184,3 +4276,55 @@ mod tests {
assert!(is_err_object_not_found(&lookup_err), "{lookup_err}");
}
}
/// Test-only request body that records how often it is polled, so admission
/// tests can prove a rejection happened before any body byte was requested.
#[cfg(test)]
pub(crate) struct PollCountingBody {
pub(crate) polls: std::sync::Arc<std::sync::atomic::AtomicUsize>,
}
#[cfg(test)]
impl PollCountingBody {
pub(crate) fn streaming_blob() -> (StreamingBlob, std::sync::Arc<std::sync::atomic::AtomicUsize>) {
let polls = std::sync::Arc::new(std::sync::atomic::AtomicUsize::new(0));
let body = StreamingBlob::new(Self {
polls: std::sync::Arc::clone(&polls),
});
(body, polls)
}
}
#[cfg(test)]
impl Stream for PollCountingBody {
type Item = Result<Bytes, StdError>;
fn poll_next(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<Option<Self::Item>> {
self.polls.fetch_add(1, std::sync::atomic::Ordering::SeqCst);
Poll::Ready(Some(Ok(Bytes::from_static(b"x"))))
}
}
#[cfg(test)]
impl ByteStream for PollCountingBody {}
#[cfg(test)]
mod oversize_single_upload_tests {
use super::*;
#[test]
fn reject_oversize_single_upload_enforces_the_single_request_ceiling() {
let ceiling = i64::try_from(rustfs_config::MAX_SINGLE_PUT_OBJECT_SIZE).expect("ceiling fits i64");
assert!(reject_oversize_single_upload(0).is_ok());
assert!(reject_oversize_single_upload(ceiling).is_ok(), "exact ceiling is allowed");
assert!(reject_oversize_single_upload(-1).is_ok(), "unknown length is left to later validation");
let err = reject_oversize_single_upload(ceiling + 1).expect_err("one byte over must be rejected");
assert_eq!(*err.code(), S3ErrorCode::EntityTooLarge);
assert_eq!(
err.message(),
Some(ApiError::error_code_to_message(&S3ErrorCode::EntityTooLarge).as_str())
);
}
}
+364
View File
@@ -0,0 +1,364 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Client inactivity is observed before decoding, but charged only while the
//! final, transformed reader is waiting. Compression can return buffered output
//! after its input returned Pending, so the transport cannot infer read demand.
use super::{LOG_COMPONENT_APP, LOG_SUBSYSTEM_OBJECT};
use crate::error::ClientBodyReadTimeout;
use bytes::Bytes;
use http_body::{Body, Frame, SizeHint};
use parking_lot::Mutex;
use rustfs_rio::{DynReader, EtagResolvable, HashReaderDetector, HashReaderMut, Index, TryGetIndex};
use std::error::Error;
use std::fmt;
use std::future::Future;
use std::pin::Pin;
use std::sync::Arc;
use std::sync::atomic::{AtomicBool, Ordering};
use std::task::{Context, Poll};
use std::time::Duration;
use tokio::io::{AsyncRead, ReadBuf};
use tokio::time::{Instant, Sleep};
const EVENT_UPLOAD_PART_BODY_READ_STALLED: &str = "upload_part_body_read_stalled";
struct ReadPolicy {
timeout: Duration,
bucket: String,
key: String,
request_id: String,
expected_decoded_bytes: u64,
}
#[derive(Default)]
struct ReadBudget {
policy: Option<ReadPolicy>,
demand: bool,
waiting_since: Option<Instant>,
waited: Duration,
raw_bytes_received: u64,
finished: bool,
}
impl ReadBudget {
fn pause(&mut self) {
if let Some(start) = self.waiting_since.take() {
self.waited = self.waited.saturating_add(start.elapsed());
}
self.demand = false;
}
}
/// Server-owned extension shared by the raw Body and the storage-facing reader.
#[derive(Clone, Default)]
pub(crate) struct BodyReadControl(Arc<SharedBudget>);
#[derive(Default)]
struct SharedBudget {
active: AtomicBool,
budget: Mutex<ReadBudget>,
}
impl BodyReadControl {
pub(crate) fn activate(
&self,
timeout: Duration,
bucket: &str,
key: &str,
request_id: &str,
expected_decoded_bytes: u64,
) -> bool {
if timeout.is_zero() {
return false;
}
let mut state = self.0.budget.lock();
if state.finished {
return false;
}
state.policy = Some(ReadPolicy {
timeout,
bucket: bucket.to_owned(),
key: key.to_owned(),
request_id: request_id.to_owned(),
expected_decoded_bytes,
});
self.0.active.store(true, Ordering::Release);
true
}
fn begin_read(&self) {
let mut state = self.0.budget.lock();
if !state.finished {
state.demand = true;
}
}
fn pause_read(&self) {
self.0.budget.lock().pause();
}
fn progress(&self, bytes: usize) {
// Other HTTP operations never activate this UploadPart policy.
if !self.0.active.load(Ordering::Acquire) {
return;
}
let mut state = self.0.budget.lock();
state.raw_bytes_received = state
.raw_bytes_received
.saturating_add(u64::try_from(bytes).unwrap_or(u64::MAX));
state.waited = Duration::ZERO;
state.waiting_since = None;
}
fn finish(&self) {
let mut state = self.0.budget.lock();
self.0.active.store(false, Ordering::Release);
state.finished = true;
state.policy = None;
state.waiting_since = None;
state.demand = false;
}
fn waiting_deadline(&self) -> Option<Instant> {
if !self.0.active.load(Ordering::Acquire) {
return None;
}
let mut state = self.0.budget.lock();
let timeout = state.policy.as_ref()?.timeout;
if !state.demand || state.finished {
return None;
}
let remaining = timeout.saturating_sub(state.waited);
let start = *state.waiting_since.get_or_insert_with(Instant::now);
// A timeout beyond the clock's representable range cannot elapse.
start.checked_add(remaining)
}
fn expire(&self) -> Option<ClientBodyReadTimeout> {
let (policy, raw_bytes_received) = {
let mut state = self.0.budget.lock();
let timeout = state.policy.as_ref()?.timeout;
if !state.demand || state.waited.saturating_add(state.waiting_since?.elapsed()) < timeout {
return None;
}
state.finished = true;
self.0.active.store(false, Ordering::Release);
state.demand = false;
state.waiting_since = None;
(state.policy.take()?, state.raw_bytes_received)
};
tracing::error!(
event = EVENT_UPLOAD_PART_BODY_READ_STALLED,
component = LOG_COMPONENT_APP,
subsystem = LOG_SUBSYSTEM_OBJECT,
state = "stall_timeout",
operation = "UploadPart",
request_id = %policy.request_id,
bucket = %policy.bucket,
key = %policy.key,
raw_bytes_received,
expected_decoded_bytes = policy.expected_decoded_bytes,
timeout_secs = policy.timeout.as_secs(),
"UploadPart request body read stalled"
);
Some(ClientBodyReadTimeout {
timeout: policy.timeout,
raw_bytes_received,
})
}
}
#[derive(Debug)]
pub(crate) enum ObservedBodyError<E> {
Transport(E),
Inactivity(ClientBodyReadTimeout),
}
impl<E: fmt::Display> fmt::Display for ObservedBodyError<E> {
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
match self {
Self::Transport(error) => error.fmt(f),
Self::Inactivity(error) => error.fmt(f),
}
}
}
impl<E: Error + 'static> Error for ObservedBodyError<E> {
fn source(&self) -> Option<&(dyn Error + 'static)> {
Some(match self {
Self::Transport(error) => error,
Self::Inactivity(error) => error,
})
}
}
/// Wraps the retained raw HTTP body, so a synthesized error never marks the
/// underlying transport complete or prevents HTTP/1 early-response draining.
pub(crate) struct ObservedBody<B> {
inner: B,
control: BodyReadControl,
timer: Option<Pin<Box<Sleep>>>,
ended: bool,
}
impl<B> ObservedBody<B> {
pub(crate) fn new(inner: B, control: BodyReadControl) -> Self {
Self {
inner,
control,
timer: None,
ended: false,
}
}
fn poll_timeout(&mut self, cx: &mut Context<'_>) -> Option<ClientBodyReadTimeout> {
let deadline = self.control.waiting_deadline()?;
let timer = self.timer.get_or_insert_with(|| Box::pin(tokio::time::sleep_until(deadline)));
if timer.deadline() != deadline {
timer.as_mut().reset(deadline);
}
if timer.as_mut().poll(cx).is_ready() {
return self.control.expire();
}
None
}
}
impl<B: Body<Data = Bytes> + Unpin> Body for ObservedBody<B> {
type Data = Bytes;
type Error = ObservedBodyError<B::Error>;
fn poll_frame(mut self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<Option<Result<Frame<Bytes>, Self::Error>>> {
if self.ended {
return Poll::Ready(None);
}
// Ignore empty data without resetting the budget. Bound work per poll
// even if a body repeatedly returns immediately-ready empty frames.
for _ in 0..32 {
match Pin::new(&mut self.inner).poll_frame(cx) {
Poll::Ready(Some(Ok(frame))) => {
if let Some(data) = frame.data_ref() {
if data.is_empty() {
if let Some(error) = self.poll_timeout(cx) {
self.ended = true;
return Poll::Ready(Some(Err(ObservedBodyError::Inactivity(error))));
}
continue;
}
self.control.progress(data.len());
}
return Poll::Ready(Some(Ok(frame)));
}
Poll::Ready(Some(Err(error))) => {
self.ended = true;
self.control.finish();
return Poll::Ready(Some(Err(ObservedBodyError::Transport(error))));
}
Poll::Ready(None) => {
self.ended = true;
self.control.finish();
return Poll::Ready(None);
}
Poll::Pending => {
if let Some(error) = self.poll_timeout(cx) {
self.ended = true;
return Poll::Ready(Some(Err(ObservedBodyError::Inactivity(error))));
}
return Poll::Pending;
}
}
}
cx.waker().wake_by_ref();
Poll::Pending
}
fn is_end_stream(&self) -> bool {
self.ended || self.inner.is_end_stream()
}
fn size_hint(&self) -> SizeHint {
if self.ended {
SizeHint::with_exact(0)
} else {
self.inner.size_hint()
}
}
}
impl<B> Drop for ObservedBody<B> {
fn drop(&mut self) {
self.control.finish();
}
}
/// Must wrap the final HashReader's inner reader after all write transforms.
/// Current erasure readers are owned by their read future/producer: canceling
/// that owner drops this reader. A future retained-reader cancellation path
/// must explicitly pause its read demand before retaining the reader.
pub(crate) struct DemandReader {
inner: DynReader,
control: BodyReadControl,
}
impl DemandReader {
pub(crate) fn new(inner: DynReader, control: BodyReadControl) -> Self {
Self { inner, control }
}
}
impl AsyncRead for DemandReader {
fn poll_read(mut self: Pin<&mut Self>, cx: &mut Context<'_>, buf: &mut ReadBuf<'_>) -> Poll<std::io::Result<()>> {
self.control.begin_read();
let result = Pin::new(&mut self.inner).poll_read(cx, buf);
if result.is_ready() {
self.control.pause_read();
}
result
}
}
impl Drop for DemandReader {
fn drop(&mut self) {
self.control.finish();
}
}
impl EtagResolvable for DemandReader {
fn is_etag_reader(&self) -> bool {
self.inner.is_etag_reader()
}
fn try_resolve_etag(&mut self) -> Option<String> {
self.inner.try_resolve_etag()
}
}
impl HashReaderDetector for DemandReader {
fn is_hash_reader(&self) -> bool {
self.inner.is_hash_reader()
}
fn as_hash_reader_mut(&mut self) -> Option<&mut dyn HashReaderMut> {
self.inner.as_hash_reader_mut()
}
}
impl TryGetIndex for DemandReader {
fn try_get_index(&self) -> Option<&Index> {
self.inner.try_get_index()
}
}
#[cfg(test)]
mod tests;
+266
View File
@@ -0,0 +1,266 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
use super::*;
use crate::app::storage_api::s3::{Body as S3Body, S3Error, S3ErrorCode, StreamingBlob};
use crate::error::ApiError;
use futures::{StreamExt, poll};
use http_body_util::StreamBody;
use std::io;
use tokio::io::AsyncReadExt;
use tokio::sync::mpsc;
use tokio_stream::wrappers::UnboundedReceiverStream;
use tokio_util::io::StreamReader;
mod protocol;
type FrameSender = mpsc::UnboundedSender<Result<Frame<Bytes>, io::Error>>;
fn raw_reader(timeout: Duration) -> (FrameSender, DynReader, BodyReadControl) {
let (sender, receiver) = mpsc::unbounded_channel();
let control = BodyReadControl::default();
control.activate(timeout, "bucket", "object", "request", 65536);
let body = ObservedBody::new(StreamBody::new(UnboundedReceiverStream::new(receiver)), control.clone());
let stream = StreamingBlob::from(S3Body::http_body_unsync(body));
let reader = rustfs_rio::wrap_reader(StreamReader::new(stream.map(|item| item.map_err(io::Error::other))));
(sender, reader, control)
}
#[tokio::test(start_paused = true)]
async fn body_stall_survives_s3s_and_io_wrapping() {
let (_sender, inner, control) = raw_reader(Duration::from_secs(300));
let mut reader = DemandReader::new(inner, control);
let mut output = Vec::new();
let mut read = Box::pin(reader.read_to_end(&mut output));
assert!(poll!(read.as_mut()).is_pending());
tokio::time::advance(Duration::from_secs(299)).await;
assert!(poll!(read.as_mut()).is_pending());
tokio::time::advance(Duration::from_secs(1)).await;
let error = read.await.expect_err("a body that remains open must time out");
let api = ApiError::from(error);
assert_eq!(api.code, S3ErrorCode::RequestTimeout);
let s3_error = S3Error::from(api);
assert_eq!(s3_error.status_code(), Some(http::StatusCode::BAD_REQUEST));
}
#[tokio::test(start_paused = true)]
async fn positive_raw_progress_can_outlast_the_inactivity_timeout() {
let (sender, inner, control) = raw_reader(Duration::from_secs(300));
let mut reader = DemandReader::new(inner, control);
let mut output = Vec::new();
let mut read = Box::pin(reader.read_to_end(&mut output));
assert!(poll!(read.as_mut()).is_pending());
let start = Instant::now();
for _ in 0..8 {
tokio::time::advance(Duration::from_secs(60)).await;
sender
.send(Ok(Frame::data(Bytes::from(vec![7; 8192]))))
.expect("body receiver");
assert!(poll!(read.as_mut()).is_pending());
}
drop(sender);
assert_eq!(read.await.expect("progressing upload"), 65536);
assert_eq!(start.elapsed(), Duration::from_secs(480));
assert_eq!(output, vec![7; 65536]);
}
#[tokio::test(start_paused = true)]
async fn empty_frames_do_not_extend_the_inactivity_budget() {
let (sender, inner, control) = raw_reader(Duration::from_secs(300));
let mut reader = DemandReader::new(inner, control);
let mut output = Vec::new();
let mut read = Box::pin(reader.read_to_end(&mut output));
assert!(poll!(read.as_mut()).is_pending());
for _ in 0..4 {
tokio::time::advance(Duration::from_secs(60)).await;
sender.send(Ok(Frame::data(Bytes::new()))).expect("body receiver");
assert!(poll!(read.as_mut()).is_pending());
}
tokio::time::advance(Duration::from_secs(60)).await;
sender.send(Ok(Frame::data(Bytes::new()))).expect("body receiver");
assert_eq!(
ApiError::from(read.await.expect_err("empty frames are not progress")).code,
S3ErrorCode::RequestTimeout
);
}
#[tokio::test(start_paused = true)]
async fn disabled_or_not_yet_read_body_has_no_inactivity_deadline() {
for timeout in [Duration::ZERO, Duration::from_secs(300)] {
let (sender, inner, control) = raw_reader(timeout);
// Covers both foreground admission and staging admission, before the
// owner ever asks its final reader for input.
tokio::time::advance(Duration::from_secs(1000)).await;
let mut reader = DemandReader::new(inner, control);
let mut output = Vec::new();
let mut read = Box::pin(reader.read_to_end(&mut output));
assert!(poll!(read.as_mut()).is_pending());
if timeout.is_zero() {
tokio::time::advance(Duration::from_secs(1000)).await;
assert!(poll!(read.as_mut()).is_pending());
}
sender
.send(Ok(Frame::data(Bytes::from_static(b"ok"))))
.expect("body receiver");
drop(sender);
assert_eq!(read.await.expect("queued or disabled body"), 2);
assert_eq!(output, b"ok");
}
}
#[tokio::test(start_paused = true)]
async fn compressed_output_pauses_raw_wait_during_storage_backpressure() {
use rustfs_utils::compress::CompressionAlgorithm;
let (sender, inner, control) = raw_reader(Duration::from_secs(300));
let compressed = rustfs_rio::CompressReader::with_block_size(inner, 8192, CompressionAlgorithm::default());
let mut reader = DemandReader::new(rustfs_rio::boxed_reader(compressed), control);
sender
.send(Ok(Frame::data(Bytes::from(vec![3; 1024]))))
.expect("body receiver");
let mut buffer = vec![0; 16384];
// CompressReader sees the partial input and then Pending, yet can return
// a complete compressed block to the storage writer.
let first = reader.read(&mut buffer).await.expect("buffered compressed block");
assert!(first > 0);
let mut compressed_bytes = buffer[..first].to_vec();
tokio::time::advance(Duration::from_secs(600)).await;
let mut read = Box::pin(reader.read(&mut buffer));
assert!(poll!(read.as_mut()).is_pending(), "storage backpressure is not a client stall");
tokio::time::advance(Duration::from_secs(299)).await;
assert!(poll!(read.as_mut()).is_pending());
sender
.send(Ok(Frame::data(Bytes::from(vec![4; 1024]))))
.expect("body receiver");
drop(sender);
let next = read.await.expect("input after storage backpressure");
compressed_bytes.extend_from_slice(&buffer[..next]);
reader
.read_to_end(&mut compressed_bytes)
.await
.expect("remaining compressed data");
let mut restored = Vec::new();
rustfs_rio::DecompressReader::new(std::io::Cursor::new(compressed_bytes), CompressionAlgorithm::default())
.read_to_end(&mut restored)
.await
.expect("roundtrip after backpressure");
assert_eq!(restored, [vec![3; 1024], vec![4; 1024]].concat());
}
#[tokio::test(start_paused = true)]
async fn read_owner_cancellation_releases_the_raw_body() {
let (sender, inner, control) = raw_reader(Duration::from_secs(300));
let mut reader = DemandReader::new(inner, control);
let owner = tokio::spawn(async move { reader.read_to_end(&mut Vec::new()).await });
tokio::task::yield_now().await;
owner.abort();
assert!(owner.await.expect_err("read owner was canceled").is_cancelled());
assert!(sender.is_closed(), "the canceled producer must release the body receiver");
}
#[tokio::test(start_paused = true)]
async fn buffered_output_pauses_but_does_not_reset_elapsed_inactivity() {
struct BufferedOutput {
inner: DynReader,
ready: Arc<AtomicBool>,
}
impl AsyncRead for BufferedOutput {
fn poll_read(mut self: Pin<&mut Self>, cx: &mut Context<'_>, buf: &mut ReadBuf<'_>) -> Poll<io::Result<()>> {
if self.ready.swap(false, Ordering::AcqRel) {
buf.put_slice(b"x");
return Poll::Ready(Ok(()));
}
Pin::new(&mut self.inner).poll_read(cx, buf)
}
}
let (_sender, inner, control) = raw_reader(Duration::from_secs(300));
let ready = Arc::new(AtomicBool::new(false));
let transform = BufferedOutput {
inner,
ready: Arc::clone(&ready),
};
let mut reader = DemandReader::new(rustfs_rio::wrap_reader(transform), control);
let mut buffer = [0; 1];
let mut read = Box::pin(reader.read(&mut buffer));
assert!(poll!(read.as_mut()).is_pending());
tokio::time::advance(Duration::from_secs(200)).await;
assert!(poll!(read.as_mut()).is_pending());
ready.store(true, Ordering::Release);
assert_eq!(read.await.expect("transform releases buffered output"), 1);
tokio::time::advance(Duration::from_secs(600)).await;
let mut read = Box::pin(reader.read(&mut buffer));
assert!(poll!(read.as_mut()).is_pending());
tokio::time::advance(Duration::from_secs(99)).await;
assert!(poll!(read.as_mut()).is_pending());
tokio::time::advance(Duration::from_secs(1)).await;
let result = poll!(read.as_mut());
let Poll::Ready(Err(error)) = result else {
panic!("remaining inactivity budget must expire immediately at 100 seconds");
};
assert_eq!(ApiError::from(error).code, S3ErrorCode::RequestTimeout);
}
#[tokio::test(start_paused = true)]
async fn final_write_reader_preserves_checksums_through_compression_and_sse() {
use crate::app::storage_api::multipart_usecase::io::{HashReader, WriteEncryption, WritePlan};
use rustfs_rio::{Checksum, ChecksumType};
use rustfs_utils::CompressionAlgorithm;
let payload = vec![7; 65536];
for plan in [
WritePlan::new().with_compression(CompressionAlgorithm::default()),
WritePlan::new().with_encryption(WriteEncryption::multipart([5; 32], [9; 12], 1)),
WritePlan::new()
.with_compression(CompressionAlgorithm::default())
.with_encryption(WriteEncryption::multipart([5; 32], [9; 12], 1)),
] {
let (sender, inner, control) = raw_reader(Duration::from_secs(300));
let checksum = Checksum::new_from_data(ChecksumType::CRC32, &payload).expect("plaintext checksum");
let mut plaintext = HashReader::from_reader(inner, 65536, 65536, None, None, false).expect("plaintext reader");
plaintext
.add_non_trailing_checksum(Some(checksum.clone()), false)
.expect("attach checksum");
let mut reader = plan.apply(plaintext, 65536).expect("write plan");
let inner = reader.take_inner();
reader.inner = rustfs_rio::boxed_reader(DemandReader::new(inner, control));
let mut output = Vec::new();
let mut read = Box::pin(reader.read_to_end(&mut output));
assert!(poll!(read.as_mut()).is_pending());
for chunk in payload.chunks(8192) {
tokio::time::advance(Duration::from_secs(60)).await;
sender
.send(Ok(Frame::data(Bytes::copy_from_slice(chunk))))
.expect("raw body progress");
assert!(poll!(read.as_mut()).is_pending());
}
drop(sender);
read.await.expect("transformed reader completes beyond one inactivity period");
assert!(!output.is_empty());
assert_eq!(reader.content_crc_type(), Some(ChecksumType::CRC32));
assert_eq!(reader.content_crc().get("CRC32"), Some(&checksum.encoded));
}
}
#[tokio::test(start_paused = true)]
async fn native_body_errors_preserve_their_original_source() {
let (sender, inner, control) = raw_reader(Duration::from_secs(300));
let mut reader = DemandReader::new(inner, control);
sender
.send(Err(io::Error::new(io::ErrorKind::TimedOut, "disk timeout")))
.expect("body receiver");
let error = reader.read_to_end(&mut Vec::new()).await.expect_err("native error");
assert_eq!(ApiError::from(error).code, S3ErrorCode::InternalError);
}
@@ -0,0 +1,280 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
use super::*;
use crate::app::storage_api::s3::{
Body as S3Body, S3, S3Config, S3Error, S3Request, S3Response, S3Result, S3Service, S3ServiceBuilder, SimpleAuth,
StaticConfigProvider, UploadPartInput, UploadPartOutput,
};
use http_body_util::BodyExt;
use std::sync::atomic::{AtomicUsize, Ordering};
#[derive(Clone, Default)]
struct Consumer {
received: Arc<AtomicUsize>,
committed: Arc<Mutex<Option<Vec<u8>>>>,
}
#[async_trait::async_trait]
impl S3 for Consumer {
async fn upload_part(&self, req: S3Request<UploadPartInput>) -> S3Result<S3Response<UploadPartOutput>> {
let expected = req.input.content_length.expect("S3S must normalize the decoded length");
let control = req
.extensions
.get::<BodyReadControl>()
.expect("HTTP control extension")
.clone();
control.activate(Duration::from_secs(300), "test-bucket", "test-key", "request", expected as u64);
let stream = req.input.body.expect("body");
let inner = rustfs_rio::wrap_reader(StreamReader::new(stream.map(|item| item.map_err(io::Error::other))));
let mut reader =
rustfs_rio::HashReader::from_stream(DemandReader::new(inner, control), expected, expected, None, None, false)
.expect("logical body reader");
reader
.add_checksum_from_s3s(&req.headers, req.trailing_headers, false)
.expect("request checksum context");
let mut output = Vec::new();
let mut buffer = [0; 8192];
loop {
let count = reader
.read(&mut buffer)
.await
.map_err(|error| S3Error::from(ApiError::from(error)))?;
if count == 0 {
break;
}
self.received.fetch_add(count, Ordering::Relaxed);
output.extend_from_slice(&buffer[..count]);
}
assert_eq!(output.len() as i64, expected, "wire length must not reach the business DTO");
*self.committed.lock() = Some(output);
Ok(S3Response::new(UploadPartOutput {
checksum_crc32: reader.content_crc().get("CRC32").cloned(),
..UploadPartOutput::default()
}))
}
}
fn service(consumer: Consumer) -> S3Service {
let mut builder = S3ServiceBuilder::new(consumer);
builder.set_auth(SimpleAuth::from_single("test-access", "test-secret"));
let mut config = S3Config::default();
config.presigned_url_max_skew_time_secs = u32::MAX;
builder.set_config(Arc::new(StaticConfigProvider::new(Arc::new(config))));
builder.build()
}
struct SignedRequest {
request: http::Request<S3Body>,
sender: FrameSender,
prefix: Bytes,
suffix: Bytes,
}
/// Constructs a real SigV4 fixture using the production crypto primitives.
/// S3S verifies both the request authorization and every signed chunk.
fn signed_request(payload: &[u8], unsigned_trailer: bool) -> SignedRequest {
use rustfs_utils::{hex_sha256, hmac_sha256};
const DATE: &str = "20130524T000000Z";
const SCOPE: &str = "20130524/us-east-1/s3/aws4_request";
let mode = if unsigned_trailer {
"STREAMING-UNSIGNED-PAYLOAD-TRAILER"
} else {
"STREAMING-AWS4-HMAC-SHA256-PAYLOAD"
};
let signed_headers = "host;x-amz-content-sha256;x-amz-date;x-amz-decoded-content-length";
let headers = format!(
"host:s3.amazonaws.com\nx-amz-content-sha256:{mode}\nx-amz-date:{DATE}\nx-amz-decoded-content-length:{}\n",
payload.len()
);
let canonical = format!("PUT\n/test-bucket/test-key\npartNumber=1&uploadId=test-upload\n{headers}\n{signed_headers}\n{mode}");
let key = hmac_sha256("AWS4test-secret", "20130524");
let key = hmac_sha256(key, "us-east-1");
let key = hmac_sha256(key, "s3");
let key = hmac_sha256(key, "aws4_request");
let digest = |data: &[u8]| hex_sha256(data, str::to_owned);
let encode = |data: [u8; 32]| hex_simd::encode_to_string(data, hex_simd::AsciiCase::Lower);
let seed = encode(hmac_sha256(
key,
format!("AWS4-HMAC-SHA256\n{DATE}\n{SCOPE}\n{}", digest(canonical.as_bytes())),
));
let chunk_signature = |previous: &str, data: &[u8]| {
encode(hmac_sha256(
key,
format!("AWS4-HMAC-SHA256-PAYLOAD\n{DATE}\n{SCOPE}\n{previous}\n{}\n{}", digest(b""), digest(data)),
))
};
let (prefix, suffix) = if unsigned_trailer {
(
format!("{:x}\r\n", payload.len()),
"\r\n0\r\nx-amz-checksum-crc32:y/Q5Jg==\r\n\r\n".to_owned(),
)
} else {
let signature = chunk_signature(&seed, payload);
(
format!("{:x};chunk-signature={signature}\r\n", payload.len()),
format!("\r\n0;chunk-signature={}\r\n\r\n", chunk_signature(&signature, b"")),
)
};
let (sender, receiver) = mpsc::unbounded_channel();
let control = BodyReadControl::default();
let body = ObservedBody::new(StreamBody::new(UnboundedReceiverStream::new(receiver)), control.clone());
let mut builder = http::Request::builder()
.method("PUT")
.uri("https://s3.amazonaws.com/test-bucket/test-key?partNumber=1&uploadId=test-upload")
.header("host", "s3.amazonaws.com")
.header("content-encoding", "aws-chunked")
.header("content-length", prefix.len() + payload.len() + suffix.len())
.header("x-amz-content-sha256", mode)
.header("x-amz-date", DATE)
.header("x-amz-decoded-content-length", payload.len())
.header(
"authorization",
format!("AWS4-HMAC-SHA256 Credential=test-access/{SCOPE}, SignedHeaders={signed_headers}, Signature={seed}"),
);
if unsigned_trailer {
builder = builder.header("x-amz-trailer", "x-amz-checksum-crc32");
}
let mut request = builder.body(S3Body::http_body_unsync(body)).expect("signed request");
request.extensions_mut().insert(control);
SignedRequest {
request,
sender,
prefix: Bytes::from(prefix),
suffix: Bytes::from(suffix),
}
}
#[tokio::test(start_paused = true)]
async fn signed_chunk_with_raw_progress_survives_eight_minutes_without_decoded_output() {
let payload = vec![7; 65536];
let SignedRequest {
request,
sender,
prefix,
suffix,
} = signed_request(&payload, false);
let consumer = Consumer::default();
let service = service(consumer.clone());
let mut call = Box::pin(service.call(request));
sender.send(Ok(Frame::data(prefix))).expect("body receiver");
assert!(poll!(call.as_mut()).is_pending());
let start = Instant::now();
for chunk in payload.chunks(8192) {
assert_eq!(
consumer.received.load(Ordering::Relaxed),
0,
"incomplete chunks must not escape signature validation"
);
tokio::time::advance(Duration::from_secs(60)).await;
sender
.send(Ok(Frame::data(Bytes::copy_from_slice(chunk))))
.expect("body receiver");
assert!(poll!(call.as_mut()).is_pending());
}
sender.send(Ok(Frame::data(suffix))).expect("body receiver");
drop(sender);
let response = call.await.expect("S3 response");
assert_eq!(response.status(), http::StatusCode::OK);
assert_eq!(start.elapsed(), Duration::from_secs(480));
assert_eq!(*consumer.committed.lock(), Some(payload));
}
#[tokio::test(start_paused = true)]
async fn decoded_length_does_not_end_waiting_for_terminator_trailer_or_raw_eof() {
for (unsigned_trailer, send_suffix) in [(false, false), (false, true), (true, false), (true, true)] {
let payload = b"123456789";
let SignedRequest {
request,
sender,
prefix,
suffix,
} = signed_request(payload, unsigned_trailer);
let consumer = Consumer::default();
let service = service(consumer.clone());
sender.send(Ok(Frame::data(prefix))).expect("prefix");
sender.send(Ok(Frame::data(Bytes::from_static(payload)))).expect("payload");
sender
.send(Ok(Frame::data(if send_suffix { suffix } else { Bytes::from_static(b"\r\n") })))
.expect("suffix");
let mut call = Box::pin(service.call(request));
assert!(poll!(call.as_mut()).is_pending());
assert_eq!(consumer.received.load(Ordering::Relaxed), payload.len());
tokio::time::advance(Duration::from_secs(300)).await;
let response = call.await.expect("S3 error response");
assert_eq!(response.status(), http::StatusCode::BAD_REQUEST);
let xml = BodyExt::collect(response.into_body()).await.expect("error XML").to_bytes();
assert!(String::from_utf8_lossy(&xml).contains("<Code>RequestTimeout</Code>"));
assert!(consumer.committed.lock().is_none());
}
}
#[tokio::test]
async fn unsigned_trailer_normalizes_length_and_signed_corruption_cannot_commit() {
for corrupt in [false, true] {
let payload = b"123456789";
let SignedRequest {
request,
sender,
prefix,
suffix,
} = signed_request(payload, !corrupt);
let consumer = Consumer::default();
let service = service(consumer.clone());
sender.send(Ok(Frame::data(prefix))).expect("prefix");
sender
.send(Ok(Frame::data(Bytes::from_static(if corrupt { b"923456789" } else { payload }))))
.expect("payload");
sender.send(Ok(Frame::data(suffix))).expect("suffix");
drop(sender);
let response = service.call(request).await.expect("S3 response");
if corrupt {
assert_ne!(response.status(), http::StatusCode::OK);
assert_eq!(consumer.received.load(Ordering::Relaxed), 0);
assert!(consumer.committed.lock().is_none());
} else {
assert_eq!(response.status(), http::StatusCode::OK);
assert_eq!(
response
.headers()
.get("x-amz-checksum-crc32")
.expect("validated response checksum"),
"y/Q5Jg=="
);
assert_eq!(*consumer.committed.lock(), Some(payload.to_vec()));
}
}
}
#[tokio::test]
async fn unsigned_trailer_with_wrong_checksum_cannot_commit() {
let SignedRequest {
request, sender, prefix, ..
} = signed_request(b"123456789", true);
let consumer = Consumer::default();
let service = service(consumer.clone());
sender.send(Ok(Frame::data(prefix))).expect("prefix");
sender
.send(Ok(Frame::data(Bytes::from_static(
b"123456789\r\n0\r\nx-amz-checksum-crc32:AAAAAA==\r\n\r\n",
))))
.expect("invalid trailer");
drop(sender);
let response = service.call(request).await.expect("S3 response");
assert_eq!(response.status(), http::StatusCode::BAD_REQUEST);
let xml = BodyExt::collect(response.into_body()).await.expect("error XML").to_bytes();
assert!(String::from_utf8_lossy(&xml).contains("<Code>BadDigest</Code>"));
assert!(consumer.committed.lock().is_none());
}
+4 -3
View File
@@ -1424,9 +1424,10 @@ mod tests {
#[test]
fn resolve_bucket_default_sse_falls_back_to_aes256_for_an_unknown_algorithm() {
// Reachable only through corrupt or hand-edited bucket metadata;
// PutBucketEncryption rejects unknown algorithms. All three call sites
// now share this single decision (backlog#1826).
// PutBucketEncryption refuses unknown algorithms, so this is reachable
// only through a configuration stored before that check or through
// hand-edited bucket metadata. All three call sites share this single
// decision (backlog#1826).
let config = bucket_sse_config_with("garbage", None);
let (sse, kms_key_id) = resolve_bucket_default_sse(Some(&config), None, None, false);
+13 -7
View File
@@ -28,9 +28,9 @@ pub(crate) fn EndpointServerPools(
/// the direct s3s surface (s3s footprint ratchet, `scripts/check_s3s_footprint.sh`).
pub(crate) mod s3 {
#[cfg(test)]
pub(crate) use s3s::S3Response;
pub(crate) use s3s::auth::SimpleAuth;
#[cfg(test)]
pub(crate) use s3s::dto::ListObjectsInput;
pub(crate) use s3s::config::{S3Config, StaticConfigProvider};
#[cfg(test)]
pub(crate) use s3s::dto::{
BucketVersioningStatus, DeleteMarkerReplication, DeleteMarkerReplicationStatus, Destination, GetObjectInput,
@@ -39,7 +39,13 @@ pub(crate) mod s3 {
ServerSideEncryptionRule, Tag, VersioningConfiguration,
};
#[cfg(test)]
pub(crate) use s3s::dto::{ListObjectsInput, StreamingBlob, UploadPartInput, UploadPartOutput};
#[cfg(test)]
pub(crate) use s3s::service::{S3Service, S3ServiceBuilder};
#[cfg(test)]
pub(crate) use s3s::xml::{Serialize as XmlSerialize, Serializer as XmlSerializer};
#[cfg(test)]
pub(crate) use s3s::{Body, S3, S3Response};
pub(crate) use s3s::{S3Error, S3ErrorCode, S3Request, S3Result};
}
@@ -266,10 +272,10 @@ pub(crate) mod access {
pub(crate) use crate::storage::storage_api::access_consumer::ReqInfo;
pub(crate) use crate::storage::storage_api::access_consumer::{
PostObjectRequestMarker, apply_bucket_generation_guard, apply_copy_source_bucket_generation_guard, authorize_request,
bucket_config_mutation_incarnation, has_bypass_governance_header, load_bucket_generation_from_store,
log_list_buckets_iam_implicit_deny, odm_read_generation, prepare_list_buckets_iam_authorization,
prepare_odm_read_generation, recursive_force_delete_is_authorized, replication_request_authorized, req_info_mut,
req_info_ref,
bucket_config_mutation_incarnation, delete_object_authorize_action, has_bypass_governance_header,
load_bucket_generation_from_store, log_list_buckets_iam_implicit_deny, odm_read_generation,
prepare_list_buckets_iam_authorization, prepare_odm_read_generation, recursive_force_delete_has_authenticated_caller,
replication_request_authorized, req_info_mut, req_info_ref,
};
}
@@ -1196,7 +1202,7 @@ pub(crate) mod object_usecase {
}
pub(crate) mod object {
pub(crate) use super::super::super::storage_contracts::{ObjectIO, ObjectOperations};
pub(crate) use super::super::super::storage_contracts::{ListOperations, ObjectIO, ObjectOperations};
}
pub(crate) mod range {
+279 -39
View File
@@ -137,6 +137,26 @@ impl std::fmt::Display for UploadLimitExceeded {
impl std::error::Error for UploadLimitExceeded {}
/// Identifies inactivity of an external client body, rather than a storage timeout.
#[derive(Debug, Clone, Copy)]
pub(crate) struct ClientBodyReadTimeout {
pub timeout: std::time::Duration,
pub raw_bytes_received: u64,
}
impl std::fmt::Display for ClientBodyReadTimeout {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(
f,
"request body made no progress for {} seconds after {} raw bytes",
self.timeout.as_secs(),
self.raw_bytes_received
)
}
}
impl std::error::Error for ClientBodyReadTimeout {}
/// Marks a server-side object/source reader failure that must not be reported as
/// a malformed client request body.
#[derive(Debug)]
@@ -424,25 +444,7 @@ fn error_chain_has_type<T>(err: &(dyn std::error::Error + 'static)) -> bool
where
T: std::error::Error + 'static,
{
if err.downcast_ref::<T>().is_some() {
return true;
}
if let Some(io_err) = err.downcast_ref::<std::io::Error>()
&& let Some(inner) = io_err.get_ref()
&& error_chain_has_type::<T>(inner)
{
return true;
}
let mut current = Some(err);
while let Some(err) = current {
if err.downcast_ref::<T>().is_some() {
return true;
}
current = err.source();
}
false
error_chain_find(err, |error| error.is::<T>().then_some(())).is_some()
}
#[derive(Clone, Copy, Debug, Eq, PartialEq)]
@@ -463,33 +465,50 @@ fn classify_s3s_body_stream_error_display(err: &(dyn std::error::Error + 'static
}
fn error_chain_s3s_body_stream_error(err: &(dyn std::error::Error + 'static)) -> Option<S3sBodyStreamError> {
if let Some(classified) = classify_s3s_body_stream_error_display(err) {
return Some(classified);
}
error_chain_find(err, classify_s3s_body_stream_error_display)
}
if let Some(io_err) = err.downcast_ref::<std::io::Error>()
&& let Some(inner) = io_err.get_ref()
&& let Some(classified) = error_chain_s3s_body_stream_error(inner)
{
return Some(classified);
}
let mut current = err.source();
while let Some(err) = current {
if let Some(classified) = classify_s3s_body_stream_error_display(err) {
/// `io::Error::source` skips its custom payload itself. Visit that payload
/// explicitly at every level, then follow its source exactly once. The bound
/// also makes cyclic or excessively deep foreign error chains safe.
fn error_chain_find<T>(
err: &(dyn std::error::Error + 'static),
mut classify: impl FnMut(&(dyn std::error::Error + 'static)) -> Option<T>,
) -> Option<T> {
let mut current = Some(err);
for _ in 0..64 {
let error = current?;
if let Some(classified) = classify(error) {
return Some(classified);
}
if let Some(io_err) = err.downcast_ref::<std::io::Error>()
&& let Some(inner) = io_err.get_ref()
&& let Some(classified) = error_chain_s3s_body_stream_error(inner)
{
return Some(classified);
}
current = err.source();
current = error
.downcast_ref::<std::io::Error>()
.and_then(std::io::Error::get_ref)
.map(|inner| inner as &(dyn std::error::Error + 'static))
.or_else(|| error.source());
}
None
}
fn error_chain_has_body_size_limit_exceeded(err: &(dyn std::error::Error + 'static)) -> bool {
error_chain_has_type::<s3s::BodySizeLimitExceeded>(err)
}
/// hyper reports a request body whose connection hit EOF before
/// `Content-Length` bytes arrived as a `Kind::Body` error carrying an
/// `UnexpectedEof` `io::Error` (its `IncompleteBody` marker is private).
/// That is a client-side short body, not a server fault.
fn is_hyper_body_eof(err: &(dyn std::error::Error + 'static)) -> bool {
err.downcast_ref::<hyper::Error>()
.and_then(|hyper_err| std::error::Error::source(hyper_err))
.and_then(|cause| cause.downcast_ref::<std::io::Error>())
.is_some_and(|io_err| io_err.kind() == std::io::ErrorKind::UnexpectedEof)
}
fn error_chain_has_hyper_body_eof(err: &(dyn std::error::Error + 'static)) -> bool {
error_chain_find(err, |error| is_hyper_body_eof(error).then_some(())).is_some()
}
impl From<ApiError> for S3Error {
fn from(err: ApiError) -> Self {
let status = custom_error_status(&err.code);
@@ -547,6 +566,30 @@ impl From<StorageError> for ApiError {
};
}
if error_chain_has_type::<ClientBodyReadTimeout>(inner) {
return ApiError {
code: S3ErrorCode::RequestTimeout,
message: ApiError::error_code_to_message(&S3ErrorCode::RequestTimeout),
source: Some(Box::new(err)),
};
}
if error_chain_has_body_size_limit_exceeded(inner) {
return ApiError {
code: S3ErrorCode::EntityTooLarge,
message: ApiError::error_code_to_message(&S3ErrorCode::EntityTooLarge),
source: Some(Box::new(err)),
};
}
if error_chain_has_hyper_body_eof(inner) {
return ApiError {
code: S3ErrorCode::IncompleteBody,
message: ApiError::error_code_to_message(&S3ErrorCode::IncompleteBody),
source: Some(Box::new(err)),
};
}
if matches!(s3s_body_stream_error, Some(S3sBodyStreamError::IncompleteBody)) {
return ApiError {
code: S3ErrorCode::IncompleteBody,
@@ -690,6 +733,30 @@ impl From<std::io::Error> for ApiError {
source: Some(Box::new(err)),
};
}
if error_chain_has_type::<ClientBodyReadTimeout>(inner) {
return ApiError {
code: S3ErrorCode::RequestTimeout,
message: ApiError::error_code_to_message(&S3ErrorCode::RequestTimeout),
source: Some(Box::new(err)),
};
}
if error_chain_has_body_size_limit_exceeded(inner) {
return ApiError {
code: S3ErrorCode::EntityTooLarge,
message: ApiError::error_code_to_message(&S3ErrorCode::EntityTooLarge),
source: Some(Box::new(err)),
};
}
if error_chain_has_hyper_body_eof(inner) {
return ApiError {
code: S3ErrorCode::IncompleteBody,
message: ApiError::error_code_to_message(&S3ErrorCode::IncompleteBody),
source: Some(Box::new(err)),
};
}
if matches!(s3s_body_stream_error, Some(S3sBodyStreamError::IncompleteBody)) {
return ApiError {
code: S3ErrorCode::IncompleteBody,
@@ -962,6 +1029,179 @@ mod tests {
}
}
#[test]
fn body_size_limit_exceeded_maps_to_entity_too_large_across_io_boundaries() {
// Shape observed in production (issue #7596):
// Custom { UnexpectedEof, Custom { Other, BodySizeLimitExceeded { size, limit } } }
let nested = || {
IoError::new(
ErrorKind::UnexpectedEof,
IoError::other(s3s::BodySizeLimitExceeded {
size: 16384,
limit: 6389,
}),
)
};
let direct: ApiError = nested().into();
assert_eq!(direct.code, S3ErrorCode::EntityTooLarge);
assert_eq!(direct.message, ApiError::error_code_to_message(&S3ErrorCode::EntityTooLarge));
let storage: ApiError = StorageError::Io(nested()).into();
assert_eq!(storage.code, S3ErrorCode::EntityTooLarge);
assert!(storage.source.is_some());
// An unrelated message that merely mentions a limit stays internal.
let other: ApiError = IoError::other(MockS3sBodyStreamError("limit exceeded for something else")).into();
assert_eq!(other.code, S3ErrorCode::InternalError);
let impostor: ApiError = IoError::other(MockS3sBodyStreamError("body size 16384 exceeds limit 6389")).into();
assert_eq!(impostor.code, S3ErrorCode::InternalError);
}
#[test]
fn client_body_timeout_survives_intermediate_io_and_storage_errors() {
#[derive(Debug)]
struct DecoderError(IoError);
impl std::fmt::Display for DecoderError {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str("decoder source failed")
}
}
impl std::error::Error for DecoderError {
fn source(&self) -> Option<&(dyn std::error::Error + 'static)> {
Some(&self.0)
}
}
let nested = || {
IoError::other(DecoderError(IoError::other(IoError::new(
ErrorKind::TimedOut,
ClientBodyReadTimeout {
timeout: std::time::Duration::from_secs(300),
raw_bytes_received: 8192,
},
))))
};
for error in [ApiError::from(nested()), ApiError::from(StorageError::Io(nested()))] {
assert_eq!(error.code, S3ErrorCode::RequestTimeout);
assert!(error_chain_has_type::<ClientBodyReadTimeout>(&error));
let s3_error = S3Error::from(error);
assert_eq!(s3_error.status_code(), Some(StatusCode::BAD_REQUEST));
}
for error in [
ApiError::from(IoError::new(ErrorKind::TimedOut, "disk read timeout")),
ApiError::from(StorageError::Io(IoError::new(ErrorKind::TimedOut, "peer timeout"))),
] {
assert_eq!(error.code, S3ErrorCode::InternalError);
}
// A server-side source wrapper has precedence even if a remote source
// has carried its own client-body marker across an I/O boundary.
let source = ServerSideSourceReadError::new("CopyObject", nested());
assert_eq!(ApiError::from(IoError::other(source)).code, S3ErrorCode::ServiceUnavailable);
}
#[test]
fn body_error_classification_bounds_cyclic_source_chains() {
#[derive(Debug)]
struct Cycle;
impl std::fmt::Display for Cycle {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str("cycle")
}
}
impl std::error::Error for Cycle {
fn source(&self) -> Option<&(dyn std::error::Error + 'static)> {
Some(self)
}
}
assert!(!error_chain_has_type::<ClientBodyReadTimeout>(&Cycle));
assert_eq!(ApiError::from(IoError::other(Cycle)).code, S3ErrorCode::InternalError);
}
/// Exercise the public error type through the actual body budget.
#[tokio::test]
async fn real_s3s_body_size_limit_error_maps_to_entity_too_large() {
use futures::StreamExt;
let real_error = || async {
let mut body = s3s::Body::from(bytes::Bytes::from_static(b"hello"));
body.set_limit(Some(4));
body.next()
.await
.expect("one frame")
.expect_err("five bytes must exceed a four-byte budget")
};
let err = real_error().await;
assert!(err.is::<s3s::BodySizeLimitExceeded>(), "unexpected body error: {err}");
let err = real_error().await;
let storage: ApiError = StorageError::Io(IoError::new(ErrorKind::UnexpectedEof, IoError::other(err))).into();
assert_eq!(storage.code, S3ErrorCode::EntityTooLarge);
assert_eq!(storage.message, ApiError::error_code_to_message(&S3ErrorCode::EntityTooLarge));
let err = real_error().await;
let direct: ApiError = IoError::other(err).into();
assert_eq!(direct.code, S3ErrorCode::EntityTooLarge);
}
/// Drive a real hyper HTTP/1 server so the test sees hyper's own body EOF
/// error (`hyper::Error(Body, UnexpectedEof, IncompleteBody)`), which has no
/// public constructor.
async fn capture_hyper_body_eof_error() -> hyper::Error {
use http_body_util::BodyExt;
use hyper::service::service_fn;
use hyper_util::rt::TokioIo;
use std::sync::{Arc, Mutex};
use tokio::io::AsyncWriteExt;
let listener = tokio::net::TcpListener::bind("127.0.0.1:0").await.expect("bind");
let addr = listener.local_addr().expect("local addr");
let captured: Arc<Mutex<Option<hyper::Error>>> = Arc::new(Mutex::new(None));
let server_slot = Arc::clone(&captured);
let server = tokio::spawn(async move {
let (stream, _) = listener.accept().await.expect("accept");
let slot = server_slot;
let service = service_fn(move |req: hyper::Request<hyper::body::Incoming>| {
let slot = Arc::clone(&slot);
async move {
let err = req.into_body().collect().await.expect_err("short body must fail");
*slot.lock().expect("slot") = Some(err);
Ok::<_, std::convert::Infallible>(hyper::Response::new(String::new()))
}
});
let _ = hyper::server::conn::http1::Builder::new()
.serve_connection(TokioIo::new(stream), service)
.await;
});
let mut client = tokio::net::TcpStream::connect(addr).await.expect("connect");
client
.write_all(b"PUT /bucket/key HTTP/1.1\r\nHost: localhost\r\nContent-Length: 100\r\n\r\nabc")
.await
.expect("write partial body");
client.shutdown().await.expect("shutdown write side");
let _ = tokio::time::timeout(std::time::Duration::from_secs(10), server).await;
let captured = captured.lock().expect("slot").take();
captured.expect("hyper body error captured")
}
#[tokio::test]
async fn hyper_body_eof_maps_to_incomplete_body_across_io_boundaries() {
let hyper_err = capture_hyper_body_eof_error().await;
assert!(is_hyper_body_eof(&hyper_err), "unexpected hyper error shape: {hyper_err:?}");
// Shape observed in production (issue #7596):
// Custom { UnexpectedEof, Custom { Other, hyper::Error(Body, UnexpectedEof, IncompleteBody) } }
let nested = IoError::new(ErrorKind::UnexpectedEof, IoError::other(hyper_err));
let storage: ApiError = StorageError::Io(nested).into();
assert_eq!(storage.code, S3ErrorCode::IncompleteBody);
assert_eq!(storage.message, ApiError::error_code_to_message(&S3ErrorCode::IncompleteBody));
let hyper_err = capture_hyper_body_eof_error().await;
let direct: ApiError = IoError::other(hyper_err).into();
assert_eq!(direct.code, S3ErrorCode::IncompleteBody);
}
#[test]
fn server_side_source_read_error_maps_to_service_unavailable_before_incomplete_body() {
let short_source = IoError::new(ErrorKind::UnexpectedEof, rustfs_rio::IncompleteBody { remaining: 17 });
+169 -18
View File
@@ -14,6 +14,7 @@
// Import HTTP server components and compression configuration
use crate::admin;
use crate::app::object::request_body::{BodyReadControl, ObservedBody};
use crate::auth::IAMAuth;
use crate::auth_keystone;
use crate::config;
@@ -158,13 +159,11 @@ static HTTP_STATUS_CLASS_METRICS: std::sync::LazyLock<[HttpStatusClassMetrics; 6
static HTTP_TRANSPORT_FAILURES_COUNTER: std::sync::LazyLock<metrics::Counter> =
std::sync::LazyLock::new(|| counter!(METRIC_HTTP_SERVER_FAILURES_TOTAL, LABEL_HTTP_STATUS_CLASS => "transport"));
const RUSTFS_S3_PUT_OBJECT_MAX_SIZE: u64 = 5 * 1024 * 1024 * 1024;
fn rustfs_s3_config() -> S3Config {
let mut s3_config = S3Config::default();
s3_config.normalize_forward_slash_path = true;
s3_config.enable_sig_v2 = true;
s3_config.put_object_max_size = Some(RUSTFS_S3_PUT_OBJECT_MAX_SIZE);
s3_config.put_object_max_size = Some(rustfs_config::MAX_SINGLE_PUT_OBJECT_SIZE);
s3_config.sig_v4_allowed_services.push("s3tables".to_string());
s3_config
}
@@ -827,13 +826,11 @@ where
impl<S, B, ResBody, ServiceError> Service<HttpRequest<B>> for EarlyResponseBodyService<S>
where
S: Service<HttpRequest<B>, Response = Response<ResBody>, Error = ServiceError>
+ Service<HttpRequest<EarlyResponseBody<B>>, Response = Response<ResBody>, Error = ServiceError>
S: Service<HttpRequest<ObservedBody<EarlyResponseBody<B>>>, Response = Response<ResBody>, Error = ServiceError>
+ Clone
+ Send
+ 'static,
<S as Service<HttpRequest<B>>>::Future: Send + 'static,
<S as Service<HttpRequest<EarlyResponseBody<B>>>>::Future: Send + 'static,
<S as Service<HttpRequest<ObservedBody<EarlyResponseBody<B>>>>>::Future: Send + 'static,
B: http_body::Body<Data = Bytes> + Send + Unpin + 'static,
B::Error: std::error::Error + Send + Sync + 'static,
ResBody: Send + 'static,
@@ -844,32 +841,32 @@ where
type Future = Pin<Box<dyn Future<Output = std::result::Result<Self::Response, Self::Error>> + Send>>;
fn poll_ready(&mut self, cx: &mut Context<'_>) -> Poll<std::result::Result<(), Self::Error>> {
match <S as Service<HttpRequest<B>>>::poll_ready(&mut self.inner, cx)? {
Poll::Ready(()) => <S as Service<HttpRequest<EarlyResponseBody<B>>>>::poll_ready(&mut self.inner, cx),
Poll::Pending => Poll::Pending,
}
self.inner.poll_ready(cx)
}
fn call(&mut self, req: HttpRequest<B>) -> Self::Future {
let version = req.version();
let preserve_on_drop = matches!(version, Version::HTTP_10 | Version::HTTP_11) && !req.body().is_end_stream();
let mut inner = self.inner.clone();
if !preserve_on_drop {
return Box::pin(async move { <S as Service<HttpRequest<B>>>::call(&mut inner, req).await });
}
let mut req = req;
let control = BodyReadControl::default();
req.extensions_mut().insert(control.clone());
let mut drain_context = EarlyResponseBodyDrainContext::from_request(&req, self.idle_timeout);
let state = Arc::new(EarlyResponseBodyState::default());
let guarded_req = req.map({
let state = Arc::clone(&state);
move |body| EarlyResponseBody::new(body, state)
move |body| ObservedBody::new(EarlyResponseBody::new(body, state), control)
});
Box::pin(async move {
let result = <S as Service<HttpRequest<EarlyResponseBody<B>>>>::call(&mut inner, guarded_req).await;
let result = inner.call(guarded_req).await;
let Some(abandoned) = state.take_abandoned() else {
return result;
};
if !preserve_on_drop {
return result;
}
match result {
Ok(mut response) => {
@@ -2458,9 +2455,10 @@ mod tests {
use crate::storage_api::server::http::ScannerScopedDirtyUsageAckEntry;
use bytes::Bytes;
use http::Request as HttpRequest;
use http::header::CONTENT_LENGTH;
use http::{HeaderMap, StatusCode};
use http_body::Frame;
use http_body_util::{Empty, Full};
use http_body_util::{BodyExt, Empty, Full};
use metrics::with_local_recorder;
use metrics_util::debugging::{DebugValue, DebuggingRecorder};
use opentelemetry::propagation::Extractor;
@@ -2887,6 +2885,159 @@ mod tests {
assert_eq!(bytes_polled.load(Ordering::Relaxed), 0);
}
#[derive(Clone, Copy)]
struct UploadPartTimeoutS3 {
timeout: Duration,
}
#[async_trait::async_trait]
impl s3s::S3 for UploadPartTimeoutS3 {
async fn upload_part(
&self,
req: s3s::S3Request<s3s::dto::UploadPartInput>,
) -> s3s::S3Result<s3s::S3Response<s3s::dto::UploadPartOutput>> {
use futures::StreamExt;
use tokio_util::io::StreamReader;
let control = req
.extensions
.get::<BodyReadControl>()
.expect("HTTP route must install the control")
.clone();
control.activate(self.timeout, "bucket", "object", "request", 1024);
let body = req.input.body.expect("upload body");
let inner = rustfs_rio::wrap_reader(StreamReader::new(body.map(|item| item.map_err(std::io::Error::other))));
let mut reader = crate::app::object::request_body::DemandReader::new(inner, control);
reader
.read_to_end(&mut Vec::new())
.await
.map_err(|error| s3s::S3Error::from(crate::error::ApiError::from(error)))?;
Ok(s3s::S3Response::new(s3s::dto::UploadPartOutput::default()))
}
async fn head_bucket(
&self,
_req: s3s::S3Request<s3s::dto::HeadBucketInput>,
) -> s3s::S3Result<s3s::S3Response<s3s::dto::HeadBucketOutput>> {
Ok(s3s::S3Response::new(s3s::dto::HeadBucketOutput::default()))
}
}
#[tokio::test(start_paused = true)]
async fn upload_part_timeout_preserves_http1_drain_and_http2_stream_drop() {
for version in [Version::HTTP_11, Version::HTTP_2] {
let (sender, body, bytes_polled, dropped) = tracked_request_body();
let request = HttpRequest::builder()
.version(version)
.method(Method::PUT)
.uri("/bucket/object?partNumber=1&uploadId=upload")
.header(CONTENT_LENGTH, "1024")
.body(body)
.expect("upload request");
let inner = s3s::service::S3ServiceBuilder::new(UploadPartTimeoutS3 {
timeout: Duration::from_secs(300),
})
.build();
let mut service = EarlyResponseBodyService::new(inner, Duration::from_secs(30));
let mut call = Box::pin(service.call(request));
assert!(futures::poll!(call.as_mut()).is_pending());
tokio::time::advance(Duration::from_secs(300)).await;
let response = call.await.expect("timeout response");
assert_eq!(response.status(), StatusCode::BAD_REQUEST);
assert_eq!(response.headers().get(CONNECTION).is_some(), version == Version::HTTP_11);
let xml = response.into_body().collect().await.expect("timeout XML").to_bytes();
assert!(String::from_utf8_lossy(&xml).contains("<Code>RequestTimeout</Code>"));
if version == Version::HTTP_11 {
assert!(
!dropped.load(Ordering::Acquire),
"synthetic errors must retain the unfinished raw transport"
);
sender
.send(Bytes::from_static(b"late payload"))
.expect("native drain receiver");
drop(sender);
tokio::task::yield_now().await;
assert_eq!(bytes_polled.load(Ordering::Relaxed), 12);
} else {
assert!(sender.is_closed(), "HTTP/2 must drop only the failed body");
assert_eq!(bytes_polled.load(Ordering::Relaxed), 0);
}
assert!(dropped.load(Ordering::Acquire));
}
}
#[tokio::test]
async fn upload_part_timeout_leaves_other_http2_streams_usable() {
let (client_io, server_io) = tokio::io::duplex(64 * 1024);
let inner = s3s::service::S3ServiceBuilder::new(UploadPartTimeoutS3 {
timeout: Duration::from_millis(500),
})
.build();
let service = EarlyResponseBodyService::new(inner, Duration::from_secs(1));
let server = tokio::spawn(async move {
hyper::server::conn::http2::Builder::new(hyper_util::rt::TokioExecutor::new())
.serve_connection(TokioIo::new(server_io), TowerToHyperService::new(service))
.await
});
let (mut client, connection) = hyper::client::conn::http2::handshake::<_, _, TrackedRequestBody>(
hyper_util::rt::TokioExecutor::new(),
TokioIo::new(client_io),
)
.await
.expect("HTTP/2 handshake");
let connection = tokio::spawn(connection);
let (_sender, body, _, _) = tracked_request_body();
let stalled = client.send_request(
HttpRequest::builder()
.method(Method::PUT)
.uri("http://localhost/bucket/object?partNumber=1&uploadId=upload")
.header(CONTENT_LENGTH, "1024")
.body(body)
.expect("stalled request"),
);
client.ready().await.expect("same connection remains ready");
let (sender, body, _, _) = tracked_request_body();
drop(sender);
let response = client
.send_request(
HttpRequest::builder()
.method(Method::HEAD)
.uri("http://localhost/bucket")
.body(body)
.expect("healthy stream"),
)
.await
.expect("healthy response");
assert_eq!(response.status(), StatusCode::OK);
assert!(!response.headers().contains_key(CONNECTION));
response.into_body().collect().await.expect("healthy stream completes");
let response = tokio::time::timeout(Duration::from_secs(5), stalled)
.await
.expect("stream timeout")
.expect("S3 response");
assert_eq!(response.status(), StatusCode::BAD_REQUEST);
let xml = response.into_body().collect().await.expect("timeout body").to_bytes();
assert!(String::from_utf8_lossy(&xml).contains("<Code>RequestTimeout</Code>"));
client.ready().await.expect("connection after failed stream");
let (sender, body, _, _) = tracked_request_body();
drop(sender);
let response = client
.send_request(
HttpRequest::builder()
.method(Method::HEAD)
.uri("http://localhost/bucket")
.body(body)
.expect("later stream"),
)
.await
.expect("later response");
assert_eq!(response.status(), StatusCode::OK);
drop(client);
connection.abort();
server.abort();
}
#[tokio::test(start_paused = true)]
async fn early_response_body_drain_releases_stalled_body_after_idle_timeout() {
let (_sender, body, _bytes_polled, dropped) = tracked_request_body();
@@ -3049,7 +3200,7 @@ mod tests {
assert!(s3_config.normalize_forward_slash_path);
assert!(s3_config.normalize_content_length);
assert!(s3_config.enable_sig_v2);
assert_eq!(s3_config.put_object_max_size, Some(RUSTFS_S3_PUT_OBJECT_MAX_SIZE));
assert_eq!(s3_config.put_object_max_size, Some(rustfs_config::MAX_SINGLE_PUT_OBJECT_SIZE));
assert!(s3_config.sig_v4_allowed_services.iter().any(|service| service == "s3"));
assert!(s3_config.sig_v4_allowed_services.iter().any(|service| service == "sts"));
assert!(s3_config.sig_v4_allowed_services.iter().any(|service| service == "s3tables"));
+82 -12
View File
@@ -1384,9 +1384,17 @@ pub(crate) fn reconcile_site_replication_bucket_targets(
continue;
};
// Only a target this pass itself derived earlier is updated in place:
// the same ARN, or the same peer under an older ARN shape (legacy
// `arn:rustfs:` / regional) recognisable by the site-replication
// service account and the same-name target bucket. An operator's
// bucket-level target that happens to point at the peer (different
// target bucket and credentials) is left alone and the site target
// is added next to it; replacing it silently orphaned the operator's
// replication rule (rustfs/backlog#2489, MinIO `getRemoteARN` parity).
if let Some(index) = targets.iter().position(|existing| {
existing.target_type == BucketTargetType::ReplicationService
&& (bucket_target_matches_peer(existing, peer) || existing.arn == target.arn)
&& (existing.arn == target.arn || is_site_replication_owned_target(existing, &target, peer, state))
}) {
let existing = targets[index].clone();
target.path = existing.path;
@@ -1414,6 +1422,24 @@ pub(crate) fn reconcile_site_replication_bucket_targets(
Ok(BucketTargets { targets })
}
/// A stored target that site replication derived for `peer` under an older
/// ARN shape: it names the same-name target bucket and carries the site
/// replication service account. Operator targets never match — their target
/// bucket or credentials differ — so reconciliation cannot take them over.
fn is_site_replication_owned_target(
existing: &BucketTarget,
derived: &BucketTarget,
peer: &PeerInfo,
state: &SiteReplicationState,
) -> bool {
bucket_target_matches_peer(existing, peer)
&& existing.target_bucket == derived.target_bucket
&& existing
.credentials
.as_ref()
.is_some_and(|credentials| credentials.access_key == state.service_account_access_key)
}
/// Whether every `site-repl-*` rule on this bucket resolves to a live remote target.
///
/// The rule set alone cannot answer this: a rule can be perfectly formed while the endpoint
@@ -1630,6 +1656,44 @@ pub(crate) fn build_site_replication_config(
}
}
/// Reload `bucket`'s metadata on every other node of this site after a
/// site-replication write. Every S3 bucket-config write does this
/// (`app::bucket_usecase::notify_bucket_metadata_reload`); the
/// site-replication writers did not, so on a multi-node site a node other
/// than the one that applied the write served the previous targets and
/// rules for up to the 15-minute refresh — a `resync start` routed to such a
/// node reported every freshly wired bucket as `Config not found` or
/// `recorded remote target no longer exists` (backlog#2367 A-5, backlog#2195
/// item 2). Best effort like the S3 path: the write is durable and the
/// refresh loop is the fallback, so an unreachable node must not fail the
/// operation that already committed.
pub(crate) async fn reload_bucket_metadata_on_peers(bucket: &str, operation: &'static str, scanner_maintenance_change: bool) {
if scanner_maintenance_change {
rustfs_scanner::record_scanner_maintenance_change(bucket);
}
let Some(notification_sys) = crate::admin::runtime_sources::current_notification_system() else {
return;
};
let result = if scanner_maintenance_change {
notification_sys.load_bucket_metadata_for_scanner_maintenance(bucket).await
} else {
notification_sys.load_bucket_metadata(bucket).await
};
if let Err(err) = result {
warn!(
event = EVENT_ADMIN_SITE_REPLICATION_STATE,
component = LOG_COMPONENT_ADMIN,
subsystem = LOG_SUBSYSTEM_SITE_REPLICATION,
bucket = %bucket,
operation,
result = "peer_metadata_reload_failed",
error = %err,
"admin site replication state"
);
}
}
/// Returns whether the bucket targets were rewritten.
pub(crate) async fn ensure_site_replication_bucket_targets_with_runtime(
bucket: &str,
state: &SiteReplicationState,
@@ -1637,7 +1701,7 @@ pub(crate) async fn ensure_site_replication_bucket_targets_with_runtime(
config: Option<&ReplicationConfiguration>,
service_account_secret_key: &str,
expected_incarnation_id: Uuid,
) -> S3Result<()> {
) -> S3Result<bool> {
let existing = match metadata_sys::list_bucket_targets(bucket).await {
Ok(targets) => targets,
Err(StorageError::ConfigNotFound) => BucketTargets::default(),
@@ -1649,7 +1713,7 @@ pub(crate) async fn ensure_site_replication_bucket_targets_with_runtime(
let updated =
reconcile_site_replication_bucket_targets(existing, bucket, state, local_peer, config, service_account_secret_key)?;
if updated.targets.is_empty() {
return Ok(());
return Ok(false);
}
let json_targets = serde_json::to_vec(&updated)
@@ -1658,12 +1722,12 @@ pub(crate) async fn ensure_site_replication_bucket_targets_with_runtime(
// client — noticeable now that startup reconciles all buckets, not just the one bucket
// an operation touched.
if json_targets == existing_json {
return Ok(());
return Ok(false);
}
metadata_sys::update_if_incarnation(bucket, BUCKET_TARGETS_FILE, json_targets, expected_incarnation_id)
.await
.map_err(ApiError::from)?;
Ok(())
Ok(true)
}
pub(crate) async fn bucket_replication_config_for_target_refresh(bucket: &str) -> S3Result<Option<ReplicationConfiguration>> {
@@ -1674,13 +1738,14 @@ pub(crate) async fn bucket_replication_config_for_target_refresh(bucket: &str) -
}
}
/// Returns whether the replication configuration was rewritten.
pub(crate) async fn ensure_site_replication_bucket_replication_config_with_runtime(
bucket: &str,
state: &SiteReplicationState,
local_peer: &PeerInfo,
service_account_secret_key: &str,
expected_incarnation_id: Uuid,
) -> S3Result<()> {
) -> S3Result<bool> {
let existing = match metadata_sys::get_replication_config(bucket).await {
Ok((existing, _)) => Some(existing),
Err(StorageError::ConfigNotFound) => None,
@@ -1689,7 +1754,7 @@ pub(crate) async fn ensure_site_replication_bucket_replication_config_with_runti
let Some(desired) = build_site_replication_config(bucket, state, local_peer, service_account_secret_key, existing.as_ref())?
else {
return Ok(());
return Ok(false);
};
// Derived rules are state owned by this site: rebuild them from the current peer
@@ -1721,7 +1786,7 @@ pub(crate) async fn ensure_site_replication_bucket_replication_config_with_runti
};
if rules == existing_rules && role == existing_role {
return Ok(());
return Ok(false);
}
let data = serialize(&ReplicationConfiguration { role, rules })
@@ -1730,7 +1795,7 @@ pub(crate) async fn ensure_site_replication_bucket_replication_config_with_runti
.await
.map_err(ApiError::from)?;
Ok(())
Ok(true)
}
pub(crate) async fn ensure_site_replication_bucket_setup_with_runtime(
@@ -1748,9 +1813,9 @@ pub(crate) async fn ensure_site_replication_bucket_setup_with_runtime_for_incarn
runtime: &SiteReplicationRuntime,
expected_incarnation_id: Uuid,
) -> S3Result<()> {
let _targets_guard = lock_bucket_targets_metadata(bucket).await;
let targets_guard = lock_bucket_targets_metadata(bucket).await;
let config = bucket_replication_config_for_target_refresh(bucket).await?;
ensure_site_replication_bucket_targets_with_runtime(
let targets_written = ensure_site_replication_bucket_targets_with_runtime(
bucket,
&runtime.state,
&runtime.local_peer,
@@ -1759,7 +1824,7 @@ pub(crate) async fn ensure_site_replication_bucket_setup_with_runtime_for_incarn
expected_incarnation_id,
)
.await?;
ensure_site_replication_bucket_replication_config_with_runtime(
let config_written = ensure_site_replication_bucket_replication_config_with_runtime(
bucket,
&runtime.state,
&runtime.local_peer,
@@ -1767,6 +1832,10 @@ pub(crate) async fn ensure_site_replication_bucket_setup_with_runtime_for_incarn
expected_incarnation_id,
)
.await?;
drop(targets_guard);
if targets_written || config_written {
reload_bucket_metadata_on_peers(bucket, "site_replication_bucket_setup", config_written).await;
}
Ok(())
}
@@ -1791,6 +1860,7 @@ pub(crate) async fn ensure_site_replication_bucket_versioning(bucket: &str) -> S
metadata_sys::update_if_incarnation(bucket, BUCKET_VERSIONING_CONFIG, bucket_versioning_xml()?, expected_incarnation_id)
.await
.map_err(ApiError::from)?;
reload_bucket_metadata_on_peers(bucket, "site_replication_bucket_versioning", false).await;
Ok(())
}
+45 -34
View File
@@ -52,7 +52,11 @@ pub(crate) struct SiteReplicationRetryEvent {
/// deletion body (if it was a deletion) recorded in
/// [`SiteReplicationState::iam_deletion_replays`]. Only then may a
/// successful deletion replay plus a stable snapshot resend settle the
/// entry; a legacy entry (or one degraded by record overflow) keeps the
/// entry. Every entry this binary creates starts recorded: the IAM
/// change hook records deletion bodies, and the other creators (the add
/// bootstrap's snapshot send, the drain's own replay) never carry a
/// deletion. A legacy entry persisted by a binary that predates recording
/// (serde default `false`), or one degraded by record overflow, keeps the
/// escalation semantics because an unrecorded deletion may hide in it.
#[serde(default, skip_serializing_if = "std::ops::Not::not")]
pub(crate) deletions_recorded: bool,
@@ -308,7 +312,12 @@ fn push_site_replication_retry_event(
updated_at: Some(OffsetDateTime::now_utc()),
edit_generation: generation,
peer_unreachable,
deletions_recorded: false,
// See the field doc: only a row persisted by an older binary is
// unrecorded. Stamping at creation is what lets an entry first
// created by the bootstrap snapshot send settle after a later
// deletion is replayed, instead of escalating forever
// (backlog#2367 A-3).
deletions_recorded: true,
});
Ok(evicted)
}
@@ -598,22 +607,7 @@ pub(crate) fn record_failed_iam_delivery(
item: &SRIAMItem,
error: &str,
) -> S3Result<()> {
let existed = state
.retry_queue
.iter()
.any(|event| retry_event_matches(event, peer, SITE_REPLICATION_RETRY_IAM_SNAPSHOT_PATH));
upsert_site_replication_retry_event(&mut state.retry_queue, peer, SITE_REPLICATION_PEER_IAM_ITEM_WIRE_PATH, error, None)?;
if !existed
&& let Some(event) = state
.retry_queue
.iter_mut()
.find(|event| retry_event_matches(event, peer, SITE_REPLICATION_RETRY_IAM_SNAPSHOT_PATH))
{
// Fresh entry: every failure it will ever collapse goes through this
// recording path, so a deletion replay plus a stable snapshot resend
// can later settle it instead of escalating.
event.deletions_recorded = true;
}
let Some(entity) = iam_item_deletion_entity(item) else {
return Ok(());
@@ -1418,6 +1412,33 @@ pub(crate) fn site_replication_retry_backoff_elapsed(event: &SiteReplicationRetr
now.unix_timestamp().saturating_sub(updated_at.unix_timestamp()) >= delay
}
/// Backoff evaluation time for the heavyweight tick: halfway to the next
/// tick. Backoffs are multiples of the tick interval, so an entry stamped δ
/// seconds after a tick is `600 − δ` old at the next one and slipped a whole
/// extra interval for every δ > 0 — a first replay landed at T+1200 rather
/// than T+600 (backlog#2367 A-1). Evaluating at the midpoint bounds the slip
/// to half an interval either way; timestamps written back stay real time.
pub(crate) fn heavyweight_retry_drain_horizon(now: OffsetDateTime) -> OffsetDateTime {
let half_interval = crate::site_replication_reconcile::RECONCILE_INTERVAL / 2;
now + time::Duration::seconds(i64::try_from(half_interval.as_secs()).unwrap_or(i64::MAX))
}
/// What the lightweight 30-second pass may act on. It replays bounded bucket
/// ops only, but probes every backed-off class: promotion is a state flip
/// the heavyweight tick then replays, so an IAM or bucket-metadata snapshot
/// owed to a peer that came back is resent at the next tick instead of
/// after its own backoff has fully elapsed (backlog#2367 A-1).
pub(crate) fn lightweight_retry_drain_partition(
state: &SiteReplicationState,
now: OffsetDateTime,
) -> (Vec<SiteReplicationRetryEvent>, Vec<SiteReplicationRetryEvent>) {
let mut actionable = actionable_site_replication_retry_events(state, now);
actionable.retain(|event| {
classify_site_replication_retry_event(event).is_some_and(|action| is_lightweight_retry_drain_action(&action))
});
(actionable, deferred_site_replication_retry_events(state, now))
}
/// The subset of the retry queue the background drain is allowed to touch.
pub(crate) fn actionable_site_replication_retry_events(
state: &SiteReplicationState,
@@ -1666,14 +1687,7 @@ async fn drain_site_replication_retry_queue_lightweight_inner() -> S3Result<()>
return Ok(());
}
let now = OffsetDateTime::now_utc();
let mut actionable = actionable_site_replication_retry_events(&runtime.state, now);
let mut deferred = deferred_site_replication_retry_events(&runtime.state, now);
actionable.retain(|event| {
classify_site_replication_retry_event(event).is_some_and(|action| is_lightweight_retry_drain_action(&action))
});
deferred.retain(|event| {
classify_site_replication_retry_event(event).is_some_and(|action| is_lightweight_retry_drain_action(&action))
});
let (actionable, deferred) = lightweight_retry_drain_partition(&runtime.state, now);
if actionable.is_empty() && deferred.is_empty() {
return Ok(());
}
@@ -1697,10 +1711,7 @@ async fn drain_site_replication_retry_queue_lightweight_inner() -> S3Result<()>
return Ok(());
}
let now = OffsetDateTime::now_utc();
let mut actionable = actionable_site_replication_retry_events(&runtime.state, now);
actionable.retain(|event| {
classify_site_replication_retry_event(event).is_some_and(|action| is_lightweight_retry_drain_action(&action))
});
let (actionable, _) = lightweight_retry_drain_partition(&runtime.state, now);
if actionable.is_empty() {
return Ok(());
}
@@ -1717,9 +1728,9 @@ pub(crate) async fn drain_site_replication_retry_queue_inner() -> S3Result<()> {
// The alert must fire even when nothing is drainable this tick —
// escalated markers are exactly the entries the drain skips.
log_site_replication_retry_liabilities(&runtime.state);
let now = OffsetDateTime::now_utc();
let actionable = actionable_site_replication_retry_events(&runtime.state, now);
let deferred = deferred_site_replication_retry_events(&runtime.state, now);
let horizon = heavyweight_retry_drain_horizon(OffsetDateTime::now_utc());
let actionable = actionable_site_replication_retry_events(&runtime.state, horizon);
let deferred = deferred_site_replication_retry_events(&runtime.state, horizon);
if actionable.is_empty() && deferred.is_empty() {
return Ok(());
}
@@ -1764,8 +1775,8 @@ pub(crate) async fn drain_site_replication_retry_queue_inner() -> S3Result<()> {
{
return Ok(());
}
let now = OffsetDateTime::now_utc();
let actionable = actionable_site_replication_retry_events(&runtime.state, now);
let horizon = heavyweight_retry_drain_horizon(OffsetDateTime::now_utc());
let actionable = actionable_site_replication_retry_events(&runtime.state, horizon);
if actionable.is_empty() {
return Ok(());
}
+235 -2
View File
@@ -98,11 +98,46 @@ async fn spawn_test_tls_server_with_response(response: &'static [u8]) -> (String
break;
}
}
stream.write_all(response).await.is_ok()
// Flush buffered TLS records and send close_notify before dropping the socket.
stream.write_all(response).await.is_ok() && stream.shutdown().await.is_ok()
});
(endpoint, ca_pem, task)
}
#[tokio::test]
async fn tls_test_server_delivers_response_and_closes_cleanly() {
use rustls_pki_types::pem::PemObject;
let (endpoint, ca_pem, server) = spawn_test_tls_server().await;
let mut roots = rustls::RootCertStore::empty();
roots
.add(rustls_pki_types::CertificateDer::from_pem_slice(ca_pem.as_bytes()).expect("parse test CA"))
.expect("trust test CA");
let config = rustls::ClientConfig::builder()
.with_root_certificates(roots)
.with_no_client_auth();
let connector = tokio_rustls::TlsConnector::from(Arc::new(config));
let socket = tokio::net::TcpStream::connect(endpoint.strip_prefix("https://").expect("TLS endpoint"))
.await
.expect("connect to TLS test server");
let mut stream = connector
.connect(rustls_pki_types::ServerName::try_from("127.0.0.1").expect("test server name"), socket)
.await
.expect("trust TLS test server");
stream
.write_all(b"GET / HTTP/1.1\r\nHost: localhost\r\nConnection: close\r\n\r\n")
.await
.expect("write test request");
stream.flush().await.expect("flush test request");
let mut response = Vec::new();
tokio::time::timeout(Duration::from_secs(5), stream.read_to_end(&mut response))
.await
.expect("TLS response must finish")
.expect("TLS test server must send close_notify before closing");
assert!(response.ends_with(b"\r\n\r\nok"));
assert!(server.await.expect("TLS test server task"));
}
#[test]
fn peer_connection_validation_accepts_supported_combinations() {
let ca = valid_test_ca_pem("peer.example.com");
@@ -845,7 +880,8 @@ fn test_record_failed_iam_delivery_records_deletions_and_flags_entry() {
record_failed_iam_delivery(&mut state, &target, &policy_delete_item("readonly"), "peer offline").expect("record failure");
assert_eq!(state.iam_deletion_replays.len(), 2);
// A legacy entry (created without recording) is never stamped.
// A legacy entry (persisted by a binary that predates recording, so it
// deserialized with the `false` default) is never stamped.
let legacy = PeerInfo {
deployment_id: "legacy-dep".to_string(),
..peer("legacy", "https://legacy.example.com")
@@ -859,6 +895,12 @@ fn test_record_failed_iam_delivery_records_deletions_and_flags_entry() {
None,
)
.expect("upsert retry event");
state
.retry_queue
.iter_mut()
.find(|event| event.peer_deployment_id == legacy.deployment_id)
.expect("legacy entry")
.deletions_recorded = false;
record_failed_iam_delivery(&mut state, &legacy, &user_delete_item("bob"), "peer offline").expect("record failure");
let legacy_event = state
.retry_queue
@@ -871,6 +913,43 @@ fn test_record_failed_iam_delivery_records_deletions_and_flags_entry() {
);
}
/// backlog#2367 A-3: an entry first created by a non-deletion failure — the
/// add bootstrap's snapshot send, or the drain's own replay — hides no
/// unrecorded deletion, so a deletion recorded later plus a stable snapshot
/// resend must settle it instead of escalating it to the permanent marker
/// that only `replicate repair` clears.
#[test]
fn test_bootstrap_created_iam_entry_settles_after_deletion_replay() {
let target = PeerInfo {
deployment_id: "remote-dep".to_string(),
..peer("remote", "https://remote.example.com")
};
let mut state = deletion_replay_state(&target);
upsert_site_replication_retry_event(
&mut state.retry_queue,
&target,
SITE_REPLICATION_PEER_IAM_ITEM_WIRE_PATH,
"peer request to https://remote.example.com failed (connect): connection refused",
None,
)
.expect("bootstrap send failure");
assert!(state.retry_queue[0].deletions_recorded, "a fresh entry carries no unrecorded deletion");
record_failed_iam_delivery(&mut state, &target, &user_delete_item("alice"), "peer offline").expect("record failure");
assert_eq!(state.retry_queue.len(), 1, "the hook failure collapses into the bootstrap entry");
assert!(state.retry_queue[0].deletions_recorded);
assert_eq!(state.iam_deletion_replays.len(), 1);
let observed = state.retry_queue[0].clone();
let replayed: Vec<String> = state.iam_deletion_replays.iter().map(|record| record.id.clone()).collect();
assert!(
settle_replayed_iam_retry_events(&mut state, &target, &observed, &replayed),
"the replayed deletion plus the snapshot resend settle the entry"
);
assert!(state.retry_queue.is_empty(), "no escalation marker may remain: {:?}", state.retry_queue);
assert!(state.iam_deletion_replays.is_empty());
}
/// Overflowing the per-peer record cap degrades the entry back to the
/// escalation semantics: the record set is no longer complete, so a replay
/// can no longer prove the peer converged.
@@ -1626,6 +1705,74 @@ fn test_deferred_retry_events_do_not_probe_fresh_application_failures() {
assert!(actionable_site_replication_retry_events(&state, now).is_empty());
}
/// backlog#2367 A-1: the lightweight pass replays bucket ops only, but
/// probes every backed-off class so a recovered peer's IAM snapshot is
/// promoted within 30 seconds instead of waiting for the heavyweight tick
/// to notice it.
#[test]
fn test_lightweight_partition_probes_snapshot_entries_but_replays_bucket_ops_only() {
let now = OffsetDateTime::from_unix_timestamp(1_700_000_000).expect("timestamp");
let mut state = SiteReplicationState::default();
state
.peers
.insert("remote".to_string(), peer("remote", "https://remote.example.com"));
let bucket_make = "/rustfs/admin/v3/site-replication/peer/bucket-ops?bucket=photos&operation=make-with-versioning";
let mut iam_unreachable = drain_event(
"remote",
SITE_REPLICATION_RETRY_IAM_SNAPSHOT_PATH,
3,
Some(now - time::Duration::seconds(30)),
);
iam_unreachable.peer_unreachable = true;
let mut bucket_unreachable = drain_event("remote", bucket_make, 3, Some(now - time::Duration::seconds(30)));
bucket_unreachable.peer_unreachable = true;
state.retry_queue = vec![
iam_unreachable,
bucket_unreachable,
// Already promoted (or never stamped): due now.
drain_event("remote", SITE_REPLICATION_RETRY_BUCKET_METADATA_SNAPSHOT_PATH, 1, None),
drain_event("remote", bucket_make, 1, None),
];
let (actionable, deferred) = lightweight_retry_drain_partition(&state, now);
let deferred_paths: Vec<&str> = deferred.iter().map(|event| event.path.as_str()).collect();
assert!(
deferred_paths.contains(&SITE_REPLICATION_RETRY_IAM_SNAPSHOT_PATH),
"the backed-off IAM snapshot must be probed by the lightweight pass: {deferred_paths:?}"
);
assert!(deferred_paths.contains(&bucket_make));
assert_eq!(
actionable.iter().map(|event| event.path.as_str()).collect::<Vec<_>>(),
vec![bucket_make],
"only the bounded bucket op is replayed by the lightweight pass"
);
}
/// backlog#2367 A-1: the heavyweight tick evaluates backoff halfway to its
/// next tick. A first failure stamped one second after a tick is 599 s old
/// at the next tick; without the horizon it slipped to the tick after.
#[test]
fn test_heavyweight_horizon_absorbs_tick_phase() {
let now = OffsetDateTime::from_unix_timestamp(1_700_000_000).expect("timestamp");
let horizon = heavyweight_retry_drain_horizon(now);
assert_eq!(horizon - now, time::Duration::seconds(300));
let elapsed_at_horizon = |secs_ago: i64| {
site_replication_retry_backoff_elapsed(
&drain_event("remote", "/p", 1, Some(now - time::Duration::seconds(secs_ago))),
horizon,
)
};
// Stamped just after the previous tick: due at this tick, not the next.
assert!(elapsed_at_horizon(599));
// Due before the next tick's midpoint: drained now rather than a whole
// interval late.
assert!(elapsed_at_horizon(301));
// Due after the midpoint: waits for the next tick.
assert!(!elapsed_at_horizon(299));
}
/// The drain settles a peer-edit success under a freshly allocated
/// generation; legacy queue entries carry `edit_generation: None` and
/// must be cleared by that generation-scoped settlement (`(Some, None)`
@@ -1831,6 +1978,92 @@ fn test_site_replication_bucket_target_replaces_tls_and_preserves_operational_fi
assert!(target.disable_proxy);
}
/// rustfs/backlog#2489: an operator's bucket-level target that points at a
/// peer (different target bucket, operator credentials) must survive site
/// replication wiring untouched, with the site target added next to it.
/// Replacing it left the operator's rule pointing at an ARN no target backs.
#[test]
fn test_reconcile_site_replication_bucket_targets_keeps_operator_target_to_peer() {
let local = PeerInfo {
deployment_id: "local".to_string(),
..peer("local", "https://local.example.com")
};
let remote = PeerInfo {
deployment_id: "remote".to_string(),
..peer("remote", "http://remote.example.com:9000")
};
let state = SiteReplicationState {
service_account_access_key: "svc".to_string(),
peers: BTreeMap::from([("local".to_string(), local.clone()), ("remote".to_string(), remote)]),
..Default::default()
};
let operator_target = BucketTarget {
arn: "arn:minio:replication::7c0c5a1e-operator:photos-dst".to_string(),
source_bucket: "photos".to_string(),
target_bucket: "photos-dst".to_string(),
endpoint: "remote.example.com:9000".to_string(),
target_type: BucketTargetType::ReplicationService,
deployment_id: "remote".to_string(),
credentials: Some(Credentials {
access_key: "operator-key".to_string(),
secret_key: "operator-secret".to_string(),
session_token: None,
expiration: None,
}),
reset_id: "bucket-level-reset".to_string(),
..Default::default()
};
let reconciled = reconcile_site_replication_bucket_targets(
BucketTargets {
targets: vec![operator_target.clone()],
},
"photos",
&state,
&local,
None,
"secret",
)
.expect("reconcile targets");
assert_eq!(reconciled.targets.len(), 2, "the site target is added next to the operator target");
let kept = &reconciled.targets[0];
assert_eq!(
(kept.arn.as_str(), kept.target_bucket.as_str(), kept.reset_id.as_str()),
("arn:minio:replication::7c0c5a1e-operator:photos-dst", "photos-dst", "bucket-level-reset"),
"the operator target is untouched"
);
assert_eq!(kept.credentials.as_ref().map(|c| c.access_key.as_str()), Some("operator-key"));
let site_target = &reconciled.targets[1];
assert_eq!(site_target.arn, "arn:minio:replication::remote:photos");
assert_eq!(site_target.target_bucket, "photos");
assert!(
site_target.reset_id.is_empty(),
"the operator's resync identity must not leak into the site target"
);
// A site target under an older ARN shape is still recognised and updated in place.
let mut legacy = site_target.clone();
legacy.arn = "arn:rustfs:replication:us-east-1:remote:photos".to_string();
legacy.bandwidth_limit = 9;
let reconciled_again = reconcile_site_replication_bucket_targets(
BucketTargets {
targets: vec![operator_target.clone(), legacy],
},
"photos",
&state,
&local,
None,
"secret",
)
.expect("reconcile targets again");
assert_eq!(reconciled_again.targets.len(), 2, "a legacy-ARN site target is updated, not duplicated");
assert_eq!(reconciled_again.targets[0].arn, operator_target.arn);
assert_eq!(
reconciled_again.targets[1].bandwidth_limit, 9,
"operator tuning of the site target carries over"
);
}
#[test]
fn test_bucket_versioning_xml_enables_versioning() {
let data = bucket_versioning_xml().expect("versioning XML should serialize");
+1 -1
View File
@@ -32,7 +32,7 @@ use tokio::time::Instant;
use tokio_util::sync::CancellationToken;
use tracing::warn;
const RECONCILE_INTERVAL: Duration = Duration::from_secs(600);
pub(crate) const RECONCILE_INTERVAL: Duration = Duration::from_secs(600);
pub(crate) const RETRY_DRAIN_INTERVAL: Duration = Duration::from_secs(30);
/// A reconciler reports its own failures; the outcome carries no value because neither
+4 -4
View File
@@ -62,10 +62,10 @@ pub(crate) async fn init_embedded_bucket_metadata_runtime(store: Arc<ECStore>, c
let buckets: Vec<String> = buckets_list.into_iter().map(|v| v.name).collect();
try_migrate_bucket_metadata(store.clone()).await;
try_migrate_bucket_metadata(store.clone()).await?;
init_on_demand_migration_runtime();
init_bucket_metadata_sys(store.clone(), buckets.clone()).await;
try_migrate_iam_config(store).await;
try_migrate_iam_config(store).await?;
spawn_bucket_resync_startup_reconcile(buckets.clone(), ctx.clone(), false);
Ok(buckets)
@@ -82,9 +82,9 @@ pub(crate) async fn init_bucket_metadata_runtime(store: Arc<ECStore>, ctx: Cance
let buckets: Vec<String> = buckets_list.into_iter().map(|v| v.name).collect();
try_migrate_bucket_metadata(store.clone()).await;
try_migrate_bucket_metadata(store.clone()).await?;
try_migrate_iam_config(store.clone()).await;
try_migrate_iam_config(store.clone()).await?;
init_on_demand_migration_runtime();
init_bucket_metadata_sys(store, buckets.clone()).await;
spawn_bucket_resync_startup_reconcile(buckets.clone(), ctx, true);
+82 -191
View File
@@ -56,6 +56,10 @@ use std::sync::Arc;
use std::sync::OnceLock;
use url::{Url, form_urlencoded};
const EVENT_OBJECT_TAG_AUTHORIZATION: &str = "object_tag_authorization";
const LOG_COMPONENT_ACCESS: &str = "storage_access";
const LOG_SUBSYSTEM_AUTHORIZATION: &str = "authorization";
#[derive(Default, Clone, Debug)]
pub(crate) struct ReqInfo {
pub cred: Option<rustfs_credentials::Credentials>,
@@ -102,9 +106,13 @@ async fn authorize_replication_only_put_headers<T>(req: &mut S3Request<T>) -> S3
Ok(())
}
pub(crate) fn recursive_force_delete_is_authorized(headers: &HeaderMap, is_owner: bool, replica_request: bool) -> bool {
pub(crate) fn recursive_force_delete_has_authenticated_caller(
headers: &HeaderMap,
authenticated: bool,
replica_request: bool,
) -> bool {
!get_header(headers, SUFFIX_FORCE_DELETE).is_some_and(|value| value.eq_ignore_ascii_case("true"))
|| is_owner
|| authenticated
|| replica_request
}
@@ -833,15 +841,12 @@ pub(crate) fn log_list_buckets_iam_implicit_deny<T>(req: &S3Request<T>) -> S3Res
Ok(())
}
/// Extra action that may be evaluated in the same authorization flow and can
/// independently require `ExistingObjectTag` conditions.
fn secondary_tag_hint_action(action: Action, version_id: Option<&str>) -> Option<Action> {
match action {
Action::S3Action(S3Action::DeleteObjectAction) if version_id.is_some() => {
Some(Action::S3Action(S3Action::DeleteObjectVersionAction))
}
_ => None,
}
pub(crate) fn delete_object_authorize_action(version_id: Option<&str>) -> Action {
Action::S3Action(if version_id.is_some() {
S3Action::DeleteObjectVersionAction
} else {
S3Action::DeleteObjectAction
})
}
/// GHSA-3ppv: select the IAM action for an object read by whether the request
@@ -1015,14 +1020,14 @@ pub async fn authorize_request<T>(req: &mut S3Request<T>, action: Action) -> S3R
deny_only: false,
};
let prepared = iam_store.prepare_auth(&action_args).await;
let mut needs_tag_from_iam = prepared.needs_existing_object_tag;
let needs_tag_from_iam = prepared.needs_existing_object_tag;
let bucket_tag_hint = if !bucket.is_empty() && !object.is_empty() {
Some(load_bucket_policy_existing_object_tag_hint(store.as_ref(), bucket.as_str(), action).await)
} else {
None
};
let mut needs_tag_from_bucket = if let Some(hint) = bucket_tag_hint.as_ref() {
let needs_tag_from_bucket = if let Some(hint) = bucket_tag_hint.as_ref() {
let bucket_args = BucketPolicyArgs {
bucket: bucket.as_str(),
action,
@@ -1037,44 +1042,18 @@ pub async fn authorize_request<T>(req: &mut S3Request<T>, action: Action) -> S3R
false
};
let secondary_action = secondary_tag_hint_action(action, version_id.as_deref());
if let Some(extra_action) = secondary_action {
let extra_args = Args {
account: &cred.access_key,
groups: &cred.groups,
action: extra_action,
bucket: bucket.as_str(),
conditions: &conditions,
is_owner,
object: object.as_str(),
claims,
deny_only: false,
};
needs_tag_from_iam |= prepared.needs_existing_object_tag_for_args(&extra_args).await;
if let Some(hint) = bucket_tag_hint.as_ref() {
let extra_bucket_args = BucketPolicyArgs {
bucket: bucket.as_str(),
action: extra_action,
is_owner,
account: cred.access_key.as_str(),
groups: &cred.groups,
conditions: &conditions,
object: object.as_str(),
};
needs_tag_from_bucket |= bucket_policy_needs_existing_object_tag_from_hint(hint, &extra_bucket_args).await;
}
}
let needs_tag = needs_tag_from_iam || needs_tag_from_bucket;
if needs_tag {
tracing::debug!(
event = EVENT_OBJECT_TAG_AUTHORIZATION,
component = LOG_COMPONENT_ACCESS,
subsystem = LOG_SUBSYSTEM_AUTHORIZATION,
anonymous = false,
bucket = %bucket,
?action,
?secondary_action,
needs_tag_from_iam,
needs_tag_from_bucket,
"authorize_request ExistingObjectTag hint requires tag conditions"
"Object tag authorization conditions required"
);
}
maybe_merge_object_tag_conditions(
@@ -1116,41 +1095,8 @@ pub async fn authorize_request<T>(req: &mut S3Request<T>, action: Action) -> S3R
return Err(denial.deny("bucket_policy_explicit_deny", action));
}
if action == Action::S3Action(S3Action::DeleteObjectAction) && version_id.is_some() {
let delete_version_args = Args {
account: &cred.access_key,
groups: &cred.groups,
action: Action::S3Action(S3Action::DeleteObjectVersionAction),
bucket: bucket.as_str(),
conditions: &conditions,
is_owner,
object: object.as_str(),
claims,
deny_only: false,
};
let delete_version_allowed = iam_store.eval_prepared(&prepared, &delete_version_args).await;
if !delete_version_allowed
&& !PolicySys::try_is_allowed_for_store(
store.as_ref(),
&BucketPolicyArgs {
bucket: bucket.as_str(),
action: Action::S3Action(S3Action::DeleteObjectVersionAction),
is_owner,
account: &cred.access_key,
groups: &cred.groups,
conditions: &conditions,
object: object.as_str(),
},
)
.await
.map_err(ApiError::from)?
{
return Err(denial.deny("delete_object_version_denied", Action::S3Action(S3Action::DeleteObjectVersionAction)));
}
}
let iam_allowed = {
let final_args = Args {
let mut final_args = Args {
account: &cred.access_key,
groups: &cred.groups,
action,
@@ -1161,7 +1107,28 @@ pub async fn authorize_request<T>(req: &mut S3Request<T>, action: Action) -> S3R
claims,
deny_only: false,
};
iam_store.eval_prepared(&prepared, &final_args).await
let allowed = iam_store.eval_prepared(&prepared, &final_args).await;
if !allowed
&& matches!(
action,
Action::S3Action(
S3Action::DeleteObjectAction
| S3Action::DeleteObjectVersionAction
| S3Action::ListBucketVersionsAction
| S3Action::BypassGovernanceRetentionAction
| S3Action::ReplicateDeleteAction
)
)
&& prepared.combined_policy_for_view().is_some()
{
// Bucket policy Allow may supplement an implicit IAM denial,
// but must not override an explicit deletion-policy Deny.
final_args.deny_only = true;
if !iam_store.eval_prepared(&prepared, &final_args).await {
return Err(denial.deny("iam_explicit_deny", action));
}
}
allowed
};
if iam_allowed {
@@ -1199,42 +1166,6 @@ pub async fn authorize_request<T>(req: &mut S3Request<T>, action: Action) -> S3R
}
return Ok(());
}
if action == Action::S3Action(S3Action::ListBucketVersionsAction) {
let list_bucket_args = Args {
account: &cred.access_key,
groups: &cred.groups,
action: Action::S3Action(S3Action::ListBucketAction),
bucket: bucket.as_str(),
conditions: &conditions,
is_owner,
object: object.as_str(),
claims,
deny_only: false,
};
let list_bucket_allowed = iam_store.eval_prepared(&prepared, &list_bucket_args).await;
if list_bucket_allowed {
return Ok(());
}
if PolicySys::try_is_allowed_for_store(
store.as_ref(),
&BucketPolicyArgs {
bucket: bucket.as_str(),
action: Action::S3Action(S3Action::ListBucketAction),
is_owner,
account: &cred.access_key,
groups: &cred.groups,
conditions: &conditions,
object: object.as_str(),
},
)
.await
.map_err(ApiError::from)?
{
return Ok(());
}
}
} else {
let default_cred = rustfs_credentials::Credentials::default();
let client_info = req.extensions.get::<ClientInfo>();
@@ -1254,7 +1185,7 @@ pub async fn authorize_request<T>(req: &mut S3Request<T>, action: Action) -> S3R
} else {
None
};
let mut needs_tag_from_bucket = if let Some(hint) = bucket_tag_hint.as_ref() {
let needs_tag_from_bucket = if let Some(hint) = bucket_tag_hint.as_ref() {
let bucket_args = BucketPolicyArgs {
bucket: bucket.as_str(),
action,
@@ -1268,27 +1199,15 @@ pub async fn authorize_request<T>(req: &mut S3Request<T>, action: Action) -> S3R
} else {
false
};
let secondary_action = secondary_tag_hint_action(action, version_id.as_deref());
if let Some(extra_action) = secondary_action
&& let Some(hint) = bucket_tag_hint.as_ref()
{
let extra_bucket_args = BucketPolicyArgs {
bucket: bucket.as_str(),
action: extra_action,
is_owner: false,
account: "",
groups: &no_groups,
conditions: &conditions,
object: object.as_str(),
};
needs_tag_from_bucket |= bucket_policy_needs_existing_object_tag_from_hint(hint, &extra_bucket_args).await;
}
if needs_tag_from_bucket {
tracing::debug!(
event = EVENT_OBJECT_TAG_AUTHORIZATION,
component = LOG_COMPONENT_ACCESS,
subsystem = LOG_SUBSYSTEM_AUTHORIZATION,
anonymous = true,
bucket = %bucket,
?action,
?secondary_action,
"anonymous authorize_request ExistingObjectTag hint requires tag conditions"
"Object tag authorization conditions required"
);
}
maybe_merge_object_tag_conditions(
@@ -1324,28 +1243,6 @@ pub async fn authorize_request<T>(req: &mut S3Request<T>, action: Action) -> S3R
}
if action != Action::S3Action(S3Action::ListAllMyBucketsAction) {
if action == Action::S3Action(S3Action::DeleteObjectAction) && version_id.is_some() {
let delete_version_allowed = PolicySys::try_is_allowed_for_store(
store.as_ref(),
&BucketPolicyArgs {
bucket: bucket.as_str(),
action: Action::S3Action(S3Action::DeleteObjectVersionAction),
is_owner: false,
account: "",
groups: &None,
conditions: &conditions,
object: object.as_str(),
},
)
.await
.map_err(ApiError::from)?;
if !delete_version_allowed {
return Err(
denial.deny("delete_object_version_denied", Action::S3Action(S3Action::DeleteObjectVersionAction))
);
}
}
let policy_allowed = PolicySys::try_is_allowed_for_store(
store.as_ref(),
&BucketPolicyArgs {
@@ -1361,10 +1258,8 @@ pub async fn authorize_request<T>(req: &mut S3Request<T>, action: Action) -> S3R
.await
.map_err(ApiError::from)?;
// A bucket policy granting s3:ListBucket also covers listing versions. This
// fallback has to feed the same post-authorization gates as the direct grant
// below, otherwise a public bucket keeps serving anonymous
// ListObjectVersions after RestrictPublicBuckets is turned on.
// A bucket policy granting s3:ListBucket also covers listing versions.
// Keep this compatibility fallback inside the same public-access gate.
let policy_allowed = policy_allowed
|| (action == Action::S3Action(S3Action::ListBucketVersionsAction)
&& PolicySys::try_is_allowed_for_store(
@@ -2187,20 +2082,18 @@ impl S3Access for FS {
req_info.bucket = Some(req.input.bucket.clone());
req_info.object = Some(req.input.key.clone());
req_info.version_id = req.input.version_id.clone();
let is_owner = req_info.is_owner;
let authenticated = req_info.is_owner || req_info.cred.is_some();
let action = delete_object_authorize_action(req_info.version_id.as_deref());
authorize_request(req, Action::S3Action(S3Action::DeleteObjectAction)).await?;
authorize_request(req, action).await?;
let replica_request = req
.headers
.get(AMZ_BUCKET_REPLICATION_STATUS)
.and_then(|value| value.to_str().ok())
.is_some_and(|value| value == ReplicationStatusType::Replica.as_str());
if !recursive_force_delete_is_authorized(&req.headers, is_owner, replica_request) {
return Err(s3_error!(
AccessDenied,
"Recursive force-delete is restricted to internal or administrative requests"
));
if !recursive_force_delete_has_authenticated_caller(&req.headers, authenticated, replica_request) {
return Err(s3_error!(AccessDenied, "Recursive force-delete requires an authenticated caller"));
}
// S3 Standard: When bypass_governance header is set, must have s3:BypassGovernanceRetention permission
@@ -3122,14 +3015,14 @@ mod tests {
PostObjectRequestMarker, ReqInfo, S3Access, StorageError, TableDataPlanePublicationGuards, apply_bucket_generation_guard,
apply_copy_source_bucket_generation_guard, authorization_conditions, bucket_policy_needs_existing_object_tag_from_hint,
bucket_website_config_authorize_action, classify_bucket_policy_raw_load_error,
complete_multipart_upload_authorize_action, get_bucket_policy_authorize_action, has_write_offset_bytes_header,
install_restore_authorization_test_hook, legal_hold_write_requested, list_parts_authorize_action,
load_bucket_policy_existing_object_tag_hint, maybe_merge_object_tag_conditions, merge_list_bucket_query_conditions,
merge_request_object_tag_conditions, owner_can_bypass_policy_deny, post_object_authorize_action,
put_bucket_policy_authorize_action, request_context_from_req, request_object_store, require_owned_reserved_table_object,
retention_write_requested, secondary_tag_hint_action, table_data_plane_admin_action, table_data_plane_content_mutation,
table_data_plane_resource_for_request, table_publication_guard_error, validate_post_object_success_controls,
versioned_read_action,
complete_multipart_upload_authorize_action, delete_object_authorize_action, get_bucket_policy_authorize_action,
has_write_offset_bytes_header, install_restore_authorization_test_hook, legal_hold_write_requested,
list_parts_authorize_action, load_bucket_policy_existing_object_tag_hint, maybe_merge_object_tag_conditions,
merge_list_bucket_query_conditions, merge_request_object_tag_conditions, owner_can_bypass_policy_deny,
post_object_authorize_action, put_bucket_policy_authorize_action, request_context_from_req, request_object_store,
require_owned_reserved_table_object, retention_write_requested, secondary_tag_hint_action, table_data_plane_admin_action,
table_data_plane_content_mutation, table_data_plane_resource_for_request, table_publication_guard_error,
validate_post_object_success_controls, versioned_read_action,
};
use crate::error::ApiError;
use crate::storage::storage_api::contract::bucket::{BucketOperations as _, DeleteBucketOptions, MakeBucketOptions};
@@ -4010,20 +3903,18 @@ mod tests {
}
#[test]
fn test_secondary_tag_hint_action_for_delete_object_version() {
assert_eq!(
secondary_tag_hint_action(Action::S3Action(S3Action::DeleteObjectAction), Some("v1")),
Some(Action::S3Action(S3Action::DeleteObjectVersionAction))
);
assert_eq!(secondary_tag_hint_action(Action::S3Action(S3Action::DeleteObjectAction), None), None);
assert_eq!(
secondary_tag_hint_action(Action::S3Action(S3Action::ListBucketVersionsAction), None),
None
);
fn delete_authorization_selects_the_addressed_version() {
assert_eq!(delete_object_authorize_action(None), Action::S3Action(S3Action::DeleteObjectAction));
for version_id in ["null", "8f418ad0-f9f4-4458-83b2-cc72bc6f1b70"] {
assert_eq!(
delete_object_authorize_action(Some(version_id)),
Action::S3Action(S3Action::DeleteObjectVersionAction)
);
}
}
#[tokio::test]
async fn test_anonymous_delete_object_with_version_requires_secondary_policy_and_tag_hint() {
async fn test_anonymous_version_delete_uses_version_policy_and_tag_hint() {
let policy: BucketPolicy = serde_json::from_str(
r#"{
"Version":"2012-10-17",
@@ -4077,16 +3968,16 @@ mod tests {
"DeleteObjectVersion should still be denied without matching ExistingObjectTag conditions"
);
let needs_tag_main = bucket_policy_needs_existing_object_tag_from_hint(&hint, &args_delete).await;
let needs_tag_secondary = bucket_policy_needs_existing_object_tag_from_hint(&hint, &args_delete_version).await;
assert!(!needs_tag_main, "DeleteObject statement itself does not require ExistingObjectTag");
let needs_tag_current = bucket_policy_needs_existing_object_tag_from_hint(&hint, &args_delete).await;
let needs_tag_version = bucket_policy_needs_existing_object_tag_from_hint(&hint, &args_delete_version).await;
assert!(!needs_tag_current, "DeleteObject statement itself does not require ExistingObjectTag");
assert!(
needs_tag_secondary,
needs_tag_version,
"DeleteObjectVersion statement requires ExistingObjectTag when version delete is evaluated"
);
assert!(
needs_tag_main || needs_tag_secondary,
"combined primary+secondary check must require tag fetch for DeleteObject(versionId)"
needs_tag_version,
"the selected version action must require tag fetch for DeleteObject(versionId)"
);
}
+230 -12
View File
@@ -220,8 +220,9 @@ pub struct SseConfiguration {
/// malformed bucket default pass the `copy_changes_encryption` guard and take
/// the metadata-only shortcut while this layer still encrypts: fresh DEK
/// metadata is committed beside the untouched plaintext blocks and the object
/// becomes unreadable. Reachable only via corrupt or hand-edited bucket
/// metadata — PutBucketEncryption rejects unknown algorithms (backlog#1826).
/// becomes unreadable. PutBucketEncryption refuses unknown algorithms, so this
/// is reachable only through a configuration stored before that check or
/// through hand-edited bucket metadata (backlog#1826).
pub(crate) fn bucket_default_write_sse(sse: &ServerSideEncryptionByDefault) -> ServerSideEncryption {
match sse.sse_algorithm.as_str() {
"AES256" => ServerSideEncryption::from_static(ServerSideEncryption::AES256),
@@ -723,6 +724,11 @@ pub(crate) fn map_get_object_reader_error(err: StorageError) -> ApiError {
let code = match resolution_error.kind() {
EncryptionResolutionErrorKind::InvalidRequest => S3ErrorCode::InvalidRequest,
EncryptionResolutionErrorKind::ServiceUnavailable => S3ErrorCode::ServiceUnavailable,
// Same code the write path returns for this key; `From<ApiError>`
// attaches the 400 that s3s cannot derive for a custom code.
EncryptionResolutionErrorKind::KeyNotFound => S3ErrorCode::Custom(crate::error::KMS_KEY_NOT_FOUND_ERROR_CODE.into()),
EncryptionResolutionErrorKind::AccessDenied => S3ErrorCode::AccessDenied,
EncryptionResolutionErrorKind::NotImplemented => S3ErrorCode::NotImplemented,
// A permanent property of the stored object, not a transient server
// fault: 5xx would invite client retry storms against an object
// this server can never decrypt.
@@ -1512,10 +1518,19 @@ fn normalize_encryption_metadata_case(
Ok(Cow::Owned(normalized))
}
/// Carry the S3-level classification of a decryption failure through the
/// ecstore boundary. Every code produced by `data_plane_kms_error` needs a
/// kind here, otherwise the read path reports it as an internal fault even
/// though the write path already reports the same KMS error to the client.
fn map_encryption_resolution_error(error: ApiError) -> EncryptionResolutionError {
let kind = match error.code {
let kind = match &error.code {
S3ErrorCode::InvalidArgument | S3ErrorCode::InvalidRequest => EncryptionResolutionErrorKind::InvalidRequest,
S3ErrorCode::ServiceUnavailable => EncryptionResolutionErrorKind::ServiceUnavailable,
S3ErrorCode::Custom(code) if &**code == crate::error::KMS_KEY_NOT_FOUND_ERROR_CODE => {
EncryptionResolutionErrorKind::KeyNotFound
}
S3ErrorCode::AccessDenied => EncryptionResolutionErrorKind::AccessDenied,
S3ErrorCode::NotImplemented => EncryptionResolutionErrorKind::NotImplemented,
_ => EncryptionResolutionErrorKind::DecryptionFailed,
};
EncryptionResolutionError::new(kind, error.message)
@@ -1530,6 +1545,19 @@ pub struct ManagedSealedKey {
}
impl EncryptionMaterial {
/// The KMS key id a write response may advertise.
///
/// `kms_key_id` is always set for managed SSE because SSE-S3 also wraps
/// its data key under the service default key, but that key is an
/// internal detail of an `AES256` object: only an `aws:kms` object names
/// a key the caller can act on, and `x-amz-server-side-encryption-aws-kms-key-id`
/// is defined only for that scheme.
pub fn response_kms_key_id(&self) -> Option<SSEKMSKeyId> {
matches!(self.sse_type, SSEType::SseKms)
.then(|| self.kms_key_id.clone())
.flatten()
}
pub fn write_encryption(&self, multipart_part_number: Option<usize>) -> super::WriteEncryption {
match (self.key_kind, multipart_part_number) {
(EncryptionKeyKind::Object, Some(part_number)) => {
@@ -2473,7 +2501,9 @@ pub async fn classify_sse_read_response(request: DecryptionRequest<'_>) -> Resul
server_side_encryption: ServerSideEncryption::from(managed_sse_public_header(sse_type).to_string()),
sse_customer_algorithm: None,
sse_customer_key_md5: None,
ssekms_key_id: Some(SSEKMSKeyId::from(kms_key_id)),
// The key id was needed above to authorize the read, but an AES256
// object's wrapping key is internal: only aws:kms objects advertise it.
ssekms_key_id: matches!(sse_type, SSEType::SseKms).then(|| SSEKMSKeyId::from(kms_key_id)),
}))
}
@@ -2751,9 +2781,21 @@ async fn apply_managed_encryption_material_inner(
}
(SSEType::SseKms, Some(kms_key_id)) => kms_key_id,
(SSEType::SseKms, None) => {
return Err(ApiError::from(StorageError::other(
"No KMS key available for managed server-side encryption (required for SSE-KMS)",
)));
// Neither the request nor the bucket default named a key and no
// service default filled in. Without a service this is the same
// outage/misconfiguration the provider check below reports, so it
// must carry the same 503/400 split rather than an untyped
// internal error; with a running service that has no default key
// the caller simply has to name one.
if runtime_sources::current_encryption_service().await.is_none() {
return Err(sse_kms_unavailable_error(kms_configured_but_unavailable().await));
}
return Err(ApiError {
code: S3ErrorCode::InvalidRequest,
message: "SSE-KMS requires a KMS key id: the request named none and the KMS service has no default key"
.to_string(),
source: None,
});
}
_ => unreachable!("managed SSE branch only supports SSE-S3 or SSE-KMS"),
};
@@ -3354,6 +3396,25 @@ fn kms_operation_error(error: rustfs_kms::KmsError) -> ApiError {
api_error
}
/// Classification for a failed data-key unwrap.
///
/// An AEAD failure stays `500` (the envelope is an integrity fault, not a
/// request a retry or a different header can fix), but the generic internal
/// error text hides the one diagnosis an operator needs: the configured
/// backend holds different key material under this key id than the one that
/// wrapped the object, typically after re-creating a key of the same name or
/// switching backends.
fn kms_unwrap_error(error: rustfs_kms::KmsError) -> ApiError {
let unwrap_rejected = matches!(error, rustfs_kms::KmsError::CryptographicError { .. });
let mut api_error = kms_operation_error(error);
if unwrap_rejected && api_error.code == S3ErrorCode::InternalError {
api_error.message = "The object's data key envelope could not be unwrapped by the configured KMS backend: the \
key material under this key id differs from the one that wrapped it, or the envelope is damaged"
.to_string();
}
api_error
}
impl KmsSseDekProvider {
/// Create a new KMS-backed provider
pub async fn new() -> Result<Self, ApiError> {
@@ -3438,7 +3499,7 @@ impl SseDekProvider for KmsSseDekProvider {
let data_key = service
.decrypt_data_key(encrypted_dek, context)
.await
.map_err(kms_operation_error)?;
.map_err(kms_unwrap_error)?;
Ok(data_key.plaintext_key)
}
@@ -3457,7 +3518,7 @@ impl SseDekProvider for KmsSseDekProvider {
let data_key = service
.decrypt_legacy_data_key(encrypted_dek)
.await
.map_err(kms_operation_error)?;
.map_err(kms_unwrap_error)?;
Ok(data_key.plaintext_key)
}
@@ -4409,6 +4470,56 @@ mod tests {
assert_eq!(super::kms_data_plane_error_class(&missing), "key_not_found");
}
/// The read path squeezes the S3 classification through ecstore's
/// resolution-error kinds; every KMS class the write path reports to the
/// client must survive that hop instead of collapsing onto `DecryptionFailed`
/// (which the S3 layer reports as `500`).
#[test]
fn encryption_resolution_kinds_preserve_kms_read_classification() {
let cases = [
(
rustfs_kms::KmsError::key_not_found("no-such-key"),
EncryptionResolutionErrorKind::KeyNotFound,
),
(rustfs_kms::KmsError::access_denied("policy"), EncryptionResolutionErrorKind::AccessDenied),
(
rustfs_kms::KmsError::unsupported_capability("local", "decrypt_legacy"),
EncryptionResolutionErrorKind::NotImplemented,
),
(
rustfs_kms::KmsError::backend_error("connection refused"),
EncryptionResolutionErrorKind::ServiceUnavailable,
),
(
rustfs_kms::KmsError::invalid_operation("key is disabled"),
EncryptionResolutionErrorKind::InvalidRequest,
),
(
rustfs_kms::KmsError::cryptographic_error("decrypt", "authentication failed"),
EncryptionResolutionErrorKind::DecryptionFailed,
),
];
for (error, expected) in cases {
let description = error.to_string();
let resolution = super::map_encryption_resolution_error(kms_operation_error(error));
assert_eq!(resolution.kind(), expected, "{description}");
}
}
/// An unwrap the backend rejects stays an internal error, but says why in
/// words an operator can act on rather than the generic 500 text.
#[test]
fn kms_unwrap_error_keeps_500_but_names_the_envelope_mismatch() {
let rejected = super::kms_unwrap_error(rustfs_kms::KmsError::cryptographic_error("decrypt", "authentication failed"));
assert_eq!(rejected.code, S3ErrorCode::InternalError);
assert!(rejected.message.contains("could not be unwrapped"), "message was {}", rejected.message);
assert_eq!(super::kms_data_plane_error_class(&rejected), "cryptographic");
let missing = super::kms_unwrap_error(rustfs_kms::KmsError::key_not_found("no-such-key"));
assert_eq!(missing.code, S3ErrorCode::Custom(crate::error::KMS_KEY_NOT_FOUND_ERROR_CODE.into()));
assert_eq!(missing.message, "KMS key not found: no-such-key");
}
#[test]
fn sse_kms_never_falls_back_to_the_local_sse_s3_provider() {
let unconfigured = super::sse_kms_unavailable_error(false);
@@ -4496,6 +4607,56 @@ mod tests {
reset_sse_dek_provider();
}
/// A bare `aws:kms` request (no key id, no bucket default) on a node with
/// no KMS has no key to resolve. It must get the same configuration
/// refusal as the keyed form, whether or not the SSE-S3 master key is
/// set, rather than an untyped internal error (backlog#2368 B4).
#[tokio::test]
async fn sse_kms_write_without_a_key_id_is_refused_like_the_keyed_form() {
let _guard = lock_sse_test_state().await;
for master_key in [None, Some(BASE64_STANDARD.encode_to_string([9u8; 32]))] {
reset_sse_dek_provider();
async_with_vars(
[
("__RUSTFS_SSE_SIMPLE_CMK", None::<String>),
("RUSTFS_SSE_S3_MASTER_KEY", master_key.clone()),
],
async {
// Entered directly: `sse_encryption` consults the bucket
// default first, which needs a bucket metadata store.
let error = apply_managed_encryption_material(
"finance",
"ledger.csv",
ServerSideEncryption::from_static(ServerSideEncryption::AWS_KMS),
None,
None,
128,
None,
)
.await
.expect_err("SSE-KMS without a key id must be refused when no KMS is running");
assert_eq!(
error.code,
S3ErrorCode::InvalidRequest,
"master_key={master_key:?}: message was {}",
error.message
);
assert!(error.message.contains("SSE-KMS requires"), "message was {}", error.message);
assert!(
!error.message.contains("RUSTFS_SSE_S3_MASTER_KEY"),
"an SSE-KMS refusal must not name the SSE-S3 master key: {}",
error.message
);
},
)
.await;
}
reset_sse_dek_provider();
}
/// The SSE-S3 local fallback itself is unchanged: refusing SSE-KMS must not
/// take the documented no-KMS deployment down with it.
#[tokio::test]
@@ -5834,7 +5995,7 @@ mod tests {
#[tokio::test]
async fn test_sse_encryption_persists_aws_kms_header_for_kms_objects() {
let metadata = encryption_material_to_metadata(&EncryptionMaterial {
let material = EncryptionMaterial {
sse_type: SSEType::SseKms,
server_side_encryption: ServerSideEncryption::from_static(ServerSideEncryption::AWS_KMS),
kms_key_id: Some("test-key".to_string()),
@@ -5847,8 +6008,10 @@ mod tests {
key_kind: EncryptionKeyKind::Direct,
managed_kms_context: None,
managed_sealed_key: None,
})
.expect("managed SSE metadata should serialize");
};
// Only an aws:kms object names its key in write responses.
assert_eq!(material.response_kms_key_id().as_deref(), Some("test-key"));
let metadata = encryption_material_to_metadata(&material).expect("managed SSE metadata should serialize");
assert_eq!(metadata.get("x-amz-server-side-encryption").map(String::as_str), Some("aws:kms"));
assert_eq!(
@@ -6025,6 +6188,9 @@ mod tests {
let metadata = encryption_material_to_metadata(&material).expect("managed SSE-S3 metadata should serialize");
assert_eq!(material.kms_key_id.as_deref(), Some("default"));
// The wrapping key stays internal: no write response may
// advertise it for an AES256 object.
assert_eq!(material.response_kms_key_id(), None);
assert_eq!(metadata.get("x-amz-server-side-encryption").map(String::as_str), Some("AES256"));
assert!(!metadata.contains_key("x-amz-server-side-encryption-aws-kms-key-id"));
assert_eq!(metadata.get(INTERNAL_ENCRYPTION_KEY_ID_HEADER).map(String::as_str), Some("default"));
@@ -7528,6 +7694,32 @@ mod tests {
assert_eq!(err.message, "KMS unavailable");
}
/// The reader wraps the resolution error in an io error exactly like
/// `readers.rs` does; a missing key must come out as the same `400
/// KMS.NotFoundException` the write path returns, not `500`.
#[test]
fn test_map_get_object_reader_error_reports_missing_kms_key_as_client_error() {
let resolution_error =
super::EncryptionResolutionError::new(EncryptionResolutionErrorKind::KeyNotFound, "KMS key not found: finance-key");
let err = map_get_object_reader_error(StorageError::other(resolution_error));
assert_eq!(err.code, S3ErrorCode::Custom(crate::error::KMS_KEY_NOT_FOUND_ERROR_CODE.into()));
assert_eq!(err.message, "KMS key not found: finance-key");
let s3_error = s3s::S3Error::from(err);
assert_eq!(s3_error.status_code(), Some(http::StatusCode::BAD_REQUEST));
let denied = map_get_object_reader_error(StorageError::other(super::EncryptionResolutionError::new(
EncryptionResolutionErrorKind::AccessDenied,
"Access Denied",
)));
assert_eq!(denied.code, S3ErrorCode::AccessDenied);
let unsupported = map_get_object_reader_error(StorageError::other(super::EncryptionResolutionError::new(
EncryptionResolutionErrorKind::NotImplemented,
"backend cannot unwrap legacy envelopes",
)));
assert_eq!(unsupported.code, S3ErrorCode::NotImplemented);
}
#[test]
fn test_map_get_object_reader_error_maps_part_missing_to_slow_down_read() {
let err = map_get_object_reader_error(StorageError::PartMissingOrCorrupt);
@@ -8083,6 +8275,32 @@ mod tests {
// Read-side response classification (single-decrypt GET path)
// ========================================================================
/// An AES256 object is read under the same key-id resolution as aws:kms
/// (authorization needs it), but the response must not advertise that
/// internal wrapping key.
#[tokio::test]
async fn classification_withholds_the_wrapping_key_for_sse_s3_reads() {
let metadata = HashMap::from([
("x-amz-server-side-encryption".to_string(), ServerSideEncryption::AES256.to_string()),
(INTERNAL_ENCRYPTION_KEY_ID_HEADER.to_string(), "service-default".to_string()),
(INTERNAL_ENCRYPTION_KEY_HEADER.to_string(), BASE64_STANDARD.encode_to_string([1u8; 16])),
(INTERNAL_ENCRYPTION_IV_HEADER.to_string(), BASE64_STANDARD.encode_to_string([2u8; 12])),
]);
let headers = super::classify_sse_read_response(DecryptionRequest {
bucket: "finance",
key: "ledger.csv",
metadata: &metadata,
sse_customer_key: None,
sse_customer_key_md5: None,
principal: None,
})
.await
.expect("sse-s3 classification should succeed")
.expect("managed metadata should classify");
assert_eq!(headers.server_side_encryption.as_str(), ServerSideEncryption::AES256);
assert_eq!(headers.ssekms_key_id, None);
}
#[tokio::test]
async fn classification_reproduces_managed_read_headers_and_audit_without_a_kms_unwrap() {
use rustfs_kms::types::{CreateKeyRequest, KeyUsage};
+11 -7
View File
@@ -119,9 +119,9 @@ pub(crate) use super::sse::{
pub(crate) mod access_consumer {
pub(crate) use super::super::access::{
PostObjectRequestMarker, ReqInfo, apply_bucket_generation_guard, apply_copy_source_bucket_generation_guard,
authorize_internal_object_request, authorize_request, bucket_config_mutation_incarnation, has_bypass_governance_header,
load_bucket_generation_from_store, log_list_buckets_iam_implicit_deny, odm_read_generation,
prepare_list_buckets_iam_authorization, prepare_odm_read_generation, recursive_force_delete_is_authorized,
authorize_internal_object_request, authorize_request, bucket_config_mutation_incarnation, delete_object_authorize_action,
has_bypass_governance_header, load_bucket_generation_from_store, log_list_buckets_iam_implicit_deny, odm_read_generation,
prepare_list_buckets_iam_authorization, prepare_odm_read_generation, recursive_force_delete_has_authenticated_caller,
replication_request_authorized, req_info_mut, req_info_ref,
};
}
@@ -1185,17 +1185,21 @@ pub(crate) fn get_global_transition_state() -> Arc<TransitionState> {
ecstore_bucket::lifecycle::bucket_lifecycle_ops::get_global_transition_state()
}
pub(crate) async fn try_migrate_bucket_metadata(store: Arc<ECStore>) {
ecstore_bucket::migration::try_migrate_bucket_metadata(store).await;
pub(crate) async fn try_migrate_bucket_metadata(store: Arc<ECStore>) -> std::io::Result<()> {
ecstore_bucket::migration::try_migrate_bucket_metadata(store)
.await
.map_err(ecstore_bucket::migration::migration_startup_error)
}
pub(crate) async fn try_migrate_iam_config(store: Arc<ECStore>) {
pub(crate) async fn try_migrate_iam_config(store: Arc<ECStore>) -> std::io::Result<()> {
// MinIO encrypts IAM identity/service-account files at rest with a key derived
// from the root credentials. Inject the IAM crate's decryption so those blobs
// are decrypted before normalization instead of being skipped as "incompatible".
let decrypt_fn: ecstore_bucket::migration::LegacyBlobDecryptFn =
Arc::new(|data: &[u8]| rustfs_iam::try_decrypt_iam_blob(data));
ecstore_bucket::migration::try_migrate_iam_config(store, Some(decrypt_fn)).await;
ecstore_bucket::migration::try_migrate_iam_config(store, Some(decrypt_fn))
.await
.map_err(ecstore_bucket::migration::migration_startup_error)
}
pub(crate) fn init_ecstore_config() {
@@ -69,6 +69,11 @@ fn warehouse_object_prefix_from_location(
"table warehouse location must be inside the table bucket".to_string(),
));
}
if is_reserved_table_object_key(object_prefix.strip_suffix('/').unwrap_or(object_prefix)) {
return Err(TableCatalogStoreError::Invalid(
"table warehouse location overlaps the reserved table catalog prefix".to_string(),
));
}
normalize_warehouse_object_prefix(object_prefix, max_prefix_depth)
}
+24
View File
@@ -5937,6 +5937,30 @@ async fn object_table_catalog_store_rejects_invalid_table_warehouse_location() {
));
}
#[test]
fn warehouse_locations_reject_the_reserved_catalog_prefix() {
for location in [
"s3://analytics/.rustfs-table",
"s3://analytics/.rustfs-table/",
"s3://analytics/.rustfs-table/warehouses/default",
] {
let table_error = validate_table_warehouse_location("analytics", location).unwrap_err();
assert!(matches!(
table_error,
TableCatalogStoreError::Invalid(message) if message.contains("reserved table catalog prefix")
));
let view_error = validate_view_warehouse_location("analytics", location).unwrap_err();
assert!(matches!(
view_error,
TableCatalogStoreError::Invalid(message) if message.contains("reserved table catalog prefix")
));
}
assert!(validate_table_warehouse_location("analytics", "s3://analytics/.rustfs-table-other/table-id").is_ok());
assert!(validate_table_warehouse_location("analytics", "s3://analytics/user/.rustfs-table/table-id").is_ok());
}
#[tokio::test]
async fn object_table_catalog_store_rejects_deep_table_warehouse_location() {
let backend = TestCatalogObjectBackend::default();
@@ -0,0 +1,326 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
#![recursion_limit = "256"]
use reqwest::StatusCode;
use rustfs::embedded::{RustFSServerBuilder, find_available_port};
use rustfs_ecstore::api::config::com::{delete_config, read_config};
use std::fs;
use std::path::{Path, PathBuf};
use std::process::Stdio;
use std::time::Duration;
use tokio::process::Command;
mod common;
const TEST_NAME: &str = "native_migration_failure_blocks_server_startup_and_repair_preserves_records";
const STAGE_ENV: &str = "RUSTFS_NATIVE_MIGRATION_TEST_STAGE";
const ROOT_ENV: &str = "RUSTFS_NATIVE_MIGRATION_TEST_ROOT";
const ADDRESS_ENV: &str = "RUSTFS_NATIVE_MIGRATION_TEST_ADDRESS";
const FAILURE_ENV: &str = "RUSTFS_NATIVE_MIGRATION_TEST_FAILURE";
const STOP_ENV: &str = "RUSTFS_NATIVE_MIGRATION_TEST_STOP";
const ACCESS_KEY: &str = "native-migration-root";
const SECRET_KEY: &str = "native-migration-root-secret";
const LEGACY_BUCKET: &str = ".minio.sys";
const TARGET_BUCKET: &str = ".rustfs.sys";
const BUCKET_METADATA: &str = "buckets/interop/.metadata.bin";
const IAM_RECORD: &str = "config/iam/groups/migration-group/members.json";
const IAM_FORMAT: &str = "config/iam/format.json";
const EXISTING_FORMAT: &[u8] = br#"{"version":1}"#;
const STARTUP_TIMEOUT: Duration = Duration::from_secs(60);
#[derive(Clone, Copy, Debug)]
enum StartupMode {
Server,
Embedded,
}
fn volumes(root: &Path) -> Vec<PathBuf> {
(1..=4).map(|index| root.join(format!("disk{index}"))).collect()
}
fn minio_bucket_metadata() -> Vec<u8> {
let hex: String = include_str!("../../crates/ecstore/tests/fixtures/minio/bucket_metadata.blob.hex")
.chars()
.filter(|ch| !ch.is_whitespace())
.collect();
hex.as_bytes()
.as_chunks::<2>()
.0
.iter()
.map(|pair| {
u8::from_str_radix(std::str::from_utf8(pair).expect("fixture hex is UTF-8"), 16).expect("valid MinIO fixture hex")
})
.collect()
}
async fn prepare_or_verify_fixture(root: &Path, seed: bool) {
let env = rustfs_test_utils::TestECStoreEnv::builder()
.base_dir(root)
.disk_count(4)
.build()
.await;
if seed {
env.make_bucket("interop", false).await;
env.make_bucket(LEGACY_BUCKET, false).await;
env.put_object_bytes(LEGACY_BUCKET, BUCKET_METADATA, minio_bucket_metadata())
.await;
env.put_object_bytes(
LEGACY_BUCKET,
IAM_RECORD,
br#"{"version":1,"status":"enabled","members":[],"updatedAt":"2026-09-10T00:00:00Z"}"#.to_vec(),
)
.await;
env.put_object_bytes(TARGET_BUCKET, IAM_FORMAT, EXISTING_FORMAT.to_vec())
.await;
// A completed record must be skipped before reading even a broken old copy.
env.put_object_bytes(LEGACY_BUCKET, IAM_FORMAT, b"do not overwrite the existing target".to_vec())
.await;
delete_config(env.ecstore.clone(), BUCKET_METADATA)
.await
.expect("leave bucket metadata pending migration");
} else {
assert_eq!(
read_config(env.ecstore.clone(), BUCKET_METADATA)
.await
.expect("migrated bucket metadata"),
minio_bucket_metadata(),
"migration must preserve the MinIO bucket settings"
);
let group: serde_json::Value = serde_json::from_slice(
&read_config(env.ecstore.clone(), IAM_RECORD)
.await
.expect("migrated IAM group"),
)
.expect("valid migrated IAM JSON");
assert_eq!(group["status"], "enabled");
assert_eq!(group["members"], serde_json::json!([]));
}
assert_eq!(
read_config(env.ecstore.clone(), IAM_FORMAT)
.await
.expect("existing IAM format"),
EXISTING_FORMAT,
"retry must not overwrite records already migrated"
);
}
async fn run_embedded_child(root: &Path) {
let address = std::env::var(ADDRESS_ENV).expect("embedded child address");
let result = RustFSServerBuilder::new()
.address(address)
.access_key(ACCESS_KEY)
.secret_key(SECRET_KEY)
.volumes(volumes(root).iter().map(|path| path.to_string_lossy().into_owned()).collect())
.build()
.await;
match result {
Ok(server) => {
let stop = PathBuf::from(std::env::var_os(STOP_ENV).expect("embedded stop path"));
while !stop.exists() {
tokio::time::sleep(Duration::from_millis(25)).await;
}
server.shutdown().await;
}
Err(error) => {
fs::write(std::env::var_os(FAILURE_ENV).expect("embedded failure path"), error.to_string())
.expect("record the actual embedded startup error");
}
}
}
fn child_command(root: &Path, stage: &str, log: &Path) -> Command {
let mut command = Command::new(std::env::current_exe().expect("integration test executable"));
command
.args(["--exact", TEST_NAME, "--nocapture"])
.env(STAGE_ENV, stage)
.env(ROOT_ENV, root);
configure_process(&mut command, log);
command
}
fn configure_process(command: &mut Command, log: &Path) {
let output = fs::File::create(log).expect("create isolated process log");
command
// These disposable erasure volumes intentionally share the test runner's disk.
.env("RUSTFS_UNSAFE_BYPASS_DISK_CHECK", "true")
.env("RUSTFS_CONSOLE_ENABLE", "false")
.env("NO_PROXY", "localhost,127.0.0.1,::1")
.env("no_proxy", "localhost,127.0.0.1,::1")
.env("RUST_LOG", "warn")
.stdin(Stdio::null())
.stdout(Stdio::from(output.try_clone().expect("clone process log")))
.stderr(Stdio::from(output))
.kill_on_drop(true);
}
async fn fixture_process(root: &Path, stage: &str) {
let log = root.join(format!("{stage}.log"));
let status = tokio::time::timeout(STARTUP_TIMEOUT, child_command(root, stage, &log).status())
.await
.expect("fixture process must finish")
.expect("run fixture process");
assert!(status.success(), "{stage} failed: {}", fs::read_to_string(log).expect("fixture log"));
}
async fn check_startup(root: &Path, mode: StartupMode, failure_record: Option<&str>, label: &str) {
let ready = failure_record.is_none();
let address = format!("127.0.0.1:{}", find_available_port().expect("free startup probe port"));
let log = root.join(format!("{label}.log"));
let failure = root.join(format!("{label}.failure"));
let stop = root.join(format!("{label}.stop"));
let mut command = match mode {
StartupMode::Server => {
let mut command = Command::new(env!("CARGO_BIN_EXE_rustfs"));
command
.args(["--address", &address, "--access-key", ACCESS_KEY, "--secret-key", SECRET_KEY])
.args(volumes(root));
configure_process(&mut command, &log);
command
}
StartupMode::Embedded => {
let mut command = child_command(root, "embedded", &log);
command
.env(ADDRESS_ENV, &address)
.env(FAILURE_ENV, &failure)
.env(STOP_ENV, &stop);
command
}
};
let mut child = command.spawn().expect("start isolated server process");
let http = reqwest::Client::builder()
.no_proxy()
.timeout(Duration::from_millis(500))
.build()
.expect("local readiness client");
let result = tokio::time::timeout(STARTUP_TIMEOUT, async {
loop {
if let Ok(response) = http.get(format!("http://{address}/health/ready")).send().await
&& response.status() == StatusCode::OK
{
assert!(ready, "{mode:?} published Ready after a migration I/O failure");
return;
}
if let Some(status) = child.try_wait().expect("poll server process") {
let details = fs::read_to_string(&log).expect("startup log");
assert!(!ready, "{mode:?} exited before Ready ({status}): {details}");
let record = failure_record.expect("failed startup has an obstructed record");
match mode {
StartupMode::Server => {
assert_eq!(status.code(), Some(1), "startup must fail: {details}");
assert_migration_io_error(&details, record);
}
StartupMode::Embedded => {
assert!(status.success(), "embedded test process failed unexpectedly: {details}");
let error = fs::read_to_string(&failure).expect("embedded startup returned an error");
assert_migration_io_error(&error, record);
}
}
return;
}
tokio::time::sleep(Duration::from_millis(25)).await;
}
})
.await;
assert!(
result.is_ok(),
"{mode:?} did not reach the expected startup outcome: {}",
fs::read_to_string(&log).expect("startup diagnostics")
);
if ready {
match mode {
StartupMode::Embedded => {
fs::write(stop, b"stop").expect("request embedded shutdown");
assert!(
tokio::time::timeout(STARTUP_TIMEOUT, child.wait())
.await
.expect("embedded shutdown completes")
.expect("wait for embedded shutdown")
.success()
);
}
StartupMode::Server => child.kill().await.expect("stop the isolated server"),
}
}
}
fn assert_migration_io_error(error: &str, record: &str) {
let lower = error.to_ascii_lowercase();
assert!(
(lower.contains("access denied")
|| lower.contains("access is denied")
|| lower.contains("not a directory")
|| lower.contains("not regular"))
&& error.contains(&format!("{TARGET_BUCKET}/{record}")),
"startup must fail because of the obstructed metadata record, not an unrelated initialization error: {error}"
);
}
async fn run_startup_cases(mode: StartupMode) {
let ordinary = tempfile::TempDir::with_prefix("rustfs-no-legacy-").expect("ordinary store");
for volume in volumes(ordinary.path()) {
fs::create_dir_all(volume).expect("ordinary volume");
}
check_startup(ordinary.path(), mode, None, "ordinary").await;
let control = tempfile::TempDir::with_prefix("rustfs-migration-control-").expect("control fixture");
fixture_process(control.path(), "seed").await;
check_startup(control.path(), mode, None, "control").await;
fixture_process(control.path(), "verify").await;
for record in [BUCKET_METADATA, IAM_RECORD] {
let target = tempfile::TempDir::with_prefix("rustfs-migration-failure-").expect("disposable migration target");
fixture_process(target.path(), "seed").await;
let blockers: Vec<_> = volumes(target.path())
.iter()
.map(|volume| volume.join(TARGET_BUCKET).join(record))
.collect();
for blocker in &blockers {
fs::create_dir_all(blocker.parent().expect("record parent")).expect("create target parent");
assert!(!blocker.exists(), "the record must still need migration");
// A non-directory target causes real filesystem I/O errors even when tests run as root.
fs::write(blocker, b"blocked migration target").expect("block only the destination record");
}
check_startup(target.path(), mode, Some(record), "blocked").await;
for blocker in blockers {
fs::remove_file(blocker).expect("repair the same partially migrated target");
}
check_startup(target.path(), mode, None, "repaired").await;
fixture_process(target.path(), "verify").await;
}
}
#[test]
fn native_migration_failure_blocks_server_startup_and_repair_preserves_records() {
// Cold processes keep failed initialization and cached metadata out of subsequent restart attempts.
common::run_embedded_test(|| async {
match std::env::var(STAGE_ENV).ok().as_deref() {
Some("seed") => {
prepare_or_verify_fixture(&PathBuf::from(std::env::var_os(ROOT_ENV).expect("fixture root")), true).await
}
Some("verify") => {
prepare_or_verify_fixture(&PathBuf::from(std::env::var_os(ROOT_ENV).expect("fixture root")), false).await
}
Some("embedded") => run_embedded_child(&PathBuf::from(std::env::var_os(ROOT_ENV).expect("fixture root"))).await,
None => run_startup_cases(StartupMode::Server).await,
Some(stage) => panic!("unknown native migration test stage: {stage}"),
}
});
}
#[test]
fn native_migration_failure_blocks_embedded_startup_and_repair_preserves_records() {
common::run_embedded_test(|| run_startup_cases(StartupMode::Embedded));
}
+253
View File
@@ -0,0 +1,253 @@
#!/usr/bin/env python3
"""Build an identified E2E server and verify it around one test invocation."""
import argparse
from contextlib import contextmanager
import hashlib
import json
import os
from pathlib import Path
import stat
import signal
import subprocess
import sys
import tempfile
ROOT = Path(__file__).resolve().parent.parent
RECEIPT_ENV = "RUSTFS_E2E_BINARY_RECEIPT"
def feature_set(value):
return sorted(set(part.strip() for part in value.split(",") if part.strip()))
def file_hash(path):
digest = hashlib.sha256()
with path.open("rb") as source:
for chunk in iter(lambda: source.read(1024 * 1024), b""):
digest.update(chunk)
return digest.hexdigest()
def source_identity():
head = subprocess.check_output(["git", "rev-parse", "HEAD"], cwd=ROOT, text=True).strip()
tracked = subprocess.check_output(["git", "ls-files", "--cached", "--others", "--exclude-standard", "-z"], cwd=ROOT)
paths = set(tracked.decode("utf-8").rstrip("\0").split("\0")) - {""}
# RustEmbed consumes ignored console assets as well as tracked Rust sources.
static_dir = ROOT / "rustfs/static"
if static_dir.is_symlink():
raise ValueError("The embedded static directory must not be a symlink")
if static_dir.is_dir():
for path in static_dir.rglob("*"):
if path.is_symlink() and path.is_dir():
raise ValueError(f"Unsupported embedded directory symlink: {path}")
if not path.is_dir():
paths.add(str(path.relative_to(ROOT)))
elif static_dir.exists():
paths.add("rustfs/static")
digest = hashlib.sha256()
digest.update(b"static-present\0" if static_dir.is_dir() else b"static-absent\0")
for name in sorted(paths):
path = ROOT / name
digest.update(name.encode("utf-8") + b"\0")
try:
metadata = path.lstat()
except FileNotFoundError:
digest.update(b"deleted\0")
continue
if stat.S_ISLNK(metadata.st_mode):
digest.update(b"symlink\0" + os.fsencode(os.readlink(path)) + b"\0")
if path.is_dir():
target = path.resolve()
if ROOT not in target.parents:
raise ValueError(f"Directory link escapes the source inventory: {name}")
# Directory aliases such as .claude/skills share already-hashed inputs.
for child in target.rglob("*"):
if child.is_dir() and not child.is_symlink():
continue
if child.is_dir() or str(child.relative_to(ROOT)) not in paths:
raise ValueError(f"Directory link contains an unrecorded input: {child}")
digest.update(b"directory\0" + str(target.relative_to(ROOT)).encode("utf-8") + b"\0")
continue
elif not stat.S_ISREG(metadata.st_mode):
raise ValueError(f"Unsupported build input: {name}")
digest.update(str(metadata.st_mode & 0o111).encode() + b"\0")
digest.update(file_hash(path).encode() + b"\0")
return {"head": head, "sha256": digest.hexdigest()}
def sidecar_path(binary):
return binary.with_name(binary.name + ".e2e.json")
def validate_target_directory(target_dir):
if target_dir == ROOT or target_dir in ROOT.parents:
raise ValueError("CARGO_TARGET_DIR must not contain the source workspace")
if ROOT in target_dir.parents:
ignored = subprocess.run(["git", "check-ignore", "--quiet", "--no-index", str(target_dir.relative_to(ROOT))], cwd=ROOT)
if ignored.returncode != 0:
raise ValueError("An in-workspace CARGO_TARGET_DIR must be Git-ignored; use target/ or an external directory")
@contextmanager
def exclusive_binary(binary):
marker = binary.with_name(binary.name + ".e2e.lock")
try:
descriptor = os.open(marker, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
except FileExistsError as error:
raise ValueError(f"Another E2E build/run owns {marker}; do not share a target directory between concurrent runs") from error
try:
identity = os.fstat(descriptor)
with os.fdopen(descriptor, "w") as lock:
lock.write(f"pid={os.getpid()}\n")
yield
finally:
current = marker.stat()
if (current.st_dev, current.st_ino) != (identity.st_dev, identity.st_ino):
raise ValueError("The E2E ownership marker changed during the command")
marker.unlink()
def terminate_command(process):
if process.poll() is not None:
return
try:
os.killpg(process.pid, signal.SIGTERM)
except ProcessLookupError:
return
try:
process.wait(timeout=5)
except subprocess.TimeoutExpired:
os.killpg(process.pid, signal.SIGKILL)
process.wait()
def build(binary, target_dir, profile, requested, all_bins):
sidecar = sidecar_path(binary)
sidecar.unlink(missing_ok=True)
before = source_identity()
command = ["cargo", "build", "--locked", "-p", "rustfs", "--target-dir", str(target_dir), "--message-format=json-render-diagnostics"]
command.extend(["--bins"] if all_bins else ["--bin", "rustfs"])
if requested:
command.extend(["--features", ",".join(requested)])
if profile == "release":
command.append("--release")
artifact = None
with subprocess.Popen(command, cwd=ROOT, stdout=subprocess.PIPE, text=True, start_new_session=True) as process:
try:
for line in process.stdout:
message = json.loads(line)
if message.get("reason") == "compiler-message":
print(message["message"].get("rendered", ""), end="", file=sys.stderr)
if message.get("reason") == "compiler-artifact" and message.get("target", {}).get("name") == "rustfs" and "bin" in message.get("target", {}).get("kind", []):
artifact = message
if process.wait() != 0:
raise ValueError("RustFS build failed; no E2E identity was recorded")
except BaseException:
terminate_command(process)
raise
if not artifact or Path(artifact.get("executable", "")).resolve() != binary:
raise ValueError("Cargo did not produce the requested RustFS executable")
if source_identity() != before:
raise ValueError("Build inputs changed during compilation; finish preparing embedded assets and rebuild in an isolated worktree")
record = {
"schema": 1,
"source": before,
"requested_features": requested,
"features": sorted(artifact["features"]),
"profile": profile,
"rustc": subprocess.check_output(["rustc", "-Vv"], text=True),
"binary_sha256": file_hash(binary),
}
sidecar.write_text(json.dumps(record, sort_keys=True) + "\n")
print(f"Built E2E server: {binary}\nIdentity: {sidecar}", file=sys.stderr)
def verify(binary, profile, requested):
record = json.loads(sidecar_path(binary).read_text())
if not isinstance(record, dict) or set(record) != {"schema", "source", "requested_features", "features", "profile", "rustc", "binary_sha256"} or type(record["schema"]) is not int or record["schema"] != 1:
raise ValueError("Missing or unsupported E2E binary identity; run the build command")
if not isinstance(record["rustc"], str) or not record["rustc"].strip():
raise ValueError("Missing E2E build toolchain identity")
if record["requested_features"] != requested or record["profile"] != profile:
raise ValueError("E2E binary build features/profile differ from this test invocation")
if not isinstance(record["features"], list) or not all(isinstance(item, str) for item in record["features"]) or not set(requested) <= set(record["features"]):
raise ValueError("Invalid resolved E2E binary features")
if record["source"] != source_identity():
raise ValueError("E2E binary was built from different inputs; rebuild before testing")
if record["binary_sha256"] != file_hash(binary):
raise ValueError("E2E binary content differs from its build identity")
return record
def run(binary, profile, requested, command):
if not command:
raise ValueError("run requires a test command after --")
override = os.environ.get("CARGO_BIN_EXE_rustfs")
if override and Path(override).resolve() != binary:
raise ValueError("CARGO_BIN_EXE_rustfs selects a different server; use --binary explicitly")
record = verify(binary, profile, requested)
metadata = binary.stat()
with tempfile.TemporaryDirectory(prefix="rustfs-e2e-receipt-") as directory:
receipt = Path(directory) / "receipt.json"
receipt.write_text(json.dumps({
"schema": 1,
"workspace": str(ROOT),
"binary": str(binary),
"size": metadata.st_size,
"modified_ns": metadata.st_mtime_ns,
"features": record["features"],
}))
env = dict(os.environ, CARGO_BIN_EXE_rustfs=str(binary), RUSTFS_BUILD_FEATURES=",".join(record["features"]))
env[RECEIPT_ENV] = str(receipt)
with subprocess.Popen(command, cwd=ROOT, env=env, start_new_session=True) as process:
try:
status = process.wait()
except (KeyboardInterrupt, SystemExit):
terminate_command(process)
raise
try:
if verify(binary, profile, requested) != record:
raise ValueError("E2E build identity changed during testing")
except (OSError, ValueError, subprocess.SubprocessError) as error:
print(f"E2E validation invalidated: {error}", file=sys.stderr)
return status if status else 1
return status
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("mode", choices=("build", "run"))
parser.add_argument("--features", default="", help="additional Cargo features; defaults remain enabled")
parser.add_argument("--profile", choices=("debug", "release"), default="debug")
parser.add_argument("--binary", type=Path, help="prebuilt server path for run")
parser.add_argument("--bins", action="store_true", help="build all RustFS binary targets, preserving the CI build matrix")
# Parse the child command separately so its options are never interpreted here.
args = sys.argv[1:]
separator = args.index("--") if "--" in args else len(args)
command = args[separator + 1:] if separator < len(args) else []
options = parser.parse_args(args[:separator])
target_dir = Path(os.environ.get("CARGO_TARGET_DIR", ROOT / "target")).resolve()
binary = (options.binary or target_dir / options.profile / ("rustfs.exe" if os.name == "nt" else "rustfs")).resolve()
try:
validate_target_directory(target_dir)
requested = feature_set(options.features)
if options.mode == "build":
binary.parent.mkdir(parents=True, exist_ok=True)
with exclusive_binary(binary):
if options.mode == "build":
if options.binary or command:
raise ValueError("build does not accept --binary or a child command")
build(binary, target_dir, options.profile, requested, options.bins)
return 0
if options.bins:
raise ValueError("--bins is a build option")
return run(binary, options.profile, requested, command)
except (OSError, ValueError, subprocess.SubprocessError) as error:
print(f"E2E prerequisite failed: {error}", file=sys.stderr)
return 1
if __name__ == "__main__":
signal.signal(signal.SIGTERM, lambda signum, frame: sys.exit(128 + signum))
raise SystemExit(main())
+12 -3
View File
@@ -14,7 +14,12 @@ NC='\033[0m' # No Color
# Default values
PROJECT_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
TARGET_DIR="$PROJECT_ROOT/target/debug"
CARGO_TARGET_DIR="${CARGO_TARGET_DIR:-$PROJECT_ROOT/target}"
if [[ "$CARGO_TARGET_DIR" != /* ]]; then
CARGO_TARGET_DIR="$PROJECT_ROOT/$CARGO_TARGET_DIR"
fi
export CARGO_TARGET_DIR
TARGET_DIR="$CARGO_TARGET_DIR/debug"
RUSTFS_BINARY="$TARGET_DIR/rustfs"
DATA_DIR="$TARGET_DIR/rustfs_test_data"
RUSTFS_PID=""
@@ -94,7 +99,7 @@ build_rustfs() {
print_info "Building RustFS..."
cd "$PROJECT_ROOT"
if ! cargo build --bin rustfs --features "$RUSTFS_BUILD_FEATURES"; then
if ! python3 scripts/e2e_binary.py build --features "$RUSTFS_BUILD_FEATURES"; then
print_error "Failed to build RustFS"
exit 1
fi
@@ -115,6 +120,10 @@ check_dependencies() {
missing_tools+=("curl")
fi
if ! command -v python3 >/dev/null 2>&1; then
missing_tools+=("python3")
fi
if ! command -v cargo >/dev/null 2>&1; then
missing_tools+=("cargo")
fi
@@ -203,7 +212,7 @@ run_tests() {
print_info "Test command: ${test_cmd[*]}"
if "${test_cmd[@]}"; then
if python3 scripts/e2e_binary.py run --features "$RUSTFS_BUILD_FEATURES" -- "${test_cmd[@]}"; then
print_success "All tests passed!"
return 0
else
+13 -12
View File
@@ -243,9 +243,10 @@ run_quick_e2e_steps() {
return
fi
run_step "e2e-reliability-disk-fault" cargo test --package e2e_test reliability_disk_fault_test -- --nocapture
run_step "e2e-heal-erasure-disk-rebuild" cargo test --package e2e_test heal_erasure_disk_rebuild_test -- --nocapture
run_step "e2e-namespace-lock-quorum" cargo test --package e2e_test namespace_lock_quorum_test -- --nocapture
run_step "build-e2e-server" python3 scripts/e2e_binary.py build
run_step "e2e-reliability-disk-fault" python3 scripts/e2e_binary.py run -- cargo test --package e2e_test reliability_disk_fault_test -- --nocapture
run_step "e2e-heal-erasure-disk-rebuild" python3 scripts/e2e_binary.py run -- cargo test --package e2e_test heal_erasure_disk_rebuild_test -- --nocapture
run_step "e2e-namespace-lock-quorum" python3 scripts/e2e_binary.py run -- cargo test --package e2e_test namespace_lock_quorum_test -- --nocapture
}
run_quick_profile() {
@@ -313,15 +314,15 @@ write_blackbox_matrix() {
{
printf 'profile\tscenario\tgate\tcommand\tfixture_env\tstatus\n'
printf 'quick\tsingle-node disk fault read/write\tblack-box\tcargo test --package e2e_test reliability_disk_fault_test -- --nocapture\tnone\t%s\n' "$e2e_status"
printf 'quick\theal degraded erasure disk rebuild\tblack-box\tcargo test --package e2e_test heal_erasure_disk_rebuild_test -- --nocapture\tnone\t%s\n' "$e2e_status"
printf 'quick\tnamespace lock quorum under EC ops\tblack-box\tcargo test --package e2e_test namespace_lock_quorum_test -- --nocapture\tnone\t%s\n' "$e2e_status"
printf 'quick\tsingle-node disk fault read/write\tblack-box\tpython3 scripts/e2e_binary.py run -- cargo test --package e2e_test reliability_disk_fault_test -- --nocapture\tnone\t%s\n' "$e2e_status"
printf 'quick\theal degraded erasure disk rebuild\tblack-box\tpython3 scripts/e2e_binary.py run -- cargo test --package e2e_test heal_erasure_disk_rebuild_test -- --nocapture\tnone\t%s\n' "$e2e_status"
printf 'quick\tnamespace lock quorum under EC ops\tblack-box\tpython3 scripts/e2e_binary.py run -- cargo test --package e2e_test namespace_lock_quorum_test -- --nocapture\tnone\t%s\n' "$e2e_status"
printf 'full\tlegacy bitrot read fixture restore\tfixture\tcargo test -p rustfs-ecstore --test legacy_bitrot_read_test -- --nocapture\tRUSTFS_LEGACY_TEST_ROOT,RUSTFS_LEGACY_TEST_DISK\t%s\n' "$legacy_status"
printf 'full\tMinIO generated encrypted read and negative restore fixture\tfixture\tcargo test -p rustfs --features rio-v2 storage::minio_generated_read_test --lib -- --ignored --nocapture\tRUSTFS_MINIO_FIXTURE_ROOT,RUSTFS_MINIO_STATIC_KMS_KEY_B64\t%s\n' "$minio_status"
printf 'full\tS3 multipart range versioning delete subset\tblack-box\tenv TESTEXPR=\"multipart or range or versioning or delete\" DEPLOY_MODE=build MAXFAIL=0 ./scripts/s3-tests/run.sh\tnone\t%s\n' "$s3_status"
printf 'destructive\tdistributed cluster concurrency\tblack-box\tcargo test --package e2e_test cluster_concurrency_test -- --nocapture\tnone\t%s\n' "$destructive_status"
printf 'destructive\tstale multipart cleanup cluster\tblack-box\tcargo test --package e2e_test stale_multipart_cleanup_cluster_test -- --nocapture\tnone\t%s\n' "$destructive_status"
printf 'destructive\tdelete marker migration semantics\tblack-box\tcargo test --package e2e_test delete_marker_migration_semantics_test -- --nocapture\tnone\t%s\n' "$destructive_status"
printf 'destructive\tdistributed cluster concurrency\tblack-box\tpython3 scripts/e2e_binary.py run -- cargo test --package e2e_test cluster_concurrency_test -- --nocapture\tnone\t%s\n' "$destructive_status"
printf 'destructive\tstale multipart cleanup cluster\tblack-box\tpython3 scripts/e2e_binary.py run -- cargo test --package e2e_test stale_multipart_cleanup_cluster_test -- --nocapture\tnone\t%s\n' "$destructive_status"
printf 'destructive\tdelete marker migration semantics\tblack-box\tpython3 scripts/e2e_binary.py run -- cargo test --package e2e_test delete_marker_migration_semantics_test -- --nocapture\tnone\t%s\n' "$destructive_status"
} >"$BLACKBOX_MATRIX"
}
@@ -566,9 +567,9 @@ run_destructive_profile() {
return
fi
run_step "e2e-cluster-concurrency" cargo test --package e2e_test cluster_concurrency_test -- --nocapture
run_step "e2e-stale-multipart-cleanup-cluster" cargo test --package e2e_test stale_multipart_cleanup_cluster_test -- --nocapture
run_step "e2e-delete-marker-migration-semantics" cargo test --package e2e_test delete_marker_migration_semantics_test -- --nocapture
run_step "e2e-cluster-concurrency" python3 scripts/e2e_binary.py run -- cargo test --package e2e_test cluster_concurrency_test -- --nocapture
run_step "e2e-stale-multipart-cleanup-cluster" python3 scripts/e2e_binary.py run -- cargo test --package e2e_test stale_multipart_cleanup_cluster_test -- --nocapture
run_step "e2e-delete-marker-migration-semantics" python3 scripts/e2e_binary.py run -- cargo test --package e2e_test delete_marker_migration_semantics_test -- --nocapture
}
run_fuzz_profile() {
+3 -7
View File
@@ -353,12 +353,7 @@ DEBUG_DIR="$TARGET_DIR/debug"
if [[ "${RUSTFS_SCANNER_HEAL_SKIP_CLEAN:-0}" != "1" ]]; then
cargo clean -p rustfs
fi
if [[ -n "$BUILD_FEATURES" ]]; then
cargo build --locked -p rustfs --bins --features "$BUILD_FEATURES"
else
cargo build --locked -p rustfs --bins
fi
printf '%s' "$BUILD_FEATURES" >"$DEBUG_DIR/rustfs.features"
"$PYTHON_BIN" "$ROOT/scripts/e2e_binary.py" build --bins --features "$BUILD_FEATURES"
LISTING_TMP="$TMP_DIR/listing.json"
NO_PROXY="${NO_PROXY:-127.0.0.1,localhost}" \
@@ -382,7 +377,8 @@ NO_PROXY="${NO_PROXY:-127.0.0.1,localhost}" \
HTTP_PROXY= \
HTTPS_PROXY= \
RUSTFS_SCANNER_HEAL_RUN_DIR="$RUN_DIR" \
cargo nextest run --profile "$PROFILE" -p e2e_test -E "$TEST_FILTER" --no-tests=fail
"$PYTHON_BIN" "$ROOT/scripts/e2e_binary.py" run --features "$BUILD_FEATURES" -- \
cargo nextest run --profile "$PROFILE" -p e2e_test -E "$TEST_FILTER" --no-tests=fail
STATUS=$?
set -e
@@ -40,7 +40,7 @@ Options:
GitHub repository used to download the release asset (default: rustfs/rustfs)
--test NAME all, mixed-version, or rollback (default: all)
--allow-dirty Allow tracked source changes while collecting evidence
--skip-build Reuse an existing target/debug/rustfs binary
--skip-build Reuse a binary built by scripts/e2e_binary.py build from the current sources
--skip-download Reuse SOURCE_DIR/rustfs instead of downloading the previous release
--plan-only Print the resolved plan without building or running tests
--dry-run Validate configuration and print the commands without running them
@@ -156,13 +156,6 @@ cargo_target_dir() {
fi
}
write_rustfs_features_stamp() {
local target_dir
target_dir="$(cargo_target_dir)"
mkdir -p "$target_dir/debug"
: > "$target_dir/debug/rustfs.features"
}
ensure_default_asset_platform() {
if [[ -n "$SOURCE_BINARY" ]]; then
return
@@ -623,8 +616,7 @@ export RUSTFS_E2E_LOG_DIR="${RUSTFS_E2E_LOG_DIR:-$RUN_DIR/server-logs}"
mkdir -p "$RUSTFS_E2E_LOG_DIR"
if [[ "$SKIP_BUILD" != 1 ]]; then
run_logged build-current cargo build --locked -p rustfs --bin rustfs
write_rustfs_features_stamp
run_logged build-current "$PYTHON_BIN" "$ROOT/scripts/e2e_binary.py" build
fi
SOURCE_REVISION="$(git rev-parse HEAD)"
@@ -641,6 +633,7 @@ for case_name in "${CASES[@]}"; do
HTTP_PROXY= \
HTTPS_PROXY= \
RUSTFS_SCANNER_HEAL_G09_EVIDENCE_DIR="$evidence_dir" \
"$PYTHON_BIN" "$ROOT/scripts/e2e_binary.py" run -- \
cargo test --locked -p e2e_test "$test_filter" -- --ignored --exact --nocapture
done
@@ -26,7 +26,7 @@ Options:
--out-dir DIR Alias for --run-dir
--test NAME all, g04, or g12 (default: all)
--allow-dirty Allow tracked source changes while collecting evidence
--skip-build Reuse an existing target/debug/rustfs binary
--skip-build Reuse a binary built by scripts/e2e_binary.py build from the current sources
--plan-only Print the resolved plan without building or running tests
--dry-run Alias for --plan-only
--self-test Run lightweight CLI and descriptor checks
@@ -99,13 +99,6 @@ cargo_target_dir() {
fi
}
write_rustfs_features_stamp() {
local target_dir
target_dir="$(cargo_target_dir)"
mkdir -p "$target_dir/debug"
: >"$target_dir/debug/rustfs.features"
}
artifact_dir_for() {
case "$1" in
g04)
@@ -585,8 +578,7 @@ SOURCE_REVISION="$(git rev-parse HEAD)"
printf '%s\n' "$SOURCE_REVISION" >"$RUN_DIR/source-revision.txt"
if [[ "$SKIP_BUILD" != 1 ]]; then
run_logged build-current cargo build --locked -p rustfs --bin rustfs
write_rustfs_features_stamp
run_logged build-current "$PYTHON_BIN" "$ROOT/scripts/e2e_binary.py" build
fi
if [[ " ${CASES[*]} " == *" g04 "* ]]; then
@@ -601,6 +593,7 @@ if [[ " ${CASES[*]} " == *" g12 "* ]]; then
NO_PROXY="${NO_PROXY:-127.0.0.1,localhost}" \
HTTP_PROXY= \
HTTPS_PROXY= \
"$PYTHON_BIN" "$ROOT/scripts/e2e_binary.py" run -- \
cargo test --locked -p e2e_test \
distributed::replication_quota_test::four_node_four_drive_hard_quota_rejects_over_limit_put \
-- --exact --nocapture
+314
View File
@@ -0,0 +1,314 @@
#!/usr/bin/env python3
"""Exercise the E2E build/run boundary without compiling RustFS."""
import json
import os
from pathlib import Path
import shutil
import signal
import subprocess
import sys
import tempfile
import unittest
class BinaryProvenanceTests(unittest.TestCase):
def setUp(self):
self.temp = tempfile.TemporaryDirectory()
self.addCleanup(self.temp.cleanup)
self.root = Path(self.temp.name)
(self.root / "scripts").mkdir()
shutil.copy(Path(__file__).with_name("e2e_binary.py"), self.root / "scripts/e2e_binary.py")
(self.root / "Cargo.toml").write_text("[workspace]\n")
(self.root / "source.rs").write_text("original source\n")
(self.root / ".gitignore").write_text("/target/\n/rustfs/static/\n")
(self.root / ".agents/skills").mkdir(parents=True)
(self.root / ".agents/skills/SKILL.md").write_text("tracked instructions\n")
(self.root / ".claude").mkdir()
(self.root / ".claude/skills").symlink_to("../.agents/skills", target_is_directory=True)
subprocess.run(["git", "init", "-q", str(self.root)], check=True)
for args in (["add", "."], ["-c", "user.name=Test", "-c", "user.email=test@example.com", "commit", "-qm", "fixture"]):
subprocess.run(["git", "-C", str(self.root), *args], check=True)
self.commands = self.root / "target/commands"
self.commands.mkdir(parents=True)
cargo = self.commands / "cargo"
cargo.write_text(f"#!{sys.executable}\n" + '''import json, os, pathlib, sys
if os.environ.get("FAKE_BUILD_FAIL"):
raise SystemExit(23)
args = sys.argv[1:]
if args[:2] == ["nextest", "run"]:
receipt = json.loads(pathlib.Path(os.environ["RUSTFS_E2E_BINARY_RECEIPT"]).read_text())
assert pathlib.Path(receipt["binary"]) == pathlib.Path(os.environ["CARGO_BIN_EXE_rustfs"]).resolve()
if os.environ.get("RUSTFS_E2E_STARTUP_CAS_BINARY"):
assert pathlib.Path(receipt["binary"]) == pathlib.Path(os.environ["RUSTFS_E2E_STARTUP_CAS_BINARY"]).resolve()
pathlib.Path("target/nextest-command.json").write_text(json.dumps(args))
raise SystemExit(int(os.environ.get("FAKE_TEST_EXIT", "0")))
target = pathlib.Path(args[args.index("--target-dir") + 1])
binary = target / ("release" if "--release" in args else "debug") / "rustfs"
binary.parent.mkdir(parents=True, exist_ok=True)
binary.write_text("#!/bin/sh\\nexit 0\\n")
binary.chmod(0o755)
features = ["default", "ftps", "webdav"]
if "--features" in args:
features.extend(args[args.index("--features") + 1].split(","))
if "full" in features:
features.extend(["sftp", "swift", "metrics-gpu", "pyroscope"])
print(json.dumps({"reason": "compiler-artifact", "target": {"name": "rustfs", "kind": ["bin"]}, "executable": str(binary), "features": sorted(set(features))}))
if os.environ.get("FAKE_BUILD_MUTATE"):
pathlib.Path("source.rs").write_text("changed during build")
''')
cargo.chmod(0o755)
rustc = self.commands / "rustc"
rustc.write_text("#!/bin/sh\nprintf 'rustc fixture\\nhost: fixture\\n'\n")
rustc.chmod(0o755)
self.env = dict(os.environ, PATH=f"{self.commands}{os.pathsep}{os.environ['PATH']}")
for name in ("CARGO_TARGET_DIR", "CARGO_BIN_EXE_rustfs", "RUSTFS_BUILD_FEATURES", "RUSTFS_E2E_BINARY_RECEIPT"):
self.env.pop(name, None)
self.binary = self.root / "target/debug/rustfs"
self.sidecar = self.binary.with_name("rustfs.e2e.json")
def invoke(self, *args, env=None):
return subprocess.run([sys.executable, str(self.root / "scripts/e2e_binary.py"), *args], cwd=self.root, env=env or self.env, text=True, capture_output=True)
def build(self, features=""):
result = self.invoke("build", "--features", features)
self.assertEqual(result.returncode, 0, result.stderr)
def run_code(self, code="pass", features="", env=None):
return self.invoke("run", "--features", features, "--", sys.executable, "-c", code, env=env)
def test_build_run_and_receipt_cleanup(self):
self.build("full,e2e-test-hooks")
result = self.run_code("import os,pathlib; print(os.environ['RUSTFS_E2E_BINARY_RECEIPT']); assert pathlib.Path(os.environ['CARGO_BIN_EXE_rustfs']).is_file(); assert 'sftp' in os.environ['RUSTFS_BUILD_FEATURES']", "e2e-test-hooks,full")
self.assertEqual(result.returncode, 0, result.stderr)
self.assertFalse(Path(result.stdout.strip()).exists(), "run receipts must not survive their command")
self.assertIn("sftp", json.loads(self.sidecar.read_text())["features"])
def test_source_changes_are_not_hidden_by_timestamps_or_head(self):
self.build()
path = self.root / "source.rs"
old = path.stat()
path.write_text("different bytes\n")
os.utime(path, ns=(old.st_atime_ns, old.st_mtime_ns))
self.assertNotEqual(self.run_code().returncode, 0)
def test_deleted_untracked_and_ignored_embedded_inputs(self):
for mutation in ("delete", "untracked", "static"):
with self.subTest(mutation=mutation):
self.build()
path = self.root / "source.rs"
if mutation == "delete":
path.unlink()
elif mutation == "untracked":
(self.root / "new.rs").write_text("new source")
else:
static = self.root / "rustfs/static"
static.mkdir(parents=True)
(static / "index.html").write_text("embedded content")
self.assertNotEqual(self.run_code().returncode, 0)
path.write_text("original source\n")
def test_wrong_binary_features_and_manifest_fail_closed(self):
self.build("sftp")
self.assertNotEqual(self.run_code(features="webdav").returncode, 0)
self.binary.write_text("old server")
self.assertNotEqual(self.run_code(features="sftp").returncode, 0)
self.sidecar.write_text("{}")
self.assertNotEqual(self.run_code(features="sftp").returncode, 0)
self.sidecar.unlink()
self.assertNotEqual(self.run_code(features="sftp").returncode, 0)
def test_build_failure_or_source_race_does_not_leave_a_receipt(self):
for failure in ("FAKE_BUILD_FAIL", "FAKE_BUILD_MUTATE"):
self.build()
result = self.invoke("build", env=dict(self.env, **{failure: "1"}))
self.assertNotEqual(result.returncode, 0)
self.assertFalse(self.sidecar.exists())
def test_child_failure_and_changes_during_run_fail(self):
self.build()
failed = self.run_code("raise SystemExit(37)")
self.assertEqual(failed.returncode, 37, failed.stderr)
for code in ("import pathlib; pathlib.Path('source.rs').write_text('changed while testing')", "import pathlib; pathlib.Path('target/debug/rustfs').write_text('different server')"):
self.build()
self.assertNotEqual(self.run_code(code).returncode, 0)
def test_override_cannot_select_an_unverified_server(self):
self.build()
result = self.run_code(env=dict(self.env, CARGO_BIN_EXE_rustfs="/some/old/server"))
self.assertNotEqual(result.returncode, 0)
def test_artifact_moves_between_clean_checkouts(self):
self.build()
with tempfile.TemporaryDirectory() as destination:
clone = Path(destination) / "clone"
subprocess.run(["git", "clone", "-q", str(self.root), str(clone)], check=True)
(clone / "target/debug").mkdir(parents=True)
shutil.copy2(self.binary, clone / "target/debug/rustfs")
shutil.copy2(self.sidecar, clone / "target/debug/rustfs.e2e.json")
result = subprocess.run([sys.executable, str(clone / "scripts/e2e_binary.py"), "run", "--", sys.executable, "-c", "pass"], cwd=clone, env=self.env, text=True, capture_output=True)
self.assertEqual(result.returncode, 0, result.stderr)
def test_ci_build_preserves_both_manifests_and_runs_the_copied_server(self):
from check_test_wiring import yaml_block
from test_security_workflow import named_steps, shell_body
source = (Path(__file__).resolve().parents[1] / ".github/workflows/ci.yml").read_text().splitlines()
build_steps = named_steps(yaml_block(source, "build-rustfs-debug-binary", 2))
run_steps = named_steps(yaml_block(source, "e2e-full", 2))
(self.root / "Cargo.lock").write_text("fixture lock\n")
subprocess.run(["git", "add", "Cargo.lock"], cwd=self.root, check=True)
subprocess.run(["git", "-c", "user.name=Test", "-c", "user.email=test@example.com", "commit", "-qm", "lock"], cwd=self.root, check=True)
copied = self.root / "target/startup-cas-input/rustfs"
env = dict(self.env, STARTUP_CAS_INPUT=str(copied.parent), RUSTFS_E2E_STARTUP_CAS_BINARY=str(copied))
for step in (build_steps["Build debug binary"], run_steps["Preserve startup CAS binary input"]):
result = subprocess.run(["bash", "-e", "-o", "pipefail", "-c", shell_body(step)], cwd=self.root, env=env, capture_output=True, text=True)
self.assertEqual(result.returncode, 0, result.stderr)
for name in ("rustfs.e2e.json", "rustfs.e2e-startup-cas-build.json"):
self.assertIn(" target/debug/" + name, build_steps["Upload debug binary"])
self.assertEqual((self.binary.parent / name).read_bytes(), (copied.parent / name).read_bytes())
manifest = json.loads(copied.with_name("rustfs.e2e-startup-cas-build.json").read_text())
self.assertEqual(manifest["argv"], ["python3", "scripts/e2e_binary.py", "build", "--bins", "--features", "e2e-test-hooks"])
self.assertTrue(manifest["clean_before"] and manifest["clean_after"])
body = next(line.removeprefix(" run: ") for line in run_steps["Run e2e full suite"] if line.startswith(" run: "))
for status in (0, 23):
result = subprocess.run(["bash", "-e", "-o", "pipefail", "-c", body], cwd=self.root, env=dict(env, FAKE_TEST_EXIT=str(status)), capture_output=True, text=True)
self.assertEqual(result.returncode, status, result.stderr)
copied.write_text("replaced preserved binary")
result = subprocess.run(["bash", "-e", "-o", "pipefail", "-c", body], cwd=self.root, env=env, capture_output=True, text=True)
self.assertNotEqual(result.returncode, 0)
def test_distributed_workflow_runs_both_filter_branches_with_receipts(self):
from check_test_wiring import yaml_block
from test_security_workflow import named_steps, shell_body
source = (Path(__file__).resolve().parents[1] / ".github/workflows/e2e-distributed.yml").read_text().splitlines()
steps = named_steps(yaml_block(source, "distributed", 2))
result = subprocess.run(["bash", "-e", "-o", "pipefail", "-c", shell_body(steps["Build rustfs binary"])], cwd=self.root, env=self.env, capture_output=True, text=True)
self.assertEqual(result.returncode, 0, result.stderr)
for selected in ("", "test(distributed::s3_basic)"):
for status in (0, 23):
result = subprocess.run(["bash", "-e", "-o", "pipefail", "-c", shell_body(steps["Run distributed 4-node e2e suite"])], cwd=self.root, env=dict(self.env, FILTER=selected, FAKE_TEST_EXIT=str(status)), capture_output=True, text=True)
self.assertEqual(result.returncode, status, result.stderr)
argv = json.loads((self.root / "target/nextest-command.json").read_text())
self.assertEqual(argv, ["nextest", "run", "--profile", "e2e-distributed", "-p", "e2e_test", *(["-E", selected] if selected else ["--no-tests=fail"])])
def test_target_directory_and_profile_are_explicit(self):
env = dict(self.env, CARGO_TARGET_DIR="target/custom")
built = self.invoke("build", "--profile", "release", env=env)
self.assertEqual(built.returncode, 0, built.stderr)
run = self.invoke("run", "--profile", "release", "--", sys.executable, "-c", "pass", env=env)
self.assertEqual(run.returncode, 0, run.stderr)
self.assertNotEqual(self.invoke("run", "--", sys.executable, "-c", "pass", env=env).returncode, 0)
def test_target_directory_cannot_hide_source_inputs(self):
for target in (str(self.root), str(self.root / "crates"), str(self.root.parent)):
with self.subTest(target=target):
result = self.invoke("build", env=dict(self.env, CARGO_TARGET_DIR=target))
self.assertNotEqual(result.returncode, 0)
self.assertIn("CARGO_TARGET_DIR", result.stderr)
tracked = self.root / "target/tracked.rs"
tracked.write_text("tracked build input")
subprocess.run(["git", "add", "-f", "target/tracked.rs"], cwd=self.root, check=True)
self.build()
tracked.write_text("changed tracked build input")
self.assertNotEqual(self.run_code().returncode, 0)
def test_unsupported_embedded_directory_links_fail_closed(self):
self.build()
destination = self.root / "target/embedded-assets"
destination.mkdir()
(destination / "index.html").write_text("untracked embedded input")
static = self.root / "rustfs/static"
static.mkdir(parents=True)
(static / "linked-assets").symlink_to(destination, target_is_directory=True)
self.assertNotEqual(self.run_code().returncode, 0)
def test_directory_aliases_cannot_hide_unrecorded_inputs(self):
self.build()
target = self.root / ".agents/skills/SKILL.md"
target.write_text("changed instructions\n")
self.assertNotEqual(self.run_code().returncode, 0)
self.build()
(target.parent / ".gitignore").write_text("hidden.rs\n")
(target.parent / "hidden.rs").write_text("ignored build input\n")
result = self.invoke("build")
self.assertNotEqual(result.returncode, 0)
self.assertIn("unrecorded input", result.stderr)
alias = self.root / ".claude/skills"
alias.unlink()
with tempfile.TemporaryDirectory() as external:
alias.symlink_to(external, target_is_directory=True)
result = self.invoke("build")
self.assertNotEqual(result.returncode, 0)
self.assertIn("escapes the source inventory", result.stderr)
def test_directory_alias_indirection_is_part_of_the_identity(self):
for name in ("first", "second"):
directory = self.root / name
directory.mkdir()
(directory / "input.rs").write_text(name)
selection = self.root / "target/selection"
selection.symlink_to(self.root / "first", target_is_directory=True)
(self.root / "source-alias").symlink_to("target/selection", target_is_directory=True)
self.build()
selection.unlink()
selection.symlink_to(self.root / "second", target_is_directory=True)
self.assertNotEqual(self.run_code().returncode, 0)
def test_existing_embedded_files_and_symlink_targets_are_hashed(self):
static = self.root / "rustfs/static"
static.mkdir(parents=True)
index = static / "index.html"
index.write_text("embedded version one")
external = self.root / "target/embedded-file"
external.write_text("linked version one")
(static / "linked.html").symlink_to(external)
self.build()
index.write_text("embedded version two")
self.assertNotEqual(self.run_code().returncode, 0)
self.build()
external.write_text("linked version two")
self.assertNotEqual(self.run_code().returncode, 0)
def test_each_run_hashes_binary_twice_and_never_calls_cargo(self):
script = self.root / "scripts/e2e_binary.py"
script.write_text(script.read_text().replace("def file_hash(path):\n", "def file_hash(path):\n if path.name == 'rustfs':\n with (ROOT / 'target/hash-count').open('a') as count:\n count.write('hash\\n')\n"))
self.build()
count = self.root / "target/hash-count"
count.write_text("")
result = self.run_code(env=dict(self.env, FAKE_BUILD_FAIL="1"))
self.assertEqual(result.returncode, 0, result.stderr)
self.assertEqual(count.read_text().splitlines(), ["hash", "hash"])
def test_concurrent_build_or_run_is_rejected(self):
self.build()
command = [sys.executable, str(self.root / "scripts/e2e_binary.py"), "run", "--", sys.executable, "-c", "print('ready', flush=True); input()"]
with subprocess.Popen(command, cwd=self.root, env=self.env, stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True) as process:
self.assertEqual(process.stdout.readline().strip(), "ready")
try:
for args in (("build", "--features", "sftp"), ("run", "--", sys.executable, "-c", "pass")):
rejected = self.invoke(*args)
self.assertNotEqual(rejected.returncode, 0)
self.assertIn("Another E2E build/run", rejected.stderr)
finally:
output, error = process.communicate("\n", timeout=10)
self.assertEqual(process.returncode, 0, error + output)
self.assertFalse(self.binary.with_name("rustfs.e2e.lock").exists())
def test_interruption_cleans_receipt_and_releases_ownership(self):
self.build()
for signum in (signal.SIGINT, signal.SIGTERM):
command = [sys.executable, str(self.root / "scripts/e2e_binary.py"), "run", "--", sys.executable, "-c", "import os; print(os.environ['RUSTFS_E2E_BINARY_RECEIPT'], flush=True); input()"]
with subprocess.Popen(command, cwd=self.root, env=self.env, stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True) as process:
receipt = Path(process.stdout.readline().strip())
self.assertTrue(receipt.is_file())
process.send_signal(signum)
process.communicate(timeout=10)
self.assertNotEqual(process.returncode, 0)
self.assertFalse(receipt.exists())
self.assertFalse(self.binary.with_name("rustfs.e2e.lock").exists())
if __name__ == "__main__":
unittest.main()

Some files were not shown because too many files have changed in this diff Show More