* fix(ecstore): stop pruning at nonempty directories (#7616)
* fix(ecstore): stop pruning at nonempty directories
* test(ecstore): release pruning fixtures before temp cleanup
(cherry picked from commit 8f150d1d8e)
* fix(heal): preserve retryable batch failures during recovery (#7642)
* fix(heal): preserve retryable batch failures during recovery
* test(heal): pin prebuilt hooks binaries in ci
(cherry picked from commit 5cd58319ed)
* fix(s3): reject oversize single PUT early and map body errors to 4xx (#7635)
* fix(s3): reject oversize single PUT early and map body errors to 4xx
A single PutObject above the 5 GiB single-request ceiling was only
rejected after the client had streamed 5 GiB into s3s's read-time body
budget, and the resulting BodySizeLimitExceeded surfaced from the erasure
writer as 500 InternalError. A body whose connection hit EOF before
Content-Length bytes arrived (hyper's IncompleteBody) was also a 500.
SDKs retry 500s, so one oversize upload was resent from offset 0 five
times.
- PutObject and UploadPart reject a declared length above
MAX_SINGLE_PUT_OBJECT_SIZE with 400 EntityTooLarge before reading the
body; the constant moves to rustfs_config so the s3s limit and the
admission check share one value.
- ApiError maps BodySizeLimitExceeded to EntityTooLarge and a hyper body
EOF to IncompleteBody across both io::Error conversions.
Fixes#7596.
* test(s3): cover UploadPart admission, aws-chunked length, real s3s limit
- Poll-counting test body proves PutObject and UploadPart reject a
declared size above the ceiling with zero body polls; exact-cap and
zero-length parts pass admission.
- A STREAMING-* aws-chunked PUT whose framed Content-Length exceeds the
cap is admitted when the decoded length is within it and rejected when
the decoded length is over it.
- The display-based BodySizeLimitExceeded matcher is checked against the
real error produced by the pinned s3s Body budget.
(cherry picked from commit 50b31bc75b)
* fix(ecstore): make directory mtime fixture portable (#7623)
* fix(ecstore): make directory mtime fixture portable
* style(ecstore): format mtime fixture assertion
---------
Co-authored-by: houseme <housemecn@gmail.com>
Co-authored-by: Zhengchao An <anzhengchao@gmail.com>
(cherry picked from commit b1cc286cac)
* fix(storage): prevent readiness after native migration failures (#7652)
* fix(storage): prevent readiness after native migration failures
* fix(storage): skip unsupported IAM records before reading
* fix(storage): use stable typed migration metadata errors
* fix(storage): include migration record in startup errors
* test(storage): cover native migration startup failures
* test(storage): use array chunks in migration fixture
---------
Co-authored-by: RJ Regenold <214054+rjregenold@users.noreply.github.com>
Co-authored-by: cxymds <cxymds@gmail.com>
(cherry picked from commit 0cbc3ffe61)
* fix(admin): expose OIDC account display fields (#7654)
Expose verified OIDC username and email claims as display-only metadata on self-account responses while preserving the virtual parent as the authorization identity.\n\nKeep rustfs-madmin public response structs unchanged by adding the optional wire fields through private handler response wrappers.
(cherry picked from commit f02bc947cd)
* fix(replication): correct peer joins and remote-state reporting (#7650)
* fix(replication): propagate verified peer deployment identities
* fix(replication): report actual remote peer state
* fix(replication): defer initial sync until all peers join
* test(replication): shut down TLS fixtures cleanly
---------
Co-authored-by: houseme <housemecn@gmail.com>
(cherry picked from commit 853ae63b6a)
* fix(s3): bound stalled UploadPart request bodies (#7659)
* fix(s3): bound stalled UploadPart request bodies
* fix(ci): preserve the S3S footprint ratchet
(cherry picked from commit 666dfd9f9f)
* fix(tables): reject reserved warehouse locations (#7671)
Co-authored-by: cxymds <cxymds@gmail.com>
(cherry picked from commit 414176c47f)
* fix(ci): bind nightly lanes to one resolved source (#7688)
(cherry picked from commit 01d8e4347f)
* fix(replication): close the pre-stable convergence gaps from backlog#2367 (#7626)
* fix(replication): total-order rule sort and honor V1 top-level Prefix
Rule matching had two defects from the pre-GA replication audit
(rustfs/backlog#2367 C-1 and C-2):
- The actionable-rule sort compared same-destination rules by priority but
answered Equal for any other pair, which is not a total order; the
standard library sort panics on such comparators once a slice exceeds the
insertion-sort threshold, so an object matching more than 20 enabled rules
across two or more targets could panic the PUT or DELETE task. Rules now
sort by priority descending with destination and id as tie-breakers, and
filter_target_arns preserves that order instead of draining a HashSet.
- A V1 rule written without a <Filter> carries its prefix at the top level;
that field was never read, so <Prefix>logs/</Prefix> matched every object.
ReplicationRuleExt::prefix now falls back to it, with a <Filter> keeping
precedence. The existing prefix fixtures were built this way and had been
asserting nothing.
* fix(admin): advertise data-usage and listen capabilities to rc
The rc client gated `rc du` and `rc watch` on a pinned contract that
matched server versions by the string prefix `1.0.0-rc.`; a server that
reports `1.0.0` no longer matches, and the dynamic `advertised` list did
not carry either name, so `rc du` against a GA server fails with an
unsupported-capability error (rustfs/backlog#2367 E-2).
Advertise `admin.data-usage` from the admin route inventory like the IAM
entries, and `listen_notification` for the bucket `?events=` extension
route the admin router dispatches. The client merges advertised entries
ahead of its pinned contract, so no version sniffing is needed.
* fix(site-replication): stop notifying the local site on remove and rotate
The pending-remove and pending-rotation notification loops skipped the
local site by endpoint only, while finalization identifies it by
deployment id or endpoint. The reconcile tick resolves the local peer from
the node's own listen address (and a handler from the request Host), so
`remove --all` dialed the site's registered endpoint, waited out the
request timeout against the lifecycle lock it was holding, and answered
`Partial: failed to notify 1 peer(s)` for a removal that had succeeded
(rustfs/backlog#2367 A-4, backlog#2195 item 3).
Both loops now iterate the peers still awaiting notification through one
helper that applies the finalization identity.
* fix(site-replication): promote and settle IAM retries without a tick of slack
Two retry-queue behaviours kept an IAM change from converging for ten to
twenty minutes after a peer came back (rustfs/backlog#2367 A-1 and A-3,
backlog#2305):
- The lightweight 30-second pass filtered its reachability probe to bucket
ops, so a backed-off IAM or bucket-metadata snapshot waited for the
600-second tick to notice the peer. It now probes every backed-off class
and still replays only bounded bucket ops; promotion is a state flip the
heavyweight tick acts on.
- Backoffs are multiples of the tick interval, so a failure stamped δ
seconds after a tick was 600 − δ old at the next tick and slipped a whole
extra interval. The heavyweight drain now evaluates backoff halfway to its
next tick.
- An IAM entry first created by a non-deletion failure (the add bootstrap's
snapshot send, the drain's own replay, an import-iam schedule) was never
stamped `deletions_recorded`, so a later recorded deletion could not
settle it and it escalated to the marker only `replicate repair` clears.
Entries created by this binary now start recorded; a row persisted by an
older binary keeps the escalation semantics.
* fix(site-replication): reload peer node caches after bucket wiring writes
Every S3 bucket-config write ends by asking the other nodes of the cluster
to reload the bucket's metadata; the site-replication writers never did.
On a multi-node site the node that ran the pairing (or applied a peer's
bucket-meta item) rewrote the bucket targets and the derived replication
rules on disk, while every other node kept serving its cached copy for up
to the 15-minute refresh. A `resync start` routed to such a node reported
every freshly wired bucket as `Config not found` and a bucket whose
operator target the pairing had replaced as `recorded remote target no
longer exists` (rustfs/backlog#2367 A-5, backlog#2195 item 2; functional
SITE-105).
Add one best-effort reload helper in the site-replication hooks and call it
after the bucket setup, versioning, peer bucket-meta apply, removed-peer
cleanup, make-with-versioning, and endpoint-refresh writes; the ensure
helpers now report whether they wrote so unchanged passes stay silent. The
resync manifest and start now read the persisted wiring instead of the
node-local cache, matching the target read the start path already did.
The new four-node e2e pairs two clusters and starts a resync through a
non-coordinator node right after pairing; it also covers an IAM user
created on a non-coordinator node converging to the peer site.
* test(e2e): cover delete-marker replication from a multi-node source
The functional suite reported delete markers created on a 3-node source
never reaching the target (rustfs/backlog#2195 item 4, REP-105). The report
was a probe defect, but the shape had no coverage: the existing
delete-marker e2e runs a single-node source. Pin it against a four-node
source replicating to a four-node peer and to a single-node target, with
the write and the delete issued through different nodes.
* ci(e2e): refresh the distributed selection for the new replication cases
Four distributed cases were added (two site-replication, two delete-marker
replication). The linux digest is derived from the last CI listing of the
lane (34 cases, matching the previous pin) plus the four new names; the
darwin digest is the local listing, which selects the same 38 cases.
(cherry picked from commit aeaba86d73)
* fix(e2e): require a verified server binary for every e2e run (#7687)
* fix(ci): share quick checks and lint workflows
* fix(ci): install actionlint from its verified release
* fix(ci): reject dependencies on required quick checks
* feat(test): verify the E2E server build and source identity
* test(e2e): register verified Darwin test membership
* test(e2e): record verified Linux receipt test membership
* test(e2e): record compiled Darwin receipt test membership
* test(e2e): record compiled Linux receipt test membership
* test(e2e): record Darwin e2e-full membership after merging main
* fix(test): route scanner/heal evidence E2E runs through the verified server binary
The evidence runners built rustfs with plain cargo and then ran e2e_test directly, which now fails without a run receipt. They build through scripts/e2e_binary.py and run the e2e_test invocations under e2e_binary.py run; the obsolete rustfs.features stamp is removed.
* docs(e2e): run server-backed e2e commands through the verified binary wrapper
* test(e2e): record Linux e2e-full membership from the branch CI listing
(cherry picked from commit 2909b1bfe1)
* fix(kms): classify KMS/SSE error contracts and SSE-S3 headers (#7697)
* fix(sse): classify bare SSE-KMS writes when no KMS is available
A `aws:kms` request without a key id, on a bucket without a default key,
returned `500 InternalError` whenever no KMS service was running: the
"no KMS key available" branch exited with an untyped storage error before
the availability classification that the keyed form already received.
Route that branch through the same split: `503 ServiceUnavailable` while
a configured KMS is stopped, `400 InvalidRequest` when KMS was never
configured, and `400 InvalidRequest` naming the missing key id when a
running KMS has no default key. `CreateMultipartUpload` shares the path.
Adds a unit test for the bare form and an e2e module that stops KMS
through the admin API, runs a master-key-only node, and runs a Local KMS
without a default key; refreshes the e2e-full selection digests.
(cherry picked from commit c3259dadc3d603a9185a5b0ad9f83dfb884e61c8)
* fix(sse): keep KMS error classes on the encrypted read path
GetObject, CopyObject and UploadPartCopy on an SSE-KMS object whose key
no longer exists answered `500 InternalError` ("KMS key not found") while
PutObject under the same key already answered `400 KMS.NotFoundException`.
The read path carries its classification through ecstore's
`EncryptionResolutionErrorKind`, which had no kind for a missing key, a
denied KMS grant or a missing backend capability, so all three folded
onto `DecryptionFailed` and the S3 layer reported an internal fault.
Add `KeyNotFound`, `AccessDenied` and `NotImplemented` kinds, map them on
both sides of the boundary, and give an envelope the configured backend
cannot unwrap a diagnosable message while keeping its `500`.
Unit tests cover the kind round trip and the reader wrapping; a new e2e
test deletes a key immediately and checks GET/Copy return 400 with
`KMS.NotFoundException` while HEAD stays 200. The e2e-full selection
digests are refreshed from the current listing (the previous digests
predated the delete-authorization tests) and the e2e `create_default_key`
helper is updated to the accepted `EncryptDecrypt` spelling.
(cherry picked from commit 2523a9814e97caea318d4ff1a51bef3a4d4445b2)
* fix(kms): classify key-management errors on the admin routes
`POST /kms/keys`, the legacy `create-key` alias and `generate-data-key`
reported every backend refusal as `500`: a blank key name (which each
backend failed on differently, the Local backend by writing a key file
with an empty stem), a name already taken, an unknown key, a disabled key
and a capability the backend lacks. `delete` and the lifecycle routes
already classified the same errors.
Refuse a blank or whitespace name in `KmsManager::create_key` before any
backend sees it, and share one `KmsError` to status mapping across
create, delete and generate-data-key (400 for validation and key state,
404 for an unknown key, 409 for a taken name, 501 for a missing
capability, 500 only for damaged material). The XML-error routes carry
the same status explicitly since s3s derives none for a custom code.
The read-only Static backend now reports create, delete and
cancel-deletion as `UnsupportedCapability`, matching its rotate and
enable/disable answers, so the admin API returns 501 for all of them.
(cherry picked from commit e33cac5493c4d9d6662e0d2980b58ba2b24a6d1b)
* fix(sse): stop SSE-S3 responses from naming the wrapping KMS key
`x-amz-server-side-encryption-aws-kms-key-id` is defined for `aws:kms`
objects only, but PutObject, CopyObject, CreateMultipartUpload and
GetObject returned it for `AES256` objects too, carrying the KMS key that
wraps the SSE-S3 data key (the service default, or the literal `default`
on a node without KMS). The write paths copied `kms_key_id` from the
encryption material unconditionally, and the single-decrypt GET
classification did the same after resolving the key for authorization.
Add `EncryptionMaterial::response_kms_key_id`, which yields the id only
for SSE-KMS, use it at the four write-response sites, and gate the GET
classification the same way. CompleteMultipartUpload and HeadObject
already omitted the header.
Unit tests pin both directions; a new e2e test covers Put/Get/Head/Copy
and CreateMultipartUpload for AES256 with an aws:kms control. The
e2e-full selection digests are refreshed from the current listing.
(cherry picked from commit 29d793a63352b0b60fd53c565e80fdbede8964bb)
* fix(s3): validate PutBucketEncryption rules before storing them
A default-encryption rule naming an unknown `SSEAlgorithm` (for example
`AES128`), a rule without `ApplyServerSideEncryptionByDefault`, an empty
rule list, or a `KMSMasterKeyID` on an `AES256` rule was stored as
written: the only algorithm check on the route decided whether to fill
in the default KMS key. `GetBucketEncryption` then advertised that
configuration while the write path encrypted header-less writes under
its `AES256` fallback, so the bucket's declared and actual schemes
disagreed. Two comments claimed the route already refused unknown
algorithms.
Validate the configuration before any of it is applied: `MalformedXML`
for a malformed rule set or unknown algorithm, `InvalidArgument` for a
key id on a non-KMS rule, and nothing stored on refusal. Correct the two
comments to describe when the AES256 fallback is still reachable.
Unit tests cover every refusal and the accepted shapes; an e2e test
checks the refusals leave the previous configuration in place. The
e2e-full selection digests are refreshed from the current listing.
(cherry picked from commit 29e4486dce41197ed93f5253cdbabc57d27a4ddb)
* test(e2e): refresh e2e-full selection for the combined KMS/SSE fixes
* test: align two unit tests with the new KMS and bucket-encryption contracts
`scheduled_deletion_carries_a_deadline_and_can_be_cancelled` still
expects the state error (`InvalidOperation`) for cancelling a key that
is not pending deletion; only the Static backend's mutations moved to
`UnsupportedCapability`. The uninitialized-store PutBucketEncryption
test now sends a well-formed AES256 rule so it reaches the store lookup
instead of the new configuration validation.
(cherry picked from commit e2e6a2535a)
* fix(site-replication): keep an operator's bucket-level target to a peer instead of taking it over (#7709)
* fix(site-replication): keep an operator's bucket-level target to a peer instead of taking it over
Site replication wired each bucket by looking for an existing replication
target "to the same peer" and rewriting the first match in place as its own
same-name target. An operator's bucket-level target that happened to point
at that site (different target bucket, operator credentials) was the first
match whenever it pre-dated the join, and the reconciler repeats the pass
every 600s, so the takeover also depended on target order afterwards. The
operator's rule then named an ARN no target backed and their bucket
replication stopped silently, while the inherited bucket-level reset id
made every site resync report the bucket as owned by another resync
(rustfs/backlog#2479, rustfs/backlog#2489).
Follow MinIO's `getRemoteARN` / `getRemoteARNForPeer` shape instead:
- Wiring updates a target in place only under the same ARN, or when it is
recognisably the site's own under an older ARN shape (same peer,
same-name target bucket, site replication service account). Anything
else gets the site target added next to it.
- The site resync manifest takes the target the derived
`site-repl-<deployment id>` rule names (same-name shape as fallback), so
an operator target to the peer neither aborts the bucket as "multiple
remote targets matched peer" nor gets resynced into.
- Peer removal prunes only targets a pruned derived rule names or the
same-name target bucket; operator targets stamped with the peer's
deployment id survive together with their rules.
Unit tests cover the three predicates. e2e
`test_site_replication_keeps_operator_bucket_target_to_peer` runs a
bucket-level replication plus `replication-reset` to the future peer, joins
the sites, and requires the operator target untouched, both paths
delivering, the site resync completing against the site target, and the
operator target and rule surviving `replicate remove --all`; without the
fix it fails at the join with the operator target gone. The repl-nightly
selection digest is refreshed for the new case.
* test(site-replication): drop a redundant clone flagged by clippy
The reconcile unit test cloned the remote peer into the state map although
the binding is not used afterwards; workspace clippy (-D warnings) rejects
that as redundant_clone.
(cherry picked from commit ecdc55fa4b)
* fix: enforce S3 permissions for recursive force deletion (#7661)
* fix: enforce S3 authorization for recursive deletion
* fix: satisfy the s3s footprint guard
* fix: restore list versions policy compatibility (#7686)
(cherry picked from commit 3fd1ce414d)
* fix(ci): repair functional defaults and chain regression checks (#7664)
* fix(ci): default functional suites to nightly packages
* test(ci): follow the fault-tolerance chain handoff
(cherry picked from commit 509a0fa90c)
* fix(ci): align security workflow tests with chain (#7679)
(cherry picked from commit d9e47d2813)
* test(ecstore): keep tier cleanup tests stable after immediate receipt queueing
Release PR #7766 made PUT/CopyObject overwrites queue the tier free-version cleanup receipt immediately, which broke two ecstore tests on release CI. In tier_overwrite_put_and_self_copy_recover_persisted_cleanup_owners the restarted store already runs expiry workers from the second iteration on, so they deleted the remote bytes before the test could assert that the commit leaves them in place; the test now fails the first remote DELETE via set_remove_failure(true) so the cleanup owner stays durable and the later restart still has to rediscover it from xl.meta (the failed remove does not bump remove_count, and the test re-enables removes before the recovery wait). In batch_transitioned_delete_post_commit_failures_roll_back_without_free_version_receipt the convergence loop now treats a transient InsufficientReadQuorum as "not yet converged", because cleanup rewrites xl.meta disk by disk and a racing read can briefly miss quorum (seen in the rio-v2 lane); any other error still panics.
* Revert "fix(ci): align security workflow tests with chain (#7679)"
This reverts commit 4544359f6d.
* Revert "fix(ci): repair functional defaults and chain regression checks (#7664)"
This reverts commit 5ba7ec0291.
* fix(ecstore): pass shard integrity to backported ingest-mode test
The stalled-reader test backported with #7659 used main's four-argument encode_with_ingest_mode, but release's signature takes an optional IntegrityBuilder, so pass None to keep the test focused on ingest-mode cleanup.
---------
Co-authored-by: Henry Guo <marshawcoco@gmail.com>
Co-authored-by: 唐小鸭 <tangtang1251@qq.com>
Co-authored-by: houseme <housemecn@gmail.com>
Co-authored-by: RJ Regenold <rregenold@teamraft.com>
Co-authored-by: RJ Regenold <214054+rjregenold@users.noreply.github.com>
Co-authored-by: cxymds <cxymds@gmail.com>
Co-authored-by: GatewayJ <835269233@qq.com>
Co-authored-by: Jason Kossis <jkossis@gmail.com>
The overnight run after #7708 proved the security verdict greps still
counted zero: real verdict lines are '\e[1;31m[FAIL]\e[0m STS-105 ...'
— the reset escape sits between the tag and the case id, and the
pattern only tolerated escapes before the tag. Allow escapes on both
sides; the fixture now emits the reset too, mirroring the real suite.
The tier case table is produced by rustfs_tier_report.py and its rows
lead with the topology column, so the case-ID-first table grep matched
nothing ('Product result: 0 passed, 0 failed'). Parse the PASS/FAIL
counts from the '## Case Summary' bullets the report always emits,
falling back to a topology-aware table grep.
First full serial pass with the green-on-case-failure semantics
(run 34693745171 / 34695021651) exposed three report-layer defects:
security: the report step referenced LOG_FILE, which is undefined in
this workflow (set -u killed the step before writing report.md), and
the verdict greps could not match the ANSI-escaped [PASS]/[FAIL] tags
in the real suite log. Point it at the artifacts suite.log, allow any
number of color escapes before the verdict tag, and count [SKIP]
lines separately (45 passed, 6 failed, 3 skipped was reported as an
unbound-variable crash).
pool: warp is stopped early (SIGINT) at the storage threshold and
only writes its final report on a clean exit, so an empty warp.log is
the expected shape of a healthy run - require its presence, not its
size. A mid-script die() abort or a FAIL step verdict must also turn
the validator red now that the run step is continue-on-error.
tier: add the standard 'Product result: N passed, M failed' summary
line computed from the case table, matching the other suites.
The contract test fixture previously injected LOG_FILE into the
environment and wrote verdict lines without ANSI escapes, which hid
both real-world defects; the fixture now mirrors the real suite
(stdout+tee with color tags) and asserts the pass/fail/skip counters.
A failing product case used to turn the whole workflow red, so the run
conclusion carried no signal beyond 'something failed' and the report
was suppressed. New semantics across the functional suites:
- Suite steps run with continue-on-error: the outcome is still recorded
for the report and the backlog issue manager (security/tier already
carried the flag).
- Generate report always publishes the full per-case table plus a
'Product result: N passed, M failed' summary, and its exit gate is
harness health: red only when the suite never reached case level (no
case verdicts), failed wholesale (zero passes, >=3 failures), or was
cancelled/skipped. performance is unchanged (parked).
- tier's structured gate no longer fails on case failures; it keeps red
for evidence-init and missing-gate-result breakdowns.
- pool/performance keep their existing red sources (install/benchmark).
Workflow contract tests updated to the new exit semantics: the security
report matrix keys green off the suite outcome, the evidence matrix
expects green for failure outcomes with recorded case rows (except
performance), the heal staged-rerun block expects the per-step table to
always publish, and run steps are now required to carry
continue-on-error.
Verified locally: actionlint clean; test_security_workflow.py 21/21.
Add a read-only descriptor ledger mode to the Scanner/Heal Linux evidence planner so release operators can separate current-head measured descriptors from old-head measured artifacts and case-level inputs before final bundle assembly.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
* ci: manage backlog issues by signal instead of per-run filing
The suite workflows used to file one backlog issue per failed run
(dedup was by run ID, which never matched), so issues accumulated
without bound. Replace the inline filing step in every suite workflow
(s3, kms, tier, storage, heal, pool, security, replication, upgrade,
performance) with a single call to
auto-testing/scripts/issue_manager.py, which:
- dedups by signal: failing cases are searched among open issues by
label (suite category + case ID); covered cases become a coalesced
comment on the existing issue, only uncovered cases file a new one
- labels new issues with functional-test, the suite category, one
label per failing case ID (lazily created), and env for
bootstrap-class failures (no cases ran, wholesale failure, or
404/ssh/clone/dpkg signatures in the log)
- closes open issues of the suite after a fully green run, citing the
run as evidence; cancelled runs never file or close anything
The step is skipped cleanly when auto-testing (private checkout) does
not contain the manager, or when PF_TESTING_GH_TOKEN is unset.
* fix(ci): satisfy actionlint and workflow contract tests for the manager step
- heal and performance workflows have no rustfs_version dispatch input;
referencing `${{ inputs.rustfs_version }}` in the manager step failed
actionlint's expression type check. Their package source now resolves
from package_url with the nightly fallback.
- scripts/test_security_workflow.py pinned the removed inline filing
step. The wiring assertions now pin the manager step (manager path +
per-suite report argument), and the evidence/stale-file tests assert
the skip contract instead: without the private auto-testing checkout
present, the step exits 0, publishes nothing, and leaves stale
evidence untouched.
Verified locally: actionlint clean, shellcheck clean,
test_security_workflow.py 21/21.
Size the scanner collector wait budget from the requested sample count and interval instead of using a fixed 120 second timeout. This preserves the full duration/60+1 telemetry sample set for measured ABBA runs.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
* test(e2e): target multi-set outage heal candidate
Require the outage write used by EC8+4 multi-set root-heal evidence to miss the same erasure index owned by the selected replacement drive. This avoids accepting a candidate from a different set and turning a valid heal into a false negative.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
* fix(e2e): satisfy G14 heal lint gates
Remove clippy-only noise from the G14 multi-set heal evidence test and align the admin route policy inventory with the registered heal catch-all route.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
* update
---------
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
* fix(e2e): record EC8+4 drive restart set size
Include the erasure set drive count in the distributed EC8+4 drive restart oracle so Scanner/Heal release evidence matches the registry contract.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
* fix(s3): tighten s3s footprint ratchet
Replace release-merge s3_error! macro calls with equivalent S3Error constructors so the s3gate migration ratchet does not grow on the PR merge tree.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
---------
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Handle cargo-nextest writing JUnit reports under the workspace target directory even when the Rust build uses CARGO_TARGET_DIR. This keeps measured Scanner/Heal evidence cases from passing the real test but failing final receipt packaging.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Keep the multi-pool evidence runner within the registry object budget while retaining deferred outage-write diagnostics for release-gate validation.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Validate the G14 multi-pool oracle field that records down-window outage PUT refusal and requires the deferred post-rejoin outage object to be accepted and checked through S3.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
test(scanner): align G14 multi-pool outage evidence
Treat the localhost multi-pool topology as whole-pool loss when the target node is down. If that topology cannot admit the outage object while the pool is offline, defer that object write until the pool rejoins and record the oracle marker instead of failing before the real crash/restart evidence runs.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Add a status mode for the Scanner/Heal Linux evidence planner so operators can identify missing or incomplete artifacts before final bundle assembly.
The status output stays plan-only and never reports release approval.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Choose the replacement disk from the target node by the presence of complete pool metadata instead of assuming the first configured drive is the scanner metadata holder. This keeps the G14 multi-set harness aligned with multi-drive EC layouts and adds clearer diagnostics when multi-pool outage writes fail closed.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Add a Scanner/Heal Linux release-evidence planner that emits a machine-readable execution manifest for the remaining measured validation lanes.
The planner can run only lightweight preflight checks and keeps plan-only output distinct from measured release evidence.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Normalize live Scanner/Heal status/outcome observations into the measured raw artifacts consumed by the G05/G06/R-D release descriptor producer.
Reject synthetic observations, incomplete required cases, and reused run/window identities before raw artifacts can enter the release bundle flow.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Harden the Scanner/Heal checkpoint restart and G14 EC evidence producers so release descriptors cannot be assembled from marked fixture, dry-run, or synthetic case inputs.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
* feat(nightly): publish packages as assets of the rolling 'nightly' release (sync from release) (#7593)
feat(nightly): publish packages as assets of the rolling 'nightly' release (#7592)
Replace the assets-branch scheme with a proper GitHub Release on
rustfs/auto-testing: a single 'nightly' release whose deb/rpm assets
are replaced in place on every build. This is the standard channel —
visible on the repo's Releases page, stable download URLs, no git
history growth (release assets live outside the repository).
- New scripts/release/publish_nightly_assets.sh: resolves-or-creates
the 'nightly' release via the REST API, deletes same-name assets,
uploads rustfs-nightly-latest.{deb,rpm}, then PATCHes the release
body with the build provenance (ref@sha, run link, sizes, SHA256).
Plain curl + python3, no gh CLI (the build fleet has none — #7586).
- The workflow step shrinks to invoking the script; full flow
exercised end-to-end against the real release with probe files
(create / upload / overwrite / download round-trip / body update).
* test(scanner): stabilize W13 release gate evidence
Treat only real raw-entry windows as replayed scanner enumeration in the restart diagnostic, so final completion rounds without raw entries are not fail-closed as raw replays.
Allow the EC8:4 multi-pool heal evidence case to select an outage object key that routes to an online pool while an entire target pool is down, preserving strict behavior for single-pool cases.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
---------
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Require Scanner/Heal G14 release bundle JSON wrappers to mirror their outer evidence and carry self-contained case artifact provenance.
Copy proof-json case artifacts into the generated G14 descriptor bundle so assembled release bundles can validate case file hashes after relocation.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Require measured G05/G06/R-D raw status, compatibility, and disposition artifacts to share run identity before producing release bundle descriptors.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Require Scanner/Heal G03 scoped ACK evidence fields to carry their own measured provenance and concrete ACK, capability, and mixed-peer observations before release-bundle gate verification.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Require Scanner/Heal release profile wrappers to carry bundled raw profile artifacts plus the matching profile cost metrics.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Replace the assets-branch scheme with a proper GitHub Release on
rustfs/auto-testing: a single 'nightly' release whose deb/rpm assets
are replaced in place on every build. This is the standard channel —
visible on the repo's Releases page, stable download URLs, no git
history growth (release assets live outside the repository).
- New scripts/release/publish_nightly_assets.sh: resolves-or-creates
the 'nightly' release via the REST API, deletes same-name assets,
uploads rustfs-nightly-latest.{deb,rpm}, then PATCHes the release
body with the build provenance (ref@sha, run link, sizes, SHA256).
Plain curl + python3, no gh CLI (the build fleet has none — #7586).
- The workflow step shrinks to invoking the script; full flow
exercised end-to-end against the real release with probe files
(create / upload / overwrite / download round-trip / body update).
Add EC8+4 multi-set and multi-pool scanner/heal release evidence coverage, including e2e registry cases, oracle checks, and a G14 descriptor assembler for measured case artifacts.
Signed-off-by: houseme <housemecn@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Configure the s3-tests harness with a local KMS key so SSE-KMS cases run in CI without relying on an external KMS service.
Also move the anonymous POST default SSE-KMS regression onto the shared local KMS test environment.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Add a W13 durable MRF evidence runner that emits measured G07, G08, and P4 JSON artifacts and release descriptors through the existing scanner/heal bundle gate.
The runner now executes the ignored MRF replay evidence test with an exact full test path, validates raw artifact kinds and gate decisions, prepares Linux tmpfs-backed ENOSPC roots for G08, and documents the Linux/long-soak boundaries.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Add a W16 Scanner/Heal evidence runner that executes the recovery-intent crash-boundary and quota-authority lanes, emits measured G04/G12 JSON artifacts, and validates the resulting single-gate release descriptors.
Wire its shell self-test into script-tests and document the release evidence entry point.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Align the Scanner/Heal release requirement registry with the release bundle gate for P4 so the closure checklist advertises MRF scale, replay cost, retained responsibility, and cleanup/GC soak evidence together.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Require Scanner/Heal release bundle JSON artifacts to carry hard-domain evidence fields for mixed-version, crash, capacity, disk-full, replica-loss, and MRF cleanup gates. This prevents a descriptor from approving a hard gate while pointing at a generic measured artifact that lacks the boundary-specific oracle fields.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Keep checksum verification output out of the command substitution that resolves the previous-release binary for the G09 runner.
Also tolerate non-GNU sha256sum in local self-tests by falling back to shasum when GNU --check support is unavailable.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Validate the selected G09 evidence lane before process-substitution case expansion so --plan-only cannot turn an unknown --test value into an empty successful plan.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
* fix(replication): close the GA blocker set from backlog#2366 (#7503)
* fix(replication): close GA blockers from backlog#2366
Implements the P1 set from the pre-GA replication audit:
- Replication rule tag filters now require every And.Tag to match, replacing
the s3s OR semantics with a local AND matcher that fails closed on a
malformed tag.
- A replicated group membership change no longer writes the group status, so
a membership update carrying the default Enabled status cannot silently
re-enable a disabled group on the peer.
- A successful IAM import schedules one collapsed full-IAM snapshot per remote
peer instead of leaving the imported entities local-only.
- A pending endpoint refresh is redriven by the heavyweight reconcile tick,
carries its own ilm-expiry override, and no longer blocks a remove that
drops every unacknowledged peer.
- Site metrics expose local replication failure totals and rolling windows;
node-level counters no longer report a constructed zero.
- set/remove-remote-target notify peer metadata caches before returning, so a
follow-up put-bucket-replication on another node sees the target.
- Adds the site-replication operations runbook, a docs index, a replication
support boundary section, and the Replication changelog section.
* fix(site-replication): resume only a locally driven endpoint refresh
The peer-side edit handler journals a pending endpoint refresh with an empty
`remote_peers` map and commits it inside the same request through
`apply_internal_peer_edit`. The reconcile tick could not tell that journal
from the coordinator's own: with no required peers it reads as complete on
sight, so the tick committed it with `edit_state` - losing the local-name
sync - and cleared it under the request that owned it, whose commit then
reported the refresh as changed and denied the coordinator the peer
acknowledgement it was waiting for.
Resume now runs only for a journal that carries the fan-out topology. A
receiver's journal stays for the coordinator to redrive with the same
refresh id, which is the path that already recovers it.
* fix(site-replication): keep an explicit disabled group status on a snapshot
Skipping the group-status write whenever an item carries members stopped a
membership change from re-enabling a disabled group, but it also silenced the
full-IAM snapshot, which always sends members together with the sender's real
status. A peer that did not have the group yet created it through
`GroupInfo::new` - enabled - so a bootstrap, a repair, or the snapshot an IAM
import now schedules handed every member of a frozen group live access there.
The madmin wire maps an unset `groupStatus` to Enabled, so only Enabled can be
a default. Disabled is always explicit and is applied again.
* fix(site-replication): schedule the import snapshot without recording a failure
`import-iam` reused the failure-recording path to queue its full-IAM
snapshot. That raises `retry_count` on every call, so three imports - the
normal shape of a bulk migration done one archive at a time - escalated a
healthy peer to `retryStats.failed` with the scheduling note shown as
`lastError`, which is exactly the signal the runbook tells operators to
repair. A full retry queue also turned a completed import into a 503.
Scheduling now only ensures the collapsed entry exists, and a failure to
schedule is logged instead of failing the request: the entities are already
imported and the reconcile pass still closes the gap.
* fix(admin): stop reporting replication failures as retries
`retries` is the minio-go counter for redeliveries, and mc prints it as such.
Filling it with the failure count claimed a redelivery that never happens: a
failed object is not retried by an event today, it waits for the scanner heal
pass. `errors` keeps the failure counters; `retries` stays zero until there is
a real redelivery to count, and the runbook now says so.
* perf(site-replication): aggregate failure windows without cloning bucket stats
`site_metrics_snapshot` went through `get_all`, which clones every bucket's
stats, and then scanned each target's sample deque twice. That deque is
bounded only by the one-hour window, so an unreachable target under load -
the case an operator polls this endpoint for - made every
`mc admin replicate status` copy the whole backlog and hold the read lock
against the failure path while doing it.
It now folds under the read lock and takes both windows in one walk. The
`max` against the serialized `last_minute` / `last_hour` snapshots is dropped:
those are stamped onto per-bucket clones elsewhere and are always zero in this
node-local cache.
* fix(site-replication): reject a conflicting ilm-expiry override on a re-run
The commit now reads the ilm-expiry override back out of the pending refresh
journal, so a second edit that asks for a different value had it dropped while
the request still reported success. Re-running without the flag keeps pinning
the recorded value - that is the documented way to redrive a stuck refresh -
but an explicit different value is now rejected instead of ignored.
* fix(admin): do not fail a remote-target write on a peer reload error
set/remove-remote-target propagated the peer metadata reload error, so a
target that was already persisted and live on this node reported a 5xx to the
client whenever one peer could not be reached. Every S3 bucket-config write
path treats that reload as best effort and only warns; these two admin
handlers now do the same, and the reason is logged with the bucket and action.
* fix(site-replication): undo every bucket a cut-short refresh rewrote
When a remove accepted on another node clears the refresh journal mid-pass,
only the bucket holding the lock at that moment had its restored target
undone. The buckets rewritten earlier in the same pass kept a target pointing
at the removed peer whenever the remove's own cleanup had already walked past
them. The undo now covers every bucket this pass rewrote, attempting all of
them so one failure does not strand the rest.
* fix(site-replication): keep replay running while an endpoint refresh is pending
A pending endpoint refresh took the whole heavyweight pass with it, so a peer
that never came back froze IAM and bucket replay to every healthy peer too -
the stall this journal's resume path was meant to end. The refresh arm now
drains the retry queue before returning; it replays per-peer deliveries
against the endpoints currently committed in state, so it is unaffected by the
edit in flight. Bucket wiring reconciliation still waits, because it rewrites
the very targets the refresh is changing, and the runbook now says so.
* test(e2e): cover the AND semantics of a two-tag replication filter
The acceptance matrix only had a single-tag rule, which matches under both AND
and OR semantics and therefore proved nothing about the filter this fix
changed. It now also carries a two-tag `And` rule - the shape
`mc replicate add --tags "k1=v1&k2=v2"` writes - and asserts that an object
with one of the two tags is not admitted while an object with both is.
No new test function, so the nightly selection digest is unchanged.
* refactor(site-replication): fold the refresh state-change error into one constructor
The endpoint-refresh work added three `s3_error!` invocation lines, which the
s3s footprint ratchet is meant to prevent. Five copies of the same
concurrent-change error now share one constructor, so the surface nets one
line smaller than main; the baseline is retightened to match.
* fix(site-replication): report a peer whose IAM snapshot waits for a repair
An escalated snapshot entry records a deletion a snapshot cannot replay, so
only a repair settles it and the marker must survive. Scheduling an import
snapshot therefore leaves that peer's entry alone - and now says so, instead
of returning success while nothing was scheduled for it.
* docs(operations): state the group-status and escalation convergence limits
Two boundaries the fixes in this branch make load-bearing: a membership change
never carries an enable, so a group disabled on one site only has to be
re-enabled there explicitly; and a peer holding an escalated IAM entry does
not receive a scheduled snapshot, including the one a bulk import schedules,
until a repair settles it.
* fix(ci): bind performance runs to selected inputs (#7512)
* fix(targets): reject trailing batch items (#7508)
* test(scanner): emit G09 release bundle gate evidence
Write a bundle-ready G09 gate descriptor from the Linux upgrade evidence runner and validate the single G09 gate with the shared release-bundle rules without approving the full Scanner/Heal release.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
---------
Co-authored-by: 唐小鸭 <tangtang1251@qq.com>
Co-authored-by: Zhengchao An <anzhengchao@gmail.com>
Co-authored-by: cui fliter <imcusg@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
* fix(replication): close the GA blocker set from backlog#2366 (#7503)
* fix(replication): close GA blockers from backlog#2366
Implements the P1 set from the pre-GA replication audit:
- Replication rule tag filters now require every And.Tag to match, replacing
the s3s OR semantics with a local AND matcher that fails closed on a
malformed tag.
- A replicated group membership change no longer writes the group status, so
a membership update carrying the default Enabled status cannot silently
re-enable a disabled group on the peer.
- A successful IAM import schedules one collapsed full-IAM snapshot per remote
peer instead of leaving the imported entities local-only.
- A pending endpoint refresh is redriven by the heavyweight reconcile tick,
carries its own ilm-expiry override, and no longer blocks a remove that
drops every unacknowledged peer.
- Site metrics expose local replication failure totals and rolling windows;
node-level counters no longer report a constructed zero.
- set/remove-remote-target notify peer metadata caches before returning, so a
follow-up put-bucket-replication on another node sees the target.
- Adds the site-replication operations runbook, a docs index, a replication
support boundary section, and the Replication changelog section.
* fix(site-replication): resume only a locally driven endpoint refresh
The peer-side edit handler journals a pending endpoint refresh with an empty
`remote_peers` map and commits it inside the same request through
`apply_internal_peer_edit`. The reconcile tick could not tell that journal
from the coordinator's own: with no required peers it reads as complete on
sight, so the tick committed it with `edit_state` - losing the local-name
sync - and cleared it under the request that owned it, whose commit then
reported the refresh as changed and denied the coordinator the peer
acknowledgement it was waiting for.
Resume now runs only for a journal that carries the fan-out topology. A
receiver's journal stays for the coordinator to redrive with the same
refresh id, which is the path that already recovers it.
* fix(site-replication): keep an explicit disabled group status on a snapshot
Skipping the group-status write whenever an item carries members stopped a
membership change from re-enabling a disabled group, but it also silenced the
full-IAM snapshot, which always sends members together with the sender's real
status. A peer that did not have the group yet created it through
`GroupInfo::new` - enabled - so a bootstrap, a repair, or the snapshot an IAM
import now schedules handed every member of a frozen group live access there.
The madmin wire maps an unset `groupStatus` to Enabled, so only Enabled can be
a default. Disabled is always explicit and is applied again.
* fix(site-replication): schedule the import snapshot without recording a failure
`import-iam` reused the failure-recording path to queue its full-IAM
snapshot. That raises `retry_count` on every call, so three imports - the
normal shape of a bulk migration done one archive at a time - escalated a
healthy peer to `retryStats.failed` with the scheduling note shown as
`lastError`, which is exactly the signal the runbook tells operators to
repair. A full retry queue also turned a completed import into a 503.
Scheduling now only ensures the collapsed entry exists, and a failure to
schedule is logged instead of failing the request: the entities are already
imported and the reconcile pass still closes the gap.
* fix(admin): stop reporting replication failures as retries
`retries` is the minio-go counter for redeliveries, and mc prints it as such.
Filling it with the failure count claimed a redelivery that never happens: a
failed object is not retried by an event today, it waits for the scanner heal
pass. `errors` keeps the failure counters; `retries` stays zero until there is
a real redelivery to count, and the runbook now says so.
* perf(site-replication): aggregate failure windows without cloning bucket stats
`site_metrics_snapshot` went through `get_all`, which clones every bucket's
stats, and then scanned each target's sample deque twice. That deque is
bounded only by the one-hour window, so an unreachable target under load -
the case an operator polls this endpoint for - made every
`mc admin replicate status` copy the whole backlog and hold the read lock
against the failure path while doing it.
It now folds under the read lock and takes both windows in one walk. The
`max` against the serialized `last_minute` / `last_hour` snapshots is dropped:
those are stamped onto per-bucket clones elsewhere and are always zero in this
node-local cache.
* fix(site-replication): reject a conflicting ilm-expiry override on a re-run
The commit now reads the ilm-expiry override back out of the pending refresh
journal, so a second edit that asks for a different value had it dropped while
the request still reported success. Re-running without the flag keeps pinning
the recorded value - that is the documented way to redrive a stuck refresh -
but an explicit different value is now rejected instead of ignored.
* fix(admin): do not fail a remote-target write on a peer reload error
set/remove-remote-target propagated the peer metadata reload error, so a
target that was already persisted and live on this node reported a 5xx to the
client whenever one peer could not be reached. Every S3 bucket-config write
path treats that reload as best effort and only warns; these two admin
handlers now do the same, and the reason is logged with the bucket and action.
* fix(site-replication): undo every bucket a cut-short refresh rewrote
When a remove accepted on another node clears the refresh journal mid-pass,
only the bucket holding the lock at that moment had its restored target
undone. The buckets rewritten earlier in the same pass kept a target pointing
at the removed peer whenever the remove's own cleanup had already walked past
them. The undo now covers every bucket this pass rewrote, attempting all of
them so one failure does not strand the rest.
* fix(site-replication): keep replay running while an endpoint refresh is pending
A pending endpoint refresh took the whole heavyweight pass with it, so a peer
that never came back froze IAM and bucket replay to every healthy peer too -
the stall this journal's resume path was meant to end. The refresh arm now
drains the retry queue before returning; it replays per-peer deliveries
against the endpoints currently committed in state, so it is unaffected by the
edit in flight. Bucket wiring reconciliation still waits, because it rewrites
the very targets the refresh is changing, and the runbook now says so.
* test(e2e): cover the AND semantics of a two-tag replication filter
The acceptance matrix only had a single-tag rule, which matches under both AND
and OR semantics and therefore proved nothing about the filter this fix
changed. It now also carries a two-tag `And` rule - the shape
`mc replicate add --tags "k1=v1&k2=v2"` writes - and asserts that an object
with one of the two tags is not admitted while an object with both is.
No new test function, so the nightly selection digest is unchanged.
* refactor(site-replication): fold the refresh state-change error into one constructor
The endpoint-refresh work added three `s3_error!` invocation lines, which the
s3s footprint ratchet is meant to prevent. Five copies of the same
concurrent-change error now share one constructor, so the surface nets one
line smaller than main; the baseline is retightened to match.
* fix(site-replication): report a peer whose IAM snapshot waits for a repair
An escalated snapshot entry records a deletion a snapshot cannot replay, so
only a repair settles it and the marker must survive. Scheduling an import
snapshot therefore leaves that peer's entry alone - and now says so, instead
of returning success while nothing was scheduled for it.
* docs(operations): state the group-status and escalation convergence limits
Two boundaries the fixes in this branch make load-bearing: a membership change
never carries an enable, so a group disabled on one site only has to be
re-enabled there explicitly; and a peer holding an escalated IAM entry does
not receive a scheduled snapshot, including the one a bulk import schedules,
until a repair settles it.
* fix(ci): bind performance runs to selected inputs (#7512)
* test(e2e): add G09 upgrade evidence runner
Add a Linux x86_64 runner that downloads the pinned previous release, builds the current RustFS binary, runs the mixed-version and rollback upgrade compatibility lanes, and verifies the required Scanner/Heal G09 raw evidence artifacts.
Document the runner and add a shell self-test for help, dry-run, SHA validation, and non-empty artifact directory guards.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
---------
Co-authored-by: 唐小鸭 <tangtang1251@qq.com>
Co-authored-by: Zhengchao An <anzhengchao@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>