* fix(heal): rebuild truncated xl.meta from healthy quorum
* fix(test): pass topology to heal overlap RPC regression
* fix(test): drive heal admission alongside partial PUT
Poll the partial PUT and its mock heal receiver together, bound their handshake, and retain the existing repair-scope assertions.
* fix(test): prepare durable MRF fixtures and Linux heal stack
* fix(test): drive tier cleanup recovery after deferred attempts
---------
Co-authored-by: Hauser <housemecn@gmail.com>
Allow a confirmed scanner root data-usage CAS write to prove its own publication when the follow-up root readback cannot provide a proof. Keep AlreadyDurable and all stale or companion paths on the existing readback-only proof boundary.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
* fix(heal): persist and retry partial-write repairs
* test(heal): pass topology to overlap RPC tests
Use the existing coordinator endpoint fixture for the three overlap-test
calls to the endpoint-aware heal control executor. This repairs the E0061
test-build failure inherited from the release base.
Reuse the heal-control endpoint fixture in the overlap receipt regression so the test matches the updated execution helper signature and selector validation boundary.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Co-authored-by: hehutu <hehutu@gmail.com>
The overnight run after #7708 proved the security verdict greps still
counted zero: real verdict lines are '\e[1;31m[FAIL]\e[0m STS-105 ...'
— the reset escape sits between the tag and the case id, and the
pattern only tolerated escapes before the tag. Allow escapes on both
sides; the fixture now emits the reset too, mirroring the real suite.
The tier case table is produced by rustfs_tier_report.py and its rows
lead with the topology column, so the case-ID-first table grep matched
nothing ('Product result: 0 passed, 0 failed'). Parse the PASS/FAIL
counts from the '## Case Summary' bullets the report always emits,
falling back to a topology-aware table grep.
First full serial pass with the green-on-case-failure semantics
(run 34693745171 / 34695021651) exposed three report-layer defects:
security: the report step referenced LOG_FILE, which is undefined in
this workflow (set -u killed the step before writing report.md), and
the verdict greps could not match the ANSI-escaped [PASS]/[FAIL] tags
in the real suite log. Point it at the artifacts suite.log, allow any
number of color escapes before the verdict tag, and count [SKIP]
lines separately (45 passed, 6 failed, 3 skipped was reported as an
unbound-variable crash).
pool: warp is stopped early (SIGINT) at the storage threshold and
only writes its final report on a clean exit, so an empty warp.log is
the expected shape of a healthy run - require its presence, not its
size. A mid-script die() abort or a FAIL step verdict must also turn
the validator red now that the run step is continue-on-error.
tier: add the standard 'Product result: N passed, M failed' summary
line computed from the case table, matching the other suites.
The contract test fixture previously injected LOG_FILE into the
environment and wrote verdict lines without ANSI escapes, which hid
both real-world defects; the fixture now mirrors the real suite
(stdout+tee with color tags) and asserts the pass/fail/skip counters.
Remove the superseded ServiceUnavailable-only PUT helper after the G14 outage PUT probe switched to the shared retry classifier.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Treat SlowDownRead as a retryable outage PUT probe response in the G14 multi-pool runner, and label terminal candidate failures with stage context.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Treat SlowDownRead as a bounded retryable deferred PUT response after the target pool rejoins, and label terminal deferred outage PUT failures with stage context.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
A failing product case used to turn the whole workflow red, so the run
conclusion carried no signal beyond 'something failed' and the report
was suppressed. New semantics across the functional suites:
- Suite steps run with continue-on-error: the outcome is still recorded
for the report and the backlog issue manager (security/tier already
carried the flag).
- Generate report always publishes the full per-case table plus a
'Product result: N passed, M failed' summary, and its exit gate is
harness health: red only when the suite never reached case level (no
case verdicts), failed wholesale (zero passes, >=3 failures), or was
cancelled/skipped. performance is unchanged (parked).
- tier's structured gate no longer fails on case failures; it keeps red
for evidence-init and missing-gate-result breakdowns.
- pool/performance keep their existing red sources (install/benchmark).
Workflow contract tests updated to the new exit semantics: the security
report matrix keys green off the suite outcome, the evidence matrix
expects green for failure outcomes with recorded case rows (except
performance), the heal staged-rerun block expects the per-step table to
always publish, and run steps are now required to carry
continue-on-error.
Verified locally: actionlint clean; test_security_workflow.py 21/21.
RUSTFS_POOL_NODE_ENDPOINTS has no secret/var configured, so the suite
ran with the workflow's inline 3-endpoint fallback and the explicit
--node-endpoints flag overrode the script default fixed in
rustfs/auto-testing#61 — every dispatch died at startup with 'must
provide at least 4 direct node endpoints' (run 34677538650). Add
rustfs-node4 to the fallback; explicit secret/var still wins.
Allow the Scanner/Heal interruption oracle to retry transient retryable GET failures after replacement recovery has converged. The readback still verifies exact object bytes and keeps a bounded timeout, so permanently unreadable objects continue to fail the case.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Add a read-only descriptor ledger mode to the Scanner/Heal Linux evidence planner so release operators can separate current-head measured descriptors from old-head measured artifacts and case-level inputs before final bundle assembly.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
* fix(storage): alias legacy meta bucket over internode rpc
Retry read-only internode RPC metadata access from legacy .minio.sys to .rustfs.sys when mixed-version peers report missing metadata during rolling upgrades.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
* test(ecstore): target durable ILM receipt quorum fixture
Use the actual durable ILM receipt object path when taking target disks offline so the test exercises receipt write quorum instead of whichever set owns the source record path.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
---------
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
* ci: manage backlog issues by signal instead of per-run filing
The suite workflows used to file one backlog issue per failed run
(dedup was by run ID, which never matched), so issues accumulated
without bound. Replace the inline filing step in every suite workflow
(s3, kms, tier, storage, heal, pool, security, replication, upgrade,
performance) with a single call to
auto-testing/scripts/issue_manager.py, which:
- dedups by signal: failing cases are searched among open issues by
label (suite category + case ID); covered cases become a coalesced
comment on the existing issue, only uncovered cases file a new one
- labels new issues with functional-test, the suite category, one
label per failing case ID (lazily created), and env for
bootstrap-class failures (no cases ran, wholesale failure, or
404/ssh/clone/dpkg signatures in the log)
- closes open issues of the suite after a fully green run, citing the
run as evidence; cancelled runs never file or close anything
The step is skipped cleanly when auto-testing (private checkout) does
not contain the manager, or when PF_TESTING_GH_TOKEN is unset.
* fix(ci): satisfy actionlint and workflow contract tests for the manager step
- heal and performance workflows have no rustfs_version dispatch input;
referencing `${{ inputs.rustfs_version }}` in the manager step failed
actionlint's expression type check. Their package source now resolves
from package_url with the nightly fallback.
- scripts/test_security_workflow.py pinned the removed inline filing
step. The wiring assertions now pin the manager step (manager path +
per-suite report argument), and the evidence/stale-file tests assert
the skip contract instead: without the private auto-testing checkout
present, the step exits 0, publishes nothing, and leaves stale
evidence untouched.
Verified locally: actionlint clean, shellcheck clean,
test_security_workflow.py 21/21.
The rustfs_version dispatch input defaulted to 1.0.0-rc.4-preview.1,
which shadowed the nightly fallback and started 404ing once that
release was deleted. Manual dispatches with no inputs now fall through
to the nightly package (same contract the pool suite already has);
passing rustfs_version or package_url still pins the build exactly as
before. Chain (repository_dispatch) runs are unaffected: the inputs
context is empty there, so they always used the nightly fallback.
Adds a fault-tolerance suite that verifies read/write behavior under
drive and node loss against the erasure-coding contract and snapshots
health-endpoint responses at every degradation tier. Scenarios are
derived from product source (default_parity_count, erasure set sizing):
- A: single-node 4 drives (EC:2, read quorum 2): hide 1/2/3 drives
- B: multi-node 4x1 (one set of 4, EC:2): stop 1/2/3 nodes
- C: multi-node 4x4 (one set of 16, EC:4, read quorum 12): 1 node down
lands exactly on the read-quorum boundary; 2 nodes down breaks it
- C2: multi-node 4x4 with RUSTFS_STORAGE_CLASS_STANDARD=EC:8 (read
quorum 8, write quorum 9, lock majority 9): 2 nodes down puts reads
inside the reported divergence window (read quorum met while the
lock majority is broken)
By default a reads-refused-despite-met-read-quorum observation is
reported as known-divergence without failing the suite; the strict
input escalates it. Chain order becomes:
upgrade -> s3 -> kms -> tier -> storage -> heal -> pool -> security ->
replication -> fault-tolerance -> performance.
Covers the findings of the 2026-09 external degradation report
(reads at read quorum, health-endpoint readiness truthfulness,
degradedReasons capture) as automated regression probes.
Size the scanner collector wait budget from the requested sample count and interval instead of using a fixed 120 second timeout. This preserves the full duration/60+1 telemetry sample set for measured ABBA runs.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
* fix(heal): reuse API error boundary for decoding failures
* Update heal.rs
Signed-off-by: houseme <housemecn@gmail.com>
* test(ecstore): synchronize batch cleanup metadata reads
Acquire object read locks while observing unversioned and explicit-version batch cleanup, then release them before waiting for progress. This prevents snapshots from spanning per-disk marker removal while retaining the existing quorum, timeout, and remote delete count assertions.
Validation: cargo fmt --all --check and git diff --check passed. The focused nextest test passed 20 stress iterations each with test-util and test-util,rio-v2, with retries disabled.
---------
Signed-off-by: houseme <housemecn@gmail.com>
Co-authored-by: houseme <housemecn@gmail.com>
* test(e2e): target multi-set outage heal candidate
Require the outage write used by EC8+4 multi-set root-heal evidence to miss the same erasure index owned by the selected replacement drive. This avoids accepting a candidate from a different set and turning a valid heal into a false negative.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
* fix(e2e): satisfy G14 heal lint gates
Remove clippy-only noise from the G14 multi-set heal evidence test and align the admin route policy inventory with the registered heal catch-all route.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
* update
---------
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
* fix(e2e): record EC8+4 drive restart set size
Include the erasure set drive count in the distributed EC8+4 drive restart oracle so Scanner/Heal release evidence matches the registry contract.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
* fix(s3): tighten s3s footprint ratchet
Replace release-merge s3_error! macro calls with equivalent S3Error constructors so the s3gate migration ratchet does not grow on the PR merge tree.
Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
---------
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Handle cargo-nextest writing JUnit reports under the workspace target directory even when the Rust build uses CARGO_TARGET_DIR. This keeps measured Scanner/Heal evidence cases from passing the real test but failing final receipt packaging.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Keep the multi-pool evidence runner within the registry object budget while retaining deferred outage-write diagnostics for release-gate validation.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Include optional outage-write diagnostics in G14 evidence oracles so deferred multi-pool outage writes bind their down-window refusals and post-rejoin acceptance.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
test(e2e): keep multi-set restart graceful
Include the EC8+4 multi-set restart scenario in the graceful interruption lane so the unclean-shutdown marker assertion matches the scenario semantics.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Use the pool-wide namespace lock client domain for every set in a pool so degraded EC8+4 multi-set writes are gated by node-level lock quorum instead of the narrower per-set endpoint host slice.
Keep namespace-lock domain deduplication tied to both the pool namespace and shared clients, and add regression coverage for three-locker degraded writes plus cross-pool domain separation.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
fix(ecstore): admit metadata snapshots at read quorum
Allow guarded bucket metadata snapshot existence checks to use read quorum so degraded erasure sets can continue Object Lock snapshot reads without weakening bucket mutation or object write quorum.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Validate the G14 multi-pool oracle field that records down-window outage PUT refusal and requires the deferred post-rejoin outage object to be accepted and checked through S3.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
test(scanner): align G14 multi-pool outage evidence
Treat the localhost multi-pool topology as whole-pool loss when the target node is down. If that topology cannot admit the outage object while the pool is offline, defer that object write until the pool rejoins and record the oracle marker instead of failing before the real crash/restart evidence runs.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Add a status mode for the Scanner/Heal Linux evidence planner so operators can identify missing or incomplete artifacts before final bundle assembly.
The status output stays plan-only and never reports release approval.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Choose the replacement disk from the target node by the presence of complete pool metadata instead of assuming the first configured drive is the scanner metadata holder. This keeps the G14 multi-set harness aligned with multi-drive EC layouts and adds clearer diagnostics when multi-pool outage writes fail closed.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Add a Scanner/Heal Linux release-evidence planner that emits a machine-readable execution manifest for the remaining measured validation lanes.
The planner can run only lightweight preflight checks and keeps plan-only output distinct from measured release evidence.
Co-authored-by: zhi22915 <qiuzgang@gmail.com>