Commit Graph

6936 Commits

Author SHA1 Message Date
Chris 791af1888f fix(connect): preserve site replication S3 key path (#8300) 2026-10-02 05:43:24 +08:00
Chris 50e6c907f9 fix(s3): bound multipart copy admission and HTTP write buffers (#8288) 2026-10-01 21:51:27 +08:00
hector 1efa86b8a3 ci: locate the auto-testing checkout for lanes that run evidence from a subdirectory (#8279)
## Related Issues

Follow-up to #8229.

## Summary of Changes

Locate the private auto-testing checkout from the lane root or the workspace root so the nested security checkout can record functional-chain evidence.

## Verification

The exact PR head b66129ab9f passed one mechanical correctness and simplicity review, nine real-Git layout and provenance checks, and sixteen existing evidence/envelope tests. The baseline sibling layout failed with git exit 128; the corrected layout succeeded while revision mismatches and missing checkouts stayed rejected. Current PR checks are completed with successful or skipped conclusions, including the aggregate.

## Impact

Both lane and private-script revision checks remain intact. No time limits, assertions, production behavior, or evidence validation requirements change. The synthetic layout checks do not execute the actual scheduled security suite; integrated main CI and release acceptance remain separate gates.

## Additional Notes

Approved review 5374559393 is bound to the exact head above. Reverting the single-file change restores the previous checkout lookup.
2026-10-01 11:24:37 +08:00
Chris a33dd896e2 fix(ci): restore E2E membership and pagination timeouts (#8281) 2026-10-01 09:00:02 +08:00
AL 6796382cee fix(scanner): validate checkpoints against global cycle fence (#8278)
## Related Issues

Related to rustfs/backlog#2701.

## Summary of Changes

Route scanner checkpoint cycle and leader validation through the global store while retaining the owning set for cache persistence, CAS revisions and publication admission.

## Verification

Two independent final-diff source reviews found no issues across correctness, concurrency and durability, test coverage, compatibility, performance and simplicity on head `8b8fe51d092090b053f552ae283960e2e306be33`. Root approval `5373624714` is bound to that exact head. Regression tests cover real two-pool routing, stale fences, post-save rejection and CAS conflicts; their reported local execution belongs to the PR author, not this merge operation. Current required CI remains pending, and this authorized admin squash does not establish CI or runtime acceptance.

## Impact

Restores checkpoint progress when global cycle and leader state differ from a set-scoped view. No format, retry, timeout, assertion or scanner-policy changes are introduced by this diff. The three prior main scanner failures remain unproved repaired.

## Additional Notes

Full validation must run on the resulting exact main revision. Reverting this patch restores the earlier set-scoped fence lookup and its checkpoint rejection behavior.
2026-10-01 08:52:25 +08:00
overtrue 99a2af8687 docs(release): validate candidates on release branch 2026-10-01 08:37:30 +08:00
Chris 2806a80c91 fix: prevent metadata lock reentry and stabilize CI fixtures (#8257)
* fix(test): establish a writable previous-release upgrade baseline

* fix(ci): reserve capacity for durable admin fixtures

* fix(ci): group durable IAM state fixtures by resource needs

* ci: run E2E doctests with the E2E dependency graph

* fix(ci): separate fixture startup from transport deadlines

* fix(ci): bound pagination after seeding and revisit restored copies

* fix(ci): make filesystem fixture timing deterministic

* fix(ecstore): avoid metadata lock reentry during internal mutations

* test: align recovery fixtures with durable ownership contracts
2026-09-30 20:57:07 +08:00
Chris 8655c38f1d fix(rio): preserve bounded retries after peer EOF (#8272) 2026-09-30 20:18:28 +08:00
Chris 88d03d8199 fix(storage): prevent metadata deadlocks and abandoned writes (#8268)
## Related Issues

N/A

## Summary of Changes

Reject system-bucket incarnation lookups before entering the pool metadata owner. Keep each selected scanner backlog publication cohort in one task so waiter cancellation cannot abandon its remaining serialized conditional writes. Bind repair and replay fixtures to persisted bucket incarnations.

## Verification

Head f6d44603fc received an approval from houseme in review 5365588011. This merge does not add a local runtime validation claim. Full main CI remains a separate publication gate.

## Impact

System metadata writes avoid recursive pool locking. Publication retains per-replica conditional writes and stops a cancelled caller from advancing to its next publication phase. Fixture changes supply the identities required by existing repair admission rules.

## Additional Notes

Reverting this change restores the previous behavior.
2026-09-30 19:49:24 +08:00
Hauser eb350ad20d fix(ecstore): purge empty recursive bucket orphans (#8261)
## Related Issues

Related to rustfs/backlog#2688.

## Summary of Changes

Reclaim metadata-less orphan directory trees after a complete, first-page, empty recursive root listing in an authoritative never-versioned bucket. Both listing implementations use the same fail-closed cleanup decision.

## Verification

Head b4a9120173 received an approval from loverustfs in review 5365460152. This merge does not add a local runtime validation claim. Full main CI remains a separate publication gate.

## Impact

Empty recursive root listings can remove otherwise unaddressable physical residue. Versioned buckets and populated listings remain excluded; unreadable disks, object metadata, unknown files, and uncommitted data prevent cleanup.

## Additional Notes

Reverting this change restores the previous orphan-directory behavior.
2026-09-30 19:24:32 +08:00
KarlEmm e573fb087b ci(helm): sync README to root of helm repository on release (#8263)
Include helm/rustfs/README.md in the uploaded helm-package artifact so
publish-helm-package updates the root README.md in rustfs/helm.

The build-helm-package job copies helm/README.md into helm/rustfs/
before running helm package, but upload-artifact previously uploaded
only helm/rustfs/*.tgz. Because both files share the helm/rustfs/
parent directory, actions/upload-artifact places both the chart
archive and README.md at the artifact root, and download-artifact
extracts README.md into the root of rustfs/helm alongside the tarball.
2026-09-30 15:57:24 +08:00
Chris 099a408d8b test(ecstore): isolate dispatch shutdown recovery fixtures (#8265) 2026-09-30 15:56:48 +08:00
Chris 5c3a2ff277 fix(ci): reserve capacity for profile cancellation probe (#8259) 2026-09-30 15:46:20 +08:00
Chris bb04355a3d fix(heal): retry contended local metadata CAS (#8267) 2026-09-30 15:45:59 +08:00
Chris fab2d7e144 fix(scanner): serialize pause backlog replica publication (#8264) 2026-09-30 15:45:37 +08:00
Hauser d60dfbb826 fix(heal): finalize durable MRF legacy responsibility lifecycle (#8254)
## Related Issues

Resolves the corrected lifecycle fixture findings in PR #8254.

## Summary of Changes

Preserve separate incarnation-bound durable MRF responsibilities and their lifecycle audit records, and align replay fixtures with their checkpoint identities.

## Verification

Two independent source reviews and changed-delta reviews are complete. The final review approved 607a0b9bb0 with no remaining supported findings. Local runtime claims were not independently reproduced for this pull request.

## Impact

Keeps storage generation fences and retained responsibility semantics. Test-capacity reservations preserve existing deadlines and assertions.

## Additional Notes

Squash merge of the currently approved fix under the authorized CI-bypass exception. Main CI and release acceptance remain required.
2026-09-30 14:02:32 +08:00
Chris 6bc2e10bc1 fix(tier): keep recovery worker after startup reload failure (#8260) 2026-09-30 12:05:26 +08:00
Chris 3268c42e00 fix(ci): stabilize registration and rolling upgrade readiness (#8244)
## Related Issues

Follow-up to #8233.

## Summary of Changes

Registration runtime fixtures timed out in the workspace CI lane while the same cases passed in the feature lanes. Reserve nextest capacity for the exact `connect_registration` binary, as already done for related inventory and drive fixtures. Keep its existing deadlines, internal concurrency, assertions and zero-retry policy; report the last watch status on failure.

Rolling upgrades could pass `ListBuckets` readiness while restarted peers still lacked write quorum. Before each mixed-version phase, probe writes through every node outside the asserted workload prefixes. A shared 30-second deadline includes requests and sleeps; only HTTP 503 with `ServiceUnavailable` is retryable, with SDK retries disabled for these probes. The actual compatibility writes, reads, multipart operations and listing assertions remain unchanged.

Add four fast regression tests to the existing PR smoke profile, with matching exclusion from the full profile. No new workflow or job is introduced.

## Verification

- `cargo nextest run --locked --profile ci -p rustfs --lib --test connect_registration --test-threads 4 -E 'binary(/^connect_registration$/) | test(=connect::diagnostics::trace_runtime::tests::local_runtime_rejects_non_private_state)' --no-tests fail --status-level pass --final-status-level fail` passed 23/23 twice in 5.088s and 5.097s: 22 macOS registration tests plus one unrelated control. JUnit intervals confirm capacity reservation; no retries or test-process leaks occurred. The Linux-only inventory case remains for CI.
- `cargo nextest run -p e2e_test --lib --profile ci -E 'test(upgrade_write_readiness_tests)' --no-fail-fast --no-tests fail` passed 4/4 twice in 1.068s and 1.072s, with zero retries. Regressions cover metadata readiness followed by write unavailability, recovery through every writer, immediate permanent-error failure despite an incoming SDK retry configuration, and deadlines for repeated 503s and stalled requests. The original failure is recorded in [the mixed-version upgrade job](https://github.com/rustfs/rustfs/actions/runs/36569700716/job/109415806473).
- `cargo fmt --all --check`, `git diff --check`, `python3 scripts/check_test_wiring.py` and compiled smoke/full membership checks passed. Smoke membership changes from 188 to 192 by adding exactly these four tests; full membership is unchanged. The expected Linux smoke digest was derived from the actual prior Linux listing plus those four platform-independent additions and still requires confirmation by this PR's CI.

Local verification covers the exact source committed in `9b2ea836317ed035a91d1fc9fcf725f70c3098e2` on main `380e98a42cb4fcd0994fed79b30c2c7605deb0bc`. An independent final-diff correctness and reliability review found no findings. Fresh Linux workspace and real mixed-version upgrade runs are required before treating the remediation as fully verified; local fake-target tests do not establish that result.

## Impact

Test scheduling and readiness only; no production behavior, API, dependency, test deadline or compatibility assertion changes. Reserving capacity serializes registration fixture processes within a nextest run. The bounded readiness probes may add startup time while peer write health converges; permanent errors still fail immediately.

## Additional Notes

Rollback by reverting this PR. Existing CI restructuring from #8233 is independent of these follow-up fixes.
2026-09-30 10:06:09 +08:00
Chris 9e33d54269 fix(release): package stable preview tags without updating channels (#8256) 2026-09-30 09:49:27 +08:00
Hauser 498080dec4 fix(storage): default new bucket durability to strict (#8253)
* fix(storage): align new bucket durability with strict defaults

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(scanner): defer failure recovery until scan completion

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-30 09:23:32 +08:00
Hauser 296854c3d2 test(heal): cover degraded MRF replay after deletion (#8249)
* test(heal): cover deleted versions during MRF replay

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* test(heal): exercise degraded MRF replay after delete

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* test(heal): route quorum errors through test storage API

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-30 09:14:02 +08:00
Jason Kossis 56468fd553 fix(scanner): resume checkpoints across cycle and leader changes (#8252) 2026-09-30 08:18:20 +08:00
Chris 50aa482136 fix(connect): add non-Unix health runtime fallback (#8247) 2026-09-30 08:10:10 +08:00
Chris e870a6d25b fix: separate CPU profile stack lifetime from accumulator (#8250) 2026-09-30 01:36:14 +08:00
Chris 5851d9eb54 fix: filter CPU profile samples to reviewed executable symbols (#8248) 2026-09-30 01:06:24 +08:00
Chris 1e7065101d feat: Capture offline CPU profiles in serving process (#8245)
feat: capture offline CPU profiles in serving process
1.0.1-preview.13
2026-09-29 22:37:10 +08:00
Chris 380e98a42c fix(test): synchronize top-disk fixture writes with sampling (#8243) 2026-09-29 21:07:45 +08:00
Chris a88225d208 fix(ci): reduce duplicate work and preserve reliable test failures (#8233)
* fix(ci): reduce duplicate work and preserve reliable test failures

* fix(ci): retain protocol evidence and repair stale test fixtures

* test(connect): honor parent deadline during API fixture readiness

* fix(ci): reserve IO capacity for state writer proofs

* test(connect): align RPC fixtures with service capture contracts

* test(connect): cover pinned service capture failures
2026-09-29 21:06:45 +08:00
Chris 1cca0dd25c feat(connect): select offline key for client performance (#8242) 2026-09-29 20:49:59 +08:00
Chris b0a54f0e5c feat(connect): capture offline network performance in server (#8240)
* feat(connect): capture offline network performance in server

* fix(connect): use storage facade in RPC test

* fix(connect): keep RPC test body behind storage facade
2026-09-29 19:41:43 +08:00
Chris 83a2d8b02d feat(connect): capture RPC activity in offline server runtime (#8237) 2026-09-29 18:16:54 +08:00
Chris f661651aad feat(connect): capture API activity in offline server runtime (#8234) 2026-09-29 17:18:26 +08:00
Chris 9298991d14 fix(scanner): restart a finished mixed sweep under its requested plan (#8231)
A bucket sweep that ends mixed clears its position and records the finishing cycle's plan as started. The next cycle requests a new plan whenever the bucket was written in between, so a fresh sweep inherited a stale started plan and ended mixed again. On a continuously written bucket no sweep could ever certify and the census kept the old root.

Start a fresh verification sweep under the requested plan when no durable position remains. Resumed sweeps keep their started plan, so a clean tail still cannot certify an old prefix.

Refs #7108
2026-09-29 16:48:58 +08:00
Chris 3c7d85420c fix(test): control clocks in Connect network fixtures (#8228) 2026-09-29 16:48:34 +08:00
Chris 6a7880c9fc feat(connect): capture offline lock metrics in server (#8230) 2026-09-29 16:29:07 +08:00
hector 0e58bd09d0 fix(ci): run security chain evidence from the lane's own checkout (#8229)
The security workflow checks the repository out into rustfs-repo/ (#7212)
but its chain evidence steps still invoke
scripts/functional_chain_evidence.py relative to the workspace root, which
on the persistent shared runner resolves to a stale checkout left by another
job. The evidence gate then compares that checkout's HEAD against the chain
workflow_sha and rejects the lane before any case runs
("lane checkout differs from chain workflow source").

Run the evidence script from the lane's own checkout so ROOT resolves to
rustfs-repo/, whose HEAD is exactly the chain-pinned workflow_sha.
2026-09-29 16:09:31 +08:00
Hauser fcb60e7322 fix(rpc): preserve walk_dir missing-path errors (#8226)
## Related Issues

rustfs/rustfs#8217 and rustfs/backlog#2683.

## Summary of Changes

Reviewed 539ed275db against d2175d1e1e. No findings. Capable callers requesting missing-path reporting recover explicit pre-output FileNotFound or VolumeNotFound errors; partial-output and other failures continue to fail the stream.

## Verification

Two independent source reviews completed across all lenses below; the frozen diff matches Git and passes diff whitespace checks. Traced the authenticated handler, bounded first-chunk preflight, cancellation ownership, HTTP status/token classification and quorum error handling. The new tests cover ordinary streaming, typed missing errors, byte preservation and mixed missing/I/O failures.

Author-reported focused tests and Docker reproduction were not rerun or independently audited in this review. The reported Docker comparison predates the final report_notfound-only gate; its scope is not distributed acceptance of the final head.

## Impact

Correctness: no findings; missing errors remain typed only before output. Security/trust: no findings; authentication, body digest and operation/status/token restrictions remain. Compatibility: no findings; legacy clean EOF remains, with the documented old-server limitation. Concurrency/durability: no findings; receiver ownership cancels an abandoned producer. Simplicity: no findings. Coverage: no blocking gap identified. Performance: no findings; preflight retains one bounded chunk and ordinary report_notfound=false streams start immediately.

## Additional Notes

This review does not establish that the patch fixes any current main CI failure. Full main CI and release acceptance remain separate gates.
2026-09-29 15:29:59 +08:00
Hauser 0a61f4a81e fix(heal): park healthy legacy MRF intents (#8225)
## Related Issues

rustfs/backlog#2682 and rustfs/rustfs#8192.

## Summary of Changes

Reviewed ca970f55ec against d2175d1e1e. No blocking code findings. The hold requires a completed healthy legacy check and an exact incarnation, lease, object, version and scope; it preserves the durable journal and rechecks on restart.

## Verification

Two independent source reviews completed, covering the lenses below. The frozen diff matches Git and passes diff whitespace checks. Reported Cargo and Docker results were not rerun or independently audited in this review.

One factual correction to the PR description: `legacy_sigkill_replay_repairs_without_releasing_unverified_responsibility` uses the default four-member fixture, with three replicas before restoration and four afterward (`mrf_partial_write_test.rs:559,567`). That named regression exercises the production manager path, but is not EC12+4. Other tests in the file use sixteen disks.

## Impact

Correctness: no findings; proofless legacy health never becomes a verified receipt. Security/trust: no findings; exact identity fences remain. Compatibility: no findings; journal encoding remains unchanged. Concurrency/durability: no findings; replay retains one checkpointed owner and replacement generations become retryable. Simplicity: no findings. Coverage: no blocking gap identified, with the test-scope correction above. Performance: no findings; held entries leave the retry index without adding per-admission queue scans.

## Additional Notes

The current main journal failure concerns DecodeFailure, which follows the separate ECDecode task path. This review does not establish that this PR fixes that failure or the scanner-cycle failure, and does not establish a passing main CI or release gate.
2026-09-29 15:28:46 +08:00
Chris d2175d1e1e fix(test): keep bucket disk faults across reconnects (#8224) 2026-09-29 14:07:30 +08:00
Chris 2019715d1a Fix fresh capacity probes for formatted local disks (#8222) 2026-09-29 12:23:06 +08:00
Chris 31f05a44af fix(test): use expect_err for offline state root rejection (#8220) 2026-09-29 11:22:16 +08:00
Chris a1724f3dbe feat(connect): capture signed offline service health (#8219) 2026-09-29 10:50:23 +08:00
Chris c62979a45b fix(connect): secure new offline state roots (#8218) 2026-09-29 10:13:58 +08:00
Chris e14eacc5ed fix(rpc): keep legacy format reads in their original namespace (#8216) 2026-09-29 09:24:39 +08:00
Chris 47af5565c7 fix(ci): preserve cluster startup failure evidence (#8213) 2026-09-29 08:50:09 +08:00
Chris f0ce628a7f feat(connect): capture offline native threads in service process (#8215) 2026-09-29 08:48:56 +08:00
Chris 32cc1eb76c Capture offline memory profile from the running service (#8214)
feat(connect): capture offline memory from the running service
2026-09-29 08:43:25 +08:00
Chris 84fe13989e fix(ecstore): retain read quota and control reserve test hedging (#8211) 2026-09-29 07:16:31 +08:00
Chris c8f42b8dca fix(ci): isolate scanner deadline and expose test failure details (#8210)
* fix(ci): isolate scanner deadline fixture and expose readiness errors

* test(ecstore): report unexpected capacity reservation errors
2026-09-29 07:16:17 +08:00
Chris c9acf01fd0 fix(tier): reread mutation intents after lock contention (#8209) 2026-09-29 07:16:01 +08:00