Commit Graph

6911 Commits

Author SHA1 Message Date
Chris 1e7065101d feat: Capture offline CPU profiles in serving process (#8245)
feat: capture offline CPU profiles in serving process
1.0.1-preview.13
2026-09-29 22:37:10 +08:00
Chris 380e98a42c fix(test): synchronize top-disk fixture writes with sampling (#8243) 2026-09-29 21:07:45 +08:00
Chris a88225d208 fix(ci): reduce duplicate work and preserve reliable test failures (#8233)
* fix(ci): reduce duplicate work and preserve reliable test failures

* fix(ci): retain protocol evidence and repair stale test fixtures

* test(connect): honor parent deadline during API fixture readiness

* fix(ci): reserve IO capacity for state writer proofs

* test(connect): align RPC fixtures with service capture contracts

* test(connect): cover pinned service capture failures
2026-09-29 21:06:45 +08:00
Chris 1cca0dd25c feat(connect): select offline key for client performance (#8242) 2026-09-29 20:49:59 +08:00
Chris b0a54f0e5c feat(connect): capture offline network performance in server (#8240)
* feat(connect): capture offline network performance in server

* fix(connect): use storage facade in RPC test

* fix(connect): keep RPC test body behind storage facade
2026-09-29 19:41:43 +08:00
Chris 83a2d8b02d feat(connect): capture RPC activity in offline server runtime (#8237) 2026-09-29 18:16:54 +08:00
Chris f661651aad feat(connect): capture API activity in offline server runtime (#8234) 2026-09-29 17:18:26 +08:00
Chris 9298991d14 fix(scanner): restart a finished mixed sweep under its requested plan (#8231)
A bucket sweep that ends mixed clears its position and records the finishing cycle's plan as started. The next cycle requests a new plan whenever the bucket was written in between, so a fresh sweep inherited a stale started plan and ended mixed again. On a continuously written bucket no sweep could ever certify and the census kept the old root.

Start a fresh verification sweep under the requested plan when no durable position remains. Resumed sweeps keep their started plan, so a clean tail still cannot certify an old prefix.

Refs #7108
2026-09-29 16:48:58 +08:00
Chris 3c7d85420c fix(test): control clocks in Connect network fixtures (#8228) 2026-09-29 16:48:34 +08:00
Chris 6a7880c9fc feat(connect): capture offline lock metrics in server (#8230) 2026-09-29 16:29:07 +08:00
hector 0e58bd09d0 fix(ci): run security chain evidence from the lane's own checkout (#8229)
The security workflow checks the repository out into rustfs-repo/ (#7212)
but its chain evidence steps still invoke
scripts/functional_chain_evidence.py relative to the workspace root, which
on the persistent shared runner resolves to a stale checkout left by another
job. The evidence gate then compares that checkout's HEAD against the chain
workflow_sha and rejects the lane before any case runs
("lane checkout differs from chain workflow source").

Run the evidence script from the lane's own checkout so ROOT resolves to
rustfs-repo/, whose HEAD is exactly the chain-pinned workflow_sha.
2026-09-29 16:09:31 +08:00
Hauser fcb60e7322 fix(rpc): preserve walk_dir missing-path errors (#8226)
## Related Issues

rustfs/rustfs#8217 and rustfs/backlog#2683.

## Summary of Changes

Reviewed 539ed275db against d2175d1e1e. No findings. Capable callers requesting missing-path reporting recover explicit pre-output FileNotFound or VolumeNotFound errors; partial-output and other failures continue to fail the stream.

## Verification

Two independent source reviews completed across all lenses below; the frozen diff matches Git and passes diff whitespace checks. Traced the authenticated handler, bounded first-chunk preflight, cancellation ownership, HTTP status/token classification and quorum error handling. The new tests cover ordinary streaming, typed missing errors, byte preservation and mixed missing/I/O failures.

Author-reported focused tests and Docker reproduction were not rerun or independently audited in this review. The reported Docker comparison predates the final report_notfound-only gate; its scope is not distributed acceptance of the final head.

## Impact

Correctness: no findings; missing errors remain typed only before output. Security/trust: no findings; authentication, body digest and operation/status/token restrictions remain. Compatibility: no findings; legacy clean EOF remains, with the documented old-server limitation. Concurrency/durability: no findings; receiver ownership cancels an abandoned producer. Simplicity: no findings. Coverage: no blocking gap identified. Performance: no findings; preflight retains one bounded chunk and ordinary report_notfound=false streams start immediately.

## Additional Notes

This review does not establish that the patch fixes any current main CI failure. Full main CI and release acceptance remain separate gates.
2026-09-29 15:29:59 +08:00
Hauser 0a61f4a81e fix(heal): park healthy legacy MRF intents (#8225)
## Related Issues

rustfs/backlog#2682 and rustfs/rustfs#8192.

## Summary of Changes

Reviewed ca970f55ec against d2175d1e1e. No blocking code findings. The hold requires a completed healthy legacy check and an exact incarnation, lease, object, version and scope; it preserves the durable journal and rechecks on restart.

## Verification

Two independent source reviews completed, covering the lenses below. The frozen diff matches Git and passes diff whitespace checks. Reported Cargo and Docker results were not rerun or independently audited in this review.

One factual correction to the PR description: `legacy_sigkill_replay_repairs_without_releasing_unverified_responsibility` uses the default four-member fixture, with three replicas before restoration and four afterward (`mrf_partial_write_test.rs:559,567`). That named regression exercises the production manager path, but is not EC12+4. Other tests in the file use sixteen disks.

## Impact

Correctness: no findings; proofless legacy health never becomes a verified receipt. Security/trust: no findings; exact identity fences remain. Compatibility: no findings; journal encoding remains unchanged. Concurrency/durability: no findings; replay retains one checkpointed owner and replacement generations become retryable. Simplicity: no findings. Coverage: no blocking gap identified, with the test-scope correction above. Performance: no findings; held entries leave the retry index without adding per-admission queue scans.

## Additional Notes

The current main journal failure concerns DecodeFailure, which follows the separate ECDecode task path. This review does not establish that this PR fixes that failure or the scanner-cycle failure, and does not establish a passing main CI or release gate.
2026-09-29 15:28:46 +08:00
Chris d2175d1e1e fix(test): keep bucket disk faults across reconnects (#8224) 2026-09-29 14:07:30 +08:00
Chris 2019715d1a Fix fresh capacity probes for formatted local disks (#8222) 2026-09-29 12:23:06 +08:00
Chris 31f05a44af fix(test): use expect_err for offline state root rejection (#8220) 2026-09-29 11:22:16 +08:00
Chris a1724f3dbe feat(connect): capture signed offline service health (#8219) 2026-09-29 10:50:23 +08:00
Chris c62979a45b fix(connect): secure new offline state roots (#8218) 2026-09-29 10:13:58 +08:00
Chris e14eacc5ed fix(rpc): keep legacy format reads in their original namespace (#8216) 2026-09-29 09:24:39 +08:00
Chris 47af5565c7 fix(ci): preserve cluster startup failure evidence (#8213) 2026-09-29 08:50:09 +08:00
Chris f0ce628a7f feat(connect): capture offline native threads in service process (#8215) 2026-09-29 08:48:56 +08:00
Chris 32cc1eb76c Capture offline memory profile from the running service (#8214)
feat(connect): capture offline memory from the running service
2026-09-29 08:43:25 +08:00
Chris 84fe13989e fix(ecstore): retain read quota and control reserve test hedging (#8211) 2026-09-29 07:16:31 +08:00
Chris c8f42b8dca fix(ci): isolate scanner deadline and expose test failure details (#8210)
* fix(ci): isolate scanner deadline fixture and expose readiness errors

* test(ecstore): report unexpected capacity reservation errors
2026-09-29 07:16:17 +08:00
Chris c9acf01fd0 fix(tier): reread mutation intents after lock contention (#8209) 2026-09-29 07:16:01 +08:00
Chris 87e9a84a4e fix(ecstore): preserve RPC status when cloning storage errors (#8207)
* fix(ecstore): preserve RPC status when cloning storage errors

* test(heal): retain clone coverage without redundant ownership
2026-09-29 07:15:46 +08:00
Chris af0ae798f8 fix(test): serialize replacement_bucket_metadata heal test under parallel load (#8212) 2026-09-29 07:07:31 +08:00
dependabot[bot] a2bdcf497b build(deps): bump the dependencies group with 14 updates (#8176) 2026-09-29 01:59:07 +08:00
Chris 5564932f38 fix(heal): retry cancelled internode RPC operations (#8205) 2026-09-29 01:58:35 +08:00
Chris 8778d55e46 fix(ci): preserve diagnostics for rio storage test failures (#8204) 2026-09-29 00:15:26 +08:00
Chris 5ebc1e480b test: stabilize heal timing and improve heartbeat replay coverage (#8150)
fix: preserve heartbeat replay and stabilize timing tests
2026-09-28 22:07:15 +08:00
Chris 4d782141d8 fix(test): add missing unavailable_drives field to readiness test initializers (#8200)
Commit 151103a609 added the  field to
StorageReadinessDetails but did not update two test initializers in
readiness.rs, causing E0063 compile errors in --all-targets builds.
2026-09-28 21:26:42 +08:00
Hauser 3bc36acf7f fix(notify): persist Docker queue stores and expose open errors (#8201)
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-28 21:24:12 +08:00
Chris 4a91c75152 fix(ci): restore checks and multipart migration retries (#8196)
* fix(ci): restore multipart lint and full E2E membership checks

* fix(ecstore): preserve multipart completion identity during migration
2026-09-28 21:22:49 +08:00
Chris 2648776be7 fix: accept catalog files in health service artifact (#8198)
* fix: accept catalog files in health service artifact

* ci: retain configured health acceptance runner
2026-09-28 20:45:26 +08:00
RustFS 9442e89f5f fix(iam): require an explicit permission for force-delete (#8154)
* fix(iam): require an explicit permission for force-delete

A force-delete header no longer inherits s3:* or consoleAdmin. Bucket
force-delete requires s3:ForceDeleteBucket whenever the header is present,
and recursive object force-delete requires s3:ForceDeleteObject. A plain
delete keeps the existing checks.

Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>

* test(e2e): keep force-delete header names static

The bucket force-delete helper must pass a static header name into the
SDK request mutator.

Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>

* test(e2e): move the force-delete header into the request mutator

The SDK request customizer requires a static header name owned by the
closure.

Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>

* fix(iam): keep force-delete out of NotAction grants

NotAction now uses plain wildcard matching, so NotAction "s3:*" still
excludes force-delete. An Allow statement grants s3:ForceDeleteObject or
s3:ForceDeleteBucket only when its Action list names the action; a
NotAction-only Allow never does. The rule applies to both IAM and bucket
policy statements.

Also build the invalid-header errors with S3Error::with_message to keep
the s3s footprint at its baseline, and fix a clippy single_match.

---------

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-09-28 19:08:33 +08:00
RustFS fc5609bbb0 fix(s3): make retried CompleteMultipartUpload idempotent (#8153)
Record the upload id on the completed object so a lost-response retry
returns that object's ETag instead of NoSuchUpload, while a different
part list or a replaced object keeps the existing errors.

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-28 19:08:10 +08:00
Chris 64bd224b3b fix(connect): accept the health-era heartbeat and measured Top windows (#8195)
Two Connect tests fail on main in every Test and Lint lane.

Heartbeat: #8155 replaced "today's capabilities minus
profile.memory.service@1" with "the pre-health set minus it". That drops
the release that advertised health.check.service@1 but not yet in-service
memory profiling, so a pending heartbeat persisted by that release now
fails validation and the runtime stops. Accept that release as its own
frozen list.

Top disk/net: execute_*_job sets max_duration_millis to the requested
duration, then the capture compared the measured sleep against it. The
timer only overshoots, so a job that ran exactly as requested was
rejected with LimitExceeded whenever the overshoot reached 1 ms. Report
the authorized window, as Top API already does; validate_capture bounds
it by the limit.
2026-09-28 19:07:39 +08:00
Chris 71f8b607ad fix(ci): restore mainline and scheduled test reliability (#8187)
* fix(ci): restore mainline and scheduled test reliability

* fix(ci): provide GitHub CLI for CPU acceptance

* ci: provide Docker for CPU service acceptance

* ci: provision Python and Docker for OIDC validation

* ci: restore hosted runners for Docker validation

* ci: use verified MinIO release packages for interop

* ci: preserve host ownership of MinIO fixtures

* fix(ci): correct diagnostic limits and isolate startup checks

* test(connect): include object CLI failure details

* test(readiness): initialize unavailable drive diagnostics
2026-09-28 19:07:29 +08:00
GatewayJ cef0c61532 fix(table-catalog): preserve data sequences during file rewrites
Merge the reviewed fix from pull request #8104.
2026-09-28 18:51:58 +08:00
Justin Bradfield fc325bcc56 fix(s3): report x-amz-mp-parts-count on GetObject with partNumber
Merge the reviewed fix from pull request #8115.
2026-09-28 18:51:48 +08:00
cxymds 4a46f5ff3d fix(s3): serve bucket websites on dedicated domains
Merge the reviewed fix from pull request #8183.
2026-09-28 18:48:21 +08:00
cxymds e0a973bf4e fix(s3): preserve prefix listing pagination and marker metadata
Merge the reviewed fix from pull request #8181.
2026-09-28 18:47:03 +08:00
hector 671238458c ci: fail the pool lane when expected step markers are missing
Merge the reviewed fix from pull request #8190.
2026-09-28 18:46:12 +08:00
Chris 151103a609 fix(ecstore): recover remote disks after transient stalls
Merge the approved fix from pull request #8149.
2026-09-28 18:40:30 +08:00
Nikita Bakun cc7d5e3f5f fix(replication): stop sending versionId query once a target rejects it
Merge the approved fix from pull request #8109.
2026-09-28 18:40:20 +08:00
yi111 d556d4dc94 fix(admin): report a missing policy or user as 404 NoSuchResource, not 500 InternalError
Merge the approved fix from pull request #8127.
2026-09-28 18:40:10 +08:00
chapman 6300fbe74f fix: map IAM not-found errors to 404 in InfoCannedPolicy and RemoveUser
Merge the approved fix from pull request #8134.
2026-09-28 18:38:20 +08:00
Chris 6ebc78d115 fix(ci): run profile acceptance on Docker runners (#8191) 2026-09-28 16:14:07 +08:00
Chris 5f4d5e8fe8 fix(ci): run CPU acceptance with Docker
Use the Docker-enabled runner required by the connected CPU service-job test.
2026-09-28 15:38:36 +08:00