Commit Graph

3540 Commits

Author SHA1 Message Date
Chris a88225d208 fix(ci): reduce duplicate work and preserve reliable test failures (#8233)
* fix(ci): reduce duplicate work and preserve reliable test failures

* fix(ci): retain protocol evidence and repair stale test fixtures

* test(connect): honor parent deadline during API fixture readiness

* fix(ci): reserve IO capacity for state writer proofs

* test(connect): align RPC fixtures with service capture contracts

* test(connect): cover pinned service capture failures
2026-09-29 21:06:45 +08:00
Chris 9298991d14 fix(scanner): restart a finished mixed sweep under its requested plan (#8231)
A bucket sweep that ends mixed clears its position and records the finishing cycle's plan as started. The next cycle requests a new plan whenever the bucket was written in between, so a fresh sweep inherited a stale started plan and ended mixed again. On a continuously written bucket no sweep could ever certify and the census kept the old root.

Start a fresh verification sweep under the requested plan when no durable position remains. Resumed sweeps keep their started plan, so a clean tail still cannot certify an old prefix.

Refs #7108
2026-09-29 16:48:58 +08:00
Hauser fcb60e7322 fix(rpc): preserve walk_dir missing-path errors (#8226)
## Related Issues

rustfs/rustfs#8217 and rustfs/backlog#2683.

## Summary of Changes

Reviewed 539ed275db against d2175d1e1e. No findings. Capable callers requesting missing-path reporting recover explicit pre-output FileNotFound or VolumeNotFound errors; partial-output and other failures continue to fail the stream.

## Verification

Two independent source reviews completed across all lenses below; the frozen diff matches Git and passes diff whitespace checks. Traced the authenticated handler, bounded first-chunk preflight, cancellation ownership, HTTP status/token classification and quorum error handling. The new tests cover ordinary streaming, typed missing errors, byte preservation and mixed missing/I/O failures.

Author-reported focused tests and Docker reproduction were not rerun or independently audited in this review. The reported Docker comparison predates the final report_notfound-only gate; its scope is not distributed acceptance of the final head.

## Impact

Correctness: no findings; missing errors remain typed only before output. Security/trust: no findings; authentication, body digest and operation/status/token restrictions remain. Compatibility: no findings; legacy clean EOF remains, with the documented old-server limitation. Concurrency/durability: no findings; receiver ownership cancels an abandoned producer. Simplicity: no findings. Coverage: no blocking gap identified. Performance: no findings; preflight retains one bounded chunk and ordinary report_notfound=false streams start immediately.

## Additional Notes

This review does not establish that the patch fixes any current main CI failure. Full main CI and release acceptance remain separate gates.
2026-09-29 15:29:59 +08:00
Hauser 0a61f4a81e fix(heal): park healthy legacy MRF intents (#8225)
## Related Issues

rustfs/backlog#2682 and rustfs/rustfs#8192.

## Summary of Changes

Reviewed ca970f55ec against d2175d1e1e. No blocking code findings. The hold requires a completed healthy legacy check and an exact incarnation, lease, object, version and scope; it preserves the durable journal and rechecks on restart.

## Verification

Two independent source reviews completed, covering the lenses below. The frozen diff matches Git and passes diff whitespace checks. Reported Cargo and Docker results were not rerun or independently audited in this review.

One factual correction to the PR description: `legacy_sigkill_replay_repairs_without_releasing_unverified_responsibility` uses the default four-member fixture, with three replicas before restoration and four afterward (`mrf_partial_write_test.rs:559,567`). That named regression exercises the production manager path, but is not EC12+4. Other tests in the file use sixteen disks.

## Impact

Correctness: no findings; proofless legacy health never becomes a verified receipt. Security/trust: no findings; exact identity fences remain. Compatibility: no findings; journal encoding remains unchanged. Concurrency/durability: no findings; replay retains one checkpointed owner and replacement generations become retryable. Simplicity: no findings. Coverage: no blocking gap identified, with the test-scope correction above. Performance: no findings; held entries leave the retry index without adding per-admission queue scans.

## Additional Notes

The current main journal failure concerns DecodeFailure, which follows the separate ECDecode task path. This review does not establish that this PR fixes that failure or the scanner-cycle failure, and does not establish a passing main CI or release gate.
2026-09-29 15:28:46 +08:00
Chris d2175d1e1e fix(test): keep bucket disk faults across reconnects (#8224) 2026-09-29 14:07:30 +08:00
Chris 2019715d1a Fix fresh capacity probes for formatted local disks (#8222) 2026-09-29 12:23:06 +08:00
Chris 47af5565c7 fix(ci): preserve cluster startup failure evidence (#8213) 2026-09-29 08:50:09 +08:00
Chris 84fe13989e fix(ecstore): retain read quota and control reserve test hedging (#8211) 2026-09-29 07:16:31 +08:00
Chris c8f42b8dca fix(ci): isolate scanner deadline and expose test failure details (#8210)
* fix(ci): isolate scanner deadline fixture and expose readiness errors

* test(ecstore): report unexpected capacity reservation errors
2026-09-29 07:16:17 +08:00
Chris c9acf01fd0 fix(tier): reread mutation intents after lock contention (#8209) 2026-09-29 07:16:01 +08:00
Chris 87e9a84a4e fix(ecstore): preserve RPC status when cloning storage errors (#8207)
* fix(ecstore): preserve RPC status when cloning storage errors

* test(heal): retain clone coverage without redundant ownership
2026-09-29 07:15:46 +08:00
dependabot[bot] a2bdcf497b build(deps): bump the dependencies group with 14 updates (#8176) 2026-09-29 01:59:07 +08:00
Chris 5564932f38 fix(heal): retry cancelled internode RPC operations (#8205) 2026-09-29 01:58:35 +08:00
Chris 8778d55e46 fix(ci): preserve diagnostics for rio storage test failures (#8204) 2026-09-29 00:15:26 +08:00
Chris 5ebc1e480b test: stabilize heal timing and improve heartbeat replay coverage (#8150)
fix: preserve heartbeat replay and stabilize timing tests
2026-09-28 22:07:15 +08:00
Hauser 3bc36acf7f fix(notify): persist Docker queue stores and expose open errors (#8201)
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-28 21:24:12 +08:00
Chris 4a91c75152 fix(ci): restore checks and multipart migration retries (#8196)
* fix(ci): restore multipart lint and full E2E membership checks

* fix(ecstore): preserve multipart completion identity during migration
2026-09-28 21:22:49 +08:00
RustFS 9442e89f5f fix(iam): require an explicit permission for force-delete (#8154)
* fix(iam): require an explicit permission for force-delete

A force-delete header no longer inherits s3:* or consoleAdmin. Bucket
force-delete requires s3:ForceDeleteBucket whenever the header is present,
and recursive object force-delete requires s3:ForceDeleteObject. A plain
delete keeps the existing checks.

Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>

* test(e2e): keep force-delete header names static

The bucket force-delete helper must pass a static header name into the
SDK request mutator.

Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>

* test(e2e): move the force-delete header into the request mutator

The SDK request customizer requires a static header name owned by the
closure.

Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>

* fix(iam): keep force-delete out of NotAction grants

NotAction now uses plain wildcard matching, so NotAction "s3:*" still
excludes force-delete. An Allow statement grants s3:ForceDeleteObject or
s3:ForceDeleteBucket only when its Action list names the action; a
NotAction-only Allow never does. The rule applies to both IAM and bucket
policy statements.

Also build the invalid-header errors with S3Error::with_message to keep
the s3s footprint at its baseline, and fix a clippy single_match.

---------

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-09-28 19:08:33 +08:00
RustFS fc5609bbb0 fix(s3): make retried CompleteMultipartUpload idempotent (#8153)
Record the upload id on the completed object so a lost-response retry
returns that object's ETag instead of NoSuchUpload, while a different
part list or a replaced object keeps the existing errors.

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-28 19:08:10 +08:00
Chris 71f8b607ad fix(ci): restore mainline and scheduled test reliability (#8187)
* fix(ci): restore mainline and scheduled test reliability

* fix(ci): provide GitHub CLI for CPU acceptance

* ci: provide Docker for CPU service acceptance

* ci: provision Python and Docker for OIDC validation

* ci: restore hosted runners for Docker validation

* ci: use verified MinIO release packages for interop

* ci: preserve host ownership of MinIO fixtures

* fix(ci): correct diagnostic limits and isolate startup checks

* test(connect): include object CLI failure details

* test(readiness): initialize unavailable drive diagnostics
2026-09-28 19:07:29 +08:00
Justin Bradfield fc325bcc56 fix(s3): report x-amz-mp-parts-count on GetObject with partNumber
Merge the reviewed fix from pull request #8115.
2026-09-28 18:51:48 +08:00
cxymds 4a46f5ff3d fix(s3): serve bucket websites on dedicated domains
Merge the reviewed fix from pull request #8183.
2026-09-28 18:48:21 +08:00
cxymds e0a973bf4e fix(s3): preserve prefix listing pagination and marker metadata
Merge the reviewed fix from pull request #8181.
2026-09-28 18:47:03 +08:00
Chris 151103a609 fix(ecstore): recover remote disks after transient stalls
Merge the approved fix from pull request #8149.
2026-09-28 18:40:30 +08:00
Nikita Bakun cc7d5e3f5f fix(replication): stop sending versionId query once a target rejects it
Merge the approved fix from pull request #8109.
2026-09-28 18:40:20 +08:00
Chris 0609e7ce14 fix(connect): admit complete release CPU symbol catalogs
Bound the release catalog to 32 MiB and 200,000 symbols based on the GNU build measurement; keep complete names and reject catalogs beyond either limit.
2026-09-28 12:09:26 +08:00
RustFS 2e014d25b3 fix: bound restarted-peer stalls on quorum reads and writes (#8152)
A restarted peer can accept a pooled connection and never send response
headers, so HttpReader::open waited past the client body timeout before
the body-stall timer or erasure hedge could run. Bound that header wait
by the stall timeout and retry the open once on a fresh connection.

After write quorum, MultiWriter still waited out the full disk stall
for a silent peer, which matches the client timeout. Give remaining
writers one second, then drop them so the caller returns.

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
2026-09-28 02:22:41 +00:00
Chris e33542b0c2 fix(scanner): recover after cycle-state persistence failures (#8148)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 16:39:41 +00:00
Chris 6d2b1c629e Generate CPU symbol catalogs from release executables (#8165) 2026-09-28 00:12:36 +08:00
RustFS de484b93e0 fix: admit cold reads and readiness at read quorum (#8156)
EC 2+2 still meets read quorum with two of four nodes up. A survivor
that had not cached a bucket returned 503 on GET, HEAD, and List, and
/health/ready left the Service once write quorum was lost. Reads and
Service membership now follow read quorum and shared locks. Writes and
/minio/health/cluster still require write quorum and exclusive locks.

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 13:58:55 +00:00
Hiroaki KAWAI 20356ab709 fix(scanner): build raw enumeration indexes incrementally (#8114)
* fix(scanner): build raw enumeration indexes incrementally

* test(scanner): iterate over restart test entries directly

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 07:49:59 +08:00
Chris cf096b3c22 fix(connect): pace diagnostic network payload sends (#8096)
* fix(connect): pace diagnostic network payload sends

* test(connect): exercise native network pacing entry

---------

Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 05:11:35 +08:00
Chris 31ac243865 feat(connect): add bounded health service job 2026-09-27 02:50:49 +08:00
Chris 8741bb77ea fix(storage): weight automatic multipart admission by part size (#8118)
* fix(storage): weight automatic multipart admission by part size

* fix(ci): restore filesystem runner capabilities and typos dependency

* fix(ci): make release guard portable and spell out part variables

* ci: restore sm-standard-2 runners for io_uring and distributed e2e jobs

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-26 22:35:39 +08:00
cxymds 212b20f890 fix(ecstore): admit cold bucket metadata at read quorum (#8120)
* fix(ecstore): admit cold bucket metadata at read quorum

Encode a durable creation-commit state in bucket metadata so committed Object Lock buckets can use read quorum for cold loads while pre-physical creation intents remain fail-closed. Drain commit-write fan-out, keep write quorum for uncommitted intents, and add regression coverage for exact, below, migrated, and partial-create quorum boundaries.

* fix(ecstore): fence bucket creation-commit persistence

Address review findings on the creation-commit proof.

Persist the proof only under the bucket metadata transaction fence: at bucket creation, and through a fenced migration that re-reads the authoritative metadata and revalidates physical presence at write quorum before writing. This stops a stale snapshot from reverting an acknowledged configuration update or outliving a delete/recreate.

Establish commitment when Object Lock is enabled on an existing bucket, inside the same configuration mutation, so a healthy cluster no longer rejects object operations with ErasureWriteQuorum.

* fix(ecstore): keep commit fence error message stable

The error(format!) ratchet requires a stable Display for quorum bucketing; use a fixed message instead of embedding the bucket name.

* test(ecstore): reuse canonical Object Lock fixture in regression tests

The s3s footprint ratchet is shrink-only. Use the existing ENABLED_OBJECT_LOCK_CONFIG static instead of naming s3s DTO types in store tests.

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-26 21:40:40 +08:00
Hauser 5cac685e9a feat(ecstore): integrate io_uring backend improvements (#8105) 2026-09-26 17:14:17 +08:00
Hauser d8bf268885 fix(scanner): bound SNSD deep scans and checkpoint cloning (#8126)
* fix(scanner): bound SNSD deep scans and checkpoint cloning

* ci: run mount-dependent jobs on hosted VMs

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
2026-09-26 06:06:50 +00:00
Hauser 286bbb8e1f chore: upgrade workspace dependency releases (#8107)
* chore: upgrade workspace dependencies

* chore: refresh dependency lockfile

* fix(rebalance): avoid waiting on active entry gate during stop preparation
2026-09-24 21:24:14 +08:00
cxymds 0913686b12 fix(heal): detect stale delete-marker metadata (#8102) 2026-09-24 11:16:02 +08:00
George Melikov df64242f64 fix(ecstore): resume scan_dir past a dash-suffixed sibling directory (#8071)
A flat (recursive) listing resumed inside "s/" re-emitted everything in
its sibling "s-x/". scan_dir dropped the entries before forward_to by
comparing directory names without their trailing slash, where "s" sorts
before "s-x", while the keys they stand for sort the other way round:
'-' (0x2d) is below '/' (0x2f), so all of "s-x/..." precedes "s/...".
The drain stopped at "s" and kept "s-x".

The next page then started with keys at or before the marker; the
listing layer filtered all of them out, found no more candidates and
answered IsTruncated=false. A bucket of 104,137 objects with backup
directories named "<id>" and "<id>-rollbacks" listed as 5,000; two such
directories of 1,200 keys each listed as 2,000.

Compare every remaining entry as the key prefix it stands for, slash
included, and keep it only when it sorts at or after forward_to or
contains it. The remainder of forward_to is taken before `current` is
trimmed, so this also holds below the bucket root; retain also drops
entries when all of them precede forward_to, which the drain never did.
2026-09-24 07:13:30 +08:00
cxymds 9ab034df62 fix(heal): preserve retryable writer failures (#8089)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-24 02:41:16 +08:00
Tyler Hillery 624463c5ba fix(s3): default CopyObject content-type to binary/octet-stream on REPLACE (#8073)
fix: default CopyObject content-type to binary/octet-stream on REPLACE
2026-09-23 18:04:05 +00:00
GatewayJ f219a0aba8 perf(ecstore): parallelize multipart I/O setup and metadata reads (#8085)
* perf(ecstore): parallelize multipart I/O setup and metadata reads

* perf(ecstore): share multipart paths and increase read concurrency

* fix(ecstore): route multipart benchmark through storage API facade
2026-09-23 22:19:23 +08:00
Hauser 7f0f4941cf chore(ci): refresh checkout pin in validated workflows (#8083)
* chore(ci): update checkout pin in validated workflows

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(ci): pin migrated checkout alerts to requested commit

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* chore(ci): pin additional checkout workflows to verified commit (#8084)

* chore(ci): pin more checkout workflows to requested commit

Update six additional workflows to the verified upstream checkout commit without changing their permissions, inputs, or triggers.

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* chore(ci): pin functional and OIDC checkout uses

Extend the verified checkout commit pin to functional-chain and OIDC workflows without changing their behavior.

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>

---------

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-23 09:27:54 +00:00
cxymds 655bb470b8 fix(heal): retry faulty storage disk errors (#8086) 2026-09-23 07:48:39 +00:00
cxymds 56a4099f99 fix(heal): recreate missing bucket volumes (#8082)
* fix(heal): recreate missing bucket volumes

* fix(heal): use typed error for empty targets
2026-09-23 14:51:24 +08:00
cxymds 1880b42169 fix(heal): preserve decode repair signals and retry startup quorum (#8067)
* fix(startup): retry transient bucket metadata quorum failures

* fix(ecstore): retain recovered shard damage for read repair

* fix(ecstore): reconnect missing read sources before GET

* test(heal): prove exact marker replay preserves history

* fix(ecstore): fence reconnect retries at the deadline

Reject expired reconnect waiters before dispatch and start cooldown on sweep completion, including timeout and cancellation. Cover exact-deadline admission independently of cooldown and preserve the original strict RPC-count regression.
2026-09-22 14:58:31 +00:00
Dae-Cheol Noh 3c6c88b2e7 fix(amqp): emit S3 events directly in notification records (#8066)
* fix(amqp): emit S3 events directly in notification records

* test(amqp): avoid queued timestamp equality race
2026-09-22 12:16:40 +00:00
Dae-Cheol Noh 1f04a12abf fix(ftps): bound upload memory with multipart streaming (#8064)
* fix(ftps): bound upload memory with multipart streaming

* fix(ftps): keep multipart upload within driver module
2026-09-22 12:15:20 +00:00
cxymds 122abfaae2 feat(integrity): add inventory, audit, and protected migration (#8065)
* feat(integrity): add inventory, audit, and protected migration

* refactor(admin): use gateway facade for integrity handlers

* fix(integrity): sort fingerprint metadata explicitly
2026-09-22 19:23:41 +08:00