Commit Graph

3515 Commits

Author SHA1 Message Date
Chris 0609e7ce14 fix(connect): admit complete release CPU symbol catalogs
Bound the release catalog to 32 MiB and 200,000 symbols based on the GNU build measurement; keep complete names and reject catalogs beyond either limit.
2026-09-28 12:09:26 +08:00
RustFS 2e014d25b3 fix: bound restarted-peer stalls on quorum reads and writes (#8152)
A restarted peer can accept a pooled connection and never send response
headers, so HttpReader::open waited past the client body timeout before
the body-stall timer or erasure hedge could run. Bound that header wait
by the stall timeout and retry the open once on a fresh connection.

After write quorum, MultiWriter still waited out the full disk stall
for a silent peer, which matches the client timeout. Give remaining
writers one second, then drop them so the caller returns.

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
2026-09-28 02:22:41 +00:00
Chris e33542b0c2 fix(scanner): recover after cycle-state persistence failures (#8148)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 16:39:41 +00:00
Chris 6d2b1c629e Generate CPU symbol catalogs from release executables (#8165) 2026-09-28 00:12:36 +08:00
RustFS de484b93e0 fix: admit cold reads and readiness at read quorum (#8156)
EC 2+2 still meets read quorum with two of four nodes up. A survivor
that had not cached a bucket returned 503 on GET, HEAD, and List, and
/health/ready left the Service once write quorum was lost. Reads and
Service membership now follow read quorum and shared locks. Writes and
/minio/health/cluster still require write quorum and exclusive locks.

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 13:58:55 +00:00
Hiroaki KAWAI 20356ab709 fix(scanner): build raw enumeration indexes incrementally (#8114)
* fix(scanner): build raw enumeration indexes incrementally

* test(scanner): iterate over restart test entries directly

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 07:49:59 +08:00
Chris cf096b3c22 fix(connect): pace diagnostic network payload sends (#8096)
* fix(connect): pace diagnostic network payload sends

* test(connect): exercise native network pacing entry

---------

Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 05:11:35 +08:00
Chris 31ac243865 feat(connect): add bounded health service job 2026-09-27 02:50:49 +08:00
Chris 8741bb77ea fix(storage): weight automatic multipart admission by part size (#8118)
* fix(storage): weight automatic multipart admission by part size

* fix(ci): restore filesystem runner capabilities and typos dependency

* fix(ci): make release guard portable and spell out part variables

* ci: restore sm-standard-2 runners for io_uring and distributed e2e jobs

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-26 22:35:39 +08:00
cxymds 212b20f890 fix(ecstore): admit cold bucket metadata at read quorum (#8120)
* fix(ecstore): admit cold bucket metadata at read quorum

Encode a durable creation-commit state in bucket metadata so committed Object Lock buckets can use read quorum for cold loads while pre-physical creation intents remain fail-closed. Drain commit-write fan-out, keep write quorum for uncommitted intents, and add regression coverage for exact, below, migrated, and partial-create quorum boundaries.

* fix(ecstore): fence bucket creation-commit persistence

Address review findings on the creation-commit proof.

Persist the proof only under the bucket metadata transaction fence: at bucket creation, and through a fenced migration that re-reads the authoritative metadata and revalidates physical presence at write quorum before writing. This stops a stale snapshot from reverting an acknowledged configuration update or outliving a delete/recreate.

Establish commitment when Object Lock is enabled on an existing bucket, inside the same configuration mutation, so a healthy cluster no longer rejects object operations with ErasureWriteQuorum.

* fix(ecstore): keep commit fence error message stable

The error(format!) ratchet requires a stable Display for quorum bucketing; use a fixed message instead of embedding the bucket name.

* test(ecstore): reuse canonical Object Lock fixture in regression tests

The s3s footprint ratchet is shrink-only. Use the existing ENABLED_OBJECT_LOCK_CONFIG static instead of naming s3s DTO types in store tests.

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-26 21:40:40 +08:00
Hauser 5cac685e9a feat(ecstore): integrate io_uring backend improvements (#8105) 2026-09-26 17:14:17 +08:00
Hauser d8bf268885 fix(scanner): bound SNSD deep scans and checkpoint cloning (#8126)
* fix(scanner): bound SNSD deep scans and checkpoint cloning

* ci: run mount-dependent jobs on hosted VMs

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
2026-09-26 06:06:50 +00:00
Hauser 286bbb8e1f chore: upgrade workspace dependency releases (#8107)
* chore: upgrade workspace dependencies

* chore: refresh dependency lockfile

* fix(rebalance): avoid waiting on active entry gate during stop preparation
2026-09-24 21:24:14 +08:00
cxymds 0913686b12 fix(heal): detect stale delete-marker metadata (#8102) 2026-09-24 11:16:02 +08:00
George Melikov df64242f64 fix(ecstore): resume scan_dir past a dash-suffixed sibling directory (#8071)
A flat (recursive) listing resumed inside "s/" re-emitted everything in
its sibling "s-x/". scan_dir dropped the entries before forward_to by
comparing directory names without their trailing slash, where "s" sorts
before "s-x", while the keys they stand for sort the other way round:
'-' (0x2d) is below '/' (0x2f), so all of "s-x/..." precedes "s/...".
The drain stopped at "s" and kept "s-x".

The next page then started with keys at or before the marker; the
listing layer filtered all of them out, found no more candidates and
answered IsTruncated=false. A bucket of 104,137 objects with backup
directories named "<id>" and "<id>-rollbacks" listed as 5,000; two such
directories of 1,200 keys each listed as 2,000.

Compare every remaining entry as the key prefix it stands for, slash
included, and keep it only when it sorts at or after forward_to or
contains it. The remainder of forward_to is taken before `current` is
trimmed, so this also holds below the bucket root; retain also drops
entries when all of them precede forward_to, which the drain never did.
2026-09-24 07:13:30 +08:00
cxymds 9ab034df62 fix(heal): preserve retryable writer failures (#8089)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-24 02:41:16 +08:00
Tyler Hillery 624463c5ba fix(s3): default CopyObject content-type to binary/octet-stream on REPLACE (#8073)
fix: default CopyObject content-type to binary/octet-stream on REPLACE
2026-09-23 18:04:05 +00:00
GatewayJ f219a0aba8 perf(ecstore): parallelize multipart I/O setup and metadata reads (#8085)
* perf(ecstore): parallelize multipart I/O setup and metadata reads

* perf(ecstore): share multipart paths and increase read concurrency

* fix(ecstore): route multipart benchmark through storage API facade
2026-09-23 22:19:23 +08:00
Hauser 7f0f4941cf chore(ci): refresh checkout pin in validated workflows (#8083)
* chore(ci): update checkout pin in validated workflows

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(ci): pin migrated checkout alerts to requested commit

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* chore(ci): pin additional checkout workflows to verified commit (#8084)

* chore(ci): pin more checkout workflows to requested commit

Update six additional workflows to the verified upstream checkout commit without changing their permissions, inputs, or triggers.

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* chore(ci): pin functional and OIDC checkout uses

Extend the verified checkout commit pin to functional-chain and OIDC workflows without changing their behavior.

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>

---------

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-23 09:27:54 +00:00
cxymds 655bb470b8 fix(heal): retry faulty storage disk errors (#8086) 2026-09-23 07:48:39 +00:00
cxymds 56a4099f99 fix(heal): recreate missing bucket volumes (#8082)
* fix(heal): recreate missing bucket volumes

* fix(heal): use typed error for empty targets
2026-09-23 14:51:24 +08:00
cxymds 1880b42169 fix(heal): preserve decode repair signals and retry startup quorum (#8067)
* fix(startup): retry transient bucket metadata quorum failures

* fix(ecstore): retain recovered shard damage for read repair

* fix(ecstore): reconnect missing read sources before GET

* test(heal): prove exact marker replay preserves history

* fix(ecstore): fence reconnect retries at the deadline

Reject expired reconnect waiters before dispatch and start cooldown on sweep completion, including timeout and cancellation. Cover exact-deadline admission independently of cooldown and preserve the original strict RPC-count regression.
2026-09-22 14:58:31 +00:00
Dae-Cheol Noh 3c6c88b2e7 fix(amqp): emit S3 events directly in notification records (#8066)
* fix(amqp): emit S3 events directly in notification records

* test(amqp): avoid queued timestamp equality race
2026-09-22 12:16:40 +00:00
Dae-Cheol Noh 1f04a12abf fix(ftps): bound upload memory with multipart streaming (#8064)
* fix(ftps): bound upload memory with multipart streaming

* fix(ftps): keep multipart upload within driver module
2026-09-22 12:15:20 +00:00
cxymds 122abfaae2 feat(integrity): add inventory, audit, and protected migration (#8065)
* feat(integrity): add inventory, audit, and protected migration

* refactor(admin): use gateway facade for integrity handlers

* fix(integrity): sort fingerprint metadata explicitly
2026-09-22 19:23:41 +08:00
cxymds d0ce2f758b fix(ecstore): bound decommission target gate contention (#8061)
* fix(ecstore): bound decommission target gate contention

Target capacity gate contention during pool decommission escalated a
per-object transient into a durable bucket pause: each contended object
failed its bucket entry, which re-ran the whole bucket listing and amplified
attempts on the same objects.

- centralize the decommission capacity failure classification so gate
  contention, benign contention and fatal failures are decided once
- retry target-gate contention inline (bounded, jittered) before a mutation
  is admitted, covering put, part, complete, new-multipart and abort
- defer contended entries to the end of the round instead of failing the
  bucket entry, and require the deferred set to drain before a set completes
- treat a missing object or version, an overwrite and a superseded upload id
  as benign contention that is neither counted as a failure nor escalated
- use full-jitter exponential backoff, capped, for decommission retries
- expose the capacity pause reason, the waiting reason and a cumulative
  pause count in the admin pool status, plus gate-retry and per-object
  attempt metrics

Related: rustfs/backlog#2644

* fix(ecstore): correct deferred replay and metadata compatibility
2026-09-22 19:20:05 +08:00
Hauser 0ff06757c8 fix(obs): support custom OTLP CA bundles (#8060)
Add RUSTFS_OBS_TLS_CA_FILE for additive private CA trust in OTLP HTTP exporters.

Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-22 13:21:44 +08:00
cxymds 1afaf4491e fix(ecstore): bound decommission capacity gate waits (#8044)
* fix(ecstore): bound decommission capacity gate waits

* fix(ecstore): keep rebalance meta save lock errors retryable

The rebalance metadata retry policy only retries when the source chain still
carries the typed lock error, but resolve_rebalance_meta_save_result collapsed
every failure into a plain string. A transient rebalance.bin write-lock timeout
therefore bypassed retry_rebalance_metadata_access entirely and surfaced as a
hard failure of the periodic and terminal rebalance metadata saves.

Wrap the failure with the data-movement stage context instead, which keeps the
original lock error reachable through rebalance_error_source so the existing
transient lock policy applies. The rendered message is unchanged.

Add a regression test that feeds the wrapped meta save lock timeout through the
retry helper and asserts the retry engages.

* chore(ecstore): refresh error format ratchet baseline

Removing the string-wrapped meta save failure drops one `::other(format!`
call site in the rebalance worker, so the shrink-only baseline must be
regenerated in the same change.
2026-09-21 20:56:26 +08:00
cxymds cd2fa50e50 fix: rename admin metrics endpoint to realtime (#8046) 2026-09-21 19:23:58 +08:00
cxymds 077b961232 fix(obs): default span events to none and gate span list by level (#8047)
JSON log layers synthesized span lifecycle events for every span and
embedded the full ancestor span list on every record. With INFO spans on
object hot paths that produced multi-megabyte lines, which journald
truncated at 48 KiB: the payload was silently dropped and the volume
itself became the dominant load.

Add RUSTFS_OBS_LOG_SPAN_EVENTS (none|close|full) with a default of none,
and keep with_span_list enabled only for verbose (debug/trace) levels so
the ancestor chain stops being re-serialized per record. The request_id
promotion path reads the span scope, not the span list, so it is
unaffected.

Refs rustfs/backlog#2642
2026-09-21 19:15:39 +08:00
cxymds e1bc731331 fix(lifecycle): reject empty expiration actions (#8045)
* fix(lifecycle): reject empty expiration actions

* style: apply rustfmt to lifecycle validation
2026-09-21 15:43:07 +08:00
cxymds bb43ef861a fix(decommission): use owned data for source capacity (#8038) 2026-09-21 15:29:30 +08:00
Chris 4047ed6c8d fix(heal): restore bucket metadata before replacement completion (#8042) 2026-09-21 12:18:06 +08:00
cxymds a7c875b48f fix(heal): preserve remote corruption and reject incomplete deep scans (#8032)
* fix(heal): reject truncated encoded shards

* test(heal): cover truncated shards across integrity modes (#8035)

Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* fix(heal): preserve remote corruption and reject incomplete deep scans

* test(admin): tolerate config lock timeout while polling

* fix(test): use admin storage error facade

---------

Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-09-21 02:33:03 +00:00
Chris 0632a64f91 fix(ecstore): allow server-side copy at normal read quorum (#8029) 2026-09-20 12:36:53 +08:00
cxymds 22243e791d test(heal): cover stale current delete marker rejoin (#8021) 2026-09-19 19:33:24 +08:00
Hauser c35de00e8a fix(scanner): protect single-disk foreground latency (#8022)
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-19 19:32:02 +08:00
cxymds c3269ad4c2 fix(heal): close tail-truncation and stale current-marker gaps (#8019)
* fix(heal): verify complete encoded shard tails

* fix(heal): prove stale current marker deletion
2026-09-19 13:41:30 +08:00
Chris 0fcd34c331 fix(ecstore): retry fleet capability probe while the notification system boots (#8017)
Startup finalizes IAM before init_notification_runtime publishes the notification system, and IAM finalization is what starts the fleet capability probe. The first probe pass therefore always failed closed and then slept the full ten-second probe interval. On a single node that left every durable capability, including the durable hard quota fence, unavailable for about ten seconds after /health already reported ok, so SetBucketQuota answered 503 durable quota capability is not confirmed across the cluster during that window.

Publishing the notification system now wakes the probe immediately through a Notify permit, with a 100ms bootstrap poll as the fallback for a wakeup that races the availability check. The probe keeps failing closed while the system is absent and logs the wait once. The four probe futures no longer carry an unreachable notification-system-unavailable arm, and the seven proof slots are revoked through one helper.

Fixes #8014
2026-09-19 13:41:16 +08:00
cxymds ff62c810b4 fix(tier): gate opaque version recovery cleanup (#8006)
* fix(tier): gate opaque version recovery cleanup

* test(tier): authorize opaque recovery fixtures

---------

Co-authored-by: Chris <anzhengchao@gmail.com>
2026-09-18 21:54:03 +08:00
cxymds 950db7412c fix(heal): inspect stale members during deep scans (#8009)
* fix(heal): inspect stale members during deep scans

* test(heal): prove convergence outcomes
2026-09-18 21:53:31 +08:00
Chris 4e7a8fabb8 fix(ecstore): repair buckets left with only an incarnation sidecar (#8008)
The legacy bucket metadata migration writes the `.bucket-incarnation`
sidecar before `.metadata.bin`. A crash or a lost namespace lease between
the two writes left the bucket with a sidecar and no metadata, and every
later load failed closed with "bucket incarnation sidecar exists without
bucket metadata", so the bucket answered 500 to every request on every
node with no repair path (rustfs/rustfs#8003).

Load the sidecar-only state as a legacy bucket so the migration runs
again. The migration and the force-create path re-read the sidecar under
the transaction lock and keep its incarnation when persisting the
metadata, so the retry never replaces an identity other nodes fenced on.
A retired incarnation is never adopted: it is residue of a deleted
bucket, and re-publishing it would let heal reclaim the new objects.
2026-09-18 18:14:00 +08:00
cxymds b0f4e66c92 fix(tier): model R2 as unversioned storage (#8004) 2026-09-18 10:12:51 +00:00
Hauser 023adb1696 fix(heal): discharge absent durable partial writes (#8000)
Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-18 16:07:18 +08:00
Hauser 0f69c073de fix(ecstore): preserve bounded listing completion reason (#7994)
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-18 15:41:44 +08:00
cxymds 7a5e1efa95 fix(heal): fail admin heal with unhealthy drives (#7996) 2026-09-18 12:58:24 +08:00
Hauser 785b44a696 fix(replication): harden delete marker HEAD convergence (#7993)
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-18 11:05:47 +08:00
Hauser 31f446bed1 fix(ecstore): preserve vectored bitrot writes (#7981)
* fix(ecstore): preserve vectored bitrot writes

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(ecstore): route test trait through contracts

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-18 08:07:37 +08:00
Hauser 717e8b9c87 fix(scanner): resume and persist checkpoints across leader handoff (#7982) 2026-09-18 06:04:35 +08:00
cxymds db4b8290e7 fix(notify): validate bucket notification config before saving (#7980) 2026-09-17 22:58:29 +08:00