Commit Graph

6855 Commits

Author SHA1 Message Date
RustFS 2e014d25b3 fix: bound restarted-peer stalls on quorum reads and writes (#8152)
A restarted peer can accept a pooled connection and never send response
headers, so HttpReader::open waited past the client body timeout before
the body-stall timer or erasure hedge could run. Bound that header wait
by the stall timeout and retry the open once on a fresh connection.

After write quorum, MultiWriter still waited out the full disk stall
for a silent peer, which matches the client timeout. Give remaining
writers one second, then drop them so the caller returns.

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
2026-09-28 02:22:41 +00:00
Chris 719000460b ci(connect): match GNU build runner in acceptance checks (#8180) 2026-09-28 09:08:25 +08:00
hector 528a368144 ci: refresh the preview-release contract for the current build workflow (#8178)
The contract check asserts exact lines in build.yml, and five recent
build-workflow PRs drifted it out of sync, so Security Audit's Workflow
Pin Report job fails for every PR and every push to main since Sep 27
17:34Z (first failure on push run 36337450094). Refresh the assertions
to the current, intentional shapes - contract intent is preserved and
in fact tightened:

- #8160 added PROFILE_ARGS to the cross/native cargo build lines.
- #8171 added 'should_publish == true' as a stricter prefix on the R2
  publication guard and the create-release / upload-release-assets /
  publish-release / update-latest-version conditions.

Verified locally: the full check (including the docker and helm
workflow contract tests) passes against main's files with these
assertions.
2026-09-28 08:54:33 +08:00
Chris d791d7b761 ci(connect): use Node 25 for acceptance workflows (#8179) 2026-09-28 08:52:32 +08:00
Chris 1a251de4f3 feat(connect): execute signed disk diagnostics in the service (#8177) 2026-09-28 08:36:49 +08:00
Chris e634df8611 fix: collect offline Top disk from the running service (#8174)
fix(connect): capture offline disk IO in the running service
2026-09-28 01:35:40 +08:00
Chris f5fbb5f4e1 ci: verify Top disk in the serving process (#8173) 2026-09-28 01:34:29 +08:00
Chris 5d2d539423 ci: default manual builds to artifacts without publishing (#8171)
Default manual builds to artifacts without publishing
2026-09-28 01:10:52 +08:00
Chris acebc9649d feat(connect): capture service runtime profiles over local IPC (#8168)
feat(connect): export service runtime profiles over local IPC
2026-09-28 00:49:45 +08:00
Chris e33542b0c2 fix(scanner): recover after cycle-state persistence failures (#8148)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 16:39:41 +00:00
Chris 2ee0ad1704 Collect bounded metrics from the serving Tokio runtime (#8167)
feat(connect): collect bounded service runtime profiles
2026-09-28 00:20:19 +08:00
Chris 4c5fc3c991 Use pinned release catalogs for Connect CPU acceptance (#8166) 2026-09-28 00:18:41 +08:00
Chris 6d2b1c629e Generate CPU symbol catalogs from release executables (#8165) 2026-09-28 00:12:36 +08:00
RustFS de484b93e0 fix: admit cold reads and readiness at read quorum (#8156)
EC 2+2 still meets read quorum with two of four nodes up. A survivor
that had not cached a bucket returned 503 on GET, HEAD, and List, and
/health/ready left the Service once write quorum was lost. Reads and
Service membership now follow read quorum and shared locks. Writes and
/minio/health/cluster still require write quorum and exclusive locks.

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 13:58:55 +00:00
hector 4f61b007da ci: build linux-aarch64 natively on sm-standard-4-arm and verify every Linux package (#8160)
Route the two linux-aarch64-* release legs to the dedicated arm64 runner
and build them natively (cross: false) instead of cross-compiling via
cargo-zigbuild on x86_64 hosts.

With native legs now executable, extend the packaged-artifact checks that
were previously gated to x86_64-unknown-linux-gnu only:

- Every Linux leg runs rustfs --version and rustfs-cli --help from its own
  package, proving the artifact matches the runner architecture.
- The packaged-console runtime smoke (server boot + console HTTP 200) also
  covers aarch64-unknown-linux-gnu, so both architectures get full runtime
  verification. Musl legs run the binary checks but skip server boot for
  now.
2026-09-27 13:40:23 +00:00
cxymds 1dd354233c test(connect): sync heartbeat capability expectations (#8161) 2026-09-27 12:22:48 +00:00
Chris 6eb14a11d0 fix(ci): run health acceptance on Docker-capable workers (#8158) 2026-09-27 14:26:28 +08:00
Chris 3667eea893 fix(ci): validate health binary without file utility (#8157) 2026-09-27 14:22:49 +08:00
hector e67bbd7bd3 ci: run the chain orchestration jobs on the smoke-testing runner (#8147)
* ci: run the chain orchestration jobs where the gh CLI exists

#8112 moved these jobs to the sm-standard-2 pool, whose images ship
without the gh CLI (a known property of the build fleet, see #7572).
Last night's chain died in prepare before any suite ran:
resolve_functional_candidate.py:34 does subprocess.run(["gh", "api"])
and got FileNotFoundError; complete-chain and the hourly
functional-chain-health canary fail the same way. On Sep 22/23 the same
jobs ran green on GitHub-hosted runners.

Move prepare, complete-chain and the health canary to ubuntu-latest.
The suite lanes stay on smoke-testing.

* ci: run the chain orchestration jobs where the gh CLI exists

#8112 moved these jobs to the sm-standard-2 pool, whose images ship
without the gh CLI (a known property of the build fleet, see #7572).
Last night's chain died in prepare before any suite ran:
resolve_functional_candidate.py:34 does subprocess.run(["gh", "api"])
and got FileNotFoundError; complete-chain and the hourly
functional-chain-health canary fail the same way. On Sep 22/23 the same
jobs ran green on GitHub-hosted runners.

Move prepare, complete-chain and the health canary to ubuntu-latest.
The suite lanes stay on smoke-testing.

* ci: point the chain orchestration jobs at the smoke-testing runner

Per maintainer decision, consolidate them onto the same runner as the
test lanes. Verified on rustfs-smoke-testing as the runner user:
gh 2.45.0, jq 1.7.

* ci: point the chain orchestration jobs at the smoke-testing runner

Per maintainer decision, consolidate them onto the same runner as the
test lanes. Verified on rustfs-smoke-testing as the runner user:
gh 2.45.0, jq 1.7.
2026-09-27 08:26:43 +08:00
Chris 8d033c85d8 fix(ci): install GitHub CLI for Connect health acceptance (#8146) 2026-09-27 08:04:54 +08:00
Chris b7420fe5f2 fix(ci): bootstrap release upload CLIs (#8144)
fix(ci): bootstrap release upload CLIs on self-hosted runners
2026-09-27 08:00:48 +08:00
Chris 5cb1c9e8cb fix(architecture): route health.rs storage imports through storage_api facade (#8143)
The health diagnostics module imported directly from
crate::storage::storage_api, bypassing the architecture migration
guardrail. Add a connect facade module to rustfs/src/storage_api.rs
and redirect the imports. Also fix a clippy::redundant-guards lint
in the coarse-flags match arm.
2026-09-27 07:53:54 +08:00
Hiroaki KAWAI 20356ab709 fix(scanner): build raw enumeration indexes incrementally (#8114)
* fix(scanner): build raw enumeration indexes incrementally

* test(scanner): iterate over restart test entries directly

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 07:49:59 +08:00
Chris 2c5c43b0e0 test(connect): expose typed drive measurement failure context (#8099)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 05:11:47 +08:00
Chris cf096b3c22 fix(connect): pace diagnostic network payload sends (#8096)
* fix(connect): pace diagnostic network payload sends

* test(connect): exercise native network pacing entry

---------

Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 05:11:35 +08:00
Chris 82024907c5 test: add health service acceptance workflow (#8140) 2026-09-27 03:55:17 +08:00
Chris e2785846d0 ci: keep gh-dependent release jobs on GitHub-hosted runners (#8137)
The sm-standard-2 fleet has no gh CLI (see #7572), so moving the release publication jobs onto it in #8112 broke every tag release: the 1.0.1-preview.12 build failed in Create GitHub Release with 'gh: command not found'. Move create-release, upload-release-assets, publish-release, cleanup-preview-releases, and package.yml resolve/package back to ubuntu-latest.
2026-09-27 03:37:23 +08:00
Chris 31ac243865 feat(connect): add bounded health service job 2026-09-27 02:50:49 +08:00
Chris 8741bb77ea fix(storage): weight automatic multipart admission by part size (#8118)
* fix(storage): weight automatic multipart admission by part size

* fix(ci): restore filesystem runner capabilities and typos dependency

* fix(ci): make release guard portable and spell out part variables

* ci: restore sm-standard-2 runners for io_uring and distributed e2e jobs

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
1.0.1-preview.12
2026-09-26 22:35:39 +08:00
cxymds 212b20f890 fix(ecstore): admit cold bucket metadata at read quorum (#8120)
* fix(ecstore): admit cold bucket metadata at read quorum

Encode a durable creation-commit state in bucket metadata so committed Object Lock buckets can use read quorum for cold loads while pre-physical creation intents remain fail-closed. Drain commit-write fan-out, keep write quorum for uncommitted intents, and add regression coverage for exact, below, migrated, and partial-create quorum boundaries.

* fix(ecstore): fence bucket creation-commit persistence

Address review findings on the creation-commit proof.

Persist the proof only under the bucket metadata transaction fence: at bucket creation, and through a fenced migration that re-reads the authoritative metadata and revalidates physical presence at write quorum before writing. This stops a stale snapshot from reverting an acknowledged configuration update or outliving a delete/recreate.

Establish commitment when Object Lock is enabled on an existing bucket, inside the same configuration mutation, so a healthy cluster no longer rejects object operations with ErasureWriteQuorum.

* fix(ecstore): keep commit fence error message stable

The error(format!) ratchet requires a stable Display for quorum bucketing; use a fixed message instead of embedding the bucket name.

* test(ecstore): reuse canonical Object Lock fixture in regression tests

The s3s footprint ratchet is shrink-only. Use the existing ENABLED_OBJECT_LOCK_CONFIG static instead of naming s3s DTO types in store tests.

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-26 21:40:40 +08:00
hector 7d3d0b6c2b ci: align the matrix contract with the 36-combination sweep (#8136)
Companion to rustfs/auto-testing#109: the 4x1 topology is removed from
MATRIX_TOPOS because the product rejects a 1-drive-per-node distributed
pool at startup (FATAL BelowMinimum). The unfiltered sweep is now 36
combinations / 300 cases; update the contract totals and comments so an
unfiltered sweep is still held to an exact expected case count.
2026-09-26 20:31:32 +08:00
hector d317bb3274 ci: run the KMS Vault lanes on ubuntu-latest again (#8133)
#8112 replaced ubuntu-latest with sm-standard-2 across workflows. Both nightly
KMS Vault lanes need a working Docker daemon; the sm-standard-2 pool does not
provide one, so both lanes die in seconds on 'Docker is not available' while
the package compiles and publishes fine - and the functional chain never fires
because its gate requires a successful build run. Restore the lanes to
ubuntu-latest, as documented in the lane comment before #8112.
2026-09-26 19:12:18 +08:00
Hauser 5cac685e9a feat(ecstore): integrate io_uring backend improvements (#8105) 2026-09-26 17:14:17 +08:00
Hauser d8bf268885 fix(scanner): bound SNSD deep scans and checkpoint cloning (#8126)
* fix(scanner): bound SNSD deep scans and checkpoint cloning

* ci: run mount-dependent jobs on hosted VMs

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
2026-09-26 06:06:50 +00:00
hector fdd753a1dd ci: run distributed e2e on ubuntu-latest to restore tmpfs mounts (#8128) 2026-09-26 12:00:26 +08:00
Chris de80e8a673 fix(ci): make shared checks portable across runners (#8124)
* fix(ci): isolate monitor argument checks from runner tools

* fix(ci): install the Typos action download dependency

* fix(ci): make release policy matching portable across awk variants
2026-09-26 10:55:06 +08:00
RustFS 4cf45e9ed2 Update bug report storage requirements: allow SAN/JBOD, recommend XFS (#8125)
docs(github): clarify supported storage in bug reports

State that local disks, SAN/Fibre Channel volumes, and JBOD are supported,
recommend XFS, and keep the existing unsupported remote/shared filesystem rule.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-09-26 10:43:33 +08:00
Chris 9f8acf688c feat(github): add issue forms with storage support checks (#8123) 2026-09-26 10:07:49 +08:00
Chris 96fc26cd6a fix(release): defer installation updates until artifacts are live (#8122) 2026-09-26 09:17:15 +08:00
hector edcc81a8fd ci: replace ubuntu-latest runner with sm-standard-2 across workflows (#8112)
* ci: replace ubuntu-latest runner with sm-standard-2 across workflows

* ci: keep scheduled-validation monitors on hosted runners

The freshness and watchdog jobs report stalled scheduled validations.
Running them on the same sm-standard-2 pool means a pool outage stalls
the monitors too, so nothing reports it.

---------

Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-09-25 10:02:45 +08:00
Hauser 286bbb8e1f chore: upgrade workspace dependency releases (#8107)
* chore: upgrade workspace dependencies

* chore: refresh dependency lockfile

* fix(rebalance): avoid waiting on active entry gate during stop preparation
2026-09-24 21:24:14 +08:00
cxymds 0913686b12 fix(heal): detect stale delete-marker metadata (#8102) 1.0.1-preview.11 2026-09-24 11:16:02 +08:00
George Melikov df64242f64 fix(ecstore): resume scan_dir past a dash-suffixed sibling directory (#8071)
A flat (recursive) listing resumed inside "s/" re-emitted everything in
its sibling "s-x/". scan_dir dropped the entries before forward_to by
comparing directory names without their trailing slash, where "s" sorts
before "s-x", while the keys they stand for sort the other way round:
'-' (0x2d) is below '/' (0x2f), so all of "s-x/..." precedes "s/...".
The drain stopped at "s" and kept "s-x".

The next page then started with keys at or before the marker; the
listing layer filtered all of them out, found no more candidates and
answered IsTruncated=false. A bucket of 104,137 objects with backup
directories named "<id>" and "<id>-rollbacks" listed as 5,000; two such
directories of 1,200 keys each listed as 2,000.

Compare every remaining entry as the key prefix it stands for, slash
included, and keep it only when it sorts at or after forward_to or
contains it. The remainder of forward_to is taken before `current` is
trimmed, so this also holds below the bucket root; retain also drops
entries when all of them precede forward_to, which the drain never did.
2026-09-24 07:13:30 +08:00
Chris d3b75e695e test(connect): add scheduler receipt acceptance workflow (#8078)
* test(connect): add scheduler receipt acceptance workflow

* chore(deps): upgrade crates and fix faster-hex advisory (#8098)

* test(connect): isolate scheduler acceptance workflow

---------

Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-24 03:34:27 +08:00
Chris dde1aa77a1 fix(connect): preserve diagnostic schedules across restart (#8092) 2026-09-24 03:28:37 +08:00
Chris be01e513be feat(connect): sample memory within the running service (#8091)
* feat(connect): sample memory within the running service

* test(connect): add official memory service acceptance

* test(connect): consume the final memory job without cloning
2026-09-24 03:25:28 +08:00
cxymds 9ab034df62 fix(heal): preserve retryable writer failures (#8089)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-24 02:41:16 +08:00
Tyler Hillery 624463c5ba fix(s3): default CopyObject content-type to binary/octet-stream on REPLACE (#8073)
fix: default CopyObject content-type to binary/octet-stream on REPLACE
2026-09-23 18:04:05 +00:00
Chris 713920dcd6 docs: update license copyright to RustFS, Inc. (#8097) 2026-09-24 00:56:08 +08:00
GatewayJ f219a0aba8 perf(ecstore): parallelize multipart I/O setup and metadata reads (#8085)
* perf(ecstore): parallelize multipart I/O setup and metadata reads

* perf(ecstore): share multipart paths and increase read concurrency

* fix(ecstore): route multipart benchmark through storage API facade
2026-09-23 22:19:23 +08:00