Commit Graph

6875 Commits

Author SHA1 Message Date
RustFS fc5609bbb0 fix(s3): make retried CompleteMultipartUpload idempotent (#8153)
Record the upload id on the completed object so a lost-response retry
returns that object's ETag instead of NoSuchUpload, while a different
part list or a replaced object keeps the existing errors.

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-28 19:08:10 +08:00
Chris 64bd224b3b fix(connect): accept the health-era heartbeat and measured Top windows (#8195)
Two Connect tests fail on main in every Test and Lint lane.

Heartbeat: #8155 replaced "today's capabilities minus
profile.memory.service@1" with "the pre-health set minus it". That drops
the release that advertised health.check.service@1 but not yet in-service
memory profiling, so a pending heartbeat persisted by that release now
fails validation and the runtime stops. Accept that release as its own
frozen list.

Top disk/net: execute_*_job sets max_duration_millis to the requested
duration, then the capture compared the measured sleep against it. The
timer only overshoots, so a job that ran exactly as requested was
rejected with LimitExceeded whenever the overshoot reached 1 ms. Report
the authorized window, as Top API already does; validate_capture bounds
it by the limit.
2026-09-28 19:07:39 +08:00
Chris 71f8b607ad fix(ci): restore mainline and scheduled test reliability (#8187)
* fix(ci): restore mainline and scheduled test reliability

* fix(ci): provide GitHub CLI for CPU acceptance

* ci: provide Docker for CPU service acceptance

* ci: provision Python and Docker for OIDC validation

* ci: restore hosted runners for Docker validation

* ci: use verified MinIO release packages for interop

* ci: preserve host ownership of MinIO fixtures

* fix(ci): correct diagnostic limits and isolate startup checks

* test(connect): include object CLI failure details

* test(readiness): initialize unavailable drive diagnostics
2026-09-28 19:07:29 +08:00
GatewayJ cef0c61532 fix(table-catalog): preserve data sequences during file rewrites
Merge the reviewed fix from pull request #8104.
2026-09-28 18:51:58 +08:00
Justin Bradfield fc325bcc56 fix(s3): report x-amz-mp-parts-count on GetObject with partNumber
Merge the reviewed fix from pull request #8115.
2026-09-28 18:51:48 +08:00
cxymds 4a46f5ff3d fix(s3): serve bucket websites on dedicated domains
Merge the reviewed fix from pull request #8183.
2026-09-28 18:48:21 +08:00
cxymds e0a973bf4e fix(s3): preserve prefix listing pagination and marker metadata
Merge the reviewed fix from pull request #8181.
2026-09-28 18:47:03 +08:00
hector 671238458c ci: fail the pool lane when expected step markers are missing
Merge the reviewed fix from pull request #8190.
2026-09-28 18:46:12 +08:00
Chris 151103a609 fix(ecstore): recover remote disks after transient stalls
Merge the approved fix from pull request #8149.
2026-09-28 18:40:30 +08:00
Nikita Bakun cc7d5e3f5f fix(replication): stop sending versionId query once a target rejects it
Merge the approved fix from pull request #8109.
2026-09-28 18:40:20 +08:00
yi111 d556d4dc94 fix(admin): report a missing policy or user as 404 NoSuchResource, not 500 InternalError
Merge the approved fix from pull request #8127.
2026-09-28 18:40:10 +08:00
chapman 6300fbe74f fix: map IAM not-found errors to 404 in InfoCannedPolicy and RemoveUser
Merge the approved fix from pull request #8134.
2026-09-28 18:38:20 +08:00
Chris 6ebc78d115 fix(ci): run profile acceptance on Docker runners (#8191) 2026-09-28 16:14:07 +08:00
Chris 5f4d5e8fe8 fix(ci): run CPU acceptance with Docker
Use the Docker-enabled runner required by the connected CPU service-job test.
2026-09-28 15:38:36 +08:00
Chris e4520f70a1 fix(ci): verify CPU acceptance identity without gh
Use authenticated curl for the pinned release run and artifact checks on sm-standard-2.
2026-09-28 15:13:03 +08:00
hector f54945a4c4 fix(ci): repair pool budgets and Connect test failures (#8155)
* ci: give the pool run a real wall-clock budget

The pool suite budgets up to 24h each for rebalance and decommission
(REBALANCE_TIMEOUT / DECOMMISSION_TIMEOUT in rustfs_pool_expand.sh),
but the workflow step capped it at 45 minutes. Four consecutive runs
died identically: steps 1-7 all PASS, then the rebalance wait was
killed at exactly 45:12 - chain 35680308150, standalone 35702803577,
chain 35757773373, chain 36283120750 - with rebalance at completed=1/4
(~21 minutes in), so a full pass has never been observed.

Make the budget an input (default 240 minutes: covers the observed
rebalance pace plus one decommission pass) and document why. The suite
keeps failing the job through its [POOL-STEP] marker adjudication.

* fix(ci): align pool job budget and timeout contract tests

* fix(connect): preserve legacy heartbeats and isolate I/O tests

---------

Signed-off-by: Hauser <housemecn@gmail.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: RustFS <hello@rustfs.com>
2026-09-28 14:26:21 +08:00
Hauser 4beac12439 fix(connect): return profile capture result directly (#8182)
* fix(connect): return profile capture result directly

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* style(connect): normalize profile capture formatting

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(tests): use async mutex for thread profiles

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* test(connect): widen diagnostic timing budgets

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-28 14:25:27 +08:00
Chris 6523ab8d9f fix(ci): verify GNU packages with CPU catalogs
Match the archive-member check to the GNU pyroscope catalog packaging condition.
2026-09-28 13:23:46 +08:00
Chris 0609e7ce14 fix(connect): admit complete release CPU symbol catalogs
Bound the release catalog to 32 MiB and 200,000 symbols based on the GNU build measurement; keep complete names and reject catalogs beyond either limit.
2026-09-28 12:09:26 +08:00
Chris afb9ddabb4 fix(helm): route Gateway API traffic to the chart service (#8145)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-28 02:26:56 +00:00
RustFS 2e014d25b3 fix: bound restarted-peer stalls on quorum reads and writes (#8152)
A restarted peer can accept a pooled connection and never send response
headers, so HttpReader::open waited past the client body timeout before
the body-stall timer or erasure hedge could run. Bound that header wait
by the stall timeout and retry the open once on a fresh connection.

After write quorum, MultiWriter still waited out the full disk stall
for a silent peer, which matches the client timeout. Give remaining
writers one second, then drop them so the caller returns.

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
2026-09-28 02:22:41 +00:00
Chris 719000460b ci(connect): match GNU build runner in acceptance checks (#8180) 2026-09-28 09:08:25 +08:00
hector 528a368144 ci: refresh the preview-release contract for the current build workflow (#8178)
The contract check asserts exact lines in build.yml, and five recent
build-workflow PRs drifted it out of sync, so Security Audit's Workflow
Pin Report job fails for every PR and every push to main since Sep 27
17:34Z (first failure on push run 36337450094). Refresh the assertions
to the current, intentional shapes - contract intent is preserved and
in fact tightened:

- #8160 added PROFILE_ARGS to the cross/native cargo build lines.
- #8171 added 'should_publish == true' as a stricter prefix on the R2
  publication guard and the create-release / upload-release-assets /
  publish-release / update-latest-version conditions.

Verified locally: the full check (including the docker and helm
workflow contract tests) passes against main's files with these
assertions.
2026-09-28 08:54:33 +08:00
Chris d791d7b761 ci(connect): use Node 25 for acceptance workflows (#8179) 2026-09-28 08:52:32 +08:00
Chris 1a251de4f3 feat(connect): execute signed disk diagnostics in the service (#8177) 2026-09-28 08:36:49 +08:00
Chris e634df8611 fix: collect offline Top disk from the running service (#8174)
fix(connect): capture offline disk IO in the running service
2026-09-28 01:35:40 +08:00
Chris f5fbb5f4e1 ci: verify Top disk in the serving process (#8173) 2026-09-28 01:34:29 +08:00
Chris 5d2d539423 ci: default manual builds to artifacts without publishing (#8171)
Default manual builds to artifacts without publishing
2026-09-28 01:10:52 +08:00
Chris acebc9649d feat(connect): capture service runtime profiles over local IPC (#8168)
feat(connect): export service runtime profiles over local IPC
2026-09-28 00:49:45 +08:00
Chris e33542b0c2 fix(scanner): recover after cycle-state persistence failures (#8148)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 16:39:41 +00:00
Chris 2ee0ad1704 Collect bounded metrics from the serving Tokio runtime (#8167)
feat(connect): collect bounded service runtime profiles
2026-09-28 00:20:19 +08:00
Chris 4c5fc3c991 Use pinned release catalogs for Connect CPU acceptance (#8166) 2026-09-28 00:18:41 +08:00
Chris 6d2b1c629e Generate CPU symbol catalogs from release executables (#8165) 2026-09-28 00:12:36 +08:00
RustFS de484b93e0 fix: admit cold reads and readiness at read quorum (#8156)
EC 2+2 still meets read quorum with two of four nodes up. A survivor
that had not cached a bucket returned 503 on GET, HEAD, and List, and
/health/ready left the Service once write quorum was lost. Reads and
Service membership now follow read quorum and shared locks. Writes and
/minio/health/cluster still require write quorum and exclusive locks.

Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 13:58:55 +00:00
hector 4f61b007da ci: build linux-aarch64 natively on sm-standard-4-arm and verify every Linux package (#8160)
Route the two linux-aarch64-* release legs to the dedicated arm64 runner
and build them natively (cross: false) instead of cross-compiling via
cargo-zigbuild on x86_64 hosts.

With native legs now executable, extend the packaged-artifact checks that
were previously gated to x86_64-unknown-linux-gnu only:

- Every Linux leg runs rustfs --version and rustfs-cli --help from its own
  package, proving the artifact matches the runner architecture.
- The packaged-console runtime smoke (server boot + console HTTP 200) also
  covers aarch64-unknown-linux-gnu, so both architectures get full runtime
  verification. Musl legs run the binary checks but skip server boot for
  now.
2026-09-27 13:40:23 +00:00
cxymds 1dd354233c test(connect): sync heartbeat capability expectations (#8161) 2026-09-27 12:22:48 +00:00
Chris 6eb14a11d0 fix(ci): run health acceptance on Docker-capable workers (#8158) 2026-09-27 14:26:28 +08:00
Chris 3667eea893 fix(ci): validate health binary without file utility (#8157) 2026-09-27 14:22:49 +08:00
hector e67bbd7bd3 ci: run the chain orchestration jobs on the smoke-testing runner (#8147)
* ci: run the chain orchestration jobs where the gh CLI exists

#8112 moved these jobs to the sm-standard-2 pool, whose images ship
without the gh CLI (a known property of the build fleet, see #7572).
Last night's chain died in prepare before any suite ran:
resolve_functional_candidate.py:34 does subprocess.run(["gh", "api"])
and got FileNotFoundError; complete-chain and the hourly
functional-chain-health canary fail the same way. On Sep 22/23 the same
jobs ran green on GitHub-hosted runners.

Move prepare, complete-chain and the health canary to ubuntu-latest.
The suite lanes stay on smoke-testing.

* ci: run the chain orchestration jobs where the gh CLI exists

#8112 moved these jobs to the sm-standard-2 pool, whose images ship
without the gh CLI (a known property of the build fleet, see #7572).
Last night's chain died in prepare before any suite ran:
resolve_functional_candidate.py:34 does subprocess.run(["gh", "api"])
and got FileNotFoundError; complete-chain and the hourly
functional-chain-health canary fail the same way. On Sep 22/23 the same
jobs ran green on GitHub-hosted runners.

Move prepare, complete-chain and the health canary to ubuntu-latest.
The suite lanes stay on smoke-testing.

* ci: point the chain orchestration jobs at the smoke-testing runner

Per maintainer decision, consolidate them onto the same runner as the
test lanes. Verified on rustfs-smoke-testing as the runner user:
gh 2.45.0, jq 1.7.

* ci: point the chain orchestration jobs at the smoke-testing runner

Per maintainer decision, consolidate them onto the same runner as the
test lanes. Verified on rustfs-smoke-testing as the runner user:
gh 2.45.0, jq 1.7.
2026-09-27 08:26:43 +08:00
Chris 8d033c85d8 fix(ci): install GitHub CLI for Connect health acceptance (#8146) 2026-09-27 08:04:54 +08:00
Chris b7420fe5f2 fix(ci): bootstrap release upload CLIs (#8144)
fix(ci): bootstrap release upload CLIs on self-hosted runners
2026-09-27 08:00:48 +08:00
Chris 5cb1c9e8cb fix(architecture): route health.rs storage imports through storage_api facade (#8143)
The health diagnostics module imported directly from
crate::storage::storage_api, bypassing the architecture migration
guardrail. Add a connect facade module to rustfs/src/storage_api.rs
and redirect the imports. Also fix a clippy::redundant-guards lint
in the coarse-flags match arm.
2026-09-27 07:53:54 +08:00
Hiroaki KAWAI 20356ab709 fix(scanner): build raw enumeration indexes incrementally (#8114)
* fix(scanner): build raw enumeration indexes incrementally

* test(scanner): iterate over restart test entries directly

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 07:49:59 +08:00
Chris 2c5c43b0e0 test(connect): expose typed drive measurement failure context (#8099)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 05:11:47 +08:00
Chris cf096b3c22 fix(connect): pace diagnostic network payload sends (#8096)
* fix(connect): pace diagnostic network payload sends

* test(connect): exercise native network pacing entry

---------

Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 05:11:35 +08:00
Chris 82024907c5 test: add health service acceptance workflow (#8140) 2026-09-27 03:55:17 +08:00
Chris e2785846d0 ci: keep gh-dependent release jobs on GitHub-hosted runners (#8137)
The sm-standard-2 fleet has no gh CLI (see #7572), so moving the release publication jobs onto it in #8112 broke every tag release: the 1.0.1-preview.12 build failed in Create GitHub Release with 'gh: command not found'. Move create-release, upload-release-assets, publish-release, cleanup-preview-releases, and package.yml resolve/package back to ubuntu-latest.
2026-09-27 03:37:23 +08:00
Chris 31ac243865 feat(connect): add bounded health service job 2026-09-27 02:50:49 +08:00
Chris 8741bb77ea fix(storage): weight automatic multipart admission by part size (#8118)
* fix(storage): weight automatic multipart admission by part size

* fix(ci): restore filesystem runner capabilities and typos dependency

* fix(ci): make release guard portable and spell out part variables

* ci: restore sm-standard-2 runners for io_uring and distributed e2e jobs

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
1.0.1-preview.12
2026-09-26 22:35:39 +08:00
cxymds 212b20f890 fix(ecstore): admit cold bucket metadata at read quorum (#8120)
* fix(ecstore): admit cold bucket metadata at read quorum

Encode a durable creation-commit state in bucket metadata so committed Object Lock buckets can use read quorum for cold loads while pre-physical creation intents remain fail-closed. Drain commit-write fan-out, keep write quorum for uncommitted intents, and add regression coverage for exact, below, migrated, and partial-create quorum boundaries.

* fix(ecstore): fence bucket creation-commit persistence

Address review findings on the creation-commit proof.

Persist the proof only under the bucket metadata transaction fence: at bucket creation, and through a fenced migration that re-reads the authoritative metadata and revalidates physical presence at write quorum before writing. This stops a stale snapshot from reverting an acknowledged configuration update or outliving a delete/recreate.

Establish commitment when Object Lock is enabled on an existing bucket, inside the same configuration mutation, so a healthy cluster no longer rejects object operations with ErasureWriteQuorum.

* fix(ecstore): keep commit fence error message stable

The error(format!) ratchet requires a stable Display for quorum bucketing; use a fixed message instead of embedding the bucket name.

* test(ecstore): reuse canonical Object Lock fixture in regression tests

The s3s footprint ratchet is shrink-only. Use the existing ENABLED_OBJECT_LOCK_CONFIG static instead of naming s3s DTO types in store tests.

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-26 21:40:40 +08:00