Commit Graph

622 Commits

Author SHA1 Message Date
Chris 641c4b3493 chore(release): merge 1.0.1 back into main and refresh installation references (#8314)
* fix(ci): include pagination regression in full E2E selection

* fix(deps): replace yanked yoke-derive release

* fix(scanner): expose pause backlog replica diagnostics (#8258)

* fix(scanner): expose pause backlog replica diagnostics

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(connect): stabilize runtime profile lease cancellation

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* test(connect): tolerate delayed schedule startup in CI

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* test(e2e): retry quota reads during usage warmup

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(ecstore): avoid meta-bucket incarnation self-deadlock

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(test): use persisted incarnation in heal fixture

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* fix(usage): reconcile stale counters after lifecycle expiration (#8108)

* fix(usage): reconcile stale counters after lifecycle expiration

* fix(usage): account lifecycle expiry during continuous writes

* test(usage): run lifecycle usage scenarios on one scanner store

* test(usage): use a Windows-representable pre-mutation offset

* fix(usage): harden expiry accounting recovery and quota checks

Borrow expiry receipt bucket names and avoid allocating a map key on cache hits. Cover cancelled receipts, durable snapshot recovery, and legacy quota admission after scanner confirmation. Use representable timestamp offsets in the quota regression.

Refs rustfs/backlog#2689

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* test(usage): recover stale persisted counts through the scanner

Seed incorrect complete usage for empty and retained-object buckets, then run the real scanner and publication consumer without further object mutations. Verify durable and admin usage over two cycles instead of writing a corrected snapshot in the test.

Refs rustfs/backlog#2689
Refs rustfs/backlog#2691

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* test(e2e): bound delimiter pagination fixture concurrency

The 120-second smoke timeout expired after 1018 of 1200 serial fixture PUTs, before LIST ran. Prepare the same objects with at most eight concurrent requests and await every PUT. Retain the timeout and strengthen exact prefix, KeyCount, empty Contents, and continuation-token assertions with phase diagnostics.

Refs rustfs/backlog#2689

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* test(upgrade): establish a persisted previous-release baseline

Seed the pinned previous-release cluster and restart it once with its data intact before replacing any node. Require every old writer to pass the strict readiness probe and preserve the seed through both mixed phases and the final current cluster. Keep InternalError fail-fast behavior and all existing compatibility deadlines and assertions.

Refs rustfs/backlog#2689
Refs rustfs/backlog#2384

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* test(ecstore): bound cancellation metadata persistence waits

Use the system-bucket incarnation boundary now supplied by main PR #8268. Bound the three cancellation waits that previously hung during pool.bin persistence, retaining their remote-generation, target-cohort, and durable-state assertions.

Refs rustfs/backlog#2697
Refs rustfs/backlog#2689

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: Chris <anzhengchao@gmail.com>
Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* test(e2e): retain startup and shutdown failure diagnostics

* fix(test): supply CPU workload for sampler regression

* fix(ci): locate security chain scripts in the workspace

* fix(usage): recover historical counters with generation fencing (#8273)

* fix(usage): recover historical counters during continued writes

Use newer converged scanner snapshots to reconcile stale absolute usage
baselines while preserving concurrent mutation and expiry receipt fences.
Cover durable publication, admin and quota reads, and legacy generations.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* refactor(usage): fence snapshots and move preserved cache entries

Apply the cached scanner generation floor before every reconciliation path
and retain it even when an older snapshot happens to match core counts.
Move preserved usage entries instead of cloning their histogram maps under
the cache lock, retaining expiry receipt identity and cancellation fences.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* test(scanner): verify checkpoint takeover and repair dispatch (#8275)

* test(scanner): cover checkpoint handoff and repair dispatch

Drive runtime budget expiry, partial-cycle persistence, leadership claims,
and stale checkpoint rejection between real disk-backed fixture scans.
Verify that a metadata repair beyond the first bounded prefix is saved in
the scanner ledger and dispatched by the MRF consumer.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* ci: isolate scanner fixtures and refresh full e2e membership

Reserve nextest capacity for the real-disk scanner publication and MRF
admission fixtures. Bind both platform membership checks to the reviewed
pagination deadline test added on main.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* refactor(scanner): consolidate checkpoint fixture lifecycle

Keep one durable control store across timeout and leadership transitions,
and inject generation advancement into the shared checkpoint scenario.
Check the actual saved metadata path so late-write rejection also proves
that existing checkpoint bytes remain intact.

Centralize MRF fixture isolation and reuse nextest process isolation when
the startup environment already satisfies the test contract.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* fix(obs): distinguish allocator counters from live memory (#8274)

* fix(obs): distinguish allocator counters from live memory

Preserve count/counter semantics and mark requested-byte attribution unavailable when live statistics or sampling are missing. Document sustained multipart memory diagnosis.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* refactor(obs): parse allocator statistics from one node

Resolve each statistic before interpreting its shape, avoiding unsupported-field tree scans and mixing data across wrapper scopes. Preserve unavailable-statistics policy and add precedence regressions.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* test(ecstore): isolate late parity recovery from metadata hedges

Use the existing object-scoped hedge timer barrier in exact-count recovery fixtures. Preserve payload and total-read assertions and verify that the omitted parity disk is read only during late refresh.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* fix(usage): combine identical snapshot retention branches

* fix(test): await HTTP sender readiness in Top RPC fixture

* [release/1.0.1] Gate multipart copy through write admission (#8284)

Gate multipart copy through write admission

Make UploadPartCopy acquire the shared foreground write admission permit before lifecycle locks or source readers so server-side multipart copy cannot bypass the same backpressure used by UploadPart. Document the shared queue semantics and add focused coverage for saturation, cancellation, lock ordering, and disabled admission.

Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* Gate multipart copy through write admission (#8283)

Make UploadPartCopy acquire the shared foreground write admission permit before lifecycle locks or source readers so server-side multipart copy cannot bypass the same backpressure used by UploadPart. Document the shared queue semantics and add focused coverage for saturation, cancellation, lock ordering, and disabled admission.

Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* fix: add UploadPart OOM validation guardrails (#8287)

* docs(release): validate candidates on release branch

* fix(scanner): validate checkpoints against global cycle fence (#8278)

## Related Issues

Related to rustfs/backlog#2701.

## Summary of Changes

Route scanner checkpoint cycle and leader validation through the global store while retaining the owning set for cache persistence, CAS revisions and publication admission.

## Verification

Two independent final-diff source reviews found no issues across correctness, concurrency and durability, test coverage, compatibility, performance and simplicity on head `8b8fe51d092090b053f552ae283960e2e306be33`. Root approval `5373624714` is bound to that exact head. Regression tests cover real two-pool routing, stale fences, post-save rejection and CAS conflicts; their reported local execution belongs to the PR author, not this merge operation. Current required CI remains pending, and this authorized admin squash does not establish CI or runtime acceptance.

## Impact

Restores checkpoint progress when global cycle and leader state differ from a set-scoped view. No format, retry, timeout, assertion or scanner-policy changes are introduced by this diff. The three prior main scanner failures remain unproved repaired.

## Additional Notes

Full validation must run on the resulting exact main revision. Reverting this patch restores the earlier set-scoped fence lookup and its checkpoint rejection behavior.

* fix(ci): restore E2E membership and pagination timeouts (#8281)

* ci: locate the auto-testing checkout for lanes that run evidence from a subdirectory (#8279)

## Related Issues

Follow-up to #8229.

## Summary of Changes

Locate the private auto-testing checkout from the lane root or the workspace root so the nested security checkout can record functional-chain evidence.

## Verification

The exact PR head b66129ab9f passed one mechanical correctness and simplicity review, nine real-Git layout and provenance checks, and sixteen existing evidence/envelope tests. The baseline sibling layout failed with git exit 128; the corrected layout succeeded while revision mismatches and missing checkouts stayed rejected. Current PR checks are completed with successful or skipped conclusions, including the aggregate.

## Impact

Both lane and private-script revision checks remain intact. No time limits, assertions, production behavior, or evidence validation requirements change. The synthetic layout checks do not execute the actual scheduled security suite; integrated main CI and release acceptance remain separate gates.

## Additional Notes

Approved review 5374559393 is bound to the exact head above. Reverting the single-file change restores the previous checkout lookup.

* fix: add UploadPart OOM validation guardrails

Add a Docker validation harness for backlog#2704 so the ordinary UploadPart
low-concurrency memory workload can be reproduced with comparable case metadata,
process/cgroup sampling, TLS and metrics toggles, cache-env controls, and write
reclaim/direct-write experiments.

Warn when operators set the unrecognized RUSTFS_OBJECT_CACHE_* variables that
appeared in the reporter compose file. The variables are reported but remain
ignored, so startup does not silently change object data cache behavior.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: overtrue <anzhengchao@gmail.com>
Co-authored-by: AL <allan.bednarowski@gmail.com>
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* fix(scanner): pass fence store to checkpoint fixture

* fix(s3): queue bucket operations and restore strict Clippy checks (#8290)

* fix(s3): queue concurrent bucket creation and deletion

Keep eight active bucket transactions and bound admission waiting to 128 requests and 30 seconds. Preserve detached transaction ownership and return Retry-After with overload responses.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(ecstore): restore strict Clippy compatibility on Rust 1.99

Use try_update without changing atomic ordering or overflow behavior. Keep
recursive storage futures boxed once at each frame and remove the redundant
async-recursion macro, including its non-recursive SQL planner use. Remove
needless closure borrows and orphaned dependency entries.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* chore: refresh dependencies and atomic update APIs

Update workspace dependencies and the lockfile. Replace deprecated
atomic fetch_update aliases with try_update while preserving closures
and memory ordering.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* improve

* fix(scanner): diagnose and verify pause backlog recovery (#8293)

* fix(scanner): diagnose and verify pause backlog recovery

Expose the retained replica snapshot and claimed membership in abnormal
admin status responses. Keep diagnostics off metrics updates and verify
single-pool recovery and conflicting-proof preservation across 24 sets.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* refactor(scanner): move replica snapshots into diagnostics

Consume the terminal admin read snapshot in a single state match and move
membership, revision, and error buffers into the response. Verify buffer
handoff and the unchanged JSON contract without altering ledger authority.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* fix(heal): wait for held legacy responsibility in replay test

* ci: remove Docker Hub description sync

* fix: emit NextPartNumberMarker only when ListParts is truncated

ListPartsInfo.next_part_number_marker was a non-optional usize that
defaulted to 0 and was only assigned when the response was truncated.
The S3 serializer then emitted it unconditionally as Some(0), causing
AWS SDK paginators to loop infinitely on part_number_marker=0 instead
of terminating.

Change the field to Option<usize> (None by default) and set it only
inside the is_truncated branch. The S3 output layer now uses
.and_then() so NextPartNumberMarker is absent when IsTruncated=false,
matching AWS S3 behavior.

Fixes #8208

(cherry picked from commit 44de803a38)

* fix(s3): honor sparse ListParts markers and verify termination

Resume part listings at the first part above the numeric marker, even
when that marker is absent. Use binary search over the sorted part
numbers and retain the existing exact-tail empty-slice path.

Add storage, XML, and real AWS SDK paginator regressions for empty and
terminal pages, sparse markers, and multipart completion integrity.

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-Authored-By: zhi22915 <qiuzgang@gmail.com>
(cherry picked from commit 673031eea1)

* fix(storage): publish delete rollback backups atomically

Stage rollback metadata outside the rollback directory and publish it only after the full write succeeds. A short write must not leave a backup that quorum rollback can rename over acknowledged version history.

Add an isolated real short-write regression and register the backported ListParts SDK test in the smoke and Linux full inventories.

* test(e2e): register paginator regression in Darwin inventory

* fix(release): install yq before Helm template checks

* chore(release): align installation references for 1.0.1

---------

Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Co-authored-by: Peder Bergan <pederbe@users.noreply.github.com>
Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: AL <allan.bednarowski@gmail.com>
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Chapman <touch65536@gmail.com>
2026-10-03 14:40:52 +08:00
Chris 87c1c851c2 ci: remove Docker Hub description sync (#8303) 2026-10-02 20:04:45 +08:00
Chris 2806a80c91 fix: prevent metadata lock reentry and stabilize CI fixtures (#8257)
* fix(test): establish a writable previous-release upgrade baseline

* fix(ci): reserve capacity for durable admin fixtures

* fix(ci): group durable IAM state fixtures by resource needs

* ci: run E2E doctests with the E2E dependency graph

* fix(ci): separate fixture startup from transport deadlines

* fix(ci): bound pagination after seeding and revisit restored copies

* fix(ci): make filesystem fixture timing deterministic

* fix(ecstore): avoid metadata lock reentry during internal mutations

* test: align recovery fixtures with durable ownership contracts
2026-09-30 20:57:07 +08:00
KarlEmm e573fb087b ci(helm): sync README to root of helm repository on release (#8263)
Include helm/rustfs/README.md in the uploaded helm-package artifact so
publish-helm-package updates the root README.md in rustfs/helm.

The build-helm-package job copies helm/README.md into helm/rustfs/
before running helm package, but upload-artifact previously uploaded
only helm/rustfs/*.tgz. Because both files share the helm/rustfs/
parent directory, actions/upload-artifact places both the chart
archive and README.md at the artifact root, and download-artifact
extracts README.md into the root of rustfs/helm alongside the tarball.
2026-09-30 15:57:24 +08:00
Chris 9e33d54269 fix(release): package stable preview tags without updating channels (#8256) 2026-09-30 09:49:27 +08:00
Chris a88225d208 fix(ci): reduce duplicate work and preserve reliable test failures (#8233)
* fix(ci): reduce duplicate work and preserve reliable test failures

* fix(ci): retain protocol evidence and repair stale test fixtures

* test(connect): honor parent deadline during API fixture readiness

* fix(ci): reserve IO capacity for state writer proofs

* test(connect): align RPC fixtures with service capture contracts

* test(connect): cover pinned service capture failures
2026-09-29 21:06:45 +08:00
hector 0e58bd09d0 fix(ci): run security chain evidence from the lane's own checkout (#8229)
The security workflow checks the repository out into rustfs-repo/ (#7212)
but its chain evidence steps still invoke
scripts/functional_chain_evidence.py relative to the workspace root, which
on the persistent shared runner resolves to a stale checkout left by another
job. The evidence gate then compares that checkout's HEAD against the chain
workflow_sha and rejects the lane before any case runs
("lane checkout differs from chain workflow source").

Run the evidence script from the lane's own checkout so ROOT resolves to
rustfs-repo/, whose HEAD is exactly the chain-pinned workflow_sha.
2026-09-29 16:09:31 +08:00
Chris 47af5565c7 fix(ci): preserve cluster startup failure evidence (#8213) 2026-09-29 08:50:09 +08:00
Chris 8778d55e46 fix(ci): preserve diagnostics for rio storage test failures (#8204) 2026-09-29 00:15:26 +08:00
Chris 2648776be7 fix: accept catalog files in health service artifact (#8198)
* fix: accept catalog files in health service artifact

* ci: retain configured health acceptance runner
2026-09-28 20:45:26 +08:00
Chris 71f8b607ad fix(ci): restore mainline and scheduled test reliability (#8187)
* fix(ci): restore mainline and scheduled test reliability

* fix(ci): provide GitHub CLI for CPU acceptance

* ci: provide Docker for CPU service acceptance

* ci: provision Python and Docker for OIDC validation

* ci: restore hosted runners for Docker validation

* ci: use verified MinIO release packages for interop

* ci: preserve host ownership of MinIO fixtures

* fix(ci): correct diagnostic limits and isolate startup checks

* test(connect): include object CLI failure details

* test(readiness): initialize unavailable drive diagnostics
2026-09-28 19:07:29 +08:00
hector 671238458c ci: fail the pool lane when expected step markers are missing
Merge the reviewed fix from pull request #8190.
2026-09-28 18:46:12 +08:00
Chris 6ebc78d115 fix(ci): run profile acceptance on Docker runners (#8191) 2026-09-28 16:14:07 +08:00
Chris 5f4d5e8fe8 fix(ci): run CPU acceptance with Docker
Use the Docker-enabled runner required by the connected CPU service-job test.
2026-09-28 15:38:36 +08:00
Chris e4520f70a1 fix(ci): verify CPU acceptance identity without gh
Use authenticated curl for the pinned release run and artifact checks on sm-standard-2.
2026-09-28 15:13:03 +08:00
hector f54945a4c4 fix(ci): repair pool budgets and Connect test failures (#8155)
* ci: give the pool run a real wall-clock budget

The pool suite budgets up to 24h each for rebalance and decommission
(REBALANCE_TIMEOUT / DECOMMISSION_TIMEOUT in rustfs_pool_expand.sh),
but the workflow step capped it at 45 minutes. Four consecutive runs
died identically: steps 1-7 all PASS, then the rebalance wait was
killed at exactly 45:12 - chain 35680308150, standalone 35702803577,
chain 35757773373, chain 36283120750 - with rebalance at completed=1/4
(~21 minutes in), so a full pass has never been observed.

Make the budget an input (default 240 minutes: covers the observed
rebalance pace plus one decommission pass) and document why. The suite
keeps failing the job through its [POOL-STEP] marker adjudication.

* fix(ci): align pool job budget and timeout contract tests

* fix(connect): preserve legacy heartbeats and isolate I/O tests

---------

Signed-off-by: Hauser <housemecn@gmail.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: RustFS <hello@rustfs.com>
2026-09-28 14:26:21 +08:00
Chris 6523ab8d9f fix(ci): verify GNU packages with CPU catalogs
Match the archive-member check to the GNU pyroscope catalog packaging condition.
2026-09-28 13:23:46 +08:00
Chris 719000460b ci(connect): match GNU build runner in acceptance checks (#8180) 2026-09-28 09:08:25 +08:00
Chris d791d7b761 ci(connect): use Node 25 for acceptance workflows (#8179) 2026-09-28 08:52:32 +08:00
Chris f5fbb5f4e1 ci: verify Top disk in the serving process (#8173) 2026-09-28 01:34:29 +08:00
Chris 5d2d539423 ci: default manual builds to artifacts without publishing (#8171)
Default manual builds to artifacts without publishing
2026-09-28 01:10:52 +08:00
Chris 2ee0ad1704 Collect bounded metrics from the serving Tokio runtime (#8167)
feat(connect): collect bounded service runtime profiles
2026-09-28 00:20:19 +08:00
Chris 4c5fc3c991 Use pinned release catalogs for Connect CPU acceptance (#8166) 2026-09-28 00:18:41 +08:00
Chris 6d2b1c629e Generate CPU symbol catalogs from release executables (#8165) 2026-09-28 00:12:36 +08:00
hector 4f61b007da ci: build linux-aarch64 natively on sm-standard-4-arm and verify every Linux package (#8160)
Route the two linux-aarch64-* release legs to the dedicated arm64 runner
and build them natively (cross: false) instead of cross-compiling via
cargo-zigbuild on x86_64 hosts.

With native legs now executable, extend the packaged-artifact checks that
were previously gated to x86_64-unknown-linux-gnu only:

- Every Linux leg runs rustfs --version and rustfs-cli --help from its own
  package, proving the artifact matches the runner architecture.
- The packaged-console runtime smoke (server boot + console HTTP 200) also
  covers aarch64-unknown-linux-gnu, so both architectures get full runtime
  verification. Musl legs run the binary checks but skip server boot for
  now.
2026-09-27 13:40:23 +00:00
Chris 6eb14a11d0 fix(ci): run health acceptance on Docker-capable workers (#8158) 2026-09-27 14:26:28 +08:00
Chris 3667eea893 fix(ci): validate health binary without file utility (#8157) 2026-09-27 14:22:49 +08:00
hector e67bbd7bd3 ci: run the chain orchestration jobs on the smoke-testing runner (#8147)
* ci: run the chain orchestration jobs where the gh CLI exists

#8112 moved these jobs to the sm-standard-2 pool, whose images ship
without the gh CLI (a known property of the build fleet, see #7572).
Last night's chain died in prepare before any suite ran:
resolve_functional_candidate.py:34 does subprocess.run(["gh", "api"])
and got FileNotFoundError; complete-chain and the hourly
functional-chain-health canary fail the same way. On Sep 22/23 the same
jobs ran green on GitHub-hosted runners.

Move prepare, complete-chain and the health canary to ubuntu-latest.
The suite lanes stay on smoke-testing.

* ci: run the chain orchestration jobs where the gh CLI exists

#8112 moved these jobs to the sm-standard-2 pool, whose images ship
without the gh CLI (a known property of the build fleet, see #7572).
Last night's chain died in prepare before any suite ran:
resolve_functional_candidate.py:34 does subprocess.run(["gh", "api"])
and got FileNotFoundError; complete-chain and the hourly
functional-chain-health canary fail the same way. On Sep 22/23 the same
jobs ran green on GitHub-hosted runners.

Move prepare, complete-chain and the health canary to ubuntu-latest.
The suite lanes stay on smoke-testing.

* ci: point the chain orchestration jobs at the smoke-testing runner

Per maintainer decision, consolidate them onto the same runner as the
test lanes. Verified on rustfs-smoke-testing as the runner user:
gh 2.45.0, jq 1.7.

* ci: point the chain orchestration jobs at the smoke-testing runner

Per maintainer decision, consolidate them onto the same runner as the
test lanes. Verified on rustfs-smoke-testing as the runner user:
gh 2.45.0, jq 1.7.
2026-09-27 08:26:43 +08:00
Chris 8d033c85d8 fix(ci): install GitHub CLI for Connect health acceptance (#8146) 2026-09-27 08:04:54 +08:00
Chris b7420fe5f2 fix(ci): bootstrap release upload CLIs (#8144)
fix(ci): bootstrap release upload CLIs on self-hosted runners
2026-09-27 08:00:48 +08:00
Chris 82024907c5 test: add health service acceptance workflow (#8140) 2026-09-27 03:55:17 +08:00
Chris e2785846d0 ci: keep gh-dependent release jobs on GitHub-hosted runners (#8137)
The sm-standard-2 fleet has no gh CLI (see #7572), so moving the release publication jobs onto it in #8112 broke every tag release: the 1.0.1-preview.12 build failed in Create GitHub Release with 'gh: command not found'. Move create-release, upload-release-assets, publish-release, cleanup-preview-releases, and package.yml resolve/package back to ubuntu-latest.
2026-09-27 03:37:23 +08:00
hector 7d3d0b6c2b ci: align the matrix contract with the 36-combination sweep (#8136)
Companion to rustfs/auto-testing#109: the 4x1 topology is removed from
MATRIX_TOPOS because the product rejects a 1-drive-per-node distributed
pool at startup (FATAL BelowMinimum). The unfiltered sweep is now 36
combinations / 300 cases; update the contract totals and comments so an
unfiltered sweep is still held to an exact expected case count.
2026-09-26 20:31:32 +08:00
hector d317bb3274 ci: run the KMS Vault lanes on ubuntu-latest again (#8133)
#8112 replaced ubuntu-latest with sm-standard-2 across workflows. Both nightly
KMS Vault lanes need a working Docker daemon; the sm-standard-2 pool does not
provide one, so both lanes die in seconds on 'Docker is not available' while
the package compiles and publishes fine - and the functional chain never fires
because its gate requires a successful build run. Restore the lanes to
ubuntu-latest, as documented in the lane comment before #8112.
2026-09-26 19:12:18 +08:00
Hauser d8bf268885 fix(scanner): bound SNSD deep scans and checkpoint cloning (#8126)
* fix(scanner): bound SNSD deep scans and checkpoint cloning

* ci: run mount-dependent jobs on hosted VMs

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
2026-09-26 06:06:50 +00:00
hector fdd753a1dd ci: run distributed e2e on ubuntu-latest to restore tmpfs mounts (#8128) 2026-09-26 12:00:26 +08:00
Chris de80e8a673 fix(ci): make shared checks portable across runners (#8124)
* fix(ci): isolate monitor argument checks from runner tools

* fix(ci): install the Typos action download dependency

* fix(ci): make release policy matching portable across awk variants
2026-09-26 10:55:06 +08:00
Chris 96fc26cd6a fix(release): defer installation updates until artifacts are live (#8122) 2026-09-26 09:17:15 +08:00
hector edcc81a8fd ci: replace ubuntu-latest runner with sm-standard-2 across workflows (#8112)
* ci: replace ubuntu-latest runner with sm-standard-2 across workflows

* ci: keep scheduled-validation monitors on hosted runners

The freshness and watchdog jobs report stalled scheduled validations.
Running them on the same sm-standard-2 pool means a pool outage stalls
the monitors too, so nothing reports it.

---------

Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-09-25 10:02:45 +08:00
Chris d3b75e695e test(connect): add scheduler receipt acceptance workflow (#8078)
* test(connect): add scheduler receipt acceptance workflow

* chore(deps): upgrade crates and fix faster-hex advisory (#8098)

* test(connect): isolate scheduler acceptance workflow

---------

Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-24 03:34:27 +08:00
Chris be01e513be feat(connect): sample memory within the running service (#8091)
* feat(connect): sample memory within the running service

* test(connect): add official memory service acceptance

* test(connect): consume the final memory job without cloning
2026-09-24 03:25:28 +08:00
Hauser 311e4306e6 chore(ci): complete action pin and HAProxy upgrades (#8090) 2026-09-23 19:20:39 +08:00
Hauser 7f0f4941cf chore(ci): refresh checkout pin in validated workflows (#8083)
* chore(ci): update checkout pin in validated workflows

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(ci): pin migrated checkout alerts to requested commit

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* chore(ci): pin additional checkout workflows to verified commit (#8084)

* chore(ci): pin more checkout workflows to requested commit

Update six additional workflows to the verified upstream checkout commit without changing their permissions, inputs, or triggers.

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* chore(ci): pin functional and OIDC checkout uses

Extend the verified checkout commit pin to functional-chain and OIDC workflows without changing their behavior.

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>

---------

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-23 09:27:54 +00:00
hector ee6de7d786 ci(package): build gnu and musl DEB/RPM variants with distinct file names (#8079)
* ci(package): build gnu and musl DEB/RPM variants with distinct file names

The Build and Release workflow produces four Linux binaries
(x86_64-gnu, aarch64-gnu, x86_64-musl, aarch64-musl), but packaging
only consumed the two gnu artifacts. Add matrix entries for the two
musl artifacts so every release ships all four DEB/RPM variants.

The libc variant is now part of the package file names, which would
otherwise collide between gnu and musl builds of the same version:

- deb: rustfs_<version>_<libc>_<arch>.deb
- rpm: rustfs-<libc>-<version>-<release>.<arch>.rpm

The dpkg Package and rpm Name stay plain "rustfs", so gnu and musl
remain mutually exclusive upgrades of one package rather than
co-installable packages fighting over /usr/bin/rustfs.

Dependency declarations now follow the linkage: gnu binaries
dynamically link glibc and keep Depends: libc6 (>= 2.31) /
glibc >= 2.31; musl binaries are statically linked and declare no
libc dependency. The libc variant is also visible in the package
description.

scripts/release/package_versions.sh gains a LIBC argument and its
contract tests cover both variants plus the invalid-libc cases.

* ci(package): align deb/rpm file names with the zip artifact naming

Rename the package file names so every release asset of one build
shares the same stem as its binary artifact, differing only by
extension:

- before: rustfs_<deb_version>_<libc>_<deb_arch>.deb
          rustfs-<libc>-<rpm_version>-<rpm_release>.<rpm_arch>.rpm
- after:  rustfs-linux-<arch>-<libc>-v<version>.deb / .rpm

e.g. rustfs-linux-x86_64-gnu-v1.0.0.zip,
     rustfs-linux-x86_64-gnu-v1.0.0.deb,
     rustfs-linux-x86_64-gnu-v1.0.0.rpm.

Non-development builds embed the raw release tag (with 'v'), like the
zips; development builds embed dev-<full sha>. The dpkg/rpm versions
(including the '~' prerelease ordering) are unchanged - they live in
the package metadata, and a side effect is that release asset names no
longer contain '~' (which GitHub normalizes to '.').

package_versions.sh now takes the target arch (x86_64|aarch64) instead
of the deb/rpm arch pair; the deb Architecture (amd64/arm64) in the
control metadata still comes from the workflow matrix. The two test
workflows that assemble deb download URLs from a release tag
(rustfs-table-test, rustfs-upgrade-test) are updated to the new name,
which also removes their '~'-to-'.' asset name workaround.
2026-09-23 03:17:47 +00:00
Chris 3f549e26f3 fix(ci): publish version-tagged preview Docker images (#8068) 2026-09-22 13:06:23 +00:00
hector 9c30cc8851 feat(ci): on-demand fault-tolerance matrix sweep workflow (#8058)
Adds rustfs-fault-tolerance-matrix.yml, a manually dispatched workflow that
runs the --matrix topology x EC outage sweep from rustfs/auto-testing (38
combinations, 320 cases). The scenario suite hardcodes --all and a 60-minute
budget, so the exhaustive sweep had no CI entry point.
2026-09-22 13:22:18 +08:00
hector bdea44f332 fix(ci): align FT evidence contract with scenario E (44 cases) (#8052)
The fault-tolerance suite now registers and runs scenario E
(rustfs/auto-testing#100): --all covers A, B, C, C2, D, E with 38 + 6
= 44 expected cases. The chain-evidence assertion still hardcoded
[A, B, C, C2, D] and 38 everywhere, which would fail every green FT
run's evidence validation once E is registered.

Bump the scenario list to include E and all count assertions from 38
to 44; step name updated to match. No other lanes reference 38.
2026-09-21 20:13:56 +08:00
hector ae6bbaef27 fix(ci): give the chain staleness probe a token that can read auto-testing (#8041)
resolve_functional_candidate.py probes rustfs/auto-testing (private) to
age the pinned functional-script revision and fall back to main HEAD
after 24h. The prepare step passed github.token, which cannot see the
private repo, so every nightly chain logged

  staleness probe failed (...exit status 1.); keeping pinned revision

and replayed the 09-14 harness. On the 09-20 nightly that harness died
on the dpkg conffile prompt in all 12 lanes (see rustfs/auto-testing#97)
because the --force-confold and other fixes never reached the chain.

Use PF_TESTING_GH_TOKEN - already required by the other steps in this
workflow - so the probe can actually run and the >24h fallback works.
2026-09-21 11:08:41 +08:00
マルコメ 42dd76d3b2 ci(docker): sync the Docker Hub overview from README.md on release (#8034)
Docker Hub's overview is a separate `full_description` field that
`docker push` never touches, so it had drifted into an 8 KB snapshot of
an old README.md that still linked to
https://docs.rustfs.com/introduction.html (now 404).

Add a `sync-dockerhub-description` job to docker.yml that runs after the
images are pushed and republishes README.md from the same commit via
peter-evans/dockerhub-description (pinned to v5.0.0). It reuses the
existing DOCKERHUB_USERNAME / DOCKERHUB_TOKEN credentials, so no new
secrets are needed. Relative links (docs/, CONTRIBUTING.md) are
rewritten to github.com URLs so they resolve on Docker Hub.

Docker Hub caps the field at 25,000 bytes and the action truncates to
fit with only a warning; README.md is at 22,879 bytes today. Read the
published overview back after the sync and fail the job if it hit the
cap, so a truncated overview cannot be published silently.

Fixes #7995

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-21 07:04:47 +08:00
hector ed2c3eaf77 fix(package): preserve service state across upgrades (#8024)
* fix(package): preserve service state across upgrades

* fix(package): match legacy DEB versions in tilde form

Published prerelease packages carry ~ in the dpkg control version
(package_versions.sh maps the SemVer prerelease - to ~), so the
legacy fallback list written with dots never matched 1.0.0~rc.x and
upgrades away from those DEBs still left the service stopped (#8011).

Fix the legacy glob to the tilde form and update the contract test,
which had enshrined the dot form. Verified on Ubuntu 24.04 systemd
containers: DEB upgrade 1.0.0~rc.5 -> 1.0.1 now keeps the service
running; 1.0.0 -> 1.0.1 and 1.0.1 -> 1.0.2 marker path still pass.

* docs(package): expand /etc/default/rustfs example template

Document the commonly used RUSTFS_* settings as commented examples in
the packaged conffile and point to docs.rustfs.com.
2026-09-20 11:27:41 +08:00