EC 2+2 still meets read quorum with two of four nodes up. A survivor
that had not cached a bucket returned 503 on GET, HEAD, and List, and
/health/ready left the Service once write quorum was lost. Reads and
Service membership now follow read quorum and shared locks. Writes and
/minio/health/cluster still require write quorum and exclusive locks.
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
Route the two linux-aarch64-* release legs to the dedicated arm64 runner
and build them natively (cross: false) instead of cross-compiling via
cargo-zigbuild on x86_64 hosts.
With native legs now executable, extend the packaged-artifact checks that
were previously gated to x86_64-unknown-linux-gnu only:
- Every Linux leg runs rustfs --version and rustfs-cli --help from its own
package, proving the artifact matches the runner architecture.
- The packaged-console runtime smoke (server boot + console HTTP 200) also
covers aarch64-unknown-linux-gnu, so both architectures get full runtime
verification. Musl legs run the binary checks but skip server boot for
now.
* ci: run the chain orchestration jobs where the gh CLI exists
#8112 moved these jobs to the sm-standard-2 pool, whose images ship
without the gh CLI (a known property of the build fleet, see #7572).
Last night's chain died in prepare before any suite ran:
resolve_functional_candidate.py:34 does subprocess.run(["gh", "api"])
and got FileNotFoundError; complete-chain and the hourly
functional-chain-health canary fail the same way. On Sep 22/23 the same
jobs ran green on GitHub-hosted runners.
Move prepare, complete-chain and the health canary to ubuntu-latest.
The suite lanes stay on smoke-testing.
* ci: run the chain orchestration jobs where the gh CLI exists
#8112 moved these jobs to the sm-standard-2 pool, whose images ship
without the gh CLI (a known property of the build fleet, see #7572).
Last night's chain died in prepare before any suite ran:
resolve_functional_candidate.py:34 does subprocess.run(["gh", "api"])
and got FileNotFoundError; complete-chain and the hourly
functional-chain-health canary fail the same way. On Sep 22/23 the same
jobs ran green on GitHub-hosted runners.
Move prepare, complete-chain and the health canary to ubuntu-latest.
The suite lanes stay on smoke-testing.
* ci: point the chain orchestration jobs at the smoke-testing runner
Per maintainer decision, consolidate them onto the same runner as the
test lanes. Verified on rustfs-smoke-testing as the runner user:
gh 2.45.0, jq 1.7.
* ci: point the chain orchestration jobs at the smoke-testing runner
Per maintainer decision, consolidate them onto the same runner as the
test lanes. Verified on rustfs-smoke-testing as the runner user:
gh 2.45.0, jq 1.7.
The health diagnostics module imported directly from
crate::storage::storage_api, bypassing the architecture migration
guardrail. Add a connect facade module to rustfs/src/storage_api.rs
and redirect the imports. Also fix a clippy::redundant-guards lint
in the coarse-flags match arm.
The sm-standard-2 fleet has no gh CLI (see #7572), so moving the release publication jobs onto it in #8112 broke every tag release: the 1.0.1-preview.12 build failed in Create GitHub Release with 'gh: command not found'. Move create-release, upload-release-assets, publish-release, cleanup-preview-releases, and package.yml resolve/package back to ubuntu-latest.
* fix(storage): weight automatic multipart admission by part size
* fix(ci): restore filesystem runner capabilities and typos dependency
* fix(ci): make release guard portable and spell out part variables
* ci: restore sm-standard-2 runners for io_uring and distributed e2e jobs
---------
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
* fix(ecstore): admit cold bucket metadata at read quorum
Encode a durable creation-commit state in bucket metadata so committed Object Lock buckets can use read quorum for cold loads while pre-physical creation intents remain fail-closed. Drain commit-write fan-out, keep write quorum for uncommitted intents, and add regression coverage for exact, below, migrated, and partial-create quorum boundaries.
* fix(ecstore): fence bucket creation-commit persistence
Address review findings on the creation-commit proof.
Persist the proof only under the bucket metadata transaction fence: at bucket creation, and through a fenced migration that re-reads the authoritative metadata and revalidates physical presence at write quorum before writing. This stops a stale snapshot from reverting an acknowledged configuration update or outliving a delete/recreate.
Establish commitment when Object Lock is enabled on an existing bucket, inside the same configuration mutation, so a healthy cluster no longer rejects object operations with ErasureWriteQuorum.
* fix(ecstore): keep commit fence error message stable
The error(format!) ratchet requires a stable Display for quorum bucketing; use a fixed message instead of embedding the bucket name.
* test(ecstore): reuse canonical Object Lock fixture in regression tests
The s3s footprint ratchet is shrink-only. Use the existing ENABLED_OBJECT_LOCK_CONFIG static instead of naming s3s DTO types in store tests.
---------
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
Companion to rustfs/auto-testing#109: the 4x1 topology is removed from
MATRIX_TOPOS because the product rejects a 1-drive-per-node distributed
pool at startup (FATAL BelowMinimum). The unfiltered sweep is now 36
combinations / 300 cases; update the contract totals and comments so an
unfiltered sweep is still held to an exact expected case count.
#8112 replaced ubuntu-latest with sm-standard-2 across workflows. Both nightly
KMS Vault lanes need a working Docker daemon; the sm-standard-2 pool does not
provide one, so both lanes die in seconds on 'Docker is not available' while
the package compiles and publishes fine - and the functional chain never fires
because its gate requires a successful build run. Restore the lanes to
ubuntu-latest, as documented in the lane comment before #8112.
* fix(scanner): bound SNSD deep scans and checkpoint cloning
* ci: run mount-dependent jobs on hosted VMs
---------
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
docs(github): clarify supported storage in bug reports
State that local disks, SAN/Fibre Channel volumes, and JBOD are supported,
recommend XFS, and keep the existing unsupported remote/shared filesystem rule.
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* ci: replace ubuntu-latest runner with sm-standard-2 across workflows
* ci: keep scheduled-validation monitors on hosted runners
The freshness and watchdog jobs report stalled scheduled validations.
Running them on the same sm-standard-2 pool means a pool outage stalls
the monitors too, so nothing reports it.
---------
Co-authored-by: overtrue <anzhengchao@gmail.com>
A flat (recursive) listing resumed inside "s/" re-emitted everything in
its sibling "s-x/". scan_dir dropped the entries before forward_to by
comparing directory names without their trailing slash, where "s" sorts
before "s-x", while the keys they stand for sort the other way round:
'-' (0x2d) is below '/' (0x2f), so all of "s-x/..." precedes "s/...".
The drain stopped at "s" and kept "s-x".
The next page then started with keys at or before the marker; the
listing layer filtered all of them out, found no more candidates and
answered IsTruncated=false. A bucket of 104,137 objects with backup
directories named "<id>" and "<id>-rollbacks" listed as 5,000; two such
directories of 1,200 keys each listed as 2,000.
Compare every remaining entry as the key prefix it stands for, slash
included, and keep it only when it sorts at or after forward_to or
contains it. The remainder of forward_to is taken before `current` is
trimmed, so this also holds below the bucket root; retain also drops
entries when all of them precede forward_to, which the drain never did.
* feat(connect): sample memory within the running service
* test(connect): add official memory service acceptance
* test(connect): consume the final memory job without cloning
* ci(package): build gnu and musl DEB/RPM variants with distinct file names
The Build and Release workflow produces four Linux binaries
(x86_64-gnu, aarch64-gnu, x86_64-musl, aarch64-musl), but packaging
only consumed the two gnu artifacts. Add matrix entries for the two
musl artifacts so every release ships all four DEB/RPM variants.
The libc variant is now part of the package file names, which would
otherwise collide between gnu and musl builds of the same version:
- deb: rustfs_<version>_<libc>_<arch>.deb
- rpm: rustfs-<libc>-<version>-<release>.<arch>.rpm
The dpkg Package and rpm Name stay plain "rustfs", so gnu and musl
remain mutually exclusive upgrades of one package rather than
co-installable packages fighting over /usr/bin/rustfs.
Dependency declarations now follow the linkage: gnu binaries
dynamically link glibc and keep Depends: libc6 (>= 2.31) /
glibc >= 2.31; musl binaries are statically linked and declare no
libc dependency. The libc variant is also visible in the package
description.
scripts/release/package_versions.sh gains a LIBC argument and its
contract tests cover both variants plus the invalid-libc cases.
* ci(package): align deb/rpm file names with the zip artifact naming
Rename the package file names so every release asset of one build
shares the same stem as its binary artifact, differing only by
extension:
- before: rustfs_<deb_version>_<libc>_<deb_arch>.deb
rustfs-<libc>-<rpm_version>-<rpm_release>.<rpm_arch>.rpm
- after: rustfs-linux-<arch>-<libc>-v<version>.deb / .rpm
e.g. rustfs-linux-x86_64-gnu-v1.0.0.zip,
rustfs-linux-x86_64-gnu-v1.0.0.deb,
rustfs-linux-x86_64-gnu-v1.0.0.rpm.
Non-development builds embed the raw release tag (with 'v'), like the
zips; development builds embed dev-<full sha>. The dpkg/rpm versions
(including the '~' prerelease ordering) are unchanged - they live in
the package metadata, and a side effect is that release asset names no
longer contain '~' (which GitHub normalizes to '.').
package_versions.sh now takes the target arch (x86_64|aarch64) instead
of the deb/rpm arch pair; the deb Architecture (amd64/arm64) in the
control metadata still comes from the workflow matrix. The two test
workflows that assemble deb download URLs from a release tag
(rustfs-table-test, rustfs-upgrade-test) are updated to the new name,
which also removes their '~'-to-'.' asset name workaround.
* fix(startup): retry transient bucket metadata quorum failures
* fix(ecstore): retain recovered shard damage for read repair
* fix(ecstore): reconnect missing read sources before GET
* test(heal): prove exact marker replay preserves history
* fix(ecstore): fence reconnect retries at the deadline
Reject expired reconnect waiters before dispatch and start cooldown on sweep completion, including timeout and cancellation. Cover exact-deadline admission independently of cooldown and preserve the original strict RPC-count regression.