Commit Graph

6820 Commits

Author SHA1 Message Date
Chris de80e8a673 fix(ci): make shared checks portable across runners (#8124)
* fix(ci): isolate monitor argument checks from runner tools

* fix(ci): install the Typos action download dependency

* fix(ci): make release policy matching portable across awk variants
2026-09-26 10:55:06 +08:00
RustFS 4cf45e9ed2 Update bug report storage requirements: allow SAN/JBOD, recommend XFS (#8125)
docs(github): clarify supported storage in bug reports

State that local disks, SAN/Fibre Channel volumes, and JBOD are supported,
recommend XFS, and keep the existing unsupported remote/shared filesystem rule.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-09-26 10:43:33 +08:00
Chris 9f8acf688c feat(github): add issue forms with storage support checks (#8123) 2026-09-26 10:07:49 +08:00
Chris 96fc26cd6a fix(release): defer installation updates until artifacts are live (#8122) 2026-09-26 09:17:15 +08:00
hector edcc81a8fd ci: replace ubuntu-latest runner with sm-standard-2 across workflows (#8112)
* ci: replace ubuntu-latest runner with sm-standard-2 across workflows

* ci: keep scheduled-validation monitors on hosted runners

The freshness and watchdog jobs report stalled scheduled validations.
Running them on the same sm-standard-2 pool means a pool outage stalls
the monitors too, so nothing reports it.

---------

Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-09-25 10:02:45 +08:00
Hauser 286bbb8e1f chore: upgrade workspace dependency releases (#8107)
* chore: upgrade workspace dependencies

* chore: refresh dependency lockfile

* fix(rebalance): avoid waiting on active entry gate during stop preparation
2026-09-24 21:24:14 +08:00
cxymds 0913686b12 fix(heal): detect stale delete-marker metadata (#8102) 1.0.1-preview.11 2026-09-24 11:16:02 +08:00
George Melikov df64242f64 fix(ecstore): resume scan_dir past a dash-suffixed sibling directory (#8071)
A flat (recursive) listing resumed inside "s/" re-emitted everything in
its sibling "s-x/". scan_dir dropped the entries before forward_to by
comparing directory names without their trailing slash, where "s" sorts
before "s-x", while the keys they stand for sort the other way round:
'-' (0x2d) is below '/' (0x2f), so all of "s-x/..." precedes "s/...".
The drain stopped at "s" and kept "s-x".

The next page then started with keys at or before the marker; the
listing layer filtered all of them out, found no more candidates and
answered IsTruncated=false. A bucket of 104,137 objects with backup
directories named "<id>" and "<id>-rollbacks" listed as 5,000; two such
directories of 1,200 keys each listed as 2,000.

Compare every remaining entry as the key prefix it stands for, slash
included, and keep it only when it sorts at or after forward_to or
contains it. The remainder of forward_to is taken before `current` is
trimmed, so this also holds below the bucket root; retain also drops
entries when all of them precede forward_to, which the drain never did.
2026-09-24 07:13:30 +08:00
Chris d3b75e695e test(connect): add scheduler receipt acceptance workflow (#8078)
* test(connect): add scheduler receipt acceptance workflow

* chore(deps): upgrade crates and fix faster-hex advisory (#8098)

* test(connect): isolate scheduler acceptance workflow

---------

Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-24 03:34:27 +08:00
Chris dde1aa77a1 fix(connect): preserve diagnostic schedules across restart (#8092) 2026-09-24 03:28:37 +08:00
Chris be01e513be feat(connect): sample memory within the running service (#8091)
* feat(connect): sample memory within the running service

* test(connect): add official memory service acceptance

* test(connect): consume the final memory job without cloning
2026-09-24 03:25:28 +08:00
cxymds 9ab034df62 fix(heal): preserve retryable writer failures (#8089)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-24 02:41:16 +08:00
Tyler Hillery 624463c5ba fix(s3): default CopyObject content-type to binary/octet-stream on REPLACE (#8073)
fix: default CopyObject content-type to binary/octet-stream on REPLACE
2026-09-23 18:04:05 +00:00
Chris 713920dcd6 docs: update license copyright to RustFS, Inc. (#8097) 2026-09-24 00:56:08 +08:00
GatewayJ f219a0aba8 perf(ecstore): parallelize multipart I/O setup and metadata reads (#8085)
* perf(ecstore): parallelize multipart I/O setup and metadata reads

* perf(ecstore): share multipart paths and increase read concurrency

* fix(ecstore): route multipart benchmark through storage API facade
2026-09-23 22:19:23 +08:00
Hauser 311e4306e6 chore(ci): complete action pin and HAProxy upgrades (#8090) 2026-09-23 19:20:39 +08:00
Hauser 7f0f4941cf chore(ci): refresh checkout pin in validated workflows (#8083)
* chore(ci): update checkout pin in validated workflows

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(ci): pin migrated checkout alerts to requested commit

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* chore(ci): pin additional checkout workflows to verified commit (#8084)

* chore(ci): pin more checkout workflows to requested commit

Update six additional workflows to the verified upstream checkout commit without changing their permissions, inputs, or triggers.

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* chore(ci): pin functional and OIDC checkout uses

Extend the verified checkout commit pin to functional-chain and OIDC workflows without changing their behavior.

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>

---------

Co-Authored-By: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-23 09:27:54 +00:00
cxymds 655bb470b8 fix(heal): retry faulty storage disk errors (#8086) 2026-09-23 07:48:39 +00:00
cxymds 56a4099f99 fix(heal): recreate missing bucket volumes (#8082)
* fix(heal): recreate missing bucket volumes

* fix(heal): use typed error for empty targets
2026-09-23 14:51:24 +08:00
hector ee6de7d786 ci(package): build gnu and musl DEB/RPM variants with distinct file names (#8079)
* ci(package): build gnu and musl DEB/RPM variants with distinct file names

The Build and Release workflow produces four Linux binaries
(x86_64-gnu, aarch64-gnu, x86_64-musl, aarch64-musl), but packaging
only consumed the two gnu artifacts. Add matrix entries for the two
musl artifacts so every release ships all four DEB/RPM variants.

The libc variant is now part of the package file names, which would
otherwise collide between gnu and musl builds of the same version:

- deb: rustfs_<version>_<libc>_<arch>.deb
- rpm: rustfs-<libc>-<version>-<release>.<arch>.rpm

The dpkg Package and rpm Name stay plain "rustfs", so gnu and musl
remain mutually exclusive upgrades of one package rather than
co-installable packages fighting over /usr/bin/rustfs.

Dependency declarations now follow the linkage: gnu binaries
dynamically link glibc and keep Depends: libc6 (>= 2.31) /
glibc >= 2.31; musl binaries are statically linked and declare no
libc dependency. The libc variant is also visible in the package
description.

scripts/release/package_versions.sh gains a LIBC argument and its
contract tests cover both variants plus the invalid-libc cases.

* ci(package): align deb/rpm file names with the zip artifact naming

Rename the package file names so every release asset of one build
shares the same stem as its binary artifact, differing only by
extension:

- before: rustfs_<deb_version>_<libc>_<deb_arch>.deb
          rustfs-<libc>-<rpm_version>-<rpm_release>.<rpm_arch>.rpm
- after:  rustfs-linux-<arch>-<libc>-v<version>.deb / .rpm

e.g. rustfs-linux-x86_64-gnu-v1.0.0.zip,
     rustfs-linux-x86_64-gnu-v1.0.0.deb,
     rustfs-linux-x86_64-gnu-v1.0.0.rpm.

Non-development builds embed the raw release tag (with 'v'), like the
zips; development builds embed dev-<full sha>. The dpkg/rpm versions
(including the '~' prerelease ordering) are unchanged - they live in
the package metadata, and a side effect is that release asset names no
longer contain '~' (which GitHub normalizes to '.').

package_versions.sh now takes the target arch (x86_64|aarch64) instead
of the deb/rpm arch pair; the deb Architecture (amd64/arm64) in the
control metadata still comes from the workflow matrix. The two test
workflows that assemble deb download URLs from a release tag
(rustfs-table-test, rustfs-upgrade-test) are updated to the new name,
which also removes their '~'-to-'.' asset name workaround.
2026-09-23 03:17:47 +00:00
cxymds 1880b42169 fix(heal): preserve decode repair signals and retry startup quorum (#8067)
* fix(startup): retry transient bucket metadata quorum failures

* fix(ecstore): retain recovered shard damage for read repair

* fix(ecstore): reconnect missing read sources before GET

* test(heal): prove exact marker replay preserves history

* fix(ecstore): fence reconnect retries at the deadline

Reject expired reconnect waiters before dispatch and start cooldown on sweep completion, including timeout and cancellation. Cover exact-deadline admission independently of cooldown and preserve the original strict RPC-count regression.
1.0.1-preview.10
2026-09-22 14:58:31 +00:00
Chris 3f549e26f3 fix(ci): publish version-tagged preview Docker images (#8068) 2026-09-22 13:06:23 +00:00
Dae-Cheol Noh 3c6c88b2e7 fix(amqp): emit S3 events directly in notification records (#8066)
* fix(amqp): emit S3 events directly in notification records

* test(amqp): avoid queued timestamp equality race
2026-09-22 12:16:40 +00:00
Dae-Cheol Noh 1f04a12abf fix(ftps): bound upload memory with multipart streaming (#8064)
* fix(ftps): bound upload memory with multipart streaming

* fix(ftps): keep multipart upload within driver module
2026-09-22 12:15:20 +00:00
cxymds 122abfaae2 feat(integrity): add inventory, audit, and protected migration (#8065)
* feat(integrity): add inventory, audit, and protected migration

* refactor(admin): use gateway facade for integrity handlers

* fix(integrity): sort fingerprint metadata explicitly
2026-09-22 19:23:41 +08:00
cxymds d0ce2f758b fix(ecstore): bound decommission target gate contention (#8061)
* fix(ecstore): bound decommission target gate contention

Target capacity gate contention during pool decommission escalated a
per-object transient into a durable bucket pause: each contended object
failed its bucket entry, which re-ran the whole bucket listing and amplified
attempts on the same objects.

- centralize the decommission capacity failure classification so gate
  contention, benign contention and fatal failures are decided once
- retry target-gate contention inline (bounded, jittered) before a mutation
  is admitted, covering put, part, complete, new-multipart and abort
- defer contended entries to the end of the round instead of failing the
  bucket entry, and require the deferred set to drain before a set completes
- treat a missing object or version, an overwrite and a superseded upload id
  as benign contention that is neither counted as a failure nor escalated
- use full-jitter exponential backoff, capped, for decommission retries
- expose the capacity pause reason, the waiting reason and a cumulative
  pause count in the admin pool status, plus gate-retry and per-object
  attempt metrics

Related: rustfs/backlog#2644

* fix(ecstore): correct deferred replay and metadata compatibility
2026-09-22 19:20:05 +08:00
hector 9c30cc8851 feat(ci): on-demand fault-tolerance matrix sweep workflow (#8058)
Adds rustfs-fault-tolerance-matrix.yml, a manually dispatched workflow that
runs the --matrix topology x EC outage sweep from rustfs/auto-testing (38
combinations, 320 cases). The scenario suite hardcodes --all and a 60-minute
budget, so the exhaustive sweep had no CI entry point.
2026-09-22 13:22:18 +08:00
Hauser 0ff06757c8 fix(obs): support custom OTLP CA bundles (#8060)
Add RUSTFS_OBS_TLS_CA_FILE for additive private CA trust in OTLP HTTP exporters.

Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-22 13:21:44 +08:00
cxymds e12100b4a1 fix(s3): emit x-amz-expiration as HTTP-date (#8057)
The expiry-date field of the x-amz-expiration response header was
rendered as RFC 3339 (e.g. 2026-10-02T00:00:00Z). AWS SDKs parse this
value with an RFC 822 parser; the Java SDK v1 rejects the ISO-8601
shape, drops the whole header, and logs a WARN from
ObjectExpirationHeaderHandler. S3 specifies HTTP-date (RFC 1123 with a
literal GMT zone), which is also what MinIO emits.

Render via the shared HTTP-date formatter (extracted alongside
format_expires_header) and pin the exact output in the unit test.
2026-09-22 11:18:59 +08:00
hector d7c42d4099 fix(ci): accept the testing-sha staleness fallback in chain evidence (#8056)
The 09-22 nightly chain failed all 12 lanes in seconds at the 'Bind
functional candidate' step:

  ValueError: private script pin differs from chain

resolve_functional_candidate.py's >24h staleness fallback (added by
#8026, made functional by #8041's token fix) legitimately sets
manifest.testing_sha to auto-testing main HEAD, but current_chain()
still required it to equal .config/functional-script-revision.txt -
a check written for the pre-fallback world where the two could never
diverge. Once the fallback finally fired, prepare produced testing_sha
21edcf4 while the pin file still holds 27e9584 and every lane aborted
before checking out the test scripts.

Drop the pin-file comparison and keep what the lane actually needs to
guarantee: testing_sha is a valid commit sha (current_chain), the lane
checked out exactly that sha (record: private_head == testing_sha, kept
as-is), and the health checker validates the same format instead of
re-reading the pin file. Tests updated: a fallback testing_sha that
differs from the pin is accepted; a non-sha testing_sha is rejected.

Verified: python3 -m unittest test_functional_chain
test_functional_chain_health -> 39 tests OK.
2026-09-22 10:25:36 +08:00
hector 4b0118520e test(ci): align FT workflow contract tests with scenario E (44 cases) (#8054)
#8052 renamed the fault-tolerance step to 'Run fault-tolerance
scenarios (A, B, C, C2, D, E)' and raised the evidence gate to 44
cases, but scripts/test_security_workflow.py still expected the old
step name and a 38-case fixture, breaking Quick Checks with three
KeyErrors and one assertion failure. Update DIRECT_TESTS and the
write_result fixture to the new contract.
2026-09-22 00:18:40 +08:00
cxymds 1afaf4491e fix(ecstore): bound decommission capacity gate waits (#8044)
* fix(ecstore): bound decommission capacity gate waits

* fix(ecstore): keep rebalance meta save lock errors retryable

The rebalance metadata retry policy only retries when the source chain still
carries the typed lock error, but resolve_rebalance_meta_save_result collapsed
every failure into a plain string. A transient rebalance.bin write-lock timeout
therefore bypassed retry_rebalance_metadata_access entirely and surfaced as a
hard failure of the periodic and terminal rebalance metadata saves.

Wrap the failure with the data-movement stage context instead, which keeps the
original lock error reachable through rebalance_error_source so the existing
transient lock policy applies. The rendered message is unchanged.

Add a regression test that feeds the wrapped meta save lock timeout through the
retry helper and asserts the retry engages.

* chore(ecstore): refresh error format ratchet baseline

Removing the string-wrapped meta save failure drops one `::other(format!`
call site in the rebalance worker, so the shrink-only baseline must be
regenerated in the same change.
1.0.1-preview.9
2026-09-21 20:56:26 +08:00
hector bdea44f332 fix(ci): align FT evidence contract with scenario E (44 cases) (#8052)
The fault-tolerance suite now registers and runs scenario E
(rustfs/auto-testing#100): --all covers A, B, C, C2, D, E with 38 + 6
= 44 expected cases. The chain-evidence assertion still hardcoded
[A, B, C, C2, D] and 38 everywhere, which would fail every green FT
run's evidence validation once E is registered.

Bump the scenario list to include E and all count assertions from 38
to 44; step name updated to match. No other lanes reference 38.
2026-09-21 20:13:56 +08:00
cxymds 8ba3309624 fix(rebalance): retry fleet proof before start activation (#8049)
Retry the transient missing fleet capability proof in both the admin start flow and node RPC activation. Return a retryable 503 with Retry-After after rollback when readiness does not converge.
2026-09-21 11:27:58 +00:00
cxymds cd2fa50e50 fix: rename admin metrics endpoint to realtime (#8046) 2026-09-21 19:23:58 +08:00
cxymds 077b961232 fix(obs): default span events to none and gate span list by level (#8047)
JSON log layers synthesized span lifecycle events for every span and
embedded the full ancestor span list on every record. With INFO spans on
object hot paths that produced multi-megabyte lines, which journald
truncated at 48 KiB: the payload was silently dropped and the volume
itself became the dominant load.

Add RUSTFS_OBS_LOG_SPAN_EVENTS (none|close|full) with a default of none,
and keep with_span_list enabled only for verbose (debug/trace) levels so
the ancestor chain stops being re-serialized per record. The request_id
promotion path reads the span scope, not the span list, so it is
unaffected.

Refs rustfs/backlog#2642
2026-09-21 19:15:39 +08:00
cxymds e1bc731331 fix(lifecycle): reject empty expiration actions (#8045)
* fix(lifecycle): reject empty expiration actions

* style: apply rustfmt to lifecycle validation
2026-09-21 15:43:07 +08:00
cxymds bb43ef861a fix(decommission): use owned data for source capacity (#8038) 2026-09-21 15:29:30 +08:00
Chris 4047ed6c8d fix(heal): restore bucket metadata before replacement completion (#8042) 2026-09-21 12:18:06 +08:00
hector ae6bbaef27 fix(ci): give the chain staleness probe a token that can read auto-testing (#8041)
resolve_functional_candidate.py probes rustfs/auto-testing (private) to
age the pinned functional-script revision and fall back to main HEAD
after 24h. The prepare step passed github.token, which cannot see the
private repo, so every nightly chain logged

  staleness probe failed (...exit status 1.); keeping pinned revision

and replayed the 09-14 harness. On the 09-20 nightly that harness died
on the dpkg conffile prompt in all 12 lanes (see rustfs/auto-testing#97)
because the --force-confold and other fixes never reached the chain.

Use PF_TESTING_GH_TOKEN - already required by the other steps in this
workflow - so the probe can actually run and the >24h fallback works.
2026-09-21 11:08:41 +08:00
cxymds a7c875b48f fix(heal): preserve remote corruption and reject incomplete deep scans (#8032)
* fix(heal): reject truncated encoded shards

* test(heal): cover truncated shards across integrity modes (#8035)

Co-authored-by: zhi22915 <qiuzgang@gmail.com>

* fix(heal): preserve remote corruption and reject incomplete deep scans

* test(admin): tolerate config lock timeout while polling

* fix(test): use admin storage error facade

---------

Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-09-21 02:33:03 +00:00
マルコメ 42dd76d3b2 ci(docker): sync the Docker Hub overview from README.md on release (#8034)
Docker Hub's overview is a separate `full_description` field that
`docker push` never touches, so it had drifted into an 8 KB snapshot of
an old README.md that still linked to
https://docs.rustfs.com/introduction.html (now 404).

Add a `sync-dockerhub-description` job to docker.yml that runs after the
images are pushed and republishes README.md from the same commit via
peter-evans/dockerhub-description (pinned to v5.0.0). It reuses the
existing DOCKERHUB_USERNAME / DOCKERHUB_TOKEN credentials, so no new
secrets are needed. Relative links (docs/, CONTRIBUTING.md) are
rewritten to github.com URLs so they resolve on Docker Hub.

Docker Hub caps the field at 25,000 bytes and the action truncates to
fit with only a warning; README.md is at 22,879 bytes today. Read the
published overview back after the sync and fail the job if it hit the
cap, so a truncated overview cannot be published silently.

Fixes #7995

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Hauser <housemecn@gmail.com>
1.0.1-preview.8
2026-09-21 07:04:47 +08:00
Chris 9aeecad4a9 docs(security): add encrypted ETag advisory lesson (#8040) 2026-09-20 23:52:47 +08:00
Hauser 3738b32df1 chore(deps): update flake.lock (#8033)
Flake lock file updates:

• Updated input 'nixpkgs':
    'github:NixOS/nixpkgs/aff8a0b' (2026-09-10)
  → 'github:NixOS/nixpkgs/a32edd7' (2026-09-17)
• Updated input 'rust-overlay':
    'github:oxalica/rust-overlay/228ecef' (2026-09-12)
  → 'github:oxalica/rust-overlay/26a71e6' (2026-09-19)
2026-09-20 14:54:45 +08:00
Hauser d997b07505 fix(get): skip futile resume for single-disk objects (#8031)
* chore: pin s3s to upstream git revision

Use the upstream s3s git repository at 50c94aeb1e5ea9ef8e7393ef6bbc87b72f6891c6 for both s3s and s3s-sigv4.

Allow the official s3s git source in cargo-deny so source checks continue to pass.

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

* fix(get): skip futile resume for single-disk objects

Co-Authored-By: heihutu <heihutu@gmail.com>

Co-Authored-By: zhi22915 <qiuzgang@gmail.com>

---------

Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-20 13:25:08 +08:00
Chris 0632a64f91 fix(ecstore): allow server-side copy at normal read quorum (#8029) 2026-09-20 12:36:53 +08:00
hector cbf6b02905 ci(chain): fall back to auto-testing main HEAD when the pinned harness goes stale (#8026)
The functional chain pins the test harness to the revision recorded in
.config/functional-script-revision.txt, which is only refreshed by
release syncs - it went stale on Sep 16 and the nightly chain has been
replaying a Sep-17-era harness ever since. Last night that meant:
the performance lane died on a dpkg conffile prompt (fix merged as
#88 but never reached the chain), the FT lane died on 'unknown option:
--results-file' (harness/workflow version mismatch), and the
STS-105 signing fix (#94) sat untested.

When the pinned revision's commit is older than 24 hours, fall back to
auto-testing main HEAD so the chain always runs the current harness.
Any staleness-probe failure keeps the pinned revision (fail-safe).
The revision file stays in place as the audit trail and the release
sync continues to manage it.
2026-09-20 11:27:56 +08:00
hector ed2c3eaf77 fix(package): preserve service state across upgrades (#8024)
* fix(package): preserve service state across upgrades

* fix(package): match legacy DEB versions in tilde form

Published prerelease packages carry ~ in the dpkg control version
(package_versions.sh maps the SemVer prerelease - to ~), so the
legacy fallback list written with dots never matched 1.0.0~rc.x and
upgrades away from those DEBs still left the service stopped (#8011).

Fix the legacy glob to the tilde form and update the contract test,
which had enshrined the dot form. Verified on Ubuntu 24.04 systemd
containers: DEB upgrade 1.0.0~rc.5 -> 1.0.1 now keeps the service
running; 1.0.0 -> 1.0.1 and 1.0.1 -> 1.0.2 marker path still pass.

* docs(package): expand /etc/default/rustfs example template

Document the commonly used RUSTFS_* settings as commented examples in
the packaged conffile and point to docs.rustfs.com.
2026-09-20 11:27:41 +08:00
cxymds 22243e791d test(heal): cover stale current delete marker rejoin (#8021) 1.0.1-preview.7 2026-09-19 19:33:24 +08:00
Hauser c35de00e8a fix(scanner): protect single-disk foreground latency (#8022)
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-09-19 19:32:02 +08:00