mirror of
https://github.com/rustfs/rustfs.git
synced 2026-10-04 12:31:36 +00:00
4e16705873
* fix(ecstore): stop pruning at nonempty directories (#7616) * fix(ecstore): stop pruning at nonempty directories * test(ecstore): release pruning fixtures before temp cleanup (cherry picked from commit8f150d1d8e) * fix(heal): preserve retryable batch failures during recovery (#7642) * fix(heal): preserve retryable batch failures during recovery * test(heal): pin prebuilt hooks binaries in ci (cherry picked from commit5cd58319ed) * fix(s3): reject oversize single PUT early and map body errors to 4xx (#7635) * fix(s3): reject oversize single PUT early and map body errors to 4xx A single PutObject above the 5 GiB single-request ceiling was only rejected after the client had streamed 5 GiB into s3s's read-time body budget, and the resulting BodySizeLimitExceeded surfaced from the erasure writer as 500 InternalError. A body whose connection hit EOF before Content-Length bytes arrived (hyper's IncompleteBody) was also a 500. SDKs retry 500s, so one oversize upload was resent from offset 0 five times. - PutObject and UploadPart reject a declared length above MAX_SINGLE_PUT_OBJECT_SIZE with 400 EntityTooLarge before reading the body; the constant moves to rustfs_config so the s3s limit and the admission check share one value. - ApiError maps BodySizeLimitExceeded to EntityTooLarge and a hyper body EOF to IncompleteBody across both io::Error conversions. Fixes #7596. * test(s3): cover UploadPart admission, aws-chunked length, real s3s limit - Poll-counting test body proves PutObject and UploadPart reject a declared size above the ceiling with zero body polls; exact-cap and zero-length parts pass admission. - A STREAMING-* aws-chunked PUT whose framed Content-Length exceeds the cap is admitted when the decoded length is within it and rejected when the decoded length is over it. - The display-based BodySizeLimitExceeded matcher is checked against the real error produced by the pinned s3s Body budget. (cherry picked from commit50b31bc75b) * fix(ecstore): make directory mtime fixture portable (#7623) * fix(ecstore): make directory mtime fixture portable * style(ecstore): format mtime fixture assertion --------- Co-authored-by: houseme <housemecn@gmail.com> Co-authored-by: Zhengchao An <anzhengchao@gmail.com> (cherry picked from commitb1cc286cac) * fix(storage): prevent readiness after native migration failures (#7652) * fix(storage): prevent readiness after native migration failures * fix(storage): skip unsupported IAM records before reading * fix(storage): use stable typed migration metadata errors * fix(storage): include migration record in startup errors * test(storage): cover native migration startup failures * test(storage): use array chunks in migration fixture --------- Co-authored-by: RJ Regenold <214054+rjregenold@users.noreply.github.com> Co-authored-by: cxymds <cxymds@gmail.com> (cherry picked from commit0cbc3ffe61) * fix(admin): expose OIDC account display fields (#7654) Expose verified OIDC username and email claims as display-only metadata on self-account responses while preserving the virtual parent as the authorization identity.\n\nKeep rustfs-madmin public response structs unchanged by adding the optional wire fields through private handler response wrappers. (cherry picked from commitf02bc947cd) * fix(replication): correct peer joins and remote-state reporting (#7650) * fix(replication): propagate verified peer deployment identities * fix(replication): report actual remote peer state * fix(replication): defer initial sync until all peers join * test(replication): shut down TLS fixtures cleanly --------- Co-authored-by: houseme <housemecn@gmail.com> (cherry picked from commit853ae63b6a) * fix(s3): bound stalled UploadPart request bodies (#7659) * fix(s3): bound stalled UploadPart request bodies * fix(ci): preserve the S3S footprint ratchet (cherry picked from commit666dfd9f9f) * fix(tables): reject reserved warehouse locations (#7671) Co-authored-by: cxymds <cxymds@gmail.com> (cherry picked from commit414176c47f) * fix(ci): bind nightly lanes to one resolved source (#7688) (cherry picked from commit01d8e4347f) * fix(replication): close the pre-stable convergence gaps from backlog#2367 (#7626) * fix(replication): total-order rule sort and honor V1 top-level Prefix Rule matching had two defects from the pre-GA replication audit (rustfs/backlog#2367 C-1 and C-2): - The actionable-rule sort compared same-destination rules by priority but answered Equal for any other pair, which is not a total order; the standard library sort panics on such comparators once a slice exceeds the insertion-sort threshold, so an object matching more than 20 enabled rules across two or more targets could panic the PUT or DELETE task. Rules now sort by priority descending with destination and id as tie-breakers, and filter_target_arns preserves that order instead of draining a HashSet. - A V1 rule written without a <Filter> carries its prefix at the top level; that field was never read, so <Prefix>logs/</Prefix> matched every object. ReplicationRuleExt::prefix now falls back to it, with a <Filter> keeping precedence. The existing prefix fixtures were built this way and had been asserting nothing. * fix(admin): advertise data-usage and listen capabilities to rc The rc client gated `rc du` and `rc watch` on a pinned contract that matched server versions by the string prefix `1.0.0-rc.`; a server that reports `1.0.0` no longer matches, and the dynamic `advertised` list did not carry either name, so `rc du` against a GA server fails with an unsupported-capability error (rustfs/backlog#2367 E-2). Advertise `admin.data-usage` from the admin route inventory like the IAM entries, and `listen_notification` for the bucket `?events=` extension route the admin router dispatches. The client merges advertised entries ahead of its pinned contract, so no version sniffing is needed. * fix(site-replication): stop notifying the local site on remove and rotate The pending-remove and pending-rotation notification loops skipped the local site by endpoint only, while finalization identifies it by deployment id or endpoint. The reconcile tick resolves the local peer from the node's own listen address (and a handler from the request Host), so `remove --all` dialed the site's registered endpoint, waited out the request timeout against the lifecycle lock it was holding, and answered `Partial: failed to notify 1 peer(s)` for a removal that had succeeded (rustfs/backlog#2367 A-4, backlog#2195 item 3). Both loops now iterate the peers still awaiting notification through one helper that applies the finalization identity. * fix(site-replication): promote and settle IAM retries without a tick of slack Two retry-queue behaviours kept an IAM change from converging for ten to twenty minutes after a peer came back (rustfs/backlog#2367 A-1 and A-3, backlog#2305): - The lightweight 30-second pass filtered its reachability probe to bucket ops, so a backed-off IAM or bucket-metadata snapshot waited for the 600-second tick to notice the peer. It now probes every backed-off class and still replays only bounded bucket ops; promotion is a state flip the heavyweight tick acts on. - Backoffs are multiples of the tick interval, so a failure stamped δ seconds after a tick was 600 − δ old at the next tick and slipped a whole extra interval. The heavyweight drain now evaluates backoff halfway to its next tick. - An IAM entry first created by a non-deletion failure (the add bootstrap's snapshot send, the drain's own replay, an import-iam schedule) was never stamped `deletions_recorded`, so a later recorded deletion could not settle it and it escalated to the marker only `replicate repair` clears. Entries created by this binary now start recorded; a row persisted by an older binary keeps the escalation semantics. * fix(site-replication): reload peer node caches after bucket wiring writes Every S3 bucket-config write ends by asking the other nodes of the cluster to reload the bucket's metadata; the site-replication writers never did. On a multi-node site the node that ran the pairing (or applied a peer's bucket-meta item) rewrote the bucket targets and the derived replication rules on disk, while every other node kept serving its cached copy for up to the 15-minute refresh. A `resync start` routed to such a node reported every freshly wired bucket as `Config not found` and a bucket whose operator target the pairing had replaced as `recorded remote target no longer exists` (rustfs/backlog#2367 A-5, backlog#2195 item 2; functional SITE-105). Add one best-effort reload helper in the site-replication hooks and call it after the bucket setup, versioning, peer bucket-meta apply, removed-peer cleanup, make-with-versioning, and endpoint-refresh writes; the ensure helpers now report whether they wrote so unchanged passes stay silent. The resync manifest and start now read the persisted wiring instead of the node-local cache, matching the target read the start path already did. The new four-node e2e pairs two clusters and starts a resync through a non-coordinator node right after pairing; it also covers an IAM user created on a non-coordinator node converging to the peer site. * test(e2e): cover delete-marker replication from a multi-node source The functional suite reported delete markers created on a 3-node source never reaching the target (rustfs/backlog#2195 item 4, REP-105). The report was a probe defect, but the shape had no coverage: the existing delete-marker e2e runs a single-node source. Pin it against a four-node source replicating to a four-node peer and to a single-node target, with the write and the delete issued through different nodes. * ci(e2e): refresh the distributed selection for the new replication cases Four distributed cases were added (two site-replication, two delete-marker replication). The linux digest is derived from the last CI listing of the lane (34 cases, matching the previous pin) plus the four new names; the darwin digest is the local listing, which selects the same 38 cases. (cherry picked from commitaeaba86d73) * fix(e2e): require a verified server binary for every e2e run (#7687) * fix(ci): share quick checks and lint workflows * fix(ci): install actionlint from its verified release * fix(ci): reject dependencies on required quick checks * feat(test): verify the E2E server build and source identity * test(e2e): register verified Darwin test membership * test(e2e): record verified Linux receipt test membership * test(e2e): record compiled Darwin receipt test membership * test(e2e): record compiled Linux receipt test membership * test(e2e): record Darwin e2e-full membership after merging main * fix(test): route scanner/heal evidence E2E runs through the verified server binary The evidence runners built rustfs with plain cargo and then ran e2e_test directly, which now fails without a run receipt. They build through scripts/e2e_binary.py and run the e2e_test invocations under e2e_binary.py run; the obsolete rustfs.features stamp is removed. * docs(e2e): run server-backed e2e commands through the verified binary wrapper * test(e2e): record Linux e2e-full membership from the branch CI listing (cherry picked from commit2909b1bfe1) * fix(kms): classify KMS/SSE error contracts and SSE-S3 headers (#7697) * fix(sse): classify bare SSE-KMS writes when no KMS is available A `aws:kms` request without a key id, on a bucket without a default key, returned `500 InternalError` whenever no KMS service was running: the "no KMS key available" branch exited with an untyped storage error before the availability classification that the keyed form already received. Route that branch through the same split: `503 ServiceUnavailable` while a configured KMS is stopped, `400 InvalidRequest` when KMS was never configured, and `400 InvalidRequest` naming the missing key id when a running KMS has no default key. `CreateMultipartUpload` shares the path. Adds a unit test for the bare form and an e2e module that stops KMS through the admin API, runs a master-key-only node, and runs a Local KMS without a default key; refreshes the e2e-full selection digests. (cherry picked from commit c3259dadc3d603a9185a5b0ad9f83dfb884e61c8) * fix(sse): keep KMS error classes on the encrypted read path GetObject, CopyObject and UploadPartCopy on an SSE-KMS object whose key no longer exists answered `500 InternalError` ("KMS key not found") while PutObject under the same key already answered `400 KMS.NotFoundException`. The read path carries its classification through ecstore's `EncryptionResolutionErrorKind`, which had no kind for a missing key, a denied KMS grant or a missing backend capability, so all three folded onto `DecryptionFailed` and the S3 layer reported an internal fault. Add `KeyNotFound`, `AccessDenied` and `NotImplemented` kinds, map them on both sides of the boundary, and give an envelope the configured backend cannot unwrap a diagnosable message while keeping its `500`. Unit tests cover the kind round trip and the reader wrapping; a new e2e test deletes a key immediately and checks GET/Copy return 400 with `KMS.NotFoundException` while HEAD stays 200. The e2e-full selection digests are refreshed from the current listing (the previous digests predated the delete-authorization tests) and the e2e `create_default_key` helper is updated to the accepted `EncryptDecrypt` spelling. (cherry picked from commit 2523a9814e97caea318d4ff1a51bef3a4d4445b2) * fix(kms): classify key-management errors on the admin routes `POST /kms/keys`, the legacy `create-key` alias and `generate-data-key` reported every backend refusal as `500`: a blank key name (which each backend failed on differently, the Local backend by writing a key file with an empty stem), a name already taken, an unknown key, a disabled key and a capability the backend lacks. `delete` and the lifecycle routes already classified the same errors. Refuse a blank or whitespace name in `KmsManager::create_key` before any backend sees it, and share one `KmsError` to status mapping across create, delete and generate-data-key (400 for validation and key state, 404 for an unknown key, 409 for a taken name, 501 for a missing capability, 500 only for damaged material). The XML-error routes carry the same status explicitly since s3s derives none for a custom code. The read-only Static backend now reports create, delete and cancel-deletion as `UnsupportedCapability`, matching its rotate and enable/disable answers, so the admin API returns 501 for all of them. (cherry picked from commit e33cac5493c4d9d6662e0d2980b58ba2b24a6d1b) * fix(sse): stop SSE-S3 responses from naming the wrapping KMS key `x-amz-server-side-encryption-aws-kms-key-id` is defined for `aws:kms` objects only, but PutObject, CopyObject, CreateMultipartUpload and GetObject returned it for `AES256` objects too, carrying the KMS key that wraps the SSE-S3 data key (the service default, or the literal `default` on a node without KMS). The write paths copied `kms_key_id` from the encryption material unconditionally, and the single-decrypt GET classification did the same after resolving the key for authorization. Add `EncryptionMaterial::response_kms_key_id`, which yields the id only for SSE-KMS, use it at the four write-response sites, and gate the GET classification the same way. CompleteMultipartUpload and HeadObject already omitted the header. Unit tests pin both directions; a new e2e test covers Put/Get/Head/Copy and CreateMultipartUpload for AES256 with an aws:kms control. The e2e-full selection digests are refreshed from the current listing. (cherry picked from commit 29d793a63352b0b60fd53c565e80fdbede8964bb) * fix(s3): validate PutBucketEncryption rules before storing them A default-encryption rule naming an unknown `SSEAlgorithm` (for example `AES128`), a rule without `ApplyServerSideEncryptionByDefault`, an empty rule list, or a `KMSMasterKeyID` on an `AES256` rule was stored as written: the only algorithm check on the route decided whether to fill in the default KMS key. `GetBucketEncryption` then advertised that configuration while the write path encrypted header-less writes under its `AES256` fallback, so the bucket's declared and actual schemes disagreed. Two comments claimed the route already refused unknown algorithms. Validate the configuration before any of it is applied: `MalformedXML` for a malformed rule set or unknown algorithm, `InvalidArgument` for a key id on a non-KMS rule, and nothing stored on refusal. Correct the two comments to describe when the AES256 fallback is still reachable. Unit tests cover every refusal and the accepted shapes; an e2e test checks the refusals leave the previous configuration in place. The e2e-full selection digests are refreshed from the current listing. (cherry picked from commit 29e4486dce41197ed93f5253cdbabc57d27a4ddb) * test(e2e): refresh e2e-full selection for the combined KMS/SSE fixes * test: align two unit tests with the new KMS and bucket-encryption contracts `scheduled_deletion_carries_a_deadline_and_can_be_cancelled` still expects the state error (`InvalidOperation`) for cancelling a key that is not pending deletion; only the Static backend's mutations moved to `UnsupportedCapability`. The uninitialized-store PutBucketEncryption test now sends a well-formed AES256 rule so it reaches the store lookup instead of the new configuration validation. (cherry picked from commite2e6a2535a) * fix(site-replication): keep an operator's bucket-level target to a peer instead of taking it over (#7709) * fix(site-replication): keep an operator's bucket-level target to a peer instead of taking it over Site replication wired each bucket by looking for an existing replication target "to the same peer" and rewriting the first match in place as its own same-name target. An operator's bucket-level target that happened to point at that site (different target bucket, operator credentials) was the first match whenever it pre-dated the join, and the reconciler repeats the pass every 600s, so the takeover also depended on target order afterwards. The operator's rule then named an ARN no target backed and their bucket replication stopped silently, while the inherited bucket-level reset id made every site resync report the bucket as owned by another resync (rustfs/backlog#2479, rustfs/backlog#2489). Follow MinIO's `getRemoteARN` / `getRemoteARNForPeer` shape instead: - Wiring updates a target in place only under the same ARN, or when it is recognisably the site's own under an older ARN shape (same peer, same-name target bucket, site replication service account). Anything else gets the site target added next to it. - The site resync manifest takes the target the derived `site-repl-<deployment id>` rule names (same-name shape as fallback), so an operator target to the peer neither aborts the bucket as "multiple remote targets matched peer" nor gets resynced into. - Peer removal prunes only targets a pruned derived rule names or the same-name target bucket; operator targets stamped with the peer's deployment id survive together with their rules. Unit tests cover the three predicates. e2e `test_site_replication_keeps_operator_bucket_target_to_peer` runs a bucket-level replication plus `replication-reset` to the future peer, joins the sites, and requires the operator target untouched, both paths delivering, the site resync completing against the site target, and the operator target and rule surviving `replicate remove --all`; without the fix it fails at the join with the operator target gone. The repl-nightly selection digest is refreshed for the new case. * test(site-replication): drop a redundant clone flagged by clippy The reconcile unit test cloned the remote peer into the state map although the binding is not used afterwards; workspace clippy (-D warnings) rejects that as redundant_clone. (cherry picked from commitecdc55fa4b) * fix: enforce S3 permissions for recursive force deletion (#7661) * fix: enforce S3 authorization for recursive deletion * fix: satisfy the s3s footprint guard * fix: restore list versions policy compatibility (#7686) (cherry picked from commit3fd1ce414d) * fix(ci): repair functional defaults and chain regression checks (#7664) * fix(ci): default functional suites to nightly packages * test(ci): follow the fault-tolerance chain handoff (cherry picked from commit509a0fa90c) * fix(ci): align security workflow tests with chain (#7679) (cherry picked from commitd9e47d2813) * test(ecstore): keep tier cleanup tests stable after immediate receipt queueing Release PR #7766 made PUT/CopyObject overwrites queue the tier free-version cleanup receipt immediately, which broke two ecstore tests on release CI. In tier_overwrite_put_and_self_copy_recover_persisted_cleanup_owners the restarted store already runs expiry workers from the second iteration on, so they deleted the remote bytes before the test could assert that the commit leaves them in place; the test now fails the first remote DELETE via set_remove_failure(true) so the cleanup owner stays durable and the later restart still has to rediscover it from xl.meta (the failed remove does not bump remove_count, and the test re-enables removes before the recovery wait). In batch_transitioned_delete_post_commit_failures_roll_back_without_free_version_receipt the convergence loop now treats a transient InsufficientReadQuorum as "not yet converged", because cleanup rewrites xl.meta disk by disk and a racing read can briefly miss quorum (seen in the rio-v2 lane); any other error still panics. * Revert "fix(ci): align security workflow tests with chain (#7679)" This reverts commit4544359f6d. * Revert "fix(ci): repair functional defaults and chain regression checks (#7664)" This reverts commit5ba7ec0291. * fix(ecstore): pass shard integrity to backported ingest-mode test The stalled-reader test backported with #7659 used main's four-argument encode_with_ingest_mode, but release's signature takes an optional IntegrityBuilder, so pass None to keep the test focused on ingest-mode cleanup. --------- Co-authored-by: Henry Guo <marshawcoco@gmail.com> Co-authored-by: 唐小鸭 <tangtang1251@qq.com> Co-authored-by: houseme <housemecn@gmail.com> Co-authored-by: RJ Regenold <rregenold@teamraft.com> Co-authored-by: RJ Regenold <214054+rjregenold@users.noreply.github.com> Co-authored-by: cxymds <cxymds@gmail.com> Co-authored-by: GatewayJ <835269233@qq.com> Co-authored-by: Jason Kossis <jkossis@gmail.com>
1253 lines
54 KiB
YAML
1253 lines
54 KiB
YAML
# Copyright 2024 RustFS Team
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
name: Continuous Integration
|
|
|
|
on:
|
|
push:
|
|
branches: [ main, release ]
|
|
paths-ignore:
|
|
- "**.md"
|
|
- "docs/**"
|
|
- "deploy/**"
|
|
- "scripts/dev_*.sh"
|
|
- "scripts/probe.sh"
|
|
- "LICENSE*"
|
|
- ".gitignore"
|
|
- ".dockerignore"
|
|
- "README*"
|
|
- "**/*.png"
|
|
- "**/*.jpg"
|
|
- "**/*.svg"
|
|
- ".github/workflows/build.yml"
|
|
- ".github/workflows/docker.yml"
|
|
- ".github/workflows/audit.yml"
|
|
- "flake.lock"
|
|
pull_request:
|
|
types: [ opened, synchronize, reopened, closed ]
|
|
branches: [ main, release ]
|
|
merge_group:
|
|
types: [ checks_requested ]
|
|
schedule:
|
|
- cron: "11 0 * * 0" # Weekly on Sunday 00:11 UTC
|
|
workflow_dispatch:
|
|
|
|
permissions:
|
|
contents: read
|
|
|
|
# Concurrency groups are scoped per event so different triggers never cancel
|
|
# each other: PR pushes cancel the previous run of that PR, main pushes keep
|
|
# latest-wins semantics among themselves, and scheduled runs always complete
|
|
# (a shared group used to let every merge kill the weekly scheduled run).
|
|
concurrency:
|
|
group: ${{ github.workflow }}-${{ github.event_name }}-${{ github.event.pull_request.number || github.ref }}
|
|
cancel-in-progress: ${{ github.event_name != 'schedule' }}
|
|
|
|
env:
|
|
CARGO_TERM_COLOR: always
|
|
RUST_BACKTRACE: 1
|
|
|
|
jobs:
|
|
|
|
cancel-closed-pr-runs:
|
|
name: Cancel Closed PR Runs
|
|
if: github.event_name == 'pull_request' && github.event.action == 'closed'
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
steps:
|
|
- name: Explain cancellation run
|
|
run: echo "PR closed; this run only cancels older runs in the same concurrency group."
|
|
|
|
classify-changes:
|
|
name: Select CI scope
|
|
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
outputs:
|
|
mode: ${{ steps.scope.outputs.mode }}
|
|
steps:
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
fetch-depth: 2
|
|
persist-credentials: false
|
|
- name: Select scope using the base revision's policy
|
|
id: scope
|
|
env:
|
|
CI_BASE_SHA: ${{ github.event.pull_request.base.sha }}
|
|
run: |
|
|
if [[ "$GITHUB_EVENT_NAME" != "pull_request" ]]; then
|
|
printf '%s\n' 'mode=full' >> "$GITHUB_OUTPUT"
|
|
elif [[ "$CI_BASE_SHA" =~ ^[0-9a-f]{40}$ ]] && git show "$CI_BASE_SHA:scripts/ci_gate.py" > "$RUNNER_TEMP/ci-gate-base.py"; then
|
|
python3 -I "$RUNNER_TEMP/ci-gate-base.py" select
|
|
else
|
|
printf '%s\n' 'mode=full' >> "$GITHUB_OUTPUT"
|
|
echo "Base CI policy unavailable; running the full matrix."
|
|
fi
|
|
|
|
typos:
|
|
name: Typos
|
|
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
steps:
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
- name: Typos check with custom config file
|
|
uses: crate-ci/typos@37bb98842b0d8c4ffebdb75301a13db0267cef89 # master
|
|
|
|
# Fail early with compile-free checks for every pull request.
|
|
quick-checks:
|
|
name: Quick Checks
|
|
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Run shared quick checks
|
|
uses: ./.github/actions/quick-checks
|
|
|
|
test-and-lint:
|
|
name: Workspace Test and Lint
|
|
if: needs.classify-changes.outputs.mode == 'full' && (github.event_name != 'pull_request' || github.event.action != 'closed')
|
|
needs: [ quick-checks, classify-changes ]
|
|
runs-on: sm-standard-4
|
|
timeout-minutes: 90
|
|
env:
|
|
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
# Checkout otherwise writes the token into .git/config, where a PR's
|
|
# own build.rs or proc-macro could read it back out.
|
|
persist-credentials: false
|
|
|
|
- name: Setup Rust environment
|
|
uses: ./.github/actions/setup
|
|
with:
|
|
rust-version: stable
|
|
# Every lane in this workflow reads its cache and none writes it.
|
|
# cache-warm.yml is the sole writer for all four keys: this workflow
|
|
# cancels superseded runs on main, and a cancelled run never reaches
|
|
# rust-cache's post step, so writing from here saved nothing (12 of 15
|
|
# consecutive main-push runs were cancelled). See rustfs/backlog#1600.
|
|
cache-shared-key: ci-dev
|
|
cache-save-if: 'false'
|
|
install-build-packaging-tools: 'false'
|
|
|
|
- name: Protect Connect test home
|
|
run: chmod go-w "$(realpath "$HOME")"
|
|
|
|
- name: Prepare test evidence
|
|
run: |
|
|
mkdir -p artifacts/test-and-lint
|
|
{
|
|
echo "run_id=${GITHUB_RUN_ID}"
|
|
echo "job=${GITHUB_JOB}"
|
|
echo "runner=${RUNNER_NAME}"
|
|
echo "started_at=$(date --utc --iso-8601=seconds)"
|
|
} > artifacts/test-and-lint/run-metadata.txt
|
|
|
|
# Clippy runs before the test pass: lint failures are the most common
|
|
# CI-only breakage and should surface in minutes, not after 20+ minutes
|
|
# of tests.
|
|
# Sampled too: clippy is the natural control arm for any CARGO_BUILD_JOBS
|
|
# experiment, since --all-targets is check-only for workspace members and
|
|
# never links the ~100 test binaries the limit exists to throttle.
|
|
- name: Run clippy lints
|
|
run: |
|
|
./scripts/ci/resource_sampler.sh start clippy
|
|
trap './scripts/ci/resource_sampler.sh stop' EXIT
|
|
cargo clippy --all-targets -- -D warnings
|
|
|
|
- name: Run nextest tests
|
|
env:
|
|
# #5394 mitigation, now under a measured experiment (backlog#1601).
|
|
#
|
|
# 2 was chosen when three concurrent workspace test links were believed
|
|
# to saturate the runner's overlay I/O and wedge Cargo until the 75m
|
|
# timeout. cgroup v2 readings from the sampler show the pod actually
|
|
# has 14 CPUs and 28GB (peak use 2.1GB), so 2 throttles compilation to
|
|
# a seventh of what is available and memory was never the constraint —
|
|
# the label name "sm-standard-4" had led everyone, including the
|
|
# original mitigation, to assume 4 cores.
|
|
#
|
|
# Raised to 3 on main pushes and manual dispatches; PRs keep 2 so the
|
|
# merge path is untouched while the experiment runs.
|
|
#
|
|
# Dispatch is included because push alone cannot supply the samples:
|
|
# this workflow cancels superseded runs on main, and only 4 of the last
|
|
# 20 push-triggered Test and Lint jobs reached a terminal state — at
|
|
# that rate ten samples would take roughly fifty merges. The
|
|
# concurrency group is scoped by event_name, so a dispatched run has
|
|
# its own group and is not cancelled by merge traffic, which makes the
|
|
# sample collectable on demand rather than by waiting.
|
|
#
|
|
# Baseline over 17 samples at 2:
|
|
# median nextest/clippy step ratio 1.95, spread 1.85-2.06. The gate-2
|
|
# criterion is that ratio dropping at least 10% (below ~1.76) with no
|
|
# 75m timeout and no run showing three consecutive samples of
|
|
# rustc/collect2/rust-lld in D state. If it does not, the conclusion is
|
|
# "this limit is not the bottleneck" — fix it back at 2 and record the
|
|
# experiment, which is a result, not a failure.
|
|
#
|
|
# Must stay step-level: rust-cache hashes CARGO/CC/CFLAGS/CXX/CMAKE/RUST
|
|
# prefixed variables from process.env into the cache key, so promoting
|
|
# this to job level would rotate every key on this lane.
|
|
CARGO_BUILD_JOBS: ${{ (github.event_name == 'push' || github.event_name == 'workflow_dispatch') && '3' || '2' }}
|
|
run: |
|
|
mkdir -p artifacts/test-and-lint
|
|
rm -f target/nextest/ci/junit.xml
|
|
./scripts/ci/resource_sampler.sh start nextest
|
|
trap './scripts/ci/resource_sampler.sh stop' EXIT
|
|
set +e
|
|
NEXTEST_HIDE_PROGRESS_BAR=1 timeout --verbose --signal=TERM --kill-after=30s 75m \
|
|
cargo nextest run --profile ci --all --exclude e2e_test \
|
|
--status-level all --final-status-level all \
|
|
2>&1 | tee artifacts/test-and-lint/nextest.log
|
|
status=${PIPESTATUS[0]}
|
|
if [[ "${status}" -eq 0 ]]; then
|
|
cargo nextest list --profile ci --all --exclude e2e_test --message-format json \
|
|
> artifacts/test-and-lint/core-test-listing.json \
|
|
&& python3 scripts/check_test_wiring.py --check-core artifacts/test-and-lint/core-test-listing.json \
|
|
&& test -s target/nextest/ci/junit.xml || status=$?
|
|
fi
|
|
{
|
|
echo "command=cargo nextest run --profile ci --all --exclude e2e_test"
|
|
echo "exit_status=${status}"
|
|
echo "finished_at=$(date --utc --iso-8601=seconds)"
|
|
echo
|
|
echo "Remaining test-related processes:"
|
|
pgrep -af 'cargo|nextest|target/.*/deps/' || true
|
|
echo
|
|
echo "Kernel OOM / kill events:"
|
|
dmesg -T 2>/dev/null | grep -iE 'oom|out of memory|killed process' | tail -20 || true
|
|
} > artifacts/test-and-lint/nextest-diagnostics.txt
|
|
exit "${status}"
|
|
|
|
- name: Run documentation tests
|
|
run: |
|
|
mkdir -p artifacts/test-and-lint
|
|
set +e
|
|
timeout --verbose --signal=TERM --kill-after=30s 15m \
|
|
cargo test --all --doc \
|
|
2>&1 | tee artifacts/test-and-lint/doctest.log
|
|
status=${PIPESTATUS[0]}
|
|
{
|
|
echo "command=cargo test --all --doc"
|
|
echo "exit_status=${status}"
|
|
echo "finished_at=$(date --utc --iso-8601=seconds)"
|
|
echo
|
|
echo "Remaining test-related processes:"
|
|
pgrep -af 'cargo|rustdoc|target/.*/deps/' || true
|
|
} > artifacts/test-and-lint/doctest-diagnostics.txt
|
|
exit "${status}"
|
|
|
|
- name: Check offline enrollment E2E root boundary
|
|
run: ./scripts/check_offline_enrollment_e2e.sh
|
|
|
|
- name: Upload test reports and diagnostics
|
|
if: always()
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: junit-test-and-lint-${{ github.run_number }}
|
|
path: |
|
|
target/nextest/ci/junit.xml
|
|
artifacts/test-and-lint
|
|
retention-days: 3
|
|
if-no-files-found: error
|
|
|
|
# rustfs/backlog#1289: fail if a seed rule's log anchor no longer exists
|
|
# verbatim in the source tree (log message drifted without updating the
|
|
# rule). Placed here where the workspace — including the la-dump-anchors
|
|
# bin — is already built by the clippy/test steps above.
|
|
- name: Check log-analyzer rule anchors
|
|
run: ./scripts/check_log_analyzer_rules.sh
|
|
|
|
# Explicit gate for migration-critical suites. These tests already ran in
|
|
# the full nextest pass above; a single filtered nextest invocation keeps
|
|
# the named gate without rebuilding or re-running them one package at a time.
|
|
#
|
|
# The gate selects tests by name substring (data_movement / rebalance /
|
|
# decommission / source_cleanup / delete_marker), so renames can silently
|
|
# thin it. The script owns the filter expression and first verifies the
|
|
# selected-test count against the committed floor in
|
|
# .config/migration-gate-floor.txt before running the gate; renames or
|
|
# removals must update that file consciously (see the script header).
|
|
# Kept on the default profile (no --profile ci): a second --profile ci run
|
|
# would clobber target/nextest/ci/junit.xml, and none of these tests are
|
|
# quarantined so they gain nothing from the ci profile's retry overrides.
|
|
- name: Run rebalance/decommission migration proofs
|
|
run: ./scripts/check_migration_gate_count.sh
|
|
|
|
# Dedicated serial lane for the ILM / lifecycle integration tests. These tests
|
|
# drive the object layer through process-global singletons (the GLOBAL_ENV
|
|
# ECStore, the global tier-config manager, background-expiry workers) and bind
|
|
# fixed ports (9002 for the scanner suite, 9003 for the app suite), so they are
|
|
# marked #[ignore] to keep them out of the default parallel `cargo nextest run
|
|
# --all` pass. Run them here explicitly with `--run-ignored ignored-only` and
|
|
# `-j1` so no two of them share global state or race for a port at once.
|
|
# serial_test's #[serial] does NOT serialize across nextest's process-per-test
|
|
# boundary (see .config/nextest.toml); `-j1` is what actually serializes them.
|
|
# See rustfs/backlog#1148 (ilm-1) and #1155.
|
|
test-ilm-integration-serial:
|
|
name: ILM Integration (serial)
|
|
if: needs.classify-changes.outputs.mode == 'full' && (github.event_name != 'pull_request' || github.event.action != 'closed')
|
|
needs: [ quick-checks, classify-changes ]
|
|
runs-on: sm-standard-4
|
|
timeout-minutes: 90
|
|
env:
|
|
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Setup Rust environment
|
|
uses: ./.github/actions/setup
|
|
with:
|
|
rust-version: stable
|
|
cache-shared-key: ci-dev
|
|
cache-save-if: 'false'
|
|
install-build-packaging-tools: 'false'
|
|
|
|
# test_transition_and_restore_flows was re-enabled by rustfs/backlog#1303:
|
|
# its "missing xl.meta on disk2" was a test-util bug (open_disk hardcoded
|
|
# disk_index 0), not an EC metadata-distribution issue.
|
|
# restore_object_usecase_reports_ongoing_conflict_and_completion was
|
|
# re-enabled by backlog#1304 (restore accepts serialize on a short CAS
|
|
# guard; the copy-back no longer holds the #4877 whole-copy-back lock,
|
|
# so the mid-restore ongoing read and fast 409 rejection it asserts are
|
|
# the implemented contract).
|
|
- name: Run ignored ILM integration tests serially
|
|
env:
|
|
# Match the measured Test and Lint link budget. The default exposed
|
|
# all 14 pod CPUs and a cold cache spent the full 80m compiling
|
|
# without starting one ILM test (main run 32982910990).
|
|
CARGO_BUILD_JOBS: ${{ (github.event_name == 'push' || github.event_name == 'workflow_dispatch') && '3' || '2' }}
|
|
run: |
|
|
mkdir -p artifacts/ilm-integration
|
|
set +e
|
|
NEXTEST_HIDE_PROGRESS_BAR=1 timeout --verbose --signal=TERM --kill-after=30s 80m \
|
|
cargo nextest run -j1 --run-ignored ignored-only \
|
|
-p rustfs-scanner -p rustfs \
|
|
-E 'binary(lifecycle_integration_test) or (package(rustfs) and test(lifecycle_transition_api_test))' \
|
|
--status-level all --final-status-level all \
|
|
2>&1 | tee artifacts/ilm-integration/nextest.log
|
|
status=${PIPESTATUS[0]}
|
|
{
|
|
echo "exit_status=${status}"
|
|
echo "finished_at=$(date --utc --iso-8601=seconds)"
|
|
echo
|
|
echo "Remaining test-related processes:"
|
|
pgrep -af 'cargo|nextest|target/.*/deps/' || true
|
|
} > artifacts/ilm-integration/diagnostics.txt
|
|
exit "${status}"
|
|
|
|
- name: Upload ILM test diagnostics
|
|
if: always()
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: ilm-integration-${{ github.run_number }}-${{ github.run_attempt }}
|
|
path: |
|
|
artifacts/ilm-integration
|
|
target/nextest/ci/junit.xml
|
|
|
|
test-and-lint-rio-v2:
|
|
name: Test and Lint (rio-v2)
|
|
if: needs.classify-changes.outputs.mode == 'full' && (github.event_name != 'pull_request' || github.event.action != 'closed')
|
|
needs: [ quick-checks, classify-changes ]
|
|
runs-on: sm-standard-4
|
|
timeout-minutes: 90
|
|
env:
|
|
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Setup Rust environment
|
|
uses: ./.github/actions/setup
|
|
with:
|
|
rust-version: stable
|
|
cache-shared-key: ci-feat-rio
|
|
cache-save-if: 'false'
|
|
install-build-packaging-tools: 'false'
|
|
|
|
- name: Protect Connect test home
|
|
run: chmod go-w "$(realpath "$HOME")"
|
|
|
|
- name: Run rio-v2 clippy lints
|
|
run: cargo clippy -p rustfs -p rustfs-ecstore --all-targets --features rio-v2 -- -D warnings
|
|
|
|
- name: Run rio-v2 feature tests
|
|
env:
|
|
# Match the main nextest lane's #5394 link-I/O guard. A cold feature
|
|
# cache otherwise fans out enough rust-lld processes to exhaust this
|
|
# job's 90-minute budget before any test starts.
|
|
CARGO_BUILD_JOBS: "2"
|
|
run: |
|
|
# --profile ci so the quarantine list (and its junit flaky markers)
|
|
# covers this leg too; the default profile is the local no-retry
|
|
# profile and silently ignored quarantined flakes here (rustfs#6703).
|
|
cargo nextest run --profile ci -p rustfs -p rustfs-ecstore --features rio-v2
|
|
cargo test -p rustfs --doc --features rio-v2
|
|
|
|
connect-short-credential-boundary:
|
|
name: Connect Short Credential Boundary
|
|
if: needs.classify-changes.outputs.mode == 'full' && (github.event_name != 'pull_request' || github.event.action != 'closed')
|
|
needs: [ quick-checks, classify-changes ]
|
|
runs-on: sm-standard-4
|
|
timeout-minutes: 60
|
|
env:
|
|
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Setup Rust environment
|
|
uses: ./.github/actions/setup
|
|
with:
|
|
rust-version: stable
|
|
cache-shared-key: ci-dev
|
|
cache-save-if: 'false'
|
|
install-build-packaging-tools: 'false'
|
|
install-test-tools: 'false'
|
|
|
|
- name: Run short credential behavior tests
|
|
env:
|
|
CARGO_BUILD_JOBS: "2"
|
|
run: |
|
|
cargo test -p rustfs --test connect_registration \
|
|
--features connect-e2e-short-credentials \
|
|
registration_enforces_build_profile_credential_lifetime -- --exact
|
|
cargo test -p rustfs --test connect_registration \
|
|
--features connect-e2e-short-credentials \
|
|
rotation_waits_for_threshold_and_stops_on_revocation -- --exact
|
|
|
|
- name: Reject short credentials in release builds
|
|
env:
|
|
CARGO_BUILD_JOBS: "2"
|
|
run: |
|
|
log="$(mktemp)"
|
|
set +e
|
|
CARGO_TERM_COLOR=never cargo check -p rustfs --release \
|
|
--features connect-e2e-short-credentials >"$log" 2>&1
|
|
status=$?
|
|
set -e
|
|
cat "$log"
|
|
expected='error: connect-e2e-short-credentials is restricted to debug builds'
|
|
summary="error: could not compile \`rustfs\` (lib) due to 1 previous error"
|
|
expected_count="$(grep -Fxc "$expected" "$log" || true)"
|
|
summary_count="$(grep -Fc "$summary" "$log" || true)"
|
|
error_count="$(grep -Ec '^error(:|\[)' "$log" || true)"
|
|
if [ "$status" -ne 101 ] || [ "$expected_count" -ne 1 ] \
|
|
|| [ "$summary_count" -ne 1 ] || [ "$error_count" -ne 2 ]; then
|
|
echo "release feature gate did not fail solely at the expected compile_error" >&2
|
|
rm -f "$log"
|
|
exit 1
|
|
fi
|
|
rm -f "$log"
|
|
|
|
test-and-lint-protocols:
|
|
name: "Test and Lint (${{ matrix.features.name }})"
|
|
if: needs.classify-changes.outputs.mode == 'full' && (github.event_name != 'pull_request' || github.event.action != 'closed')
|
|
needs: [ quick-checks, classify-changes ]
|
|
runs-on: sm-standard-4
|
|
timeout-minutes: 90
|
|
strategy:
|
|
# On a PR, one failing protocol leg is enough to know the PR is not ready,
|
|
# so stop the sibling leg instead of paying another ~40 minutes for it.
|
|
# Everywhere else (main pushes, the merge queue, the weekly schedule) keep
|
|
# the full signal: there we want to know whether swift AND sftp are broken,
|
|
# not just whichever failed first. This is the only part of the early-stop
|
|
# work that also covers fork PRs, since it needs no token.
|
|
fail-fast: ${{ github.event_name == 'pull_request' }}
|
|
matrix:
|
|
features:
|
|
- name: swift
|
|
flags: "--features swift"
|
|
- name: sftp
|
|
flags: "--features sftp"
|
|
env:
|
|
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Setup Rust environment
|
|
uses: ./.github/actions/setup
|
|
with:
|
|
rust-version: stable
|
|
cache-shared-key: ci-feat-proto
|
|
cache-save-if: 'false'
|
|
install-build-packaging-tools: 'false'
|
|
|
|
- name: Protect Connect test home
|
|
run: chmod go-w "$(realpath "$HOME")"
|
|
|
|
- name: Run clippy with ${{ matrix.features.name }}
|
|
run: |
|
|
cargo clippy -p rustfs -p rustfs-protocols --all-targets ${{ matrix.features.flags }} -- -D warnings
|
|
|
|
- name: Run tests with ${{ matrix.features.name }}
|
|
env:
|
|
# Keep feature-test linking under the same bounded concurrency as the
|
|
# main nextest lane; Clippy is metadata-only and needs no such limit.
|
|
CARGO_BUILD_JOBS: "2"
|
|
run: |
|
|
# --profile ci so the quarantine list (and its junit flaky markers)
|
|
# covers this leg too; the default profile is the local no-retry
|
|
# profile and silently ignored quarantined flakes here (rustfs#6703).
|
|
cargo nextest run --profile ci -p rustfs -p rustfs-protocols ${{ matrix.features.flags }}
|
|
|
|
build-rustfs-debug-binary:
|
|
name: Build RustFS Debug Binary
|
|
if: needs.classify-changes.outputs.mode == 'full' && (github.event_name != 'pull_request' || github.event.action != 'closed')
|
|
needs: [ quick-checks, classify-changes ]
|
|
runs-on: sm-standard-4
|
|
timeout-minutes: 30
|
|
env:
|
|
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Setup Rust environment
|
|
uses: ./.github/actions/setup
|
|
with:
|
|
rust-version: stable
|
|
cache-shared-key: ci-dev
|
|
cache-save-if: 'false'
|
|
install-build-packaging-tools: 'false'
|
|
|
|
- name: Build debug binary
|
|
run: |
|
|
python3 - <<'PYBUILD'
|
|
import hashlib
|
|
import json
|
|
import os
|
|
import pathlib
|
|
import subprocess
|
|
|
|
def git(*args):
|
|
return subprocess.check_output(["git", *args], text=True).strip()
|
|
|
|
def sha256(path):
|
|
digest = hashlib.sha256()
|
|
with pathlib.Path(path).open("rb") as source:
|
|
for chunk in iter(lambda: source.read(1024 * 1024), b""):
|
|
digest.update(chunk)
|
|
return digest.hexdigest()
|
|
|
|
argv = ["python3", "scripts/e2e_binary.py", "build", "--bins", "--features", "e2e-test-hooks"]
|
|
commit, tree = git("rev-parse", "HEAD"), git("rev-parse", "HEAD^{tree}")
|
|
clean_before = not git("status", "--porcelain", "--untracked-files=normal")
|
|
if not clean_before:
|
|
raise SystemExit("hooks binary requires a clean build checkout")
|
|
lock_sha256 = sha256("Cargo.lock")
|
|
lock_git_blob = git("hash-object", "Cargo.lock")
|
|
rustc = subprocess.check_output(["rustc", "-vV"], text=True)
|
|
host = next(line.removeprefix("host: ") for line in rustc.splitlines() if line.startswith("host: "))
|
|
if os.environ.get("CARGO_BUILD_TARGET") or pathlib.Path(os.environ.get("CARGO_TARGET_DIR", "target")).resolve() != pathlib.Path("target").resolve():
|
|
raise SystemExit("this artifact requires the native target/debug output")
|
|
subprocess.run(argv, check=True)
|
|
clean_after = not git("status", "--porcelain", "--untracked-files=normal")
|
|
if not clean_after or commit != git("rev-parse", "HEAD") or tree != git("rev-parse", "HEAD^{tree}") or lock_sha256 != sha256("Cargo.lock"):
|
|
raise SystemExit("hooks binary source changed while building")
|
|
manifest = {
|
|
"schema": 1, "commit": commit, "tree": tree,
|
|
"clean_before": clean_before, "clean_after": clean_after,
|
|
"lock_sha256": lock_sha256, "lock_git_blob": lock_git_blob,
|
|
"argv": argv, "profile": "debug", "target": host,
|
|
"features": ["e2e-test-hooks"],
|
|
"rustc_verbose": rustc,
|
|
"build_flags": {key: os.environ[key] for key in ("RUSTFLAGS", "CARGO_ENCODED_RUSTFLAGS", "CARGO_BUILD_TARGET", "CARGO_TARGET_DIR", "RUSTUP_TOOLCHAIN") if key in os.environ},
|
|
"binary_sha256": sha256("target/debug/rustfs"),
|
|
}
|
|
pathlib.Path("target/debug/rustfs.e2e-startup-cas-build.json").write_text(json.dumps(manifest, indent=2) + "\n")
|
|
PYBUILD
|
|
|
|
- name: Upload debug binary
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: rustfs-debug-binary
|
|
path: |
|
|
target/debug/rustfs
|
|
target/debug/rustfs.e2e.json
|
|
target/debug/rustfs.e2e-startup-cas-build.json
|
|
if-no-files-found: error
|
|
retention-days: 1
|
|
|
|
build-rustfs-debug-binary-rio-v2:
|
|
name: Build RustFS Debug Binary (rio-v2)
|
|
# Dormant rio-v2 variant (rustfs/backlog#1835): the feature ships in no
|
|
# default build, so this full-suite lane runs only on the weekly schedule
|
|
# and manual dispatch. Per-PR cfg-seam coverage stays with
|
|
# test-and-lint-rio-v2. Lifecycle and the promote-or-delete condition:
|
|
# docs/architecture/minio-file-format-compat.md ("rio-v2 variant lifecycle").
|
|
if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
|
|
needs: [ quick-checks ]
|
|
runs-on: sm-standard-4
|
|
timeout-minutes: 30
|
|
env:
|
|
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Setup Rust environment
|
|
uses: ./.github/actions/setup
|
|
with:
|
|
rust-version: stable
|
|
cache-shared-key: ci-feat-rio
|
|
cache-save-if: 'false'
|
|
install-build-packaging-tools: 'false'
|
|
|
|
- name: Build debug binary with rio-v2
|
|
run: python3 scripts/e2e_binary.py build --bins --features rio-v2,e2e-test-hooks
|
|
|
|
- name: Upload debug binary
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: rustfs-debug-binary-rio-v2
|
|
path: |
|
|
target/debug/rustfs
|
|
target/debug/rustfs.e2e.json
|
|
if-no-files-found: error
|
|
retention-days: 1
|
|
|
|
uring-integration:
|
|
name: io_uring Integration (real)
|
|
# The pull_request trigger includes `closed` purely so the concurrency
|
|
# group cancels in-flight runs of a closed PR; every other job opts out of
|
|
# that run with this guard (or is skipped through its `needs` chain). This
|
|
# job had neither, so each closed/merged PR really ran the whole io_uring
|
|
# suite (measured 4m17s / 7m19s / 7m31s on runs 30678272341 / 30678117601 /
|
|
# 30662728539) and kept the cancellation run in progress for minutes.
|
|
if: needs.classify-changes.outputs.mode == 'full' && (github.event_name != 'pull_request' || github.event.action != 'closed')
|
|
needs: [ quick-checks, classify-changes ]
|
|
# GitHub-hosted ubuntu-latest runs a recent kernel with io_uring and, unlike
|
|
# a container, applies no seccomp filter that would block io_uring_setup — so
|
|
# the probe succeeds and the tests exercise the real UringBackend/FdCache/
|
|
# latch paths instead of the StdBackend fallback (rustfs/backlog#1179). The
|
|
# self-hosted sm-standard runners cannot guarantee this.
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 30
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Setup Rust environment
|
|
uses: ./.github/actions/setup
|
|
with:
|
|
rust-version: stable
|
|
# Keeps its own key rather than joining ci-dev. rust-cache's key is
|
|
# built from runner.os/arch plus rustc and lockfile fingerprints — it
|
|
# does NOT include the runner label or image. ubuntu-latest and
|
|
# sm-standard-4 are therefore indistinguishable to it, so sharing a key
|
|
# would let two different system images overwrite each other's
|
|
# artifacts, and would make a 2-core hosted runner unpack ci-dev's ~3GB
|
|
# instead of this lane's ~1.3GB. cache-warm.yml warms this key on
|
|
# ubuntu-latest for the same reason.
|
|
cache-shared-key: ci-uring
|
|
cache-save-if: 'false'
|
|
install-build-packaging-tools: 'false'
|
|
|
|
# ext4 supports O_DIRECT; the runner's default TMPDIR may sit on tmpfs or
|
|
# overlayfs, where open(O_DIRECT) returns EINVAL/EOPNOTSUPP and the native
|
|
# read_at_direct path silently latches off to the aligned StdBackend
|
|
# fallback. A dedicated ext4 loopback guarantees the native O_DIRECT read
|
|
# path is actually exercised, not just the fallback (rustfs/backlog#1220).
|
|
- name: Prepare ext4 loopback for O_DIRECT coverage
|
|
run: |
|
|
set -euo pipefail
|
|
dd if=/dev/zero of=/tmp/rustfs-odirect.img bs=1M count=1024
|
|
mkfs.ext4 -q -F /tmp/rustfs-odirect.img
|
|
sudo mkdir -p /mnt/rustfs-odirect
|
|
sudo mount -o loop /tmp/rustfs-odirect.img /mnt/rustfs-odirect
|
|
sudo chmod 1777 /mnt/rustfs-odirect
|
|
mount | grep rustfs-odirect
|
|
|
|
# RUSTFS_URING_TESTS_MUST_RUN=1 makes the io_uring tests fail instead of
|
|
# silently skipping if io_uring is unavailable, so this leg can never pass
|
|
# vacuously. RUSTFS_IO_URING_READ_ENABLE=true selects the UringBackend.
|
|
# TMPDIR points every tempfile-based test fixture at the ext4 loopback so
|
|
# eligible reads keep O_DIRECT and the native path is covered end to end.
|
|
- name: Run ecstore io_uring integration tests on real io_uring
|
|
env:
|
|
RUSTFS_IO_URING_READ_ENABLE: "true"
|
|
RUSTFS_URING_TESTS_MUST_RUN: "1"
|
|
TMPDIR: /mnt/rustfs-odirect
|
|
# --lib narrows what gets compiled, not what gets run: every selected
|
|
# test lives in the lib target. The 7 integration binaries under
|
|
# crates/ecstore/tests/ each reported "running 0 tests" here, so they
|
|
# were compiled and linked for nothing.
|
|
#
|
|
# The `uring_` filter must stay exactly as it is. libtest matches on
|
|
# substring, so it also selects names containing `during_` — 6 of the 18
|
|
# selected tests are such incidental matches. Narrowing the filter to
|
|
# `io_uring` would silently drop them, which is a coverage change.
|
|
# scripts/check_uring_lane_lib_only.sh guards the --lib precondition.
|
|
run: cargo test -p rustfs-ecstore --lib uring_ -- --test-threads=1 --nocapture
|
|
|
|
e2e-tests:
|
|
name: End-to-End Tests
|
|
needs: [ build-rustfs-debug-binary ]
|
|
runs-on: sm-standard-2
|
|
timeout-minutes: 30
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
# Full setup with dependency caching: the smoke-suite step below
|
|
# compiles the e2e_test crate, which pulls in most of the workspace.
|
|
# Without a restored cache this was a cold multi-GB build on every run.
|
|
- name: Setup Rust environment
|
|
uses: ./.github/actions/setup
|
|
with:
|
|
rust-version: stable
|
|
cache-shared-key: ci-dev
|
|
cache-save-if: 'false'
|
|
install-build-packaging-tools: 'false'
|
|
|
|
- name: Set up Python
|
|
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6.3.0
|
|
with:
|
|
python-version: "3.12"
|
|
|
|
- name: Install awscurl
|
|
run: |
|
|
python3 -m pip install --user --upgrade pip "awscurl==0.44"
|
|
echo "AWSCURL_PATH=$HOME/.local/bin/awscurl" >> "$GITHUB_ENV"
|
|
|
|
- name: Verify awscurl
|
|
run: test -x "$AWSCURL_PATH"
|
|
|
|
# Download after the cache restore so the freshly built binary from the
|
|
# build job always wins over anything restored into target/debug.
|
|
- name: Download debug binary
|
|
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
|
|
with:
|
|
name: rustfs-debug-binary
|
|
path: target/debug
|
|
|
|
- name: Make binary executable
|
|
run: chmod +x ./target/debug/rustfs
|
|
|
|
# Build the e2e test graph once. The archive is reused by the smoke
|
|
# selection guard, security exact-count check, and run below, avoiding a
|
|
# second compile of the same e2e_test target on cold runners (backlog#1645).
|
|
- name: Archive e2e smoke test binaries
|
|
env:
|
|
NEXTEST_ARCHIVE: ${{ runner.temp }}/rustfs-e2e-smoke.tar.zst
|
|
NEXTEST_LISTING: ${{ runner.temp }}/rustfs-e2e-smoke-list.json
|
|
run: |
|
|
cargo nextest archive --profile e2e-smoke -p e2e_test --archive-file "${NEXTEST_ARCHIVE}"
|
|
cargo nextest list --profile e2e-smoke --archive-file "${NEXTEST_ARCHIVE}" --message-format json > "${NEXTEST_LISTING}"
|
|
python3 ./scripts/check_test_wiring.py --check-profile e2e-smoke "${NEXTEST_LISTING}"
|
|
./scripts/check_security_smoke_count.sh check "${NEXTEST_LISTING}"
|
|
|
|
# PR smoke subset of the in-repo e2e suite (backlog#1149 ci-4). The
|
|
# profile.e2e-smoke default-filter in .config/nextest.toml is the single
|
|
# wiring mechanism for e2e tests in CI — extend that filter instead of
|
|
# adding new e2e jobs here. Each test spawns its own rustfs server on a
|
|
# random port and reuses the downloaded debug binary above.
|
|
- name: Run e2e smoke suite
|
|
env:
|
|
NEXTEST_ARCHIVE: ${{ runner.temp }}/rustfs-e2e-smoke.tar.zst
|
|
RUSTFS_E2E_LOG_DIR: ${{ runner.temp }}/rustfs-e2e-smoke-logs
|
|
run: |
|
|
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-smoke --archive-file "${NEXTEST_ARCHIVE}" \
|
|
--status-level all --final-status-level all --failure-output final
|
|
|
|
- name: Upload e2e smoke diagnostics
|
|
if: failure()
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: e2e-smoke-diagnostics-${{ github.run_number }}
|
|
path: |
|
|
${{ runner.temp }}/rustfs-e2e-smoke-logs/
|
|
${{ runner.temp }}/rustfs-e2e-smoke-list.json
|
|
if-no-files-found: warn
|
|
|
|
- name: Upload e2e smoke JUnit report
|
|
if: always()
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: e2e-smoke-junit-${{ github.run_number }}
|
|
path: target/nextest/e2e-smoke/junit.xml
|
|
if-no-files-found: warn
|
|
|
|
- name: Install s3s-e2e test tool
|
|
uses: taiki-e/cache-cargo-install-action@7447f04c51f2ba27ca35e7f1e28fab848c5b3ba7 # v2
|
|
with:
|
|
tool: s3s-e2e
|
|
git: https://github.com/s3s-project/s3s.git
|
|
rev: 62cb4a71dd759a6ec56b64c4c42fcc183a2c6a52
|
|
|
|
- name: Run end-to-end tests
|
|
run: |
|
|
s3s-e2e --version
|
|
RUN_ROOT="${RUNNER_TEMP}/rustfs-e2e-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}-${GITHUB_JOB}"
|
|
mkdir -p "${RUN_ROOT}"
|
|
RUSTFS_TEST_PORT="$(python3 -c 'import socket; s=socket.socket(); s.bind(("127.0.0.1", 0)); print(s.getsockname()[1]); s.close()')"
|
|
RUSTFS_TEST_PORT="${RUSTFS_TEST_PORT}" \
|
|
RUSTFS_TEST_LOG="${RUN_ROOT}/rustfs.log" \
|
|
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- ./scripts/e2e-run.sh ./target/debug/rustfs "${RUN_ROOT}/data"
|
|
|
|
- name: Upload test logs
|
|
if: failure()
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: e2e-test-logs-${{ github.run_number }}
|
|
path: ${{ runner.temp }}/rustfs-e2e-*/rustfs.log
|
|
retention-days: 3
|
|
|
|
e2e-full:
|
|
name: End-to-End Tests (full merge gate)
|
|
# Merge gate only (backlog#1149 ci-5): the never-automated user-visible
|
|
# suites — KMS, object_lock, multipart_auth, quota, checksum, encryption,
|
|
# security-boundary, ... — via the e2e-full nextest profile. Too heavy for
|
|
# every PR, so it is gated to main/release pushes, the merge queue, and manual
|
|
# dispatch. protocols / the 7 cluster suites / replication / #[ignore] are
|
|
# owned by other lanes (see .config/nextest.toml profile.e2e-full).
|
|
if: >-
|
|
github.event_name == 'workflow_dispatch' ||
|
|
github.event_name == 'merge_group' ||
|
|
(github.event_name == 'push' &&
|
|
(github.ref == 'refs/heads/main' || github.ref == 'refs/heads/release'))
|
|
needs: [ build-rustfs-debug-binary ]
|
|
runs-on: sm-standard-2
|
|
timeout-minutes: 55
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Setup Rust environment
|
|
uses: ./.github/actions/setup
|
|
with:
|
|
rust-version: stable
|
|
cache-shared-key: ci-dev
|
|
cache-save-if: 'false'
|
|
install-build-packaging-tools: 'false'
|
|
|
|
- name: Install network fault-injection tools
|
|
run: |
|
|
sudo apt-get install -y iptables
|
|
sudo -n iptables --version
|
|
# The endpoint-blackhole heal scenario needs CAP_NET_ADMIN. Containerised
|
|
# runners can run iptables but not touch the rule set; the test then logs
|
|
# a skip instead of failing, so surface that here where it is visible.
|
|
if ! sudo -n iptables -w 5 -S OUTPUT >/dev/null 2>&1; then
|
|
echo "::warning::iptables cannot read the OUTPUT chain on this runner (no CAP_NET_ADMIN); the endpoint-blackhole heal scenario will be skipped"
|
|
fi
|
|
|
|
- name: Set up Python
|
|
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6.3.0
|
|
with:
|
|
python-version: "3.12"
|
|
|
|
- name: Install awscurl
|
|
run: |
|
|
python3 -m pip install --user --upgrade pip "awscurl==0.44"
|
|
echo "AWSCURL_PATH=$HOME/.local/bin/awscurl" >> "$GITHUB_ENV"
|
|
|
|
- name: Verify awscurl
|
|
run: test -x "$AWSCURL_PATH"
|
|
|
|
- name: Install mc
|
|
env:
|
|
MC_VERSION: RELEASE.2025-08-13T08-35-41Z
|
|
MC_SHA256: 01f866e9c5f9b87c2b09116fa5d7c06695b106242d829a8bb32990c00312e891
|
|
run: |
|
|
MC_BINARY="mc.linux-amd64.${MC_VERSION}"
|
|
curl -fsSLo "$RUNNER_TEMP/mc" "https://github.com/minio/mc/releases/download/${MC_VERSION}/${MC_BINARY}"
|
|
echo "${MC_SHA256} $RUNNER_TEMP/mc" | sha256sum --check --status
|
|
chmod +x "$RUNNER_TEMP/mc"
|
|
echo "$RUNNER_TEMP" >> "$GITHUB_PATH"
|
|
|
|
- name: Verify mc
|
|
run: mc --version
|
|
|
|
- name: Install Vault
|
|
run: |
|
|
VAULT_VERSION="1.17.6"
|
|
VAULT_ARCHIVE="vault_${VAULT_VERSION}_linux_amd64.zip"
|
|
curl -fsSLo "$RUNNER_TEMP/$VAULT_ARCHIVE" "https://releases.hashicorp.com/vault/${VAULT_VERSION}/${VAULT_ARCHIVE}"
|
|
echo "0cddc1fbbb88583b5ba5b845f9f8fae47c6fb39a6d48cd543c6ba6fd3ac1a669 $RUNNER_TEMP/$VAULT_ARCHIVE" | sha256sum --check --status
|
|
unzip -q "$RUNNER_TEMP/$VAULT_ARCHIVE" -d "$RUNNER_TEMP/vault-bin"
|
|
echo "RUSTFS_TEST_VAULT_BIN=$RUNNER_TEMP/vault-bin/vault" >> "$GITHUB_ENV"
|
|
|
|
- name: Verify Vault
|
|
run: |
|
|
"$RUSTFS_TEST_VAULT_BIN" version
|
|
|
|
# Download after the cache restore so the freshly built binary from the
|
|
# build job always wins over anything restored into target/debug.
|
|
- name: Download debug binary
|
|
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
|
|
with:
|
|
name: rustfs-debug-binary
|
|
path: target/debug
|
|
|
|
- name: Make binary executable
|
|
run: chmod +x ./target/debug/rustfs
|
|
|
|
- name: Preserve startup CAS binary input
|
|
env:
|
|
STARTUP_CAS_INPUT: ${{ runner.temp }}/rustfs-startup-cas-input
|
|
run: |
|
|
python3 - <<'PYINPUT'
|
|
import hashlib
|
|
import json
|
|
import os
|
|
import pathlib
|
|
import shutil
|
|
import subprocess
|
|
|
|
source = pathlib.Path("target/debug/rustfs")
|
|
manifest_path = source.with_name("rustfs.e2e-startup-cas-build.json")
|
|
manifest = json.loads(manifest_path.read_text())
|
|
target = pathlib.Path(os.environ["STARTUP_CAS_INPUT"])
|
|
target.mkdir(parents=True, exist_ok=True)
|
|
binary = target / "rustfs"
|
|
shutil.copy2(source, binary)
|
|
digest = hashlib.sha256()
|
|
with binary.open("rb") as stream:
|
|
for chunk in iter(lambda: stream.read(1024 * 1024), b""):
|
|
digest.update(chunk)
|
|
commit = subprocess.check_output(["git", "rev-parse", "HEAD"], text=True).strip()
|
|
if manifest["binary_sha256"] != digest.hexdigest() or manifest["commit"] != commit:
|
|
raise SystemExit("downloaded hooks binary identity mismatch")
|
|
shutil.copy2(manifest_path, target / manifest_path.name)
|
|
shutil.copy2(source.with_name("rustfs.e2e.json"), target / "rustfs.e2e.json")
|
|
binary.chmod(0o755)
|
|
PYINPUT
|
|
|
|
- name: Verify e2e full membership
|
|
env:
|
|
NEXTEST_LISTING: ${{ runner.temp }}/rustfs-e2e-full-list.json
|
|
run: |
|
|
cargo nextest list --profile e2e-full -p e2e_test --message-format json > "${NEXTEST_LISTING}"
|
|
python3 ./scripts/check_test_wiring.py --check-profile e2e-full "${NEXTEST_LISTING}"
|
|
|
|
# Full single-node e2e lane (backlog#1149 ci-5). The e2e-full
|
|
# default-filter in .config/nextest.toml is the single wiring mechanism —
|
|
# extend that filter, never add ad-hoc e2e jobs here. Reuses the downloaded
|
|
# debug binary; each test spawns its own rustfs server on a random port.
|
|
- name: Run e2e full suite
|
|
env:
|
|
CARGO_BIN_EXE_rustfs: ${{ runner.temp }}/rustfs-startup-cas-input/rustfs
|
|
RUSTFS_E2E_STARTUP_CAS_BINARY: ${{ runner.temp }}/rustfs-startup-cas-input/rustfs
|
|
RUSTFS_E2E_STARTUP_CAS_BUILD_MANIFEST: ${{ runner.temp }}/rustfs-startup-cas-input/rustfs.e2e-startup-cas-build.json
|
|
RUSTFS_E2E_STARTUP_CAS_ARTIFACT_DIR: ${{ runner.temp }}/rustfs-startup-cas-evidence
|
|
run: python3 scripts/e2e_binary.py run --binary "$RUSTFS_E2E_STARTUP_CAS_BINARY" --features e2e-test-hooks -- cargo nextest run --profile e2e-full -p e2e_test
|
|
|
|
- name: Upload junit
|
|
if: always()
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: e2e-full-junit-${{ github.run_number }}
|
|
path: |
|
|
target/nextest/e2e-full/junit.xml
|
|
${{ runner.temp }}/rustfs-e2e-full-list.json
|
|
retention-days: 7
|
|
|
|
- name: Upload startup CAS evidence
|
|
if: always()
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: fresh-startup-cas-evidence-${{ github.run_number }}
|
|
path: |
|
|
${{ runner.temp }}/rustfs-startup-cas-evidence
|
|
${{ runner.temp }}/rustfs-startup-cas-input/rustfs.e2e-startup-cas-build.json
|
|
if-no-files-found: warn
|
|
retention-days: 7
|
|
|
|
e2e-tests-rio-v2:
|
|
name: End-to-End Tests (rio-v2)
|
|
# Inherits the schedule/dispatch-only gate through needs: on every other
|
|
# event build-rustfs-debug-binary-rio-v2 is skipped, so this job skips
|
|
# with it (see the dormant-variant comment on that job).
|
|
needs: [ build-rustfs-debug-binary-rio-v2 ]
|
|
runs-on: sm-standard-2
|
|
timeout-minutes: 30
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Clean up previous test run
|
|
run: |
|
|
rm -rf /tmp/rustfs
|
|
rm -f /tmp/rustfs.log
|
|
|
|
- name: Download debug binary
|
|
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
|
|
with:
|
|
name: rustfs-debug-binary-rio-v2
|
|
path: target/debug
|
|
|
|
- name: Make binary executable
|
|
run: chmod +x ./target/debug/rustfs
|
|
|
|
- name: Setup Rust toolchain for s3s-e2e installation
|
|
uses: dtolnay/rust-toolchain@29eef336d9b2848a0b548edc03f92a220660cdb8 # stable
|
|
|
|
# The sm-standard-* custom runner images (introduced in #4884) ship no C
|
|
# toolchain, unlike GitHub-hosted ubuntu-latest. Installing s3s-e2e below
|
|
# compiles it from source on a cache miss, and build scripts need cc.
|
|
- name: Install build tools for s3s-e2e compilation
|
|
run: |
|
|
sudo apt-get update
|
|
sudo apt-get install -y build-essential cmake pkg-config libssl-dev
|
|
|
|
- name: Install s3s-e2e test tool
|
|
uses: taiki-e/cache-cargo-install-action@7447f04c51f2ba27ca35e7f1e28fab848c5b3ba7 # v2
|
|
with:
|
|
tool: s3s-e2e
|
|
git: https://github.com/s3s-project/s3s.git
|
|
rev: 62cb4a71dd759a6ec56b64c4c42fcc183a2c6a52
|
|
|
|
- name: Run end-to-end tests
|
|
run: |
|
|
s3s-e2e --version
|
|
python3 scripts/e2e_binary.py run --features rio-v2,e2e-test-hooks -- ./scripts/e2e-run.sh ./target/debug/rustfs /tmp/rustfs
|
|
|
|
- name: Upload test logs
|
|
if: failure()
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: e2e-test-logs-rio-v2-${{ github.run_number }}
|
|
path: /tmp/rustfs.log
|
|
retention-days: 3
|
|
|
|
s3-implemented-tests:
|
|
name: S3 Implemented Tests
|
|
needs: [ build-rustfs-debug-binary, e2e-tests ]
|
|
runs-on: sm-standard-4
|
|
timeout-minutes: 60
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Download debug binary
|
|
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
|
|
with:
|
|
name: rustfs-debug-binary
|
|
path: target/debug
|
|
|
|
- name: Make binary executable
|
|
run: chmod +x ./target/debug/rustfs
|
|
|
|
- name: Run implemented s3-tests
|
|
run: |
|
|
RUN_ROOT="${RUNNER_TEMP}/rustfs-s3tests-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}-${GITHUB_JOB}"
|
|
mkdir -p "${RUN_ROOT}"
|
|
S3_PORT="$(python3 -c 'import socket; s=socket.socket(); s.bind(("127.0.0.1", 0)); print(s.getsockname()[1]); s.close()')"
|
|
DEPLOY_MODE=binary \
|
|
RUSTFS_BINARY=./target/debug/rustfs \
|
|
TEST_MODE=single \
|
|
MAXFAIL=0 \
|
|
XDIST=4 \
|
|
S3_PORT="${S3_PORT}" \
|
|
DATA_ROOT="${RUN_ROOT}" \
|
|
S3TESTS_CONF=artifacts/s3tests-single/s3tests.conf \
|
|
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- ./scripts/s3-tests/run.sh
|
|
|
|
- name: Upload s3 test artifacts
|
|
if: always()
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: s3tests-implemented-${{ github.run_number }}
|
|
path: artifacts/s3tests-single/**
|
|
if-no-files-found: ignore
|
|
retention-days: 3
|
|
|
|
# Dedicated lifecycle behavior lane (backlog#1148 ilm-10, master plan #1155).
|
|
#
|
|
# These ceph/s3-tests cases assert that objects/versions/uploads are actually
|
|
# removed by the background scanner and the stale-multipart cleanup loop, so
|
|
# they need a debug-accelerated day plus an enabled scanner. That configuration
|
|
# cannot be applied to the s3-implemented-tests lane above because a global
|
|
# RUSTFS_ILM_DEBUG_DAY_SECS also shrinks the x-amz-expiration header those tests
|
|
# assert on (test_lifecycle_expiration_header_*). Hence a separate server here.
|
|
#
|
|
# RUSTFS_ILM_DEBUG_DAY_SECS MUST stay equal to the s3-tests lc_debug_interval
|
|
# default (10s). test_lifecycle_expiration polls the Days=1 rule with a
|
|
# 4*lc_interval window while the Days=5 rule fires at 5*debug_day; observing
|
|
# the intermediate 4-object plateau requires 4*lc_interval < 5*debug_day, i.e.
|
|
# debug_day > 8. A smaller value lets the Days=5 rule fire inside the Days=1
|
|
# poll window so the count collapses straight to 2 and the test fails (#4764).
|
|
#
|
|
# RUSTFS_DATA_USAGE_UPDATE_DIR_CYCLES=1 is also required. A compacted directory
|
|
# is only re-descended (and its objects re-evaluated for ILM) once every
|
|
# DATA_USAGE_UPDATE_DIR_CYCLES scanner cycles (default 16). At the accelerated
|
|
# RUSTFS_SCANNER_CYCLE=2 that is ~32s between evaluations, so a Days=1 object
|
|
# that becomes due at debug_day (10s) is not actually expired until the next
|
|
# ~32s boundary (~42s), landing just past the 4*lc_interval (40s) poll window
|
|
# and making the count stall at 6 (#4767). Forcing every-cycle re-descent
|
|
# evaluates ILM within ~2s of the due time, well inside the poll window.
|
|
s3-lifecycle-behavior-tests:
|
|
name: S3 Lifecycle Behavior Tests
|
|
# Also gated on e2e-tests, matching s3-implemented-tests: when the e2e smoke
|
|
# suite is already red this lane cannot tell us anything new, and it holds a
|
|
# sm-standard-4 for up to 30 minutes doing so. Both lanes only download the
|
|
# prebuilt debug binary (no cargo build), and s3-implemented-tests — which
|
|
# already waits on e2e-tests — finishes later anyway, so a green PR's total
|
|
# wall clock is unchanged.
|
|
needs: [ build-rustfs-debug-binary, e2e-tests ]
|
|
runs-on: sm-standard-4
|
|
timeout-minutes: 30
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Download debug binary
|
|
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
|
|
with:
|
|
name: rustfs-debug-binary
|
|
path: target/debug
|
|
|
|
- name: Make binary executable
|
|
run: chmod +x ./target/debug/rustfs
|
|
|
|
- name: Run lifecycle behavior s3-tests
|
|
run: |
|
|
RUN_ROOT="${RUNNER_TEMP}/rustfs-s3tests-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}-${GITHUB_JOB}"
|
|
mkdir -p "${RUN_ROOT}"
|
|
S3_PORT="$(python3 -c 'import socket; s=socket.socket(); s.bind(("127.0.0.1", 0)); print(s.getsockname()[1]); s.close()')"
|
|
DEPLOY_MODE=binary \
|
|
RUSTFS_BINARY=./target/debug/rustfs \
|
|
TEST_MODE=single \
|
|
MAXFAIL=0 \
|
|
XDIST=0 \
|
|
IMPLEMENTED_TESTS_FILE=scripts/s3-tests/lifecycle_behavior_tests.txt \
|
|
RUSTFS_ILM_DEBUG_DAY_SECS=10 \
|
|
RUSTFS_ILM_PROCESS_TIME=1 \
|
|
RUSTFS_SCANNER_ENABLED=true \
|
|
RUSTFS_SCANNER_CYCLE=2 \
|
|
RUSTFS_DATA_USAGE_UPDATE_DIR_CYCLES=1 \
|
|
RUSTFS_SCANNER_START_DELAY_SECS=0 \
|
|
RUSTFS_API_STALE_UPLOADS_CLEANUP_INTERVAL=2s \
|
|
S3_PORT="${S3_PORT}" \
|
|
DATA_ROOT="${RUN_ROOT}" \
|
|
S3TESTS_CONF=artifacts/s3tests-single/s3tests.conf \
|
|
python3 scripts/e2e_binary.py run --features e2e-test-hooks -- ./scripts/s3-tests/run.sh
|
|
|
|
- name: Upload s3 test artifacts
|
|
if: always()
|
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
|
with:
|
|
name: s3tests-lifecycle-behavior-${{ github.run_number }}
|
|
path: artifacts/s3tests-single/**
|
|
if-no-files-found: ignore
|
|
retention-days: 3
|
|
|
|
required-checks:
|
|
name: Test and Lint
|
|
if: always() && (github.event_name != 'pull_request' || github.event.action != 'closed')
|
|
needs:
|
|
- classify-changes
|
|
- typos
|
|
- quick-checks
|
|
- test-and-lint
|
|
- test-ilm-integration-serial
|
|
- test-and-lint-rio-v2
|
|
- connect-short-credential-boundary
|
|
- test-and-lint-protocols
|
|
- build-rustfs-debug-binary
|
|
- uring-integration
|
|
- e2e-tests
|
|
- s3-implemented-tests
|
|
- s3-lifecycle-behavior-tests
|
|
- build-rustfs-debug-binary-rio-v2
|
|
- e2e-tests-rio-v2
|
|
- e2e-full
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
steps:
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
- name: Require the expected result of every CI lane
|
|
env:
|
|
CI_NEEDS: ${{ toJSON(needs) }}
|
|
shell: bash
|
|
run: python3 scripts/ci_gate.py verify
|
|
|
|
alert-on-failure:
|
|
name: Alert on scheduled failure
|
|
needs:
|
|
- classify-changes
|
|
- connect-short-credential-boundary
|
|
- required-checks
|
|
- typos
|
|
- quick-checks
|
|
- test-and-lint
|
|
- test-ilm-integration-serial
|
|
- test-and-lint-rio-v2
|
|
- test-and-lint-protocols
|
|
- build-rustfs-debug-binary
|
|
- build-rustfs-debug-binary-rio-v2
|
|
- uring-integration
|
|
- e2e-tests
|
|
- e2e-full
|
|
- e2e-tests-rio-v2
|
|
- s3-implemented-tests
|
|
- s3-lifecycle-behavior-tests
|
|
if: >-
|
|
always() && github.event_name == 'schedule' &&
|
|
(contains(needs.*.result, 'failure') || contains(needs.*.result, 'cancelled'))
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
permissions:
|
|
contents: read
|
|
issues: write
|
|
steps:
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
|
with:
|
|
persist-credentials: false
|
|
- name: Open or update failure-tracking issue
|
|
uses: ./.github/actions/schedule-failure-issue
|
|
with:
|
|
github-token: ${{ secrets.GITHUB_TOKEN }}
|