* fix(auth): reject unsigned x-amz headers on header-signed SigV4 requests A SigV4 request authenticated with an Authorization header only binds the headers named in its SignedHeaders list, but RustFS acted on every x-amz-* header that arrived, so a replayed header-signed PutObject carrying an unsigned x-amz-copy-source became a CopyObject run as the signer that could copy any object the signer can read (GHSA-xm99-m3gq-83g8). The presigned form was already closed by GHSA-g8w9-qw9q-fghr. reject_unsigned_amz_headers_on_sigv4_request now guards S3Access::check and S3Router::check_access: every SigV4 signed-header list the request carries must cover every x-amz-* header, the Authorization header is parsed with the verifier's own s3s-sigv4 parser and compared case-insensitively, the algorithm token is pinned to AWS4-HMAC-SHA256 because the upstream header path accepts any token, and the exempt set mirrors the upstream s3s fix (x-amz-content-sha256, x-amz-decoded-content-length, x-amz-trailer, x-amz-checksum-algorithm) plus x-amz-cf-id. Adds ghsa_xm99 unit, router and e2e regressions, raises the security smoke floor to 28, and records the advisory in docs/testing/security-regressions.md and CHANGELOG.md. * chore(deps): switch s3s and s3s-sigv4 to crates.io 0.16.0 * fix(server): enforce the SigV4 header guard ahead of s3s dispatch s3s 0.16 verifies the claimed algorithm as the first step of its own signature flow, so a request whose Authorization header swaps the AWS4-HMAC-SHA256 token was answered with 501 NotImplemented before RustFS's access layer could rule on the unsigned x-amz-copy-source (GHSA-xm99-m3gq-83g8). Add the SigV4HeaderGuardLayer as the innermost external stack layer, running reject_unsigned_amz_headers_on_sigv4_request in front of s3s and serializing its rejections as the same AccessDenied S3 error document the access layer produces. --------- Co-authored-by: Hauser <housemecn@gmail.com>
RustFS Testing
Use this when: you need to pick a test layer for a change, name a test so a gate keeps selecting it, understand why #[serial] does nothing under nextest, or handle a flaky test.
Source of truth: .config/nextest.toml (profiles, test-groups, quarantine), .config/make/tests.mak (make test), .github/workflows/*.yml (what runs when; matrix in ci-gates.md).
Test taxonomy
Pick the lowest layer that can prove the change; add a higher-layer test only when the behaviour is not observable below it.
| Layer | What it covers | Entry command | When it runs (details: ci-gates.md) |
|---|---|---|---|
| Unit & crate integration | Per-crate logic and in-process integration tests | cargo nextest run --all --exclude e2e_test (or -p <crate>); make test wraps it |
Every PR, required (Test and Lint, ci profile) |
| ecstore black-box | Erasure-coded read/write/recovery validation; profiles quick / full / destructive / fuzz |
scripts/run_ecstore_validation_suite.sh --profile quick |
Local and release validation only; not wired into any workflow. Contract: ecstore-validation-suite-design.md |
e2e (e2e_test crate) |
A real rustfs binary per test, driven over the S3, admin, and protocol APIs |
python3 scripts/e2e_binary.py build --features e2e-test-hooks, then python3 scripts/e2e_binary.py run --features e2e-test-hooks -- cargo nextest run --profile e2e-smoke -p e2e_test |
PR: e2e-smoke (report-only); merge queue / main push: e2e-full; nightly: e2e-repl-nightly, e2e-nightly, e2e-protocols, e2e-distributed. Guide: crates/e2e_test/README.md; 4-node 4-disk map: distributed-e2e.md |
| Outbound target matrix | Replication of every object shape (empty, plain, retention, legal hold, multipart) against every remote-target failure mode the fake target models; an explicit expectation table pins known-red cells to an open issue | python3 scripts/e2e_binary.py build, then python3 scripts/e2e_binary.py run -- cargo nextest run -p e2e_test -E 'test(/^replication_target_matrix_test::/)' |
With e2e-repl-nightly; required locally for any change to outbound client defaults (SOP: docs/postmortems/2026-09-03-replication-checksum-default-regression.md) |
| s3s-e2e conformance | External S3 conformance tool against a live server | ./scripts/e2e-run.sh ./target/debug/rustfs <data-dir> |
PR, report-only (second half of the End-to-End Tests job) |
| S3 compatibility | ceph/s3-tests (boto3; allow-list scripts/s3-tests/implemented_tests.txt) and MinIO mint |
scripts/s3-tests/run.sh; mint via .github/workflows/mint.yml |
s3-tests: PR report-only plus a weekly full sweep; mint: weekly, report-only |
| Chaos / fault-injection | Single-node disk fault injection (crates/e2e_test/src/chaos.rs, crates/e2e_test/src/fault_proxy.rs) plus the 4-node kill/fresh-drive/blackhole cases in crates/e2e_test/src/distributed/chaos_test.rs |
Part of the e2e crate (e2e-reliability and e2e-distributed) |
Reliability cases with e2e-full; 4-node chaos on storage-sensitive PRs and nightly via e2e-distributed |
| Fuzz | cargo-fuzz targets over untrusted parsing surfaces; isolated sub-workspace under fuzz/ |
./scripts/fuzz/run.sh (see fuzz/README.md) |
PR smoke on the paths listed in .github/workflows/fuzz.yml, plus nightly corpus |
| Benchmarks | Criterion benches under each crate's benches/ |
cargo bench -p <crate> |
On demand; never a gate |
Every script named above is indexed with status and wiring in scripts/README.md. Fixed GHSA advisories map to named regression tests in security-regressions.md.
The scanner checkpoint fixture diagnoses retained subtree coverage across budget interruption, persistence, reload, and plan invalidation.
The scanner cache cost profile separates clone, subtree copy, encoding, and counted save costs without changing production cache behavior.
The Pool layout compatibility reference defines the topology and EC regression matrix for single-drive, single-node multi-drive, and multi-node expansion pools.
Naming conventions
Reserved test-name substrings (migration gate)
scripts/check_migration_gate_count.sh (runs in Test and Lint) selects migration-critical tests by name substring and fails when the count drops below .config/migration-gate-floor.txt. A rename that drops a substring silently thins the gate, so these substrings are reserved:
| Substring | Guards |
|---|---|
data_movement |
Cross-pool / cross-set object data-movement proofs |
rebalance |
Pool rebalance correctness |
decommission |
Pool decommission correctness |
source_cleanup |
Post-migration source cleanup |
delete_marker |
Delete-marker handling across migration |
A deliberate reduction lowers the floor in the same PR. The list above mirrors the script; change both together.
General naming
- Name a regression test after what it pins (issue or advisory number, or the invariant) so
rgfinds the guard for a past bug. - e2e lane membership is selected by test-name patterns in
.config/nextest.toml(for example_real_dual_node/_real_single_noderoute replication tests to the nightly lane). Follow the existing marker of the suite you extend. - Symbol naming follows the Rust API Guidelines (see
AGENTS.md).
nextest and #[serial]
cargo-nextest is the runner: make test requires it and CI installs it. nextest runs every test in its own process, so serial_test's in-process #[serial] mutex does not serialize tests against each other; it only affects the plain cargo test fallback. Cross-test serialization under nextest comes from a [test-groups] entry with max-threads = 1 in .config/nextest.toml (for example ecstore-serial-flaky, e2e-reliability) or from a -j 1 lane. Prefer making tests self-isolating (per-test instance context, random port, own temp dir) over adding serialization. RUSTFS_ALLOW_CARGO_TEST_FALLBACK=1 make test runs plain cargo test; its results are not authoritative because [test-groups] do not apply.
Time-driven tests use paused tokio time (start_paused plus tokio::time::advance) or explicit synchronization instead of fixed sleep windows.
Profiles
All profiles are defined in .config/nextest.toml; its block comments hold the filters and rationale.
| Profile | Role |
|---|---|
default |
Local runs; never retries |
ci |
PR gate for everything except e2e_test; global retries = 0 plus the quarantine list |
e2e-smoke |
PR subset of e2e_test |
e2e-full |
Merge-queue / main-push single-node e2e lane |
e2e-repl-nightly |
Nightly slow / cross-process replication lane |
e2e-nightly |
Nightly serial multi-process cluster fault lane |
e2e-distributed |
Storage-sensitive PR and nightly 4-node 4-disk S3 / lock / versioning / replication / quota / expand / decommission / rebalance / site-replication / chaos / upgrade (history + IAM AK/SK) lane |
e2e-protocols |
Nightly fixed-port FTPS/SFTP/WebDAV lane, run with -j 1 |
Membership of each e2e profile is pinned by a digest in .config/e2e-<profile>-selection.txt and checked by scripts/check_test_wiring.py --check-profile <profile> before the lane runs. To list what a profile selects on your platform (the result is platform-dependent because some modules are linux-only):
cargo nextest list -p e2e_test --profile e2e-smoke --message-format json \
| jq -r '.["rust-suites"][].testcases | to_entries[] | select(.value["filter-match"].status == "matches") | .key | split("::")[0]' \
| sort | uniq -c
Flake policy
A flaky test fails non-deterministically without a corresponding code change. Retry semantics live in .config/nextest.toml: default never retries, ci has global retries = 0, and only quarantined tests get retries = 2 under ci. A quarantined test that passes on retry is marked flaky in target/nextest/ci/junit.xml (uploaded as a CI artifact); that marker, not a green check, is how a live flake stays visible.
- Discover — a non-deterministic failure (CI or local) or a
flakyJUnit marker. - Open an issue within 24h describing symptom, suspected cause, and affected suite. No silent re-runs.
- Quarantine — add a
[[profile.ci.overrides]]entry withretries = 2and a comment linking exactly one OPEN issue. The current quarantine list is the[[profile.ci.overrides]]block in.config/nextest.toml. - Fix or delete within 30 days — make the test robust and remove the entry, or delete the test. An entry without a live OPEN issue link is a policy violation.
Coverage
.github/workflows/coverage.ymlmeasures workspace line coverage on its schedule and on manual dispatch:cargo llvm-cov nextest --workspace --exclude e2e_testunder theciprofile, the same scope as the PR gate. The per-crate table lands in the job summary; lcov and JSON exports are uploaded as an artifact (retention set in the workflow).- PRs touching the paths listed in
coverage.ymlalso run a report-only comparison against.config/coverage-baselines.tomlviascripts/check_security_coverage.py: a regression is recorded in the summary without failing the job; missing or malformed coverage evidence fails closed. make coverage(.config/make/coverage.mak) is the local equivalent; it writestarget/llvm-cov/lcov.infoandcoverage.jsonand prints the same table viascripts/coverage_per_crate.py.- Not measured: doctests (
ci.ymlruns them uninstrumented) and thee2e_testcrate. - A baseline change needs a linked coverage run and a reviewed explanation in the PR.