test(ci): add strict nextest ci profile with quarantine + flake policy Formalize the existing ecstore-serial-flaky mechanism into a strict CI gate (ci-10, absorbs infra-15; backlog#1149). - .config/nextest.toml: add [profile.ci] with global retries=0 (never mask a new race's first occurrence), fail-fast=false, and JUnit output at target/nextest/ci/junit.xml. Add a quarantine section where flaky tests get retries=2 under the ci profile only; each entry links one OPEN issue. First members are the two backlog#937 ecstore groups (concurrent_resend_same_part_commits_one_generation and store::bucket::tests::bucket_delete_*), which keep their existing ecstore-serial-flaky test-group serialization. Local default profile still never retries. - .github/workflows/ci.yml: run the main test step with --profile ci and upload the JUnit report (if: always(), 3-day retention, run-number in name). The migration-proof step stays on the default profile to avoid clobbering the ci JUnit artifact (its tests are not quarantined). - docs/testing/README.md: new skeleton (owned by backlog#1153 infra-11) holding the flake policy: discover -> open issue within 24h -> quarantine with issue link -> fix or delete within 30 days. AGENTS.md points to it. Refs: rustfs/backlog#1149, rustfs/backlog#937, rustfs/backlog#1155
2.6 KiB
RustFS Testing
Owner: backlog#1153 (infra-11). This file is the authoritative home for test taxonomy, naming conventions, and serial/quarantine rules. It is currently a skeleton — infra-11 fills in the remaining sections. The event × budget × required matrix lives in
docs/testing/ci-gates.md(ci-15); this file links to it rather than duplicating counts.
Test taxonomy
TODO (infra-11): unit / e2e-smoke / e2e-full / e2e-nightly / protocols / s3-tests / fuzz / perf, with naming conventions.
Serial groups & CI profiles
TODO (infra-11): document the [test-groups] mechanism and the default vs
ci nextest profiles. See .config/nextest.toml.
Flake policy
A flaky test is one that fails non-deterministically without a corresponding code change. Flakes erode trust in the gate and block tightening required checks, so they are handled on a strict, time-boxed loop.
Retry semantics (source of truth: .config/nextest.toml):
- The local
defaultprofile never retries. A red test on your machine is a real failure to investigate, not noise to paper over. - The CI
ciprofile runs with globalretries = 0. A new race must fail on its first occurrence so the first crime scene is never masked. - Only tests on the quarantine list get
retries = 2, and only under theciprofile. Each quarantine entry MUST link exactly one OPEN issue. - JUnit flaky markers are the observable. A quarantined test that passes
only after a retry is marked
flakyintarget/nextest/ci/junit.xml(uploaded as a CI artifact). That marker — not a green check — is how we see a flake is still live.
Lifecycle of a flake:
- Discover — a test fails non-deterministically (CI or local), or shows a
flakymarker in the JUnit report. - Open an issue within 24h — file/track an issue describing the flake (symptom, suspected cause, affected suite). No silent re-runs.
- Quarantine — add the test to the quarantine override block in
.config/nextest.tomlwith a comment linking that OPEN issue. This grantsretries = 2under CI so the flake stops reddening unrelated PRs, while theflakymarker keeps it visible. - Fix or delete within 30 days — make the test robust (then remove the quarantine entry) or delete the test. A quarantine entry may not outlive its fix window; an entry without a live OPEN issue link is a policy violation.
First quarantine members: the two backlog#937 ecstore groups
(concurrent_resend_same_part_commits_one_generation and
store::bucket::tests::bucket_delete_*).