mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-29 08:27:06 +00:00
a199312e45
* test(heal): add node-outage heal E2E script and workflow RustFS heal test on the 3x4 cluster (3 nodes x 4 disks, same RUSTFS_VOLUMES expression on every node): write data with warp, stop the outage node mid-write, restart it, start cluster heal via the admin API, and pass only when the heal task finishes with 0 failures AND the outage node's disk usage reaches the target. Includes the GitHub Actions workflow (smoke-testing runner, nightly deb by default) and a README. Validated end-to-end on the test environment: 40/40/16 GiB before heal -> 40/40/40 GiB after heal, summary=finished. The script also writes RUSTFS_HEAL_TASK_TIMEOUT_SECS (default 6h) into the node config because the server default (5 min) is far too short for healing tens of GiB. * ci(pool-test): chain heal regression after the pool test The pool-expansion workflow is now triggered by the Nightly GNU Build (workflow_run, replacing the schedule) and runs two sequential jobs on the shared test environment: 1. pool-expansion-test (existing) — skipped if the nightly build failed. 2. heal-test — runs after the pool test regardless of its outcome (if: always()): a pool failure makes the run red but does not block the heal regression. Runs the heal script (reset -> install/start 3x4 -> write/outage -> heal -> verify -> reset). * test(heal): address review — camelCase progress, fail-closed, workflow hygiene - Heal progress fields are camelCase in the API (objectsScanned/objectsHealed/ objectsFailed/progressPercentage); read them with a snake_case fallback and distinguish null (absent) progress from zero, logging null as evidence (rustfs/backlog#2035) instead of silently coercing. - Fail closed in step 3: the outage node must actually be inactive after stop, the write target must be reached, and an unobserved outage or incomplete write fails the test instead of warning. - Step 4 waits (bounded) for the cluster to report an active pool after the outage-node restart instead of swallowing the verification error. - Heal start fails fast on 400/403 (deterministic request/auth problems) and only retries transient server errors. - Disable the background scanner (RUSTFS_HEAL_AUTO_HEAL_ENABLE=false) so the explicit heal is the only repair mechanism and the outage is observable. - Workflows: heal and pool share one concurrency group; workflow_run requires an exact successful nightly conclusion; checkout is pinned to the triggering SHA; comma-separated step args are quoted (actionlint SC2054).