Files
rustfs/scripts/test/rustfs_heal_test.md
T
hector a199312e45 test(heal): add node-outage heal E2E script and workflow (#6733)
* test(heal): add node-outage heal E2E script and workflow

RustFS heal test on the 3x4 cluster (3 nodes x 4 disks, same
RUSTFS_VOLUMES expression on every node): write data with warp, stop the
outage node mid-write, restart it, start cluster heal via the admin API,
and pass only when the heal task finishes with 0 failures AND the outage
node's disk usage reaches the target.

Includes the GitHub Actions workflow (smoke-testing runner, nightly deb by
default) and a README. Validated end-to-end on the test environment:
40/40/16 GiB before heal -> 40/40/40 GiB after heal, summary=finished.

The script also writes RUSTFS_HEAL_TASK_TIMEOUT_SECS (default 6h) into the
node config because the server default (5 min) is far too short for
healing tens of GiB.

* ci(pool-test): chain heal regression after the pool test

The pool-expansion workflow is now triggered by the Nightly GNU Build
(workflow_run, replacing the schedule) and runs two sequential jobs on the
shared test environment:

1. pool-expansion-test (existing) — skipped if the nightly build failed.
2. heal-test — runs after the pool test regardless of its outcome
   (if: always()): a pool failure makes the run red but does not block the
   heal regression. Runs the heal script (reset -> install/start 3x4 ->
   write/outage -> heal -> verify -> reset).

* test(heal): address review — camelCase progress, fail-closed, workflow hygiene

- Heal progress fields are camelCase in the API (objectsScanned/objectsHealed/
  objectsFailed/progressPercentage); read them with a snake_case fallback and
  distinguish null (absent) progress from zero, logging null as evidence
  (rustfs/backlog#2035) instead of silently coercing.
- Fail closed in step 3: the outage node must actually be inactive after stop,
  the write target must be reached, and an unobserved outage or incomplete
  write fails the test instead of warning.
- Step 4 waits (bounded) for the cluster to report an active pool after the
  outage-node restart instead of swallowing the verification error.
- Heal start fails fast on 400/403 (deterministic request/auth problems) and
  only retries transient server errors.
- Disable the background scanner (RUSTFS_HEAL_AUTO_HEAL_ENABLE=false) so the
  explicit heal is the only repair mechanism and the outage is observable.
- Workflows: heal and pool share one concurrency group; workflow_run requires
  an exact successful nightly conclusion; checkout is pinned to the triggering
  SHA; comma-separated step args are quoted (actionlint SC2054).
2026-08-27 18:27:13 +08:00

5.0 KiB

RustFS Heal Test

Node-outage heal test driven by scripts/test/rustfs_heal_test.sh, based on the Obsidian note "RustFS Heal 测试步骤". Uses the same 3-node test environment as the pool expansion test (vm000 vm001 vm002).

All status checks talk to the RustFS admin API directly (SigV4-signed, jq assertions), no rc required.

What it does

  1. Downloads the .deb package on all nodes (release tag or a direct URL such as the nightly/R2 package).
  2. Installs it, writes the 3x4 config (http://rustfs-node{1...3}:9000/data/rustfs{1...4}/mnmd), starts all three nodes simultaneously, verifies the cluster is up.
  3. Writes data with warp while monitoring disk usage on the surviving nodes (df -B1G | grep /data/rustfs):
    • when both surviving nodes reach STOP_NODE_AT_GB (default 15 GiB), stop the outage node (vm002, OUTAGE_NODE_INDEX=2);
    • keep writing until both surviving nodes reach WARP_STOP_AT_GB (default 40 GiB), then stop warp.
  4. Restarts the outage node.
  5. Starts cluster heal: POST /rustfs/admin/v3/heal/ with body {"recursive":true} (retried, returns a clientToken).
  6. Monitors the heal task via POST /rustfs/admin/v3/heal/?clientToken=<token> until the summary is a terminal success (finished/completed), objects_failed == 0, and the outage node's disk usage reaches HEAL_TARGET_GB (default 40 GiB).
  7. Result analysis: heal stats (scanned/healed/failed), per-node disk usage, pass/fail verdict.

Success requires both the heal API completion (the server's scan/repair verdict) and the outage node's disk reaching the target.

Self-hosted runner prerequisites

  • Register the admin host (e.g. heal) as a runner with the smoke-testing label.
  • Install jq, openssl, curl and warp on the runner. rc is not required.
  • The runner user must be able to SSH to vm000/vm001/vm002 without a password prompt; nodes need passwordless sudo for the SSH user and resolvable rustfs-node* hostnames.
  • Admin API credentials need the admin:server-info, admin:heal and admin:rebalance actions.

Configuration

Same repository secrets/variables as the pool expansion workflow:

Kind Name Purpose
Secret RUSTFS_ACCESS_KEY RustFS access key (default rustfs@test)
Secret RUSTFS_SECRET_KEY RustFS secret key (default rustfs@test)
Var RUSTFS_API_ENDPOINT Admin API endpoint, e.g. http://127.0.0.1:9000 (RUSTFS_RC_ENDPOINT fallback)
Var RUSTFS_NODES vm000 vm001 vm002
Var RUSTFS_SSH_USER azureuser
Var RUSTFS_NIGHTLY_PACKAGE_URL Default nightly deb URL (defaults to the R2 latest alias)

Workflow inputs

Input Default Meaning
package_url nightly Direct .deb URL; empty = latest nightly
stop_node_gb 15 Stop outage node at N GiB on survivors
warp_stop_gb 40 Stop warp at N GiB on survivors
heal_target_gb 40 Outage node must reach N GiB after heal
cleanup_before true Reset nodes before the test
cleanup_after true Reset nodes after the test

⚠️ --reset purges the rustfs package and deletes the data directories on all nodes. Only run against a dedicated test environment.

Manual usage

./scripts/test/rustfs_heal_test.sh --all -y \
  --package-url https://dl.rustfs.com/artifacts/rustfs/packages/nightly/rustfs-nightly-latest.deb \
  --endpoint http://127.0.0.1:9000

./scripts/test/rustfs_heal_test.sh --steps 5,6,7
./scripts/test/rustfs_heal_test.sh --reset -y

Known issues

  • Nightly builds gate pool/rebalance activation on a live fleet capability proof (rustfs/backlog#2031); the script retries heal/rebalance starts and prints a hint when the signature appears.
  • The cluster-level GET /rustfs/admin/v3/background-heal/status aggregator returns 501 in the single-pool 3x4 topology (no notification system), so the script monitors the started heal task via its clientToken instead.
  • The heal task may report progress: null while running; the script logs this as evidence (rustfs/backlog#2035) rather than coercing it to zero, and reads the canonical camelCase progress fields (objectsScanned/objectsHealed/objectsFailed/progressPercentage) with a snake_case fallback.
  • The server-side per-task heal timeout defaults to 5 minutes; the script writes RUSTFS_HEAL_TASK_TIMEOUT_SECS=21600 (6h) into the node config so a multi-tens-of-GiB heal can finish. The background scanner is disabled (RUSTFS_HEAL_AUTO_HEAL_ENABLE=false) so the explicit heal is the only repair mechanism and the outage effect stays observable.