mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-29 00:17:11 +00:00
a199312e45
* test(heal): add node-outage heal E2E script and workflow RustFS heal test on the 3x4 cluster (3 nodes x 4 disks, same RUSTFS_VOLUMES expression on every node): write data with warp, stop the outage node mid-write, restart it, start cluster heal via the admin API, and pass only when the heal task finishes with 0 failures AND the outage node's disk usage reaches the target. Includes the GitHub Actions workflow (smoke-testing runner, nightly deb by default) and a README. Validated end-to-end on the test environment: 40/40/16 GiB before heal -> 40/40/40 GiB after heal, summary=finished. The script also writes RUSTFS_HEAL_TASK_TIMEOUT_SECS (default 6h) into the node config because the server default (5 min) is far too short for healing tens of GiB. * ci(pool-test): chain heal regression after the pool test The pool-expansion workflow is now triggered by the Nightly GNU Build (workflow_run, replacing the schedule) and runs two sequential jobs on the shared test environment: 1. pool-expansion-test (existing) — skipped if the nightly build failed. 2. heal-test — runs after the pool test regardless of its outcome (if: always()): a pool failure makes the run red but does not block the heal regression. Runs the heal script (reset -> install/start 3x4 -> write/outage -> heal -> verify -> reset). * test(heal): address review — camelCase progress, fail-closed, workflow hygiene - Heal progress fields are camelCase in the API (objectsScanned/objectsHealed/ objectsFailed/progressPercentage); read them with a snake_case fallback and distinguish null (absent) progress from zero, logging null as evidence (rustfs/backlog#2035) instead of silently coercing. - Fail closed in step 3: the outage node must actually be inactive after stop, the write target must be reached, and an unobserved outage or incomplete write fails the test instead of warning. - Step 4 waits (bounded) for the cluster to report an active pool after the outage-node restart instead of swallowing the verification error. - Heal start fails fast on 400/403 (deterministic request/auth problems) and only retries transient server errors. - Disable the background scanner (RUSTFS_HEAL_AUTO_HEAL_ENABLE=false) so the explicit heal is the only repair mechanism and the outage is observable. - Workflows: heal and pool share one concurrency group; workflow_run requires an exact successful nightly conclusion; checkout is pinned to the triggering SHA; comma-separated step args are quoted (actionlint SC2054).
5.0 KiB
5.0 KiB
RustFS Heal Test
Node-outage heal test driven by
scripts/test/rustfs_heal_test.sh, based on the
Obsidian note "RustFS Heal 测试步骤". Uses the same 3-node test environment as
the pool expansion test (vm000 vm001 vm002).
All status checks talk to the RustFS admin API directly (SigV4-signed,
jq assertions), no rc required.
What it does
- Downloads the
.debpackage on all nodes (release tag or a direct URL such as the nightly/R2 package). - Installs it, writes the 3x4 config
(
http://rustfs-node{1...3}:9000/data/rustfs{1...4}/mnmd), starts all three nodes simultaneously, verifies the cluster is up. - Writes data with
warpwhile monitoring disk usage on the surviving nodes (df -B1G | grep /data/rustfs):- when both surviving nodes reach
STOP_NODE_AT_GB(default 15 GiB), stop the outage node (vm002,OUTAGE_NODE_INDEX=2); - keep writing until both surviving nodes reach
WARP_STOP_AT_GB(default 40 GiB), then stop warp.
- when both surviving nodes reach
- Restarts the outage node.
- Starts cluster heal:
POST /rustfs/admin/v3/heal/with body{"recursive":true}(retried, returns aclientToken). - Monitors the heal task via
POST /rustfs/admin/v3/heal/?clientToken=<token>until the summary is a terminal success (finished/completed),objects_failed == 0, and the outage node's disk usage reachesHEAL_TARGET_GB(default 40 GiB). - Result analysis: heal stats (scanned/healed/failed), per-node disk usage, pass/fail verdict.
Success requires both the heal API completion (the server's scan/repair verdict) and the outage node's disk reaching the target.
Self-hosted runner prerequisites
- Register the admin host (e.g.
heal) as a runner with thesmoke-testinglabel. - Install
jq,openssl,curlandwarpon the runner.rcis not required. - The runner user must be able to SSH to
vm000/vm001/vm002without a password prompt; nodes need passwordlesssudofor the SSH user and resolvablerustfs-node*hostnames. - Admin API credentials need the
admin:server-info,admin:healandadmin:rebalanceactions.
Configuration
Same repository secrets/variables as the pool expansion workflow:
| Kind | Name | Purpose |
|---|---|---|
| Secret | RUSTFS_ACCESS_KEY |
RustFS access key (default rustfs@test) |
| Secret | RUSTFS_SECRET_KEY |
RustFS secret key (default rustfs@test) |
| Var | RUSTFS_API_ENDPOINT |
Admin API endpoint, e.g. http://127.0.0.1:9000 (RUSTFS_RC_ENDPOINT fallback) |
| Var | RUSTFS_NODES |
vm000 vm001 vm002 |
| Var | RUSTFS_SSH_USER |
azureuser |
| Var | RUSTFS_NIGHTLY_PACKAGE_URL |
Default nightly deb URL (defaults to the R2 latest alias) |
Workflow inputs
| Input | Default | Meaning |
|---|---|---|
package_url |
nightly | Direct .deb URL; empty = latest nightly |
stop_node_gb |
15 |
Stop outage node at N GiB on survivors |
warp_stop_gb |
40 |
Stop warp at N GiB on survivors |
heal_target_gb |
40 |
Outage node must reach N GiB after heal |
cleanup_before |
true |
Reset nodes before the test |
cleanup_after |
true |
Reset nodes after the test |
⚠️
--resetpurges therustfspackage and deletes the data directories on all nodes. Only run against a dedicated test environment.
Manual usage
./scripts/test/rustfs_heal_test.sh --all -y \
--package-url https://dl.rustfs.com/artifacts/rustfs/packages/nightly/rustfs-nightly-latest.deb \
--endpoint http://127.0.0.1:9000
./scripts/test/rustfs_heal_test.sh --steps 5,6,7
./scripts/test/rustfs_heal_test.sh --reset -y
Known issues
- Nightly builds gate pool/rebalance activation on a live fleet capability proof (rustfs/backlog#2031); the script retries heal/rebalance starts and prints a hint when the signature appears.
- The cluster-level
GET /rustfs/admin/v3/background-heal/statusaggregator returns 501 in the single-pool 3x4 topology (no notification system), so the script monitors the started heal task via itsclientTokeninstead. - The heal task may report
progress: nullwhile running; the script logs this as evidence (rustfs/backlog#2035) rather than coercing it to zero, and reads the canonical camelCase progress fields (objectsScanned/objectsHealed/objectsFailed/progressPercentage) with a snake_case fallback. - The server-side per-task heal timeout defaults to 5 minutes; the script
writes
RUSTFS_HEAL_TASK_TIMEOUT_SECS=21600(6h) into the node config so a multi-tens-of-GiB heal can finish. The background scanner is disabled (RUSTFS_HEAL_AUTO_HEAL_ENABLE=false) so the explicit heal is the only repair mechanism and the outage effect stays observable.