test(heal): add node-outage heal E2E script and workflow (#6733)

* test(heal): add node-outage heal E2E script and workflow

RustFS heal test on the 3x4 cluster (3 nodes x 4 disks, same
RUSTFS_VOLUMES expression on every node): write data with warp, stop the
outage node mid-write, restart it, start cluster heal via the admin API,
and pass only when the heal task finishes with 0 failures AND the outage
node's disk usage reaches the target.

Includes the GitHub Actions workflow (smoke-testing runner, nightly deb by
default) and a README. Validated end-to-end on the test environment:
40/40/16 GiB before heal -> 40/40/40 GiB after heal, summary=finished.

The script also writes RUSTFS_HEAL_TASK_TIMEOUT_SECS (default 6h) into the
node config because the server default (5 min) is far too short for
healing tens of GiB.

* ci(pool-test): chain heal regression after the pool test

The pool-expansion workflow is now triggered by the Nightly GNU Build
(workflow_run, replacing the schedule) and runs two sequential jobs on the
shared test environment:

1. pool-expansion-test (existing) — skipped if the nightly build failed.
2. heal-test — runs after the pool test regardless of its outcome
   (if: always()): a pool failure makes the run red but does not block the
   heal regression. Runs the heal script (reset -> install/start 3x4 ->
   write/outage -> heal -> verify -> reset).

* test(heal): address review — camelCase progress, fail-closed, workflow hygiene

- Heal progress fields are camelCase in the API (objectsScanned/objectsHealed/
  objectsFailed/progressPercentage); read them with a snake_case fallback and
  distinguish null (absent) progress from zero, logging null as evidence
  (rustfs/backlog#2035) instead of silently coercing.
- Fail closed in step 3: the outage node must actually be inactive after stop,
  the write target must be reached, and an unobserved outage or incomplete
  write fails the test instead of warning.
- Step 4 waits (bounded) for the cluster to report an active pool after the
  outage-node restart instead of swallowing the verification error.
- Heal start fails fast on 400/403 (deterministic request/auth problems) and
  only retries transient server errors.
- Disable the background scanner (RUSTFS_HEAL_AUTO_HEAL_ENABLE=false) so the
  explicit heal is the only repair mechanism and the outage is observable.
- Workflows: heal and pool share one concurrency group; workflow_run requires
  an exact successful nightly conclusion; checkout is pinned to the triggering
  SHA; comma-separated step args are quoted (actionlint SC2054).
This commit is contained in:
hector
2026-08-27 18:27:13 +08:00
committed by GitHub
parent 3420006762
commit a199312e45
4 changed files with 1427 additions and 5 deletions
+104
View File
@@ -0,0 +1,104 @@
# RustFS Heal Test
Node-outage heal test driven by
[`scripts/test/rustfs_heal_test.sh`](rustfs_heal_test.sh), based on the
Obsidian note "RustFS Heal 测试步骤". Uses the same 3-node test environment as
the pool expansion test (`vm000 vm001 vm002`).
All status checks talk to the RustFS admin API directly (SigV4-signed,
`jq` assertions), no `rc` required.
## What it does
1. Downloads the `.deb` package on all nodes (release tag or a direct URL such
as the nightly/R2 package).
2. Installs it, writes the 3x4 config
(`http://rustfs-node{1...3}:9000/data/rustfs{1...4}/mnmd`), starts all
three nodes simultaneously, verifies the cluster is up.
3. Writes data with `warp` while monitoring disk usage on the surviving nodes
(`df -B1G | grep /data/rustfs`):
- when both surviving nodes reach `STOP_NODE_AT_GB` (default 15 GiB), stop
the outage node (`vm002`, `OUTAGE_NODE_INDEX=2`);
- keep writing until both surviving nodes reach `WARP_STOP_AT_GB`
(default 40 GiB), then stop warp.
4. Restarts the outage node.
5. Starts cluster heal: `POST /rustfs/admin/v3/heal/` with body
`{"recursive":true}` (retried, returns a `clientToken`).
6. Monitors the heal task via `POST /rustfs/admin/v3/heal/?clientToken=<token>`
until the summary is a terminal success (`finished`/`completed`),
`objects_failed == 0`, **and** the outage node's disk usage reaches
`HEAL_TARGET_GB` (default 40 GiB).
7. Result analysis: heal stats (scanned/healed/failed), per-node disk usage,
pass/fail verdict.
Success requires **both** the heal API completion (the server's scan/repair
verdict) and the outage node's disk reaching the target.
## Self-hosted runner prerequisites
- Register the admin host (e.g. `heal`) as a runner with the
`smoke-testing` label.
- Install `jq`, `openssl`, `curl` and `warp` on the runner. `rc` is **not**
required.
- The runner user must be able to SSH to `vm000/vm001/vm002` without a
password prompt; nodes need passwordless `sudo` for the SSH user and
resolvable `rustfs-node*` hostnames.
- Admin API credentials need the `admin:server-info`, `admin:heal` and
`admin:rebalance` actions.
## Configuration
Same repository secrets/variables as the pool expansion workflow:
| Kind | Name | Purpose |
| ------ | --------------------- | ---------------------------------------------- |
| Secret | `RUSTFS_ACCESS_KEY` | RustFS access key (default `rustfs@test`) |
| Secret | `RUSTFS_SECRET_KEY` | RustFS secret key (default `rustfs@test`) |
| Var | `RUSTFS_API_ENDPOINT` | Admin API endpoint, e.g. `http://127.0.0.1:9000` (`RUSTFS_RC_ENDPOINT` fallback) |
| Var | `RUSTFS_NODES` | `vm000 vm001 vm002` |
| Var | `RUSTFS_SSH_USER` | `azureuser` |
| Var | `RUSTFS_NIGHTLY_PACKAGE_URL` | Default nightly deb URL (defaults to the R2 `latest` alias) |
## Workflow inputs
| Input | Default | Meaning |
| ---------------- | ------- | ----------------------------------------- |
| `package_url` | nightly | Direct `.deb` URL; empty = latest nightly |
| `stop_node_gb` | `15` | Stop outage node at N GiB on survivors |
| `warp_stop_gb` | `40` | Stop warp at N GiB on survivors |
| `heal_target_gb` | `40` | Outage node must reach N GiB after heal |
| `cleanup_before` | `true` | Reset nodes before the test |
| `cleanup_after` | `true` | Reset nodes after the test |
> ⚠️ `--reset` purges the `rustfs` package and deletes the data directories on
> all nodes. Only run against a dedicated test environment.
## Manual usage
```bash
./scripts/test/rustfs_heal_test.sh --all -y \
--package-url https://dl.rustfs.com/artifacts/rustfs/packages/nightly/rustfs-nightly-latest.deb \
--endpoint http://127.0.0.1:9000
./scripts/test/rustfs_heal_test.sh --steps 5,6,7
./scripts/test/rustfs_heal_test.sh --reset -y
```
## Known issues
- Nightly builds gate pool/rebalance activation on a live fleet capability
proof (rustfs/backlog#2031); the script retries heal/rebalance starts and
prints a hint when the signature appears.
- The cluster-level `GET /rustfs/admin/v3/background-heal/status` aggregator
returns 501 in the single-pool 3x4 topology (no notification system), so the
script monitors the started heal task via its `clientToken` instead.
- The heal task may report `progress: null` while running; the script logs
this as evidence (rustfs/backlog#2035) rather than coercing it to zero, and
reads the canonical camelCase progress fields
(`objectsScanned`/`objectsHealed`/`objectsFailed`/`progressPercentage`) with
a snake_case fallback.
- The server-side per-task heal timeout defaults to 5 minutes; the script
writes `RUSTFS_HEAL_TASK_TIMEOUT_SECS=21600` (6h) into the node config so a
multi-tens-of-GiB heal can finish. The background scanner is disabled
(`RUSTFS_HEAL_AUTO_HEAL_ENABLE=false`) so the explicit heal is the only
repair mechanism and the outage effect stays observable.