mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-29 08:27:06 +00:00
2e6c820f53
* test(heal): relative disk target and fail fast on terminal-but-short The absolute 40 GiB heal target was calibrated to the background scanner (auto-heal), which is now disabled for determinism; with only the explicit heal the recovered node lands at ~36 GiB for 40 GiB survivors. Make the success criterion relative: the outage node must reach at least 90% of the least-used surviving node (absolute HEAL_TARGET_GB floor optional, default 0 = relative only). Also fail fast when the heal task reaches a terminal success but the disk target is not met (previously the monitor kept polling until timeout), and drop the misleading 'progress absent' warning on the final (cleaned) task response — mid-run progress is reported correctly. Validated live: heal summary=finished, 0 failed, vm000/vm001=40GB, vm002=40GB (target 36GB), test PASSED. * test(heal): gate success on server verdict + data read-back, drop disk GB gate The per-node disk-usage target (40 GiB / 90% of survivors) is not a code-level invariant: EC distributes different shards per node, so the final GB per node depends on the layout, not on heal correctness. Gate the test on what the server actually verifies: - Heal task terminal success (finished/completed) with objectsFailed == 0 (the server's per-object scan/repair verdict). - S3 read-back verification: list the test bucket and GET a sample of objects, requiring HTTP 200 for every read (end-to-end proof the data is still reconstructable after repair). The GET uses a discard mode so binary bodies are not captured (no null-byte warnings / SIGPIPE). Per-node disk usage stays in the output as observability (with a warning if the outage node gained no usage), not as the pass/fail gate. Removes the heal_target_gb input and the relative-target logic. Validated live: heal summary=finished, 0 failed, 20/20 objects read back, vm002_used=40GB, PASS.
107 lines
5.2 KiB
Markdown
107 lines
5.2 KiB
Markdown
# RustFS Heal Test
|
|
|
|
Node-outage heal test driven by
|
|
[`scripts/test/rustfs_heal_test.sh`](rustfs_heal_test.sh), based on the
|
|
Obsidian note "RustFS Heal 测试步骤". Uses the same 3-node test environment as
|
|
the pool expansion test (`vm000 vm001 vm002`).
|
|
|
|
All status checks talk to the RustFS admin API directly (SigV4-signed,
|
|
`jq` assertions), no `rc` required.
|
|
|
|
## What it does
|
|
|
|
1. Downloads the `.deb` package on all nodes (release tag or a direct URL such
|
|
as the nightly/R2 package).
|
|
2. Installs it, writes the 3x4 config
|
|
(`http://rustfs-node{1...3}:9000/data/rustfs{1...4}/mnmd`), starts all
|
|
three nodes simultaneously, verifies the cluster is up.
|
|
3. Writes data with `warp` while monitoring disk usage on the surviving nodes
|
|
(`df -B1G | grep /data/rustfs`):
|
|
- when both surviving nodes reach `STOP_NODE_AT_GB` (default 15 GiB), stop
|
|
the outage node (`vm002`, `OUTAGE_NODE_INDEX=2`);
|
|
- keep writing until both surviving nodes reach `WARP_STOP_AT_GB`
|
|
(default 40 GiB), then stop warp.
|
|
4. Restarts the outage node.
|
|
5. Starts cluster heal: `POST /rustfs/admin/v3/heal/` with body
|
|
`{"recursive":true}` (retried, returns a `clientToken`).
|
|
6. Monitors the heal task via `POST /rustfs/admin/v3/heal/?clientToken=<token>`
|
|
until the server verdict is a terminal success (`finished`/`completed`) with
|
|
`objects_failed == 0`.
|
|
7. Result analysis: heal stats (scanned/healed/failed), an **S3 read-back
|
|
verification** of the written objects (list the test bucket and GET a
|
|
sample — every read must succeed), per-node disk usage (observability),
|
|
pass/fail verdict.
|
|
|
|
Success is the server's own scan/repair verdict (heal finished, 0 failed)
|
|
**plus** an end-to-end data read-back; per-node disk usage is logged as
|
|
observability, not a pass gate (EC distributes different shards per node, so a
|
|
fixed per-node GB target is not a meaningful invariant).
|
|
|
|
## Self-hosted runner prerequisites
|
|
|
|
- Register the admin host (e.g. `heal`) as a runner with the
|
|
`smoke-testing` label.
|
|
- Install `jq`, `openssl`, `curl` and `warp` on the runner. `rc` is **not**
|
|
required.
|
|
- The runner user must be able to SSH to `vm000/vm001/vm002` without a
|
|
password prompt; nodes need passwordless `sudo` for the SSH user and
|
|
resolvable `rustfs-node*` hostnames.
|
|
- Admin API credentials need the `admin:server-info`, `admin:heal` and
|
|
`admin:rebalance` actions.
|
|
|
|
## Configuration
|
|
|
|
Same repository secrets/variables as the pool expansion workflow:
|
|
|
|
| Kind | Name | Purpose |
|
|
| ------ | --------------------- | ---------------------------------------------- |
|
|
| Secret | `RUSTFS_ACCESS_KEY` | RustFS access key (default `rustfs@test`) |
|
|
| Secret | `RUSTFS_SECRET_KEY` | RustFS secret key (default `rustfs@test`) |
|
|
| Var | `RUSTFS_API_ENDPOINT` | Admin API endpoint, e.g. `http://127.0.0.1:9000` (`RUSTFS_RC_ENDPOINT` fallback) |
|
|
| Var | `RUSTFS_NODES` | `vm000 vm001 vm002` |
|
|
| Var | `RUSTFS_SSH_USER` | `azureuser` |
|
|
| Var | `RUSTFS_NIGHTLY_PACKAGE_URL` | Default nightly deb URL (defaults to the R2 `latest` alias) |
|
|
|
|
## Workflow inputs
|
|
|
|
| Input | Default | Meaning |
|
|
| ---------------- | ------- | ----------------------------------------- |
|
|
| `package_url` | nightly | Direct `.deb` URL; empty = latest nightly |
|
|
| `stop_node_gb` | `15` | Stop outage node at N GiB on survivors |
|
|
| `warp_stop_gb` | `40` | Stop warp at N GiB on survivors |
|
|
| `cleanup_before` | `true` | Reset nodes before the test |
|
|
| `cleanup_after` | `true` | Reset nodes after the test |
|
|
|
|
> ⚠️ `--reset` purges the `rustfs` package and deletes the data directories on
|
|
> all nodes. Only run against a dedicated test environment.
|
|
|
|
## Manual usage
|
|
|
|
```bash
|
|
./scripts/test/rustfs_heal_test.sh --all -y \
|
|
--package-url https://dl.rustfs.com/artifacts/rustfs/packages/nightly/rustfs-nightly-latest.deb \
|
|
--endpoint http://127.0.0.1:9000
|
|
|
|
./scripts/test/rustfs_heal_test.sh --steps 5,6,7
|
|
./scripts/test/rustfs_heal_test.sh --reset -y
|
|
```
|
|
|
|
## Known issues
|
|
|
|
- Nightly builds gate pool/rebalance activation on a live fleet capability
|
|
proof (rustfs/backlog#2031); the script retries heal/rebalance starts and
|
|
prints a hint when the signature appears.
|
|
- The cluster-level `GET /rustfs/admin/v3/background-heal/status` aggregator
|
|
returns 501 in the single-pool 3x4 topology (no notification system), so the
|
|
script monitors the started heal task via its `clientToken` instead.
|
|
- The heal task may report `progress: null` while running; the script logs
|
|
this as evidence (rustfs/backlog#2035) rather than coercing it to zero, and
|
|
reads the canonical camelCase progress fields
|
|
(`objectsScanned`/`objectsHealed`/`objectsFailed`/`progressPercentage`) with
|
|
a snake_case fallback.
|
|
- The server-side per-task heal timeout defaults to 5 minutes; the script
|
|
writes `RUSTFS_HEAL_TASK_TIMEOUT_SECS=21600` (6h) into the node config so a
|
|
multi-tens-of-GiB heal can finish. The background scanner is disabled
|
|
(`RUSTFS_HEAL_AUTO_HEAL_ENABLE=false`) so the explicit heal is the only
|
|
repair mechanism and the outage effect stays observable.
|