mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-29 00:17:11 +00:00
test(heal): relative disk target and fail fast on terminal-but-short (#6748)
* test(heal): relative disk target and fail fast on terminal-but-short The absolute 40 GiB heal target was calibrated to the background scanner (auto-heal), which is now disabled for determinism; with only the explicit heal the recovered node lands at ~36 GiB for 40 GiB survivors. Make the success criterion relative: the outage node must reach at least 90% of the least-used surviving node (absolute HEAL_TARGET_GB floor optional, default 0 = relative only). Also fail fast when the heal task reaches a terminal success but the disk target is not met (previously the monitor kept polling until timeout), and drop the misleading 'progress absent' warning on the final (cleaned) task response — mid-run progress is reported correctly. Validated live: heal summary=finished, 0 failed, vm000/vm001=40GB, vm002=40GB (target 36GB), test PASSED. * test(heal): gate success on server verdict + data read-back, drop disk GB gate The per-node disk-usage target (40 GiB / 90% of survivors) is not a code-level invariant: EC distributes different shards per node, so the final GB per node depends on the layout, not on heal correctness. Gate the test on what the server actually verifies: - Heal task terminal success (finished/completed) with objectsFailed == 0 (the server's per-object scan/repair verdict). - S3 read-back verification: list the test bucket and GET a sample of objects, requiring HTTP 200 for every read (end-to-end proof the data is still reconstructable after repair). The GET uses a discard mode so binary bodies are not captured (no null-byte warnings / SIGPIPE). Per-node disk usage stays in the output as observability (with a warning if the outage node gained no usage), not as the pass/fail gate. Removes the heal_target_gb input and the relative-target logic. Validated live: heal summary=finished, 0 failed, 20/20 objects read back, vm002_used=40GB, PASS.
This commit is contained in:
@@ -25,14 +25,17 @@ All status checks talk to the RustFS admin API directly (SigV4-signed,
|
||||
5. Starts cluster heal: `POST /rustfs/admin/v3/heal/` with body
|
||||
`{"recursive":true}` (retried, returns a `clientToken`).
|
||||
6. Monitors the heal task via `POST /rustfs/admin/v3/heal/?clientToken=<token>`
|
||||
until the summary is a terminal success (`finished`/`completed`),
|
||||
`objects_failed == 0`, **and** the outage node's disk usage reaches
|
||||
`HEAL_TARGET_GB` (default 40 GiB).
|
||||
7. Result analysis: heal stats (scanned/healed/failed), per-node disk usage,
|
||||
until the server verdict is a terminal success (`finished`/`completed`) with
|
||||
`objects_failed == 0`.
|
||||
7. Result analysis: heal stats (scanned/healed/failed), an **S3 read-back
|
||||
verification** of the written objects (list the test bucket and GET a
|
||||
sample — every read must succeed), per-node disk usage (observability),
|
||||
pass/fail verdict.
|
||||
|
||||
Success requires **both** the heal API completion (the server's scan/repair
|
||||
verdict) and the outage node's disk reaching the target.
|
||||
Success is the server's own scan/repair verdict (heal finished, 0 failed)
|
||||
**plus** an end-to-end data read-back; per-node disk usage is logged as
|
||||
observability, not a pass gate (EC distributes different shards per node, so a
|
||||
fixed per-node GB target is not a meaningful invariant).
|
||||
|
||||
## Self-hosted runner prerequisites
|
||||
|
||||
@@ -66,7 +69,6 @@ Same repository secrets/variables as the pool expansion workflow:
|
||||
| `package_url` | nightly | Direct `.deb` URL; empty = latest nightly |
|
||||
| `stop_node_gb` | `15` | Stop outage node at N GiB on survivors |
|
||||
| `warp_stop_gb` | `40` | Stop warp at N GiB on survivors |
|
||||
| `heal_target_gb` | `40` | Outage node must reach N GiB after heal |
|
||||
| `cleanup_before` | `true` | Reset nodes before the test |
|
||||
| `cleanup_after` | `true` | Reset nodes after the test |
|
||||
|
||||
|
||||
Reference in New Issue
Block a user