test(heal): relative disk target and fail fast on terminal-but-short (#6748)

* test(heal): relative disk target and fail fast on terminal-but-short

The absolute 40 GiB heal target was calibrated to the background scanner
(auto-heal), which is now disabled for determinism; with only the explicit
heal the recovered node lands at ~36 GiB for 40 GiB survivors. Make the
success criterion relative: the outage node must reach at least 90% of the
least-used surviving node (absolute HEAL_TARGET_GB floor optional, default
0 = relative only).

Also fail fast when the heal task reaches a terminal success but the disk
target is not met (previously the monitor kept polling until timeout), and
drop the misleading 'progress absent' warning on the final (cleaned) task
response — mid-run progress is reported correctly.

Validated live: heal summary=finished, 0 failed, vm000/vm001=40GB,
vm002=40GB (target 36GB), test PASSED.

* test(heal): gate success on server verdict + data read-back, drop disk GB gate

The per-node disk-usage target (40 GiB / 90% of survivors) is not a
code-level invariant: EC distributes different shards per node, so the
final GB per node depends on the layout, not on heal correctness. Gate the
test on what the server actually verifies:

- Heal task terminal success (finished/completed) with objectsFailed == 0
  (the server's per-object scan/repair verdict).
- S3 read-back verification: list the test bucket and GET a sample of
  objects, requiring HTTP 200 for every read (end-to-end proof the data is
  still reconstructable after repair). The GET uses a discard mode so
  binary bodies are not captured (no null-byte warnings / SIGPIPE).

Per-node disk usage stays in the output as observability (with a warning if
the outage node gained no usage), not as the pass/fail gate. Removes the
heal_target_gb input and the relative-target logic.

Validated live: heal summary=finished, 0 failed, 20/20 objects read back,
vm002_used=40GB, PASS.
This commit is contained in:
hector
2026-08-27 22:25:17 +08:00
committed by GitHub
parent d48dda5bdc
commit 2e6c820f53
4 changed files with 75 additions and 38 deletions
+9 -7
View File
@@ -25,14 +25,17 @@ All status checks talk to the RustFS admin API directly (SigV4-signed,
5. Starts cluster heal: `POST /rustfs/admin/v3/heal/` with body
`{"recursive":true}` (retried, returns a `clientToken`).
6. Monitors the heal task via `POST /rustfs/admin/v3/heal/?clientToken=<token>`
until the summary is a terminal success (`finished`/`completed`),
`objects_failed == 0`, **and** the outage node's disk usage reaches
`HEAL_TARGET_GB` (default 40 GiB).
7. Result analysis: heal stats (scanned/healed/failed), per-node disk usage,
until the server verdict is a terminal success (`finished`/`completed`) with
`objects_failed == 0`.
7. Result analysis: heal stats (scanned/healed/failed), an **S3 read-back
verification** of the written objects (list the test bucket and GET a
sample — every read must succeed), per-node disk usage (observability),
pass/fail verdict.
Success requires **both** the heal API completion (the server's scan/repair
verdict) and the outage node's disk reaching the target.
Success is the server's own scan/repair verdict (heal finished, 0 failed)
**plus** an end-to-end data read-back; per-node disk usage is logged as
observability, not a pass gate (EC distributes different shards per node, so a
fixed per-node GB target is not a meaningful invariant).
## Self-hosted runner prerequisites
@@ -66,7 +69,6 @@ Same repository secrets/variables as the pool expansion workflow:
| `package_url` | nightly | Direct `.deb` URL; empty = latest nightly |
| `stop_node_gb` | `15` | Stop outage node at N GiB on survivors |
| `warp_stop_gb` | `40` | Stop warp at N GiB on survivors |
| `heal_target_gb` | `40` | Outage node must reach N GiB after heal |
| `cleanup_before` | `true` | Reset nodes before the test |
| `cleanup_after` | `true` | Reset nodes after the test |