* test(heal): relative disk target and fail fast on terminal-but-short
The absolute 40 GiB heal target was calibrated to the background scanner
(auto-heal), which is now disabled for determinism; with only the explicit
heal the recovered node lands at ~36 GiB for 40 GiB survivors. Make the
success criterion relative: the outage node must reach at least 90% of the
least-used surviving node (absolute HEAL_TARGET_GB floor optional, default
0 = relative only).
Also fail fast when the heal task reaches a terminal success but the disk
target is not met (previously the monitor kept polling until timeout), and
drop the misleading 'progress absent' warning on the final (cleaned) task
response — mid-run progress is reported correctly.
Validated live: heal summary=finished, 0 failed, vm000/vm001=40GB,
vm002=40GB (target 36GB), test PASSED.
* test(heal): gate success on server verdict + data read-back, drop disk GB gate
The per-node disk-usage target (40 GiB / 90% of survivors) is not a
code-level invariant: EC distributes different shards per node, so the
final GB per node depends on the layout, not on heal correctness. Gate the
test on what the server actually verifies:
- Heal task terminal success (finished/completed) with objectsFailed == 0
(the server's per-object scan/repair verdict).
- S3 read-back verification: list the test bucket and GET a sample of
objects, requiring HTTP 200 for every read (end-to-end proof the data is
still reconstructable after repair). The GET uses a discard mode so
binary bodies are not captured (no null-byte warnings / SIGPIPE).
Per-node disk usage stays in the output as observability (with a warning if
the outage node gained no usage), not as the pass/fail gate. Removes the
heal_target_gb input and the relative-target logic.
Validated live: heal summary=finished, 0 failed, 20/20 objects read back,
vm002_used=40GB, PASS.
* test(heal): add node-outage heal E2E script and workflow
RustFS heal test on the 3x4 cluster (3 nodes x 4 disks, same
RUSTFS_VOLUMES expression on every node): write data with warp, stop the
outage node mid-write, restart it, start cluster heal via the admin API,
and pass only when the heal task finishes with 0 failures AND the outage
node's disk usage reaches the target.
Includes the GitHub Actions workflow (smoke-testing runner, nightly deb by
default) and a README. Validated end-to-end on the test environment:
40/40/16 GiB before heal -> 40/40/40 GiB after heal, summary=finished.
The script also writes RUSTFS_HEAL_TASK_TIMEOUT_SECS (default 6h) into the
node config because the server default (5 min) is far too short for
healing tens of GiB.
* ci(pool-test): chain heal regression after the pool test
The pool-expansion workflow is now triggered by the Nightly GNU Build
(workflow_run, replacing the schedule) and runs two sequential jobs on the
shared test environment:
1. pool-expansion-test (existing) — skipped if the nightly build failed.
2. heal-test — runs after the pool test regardless of its outcome
(if: always()): a pool failure makes the run red but does not block the
heal regression. Runs the heal script (reset -> install/start 3x4 ->
write/outage -> heal -> verify -> reset).
* test(heal): address review — camelCase progress, fail-closed, workflow hygiene
- Heal progress fields are camelCase in the API (objectsScanned/objectsHealed/
objectsFailed/progressPercentage); read them with a snake_case fallback and
distinguish null (absent) progress from zero, logging null as evidence
(rustfs/backlog#2035) instead of silently coercing.
- Fail closed in step 3: the outage node must actually be inactive after stop,
the write target must be reached, and an unobserved outage or incomplete
write fails the test instead of warning.
- Step 4 waits (bounded) for the cluster to report an active pool after the
outage-node restart instead of swallowing the verification error.
- Heal start fails fast on 400/403 (deterministic request/auth problems) and
only retries transient server errors.
- Disable the background scanner (RUSTFS_HEAL_AUTO_HEAL_ENABLE=false) so the
explicit heal is the only repair mechanism and the outage is observable.
- Workflows: heal and pool share one concurrency group; workflow_run requires
an exact successful nightly conclusion; checkout is pinned to the triggering
SHA; comma-separated step args are quoted (actionlint SC2054).