mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-29 16:37:07 +00:00
test(heal): add node-outage heal E2E script and workflow (#6733)
* test(heal): add node-outage heal E2E script and workflow RustFS heal test on the 3x4 cluster (3 nodes x 4 disks, same RUSTFS_VOLUMES expression on every node): write data with warp, stop the outage node mid-write, restart it, start cluster heal via the admin API, and pass only when the heal task finishes with 0 failures AND the outage node's disk usage reaches the target. Includes the GitHub Actions workflow (smoke-testing runner, nightly deb by default) and a README. Validated end-to-end on the test environment: 40/40/16 GiB before heal -> 40/40/40 GiB after heal, summary=finished. The script also writes RUSTFS_HEAL_TASK_TIMEOUT_SECS (default 6h) into the node config because the server default (5 min) is far too short for healing tens of GiB. * ci(pool-test): chain heal regression after the pool test The pool-expansion workflow is now triggered by the Nightly GNU Build (workflow_run, replacing the schedule) and runs two sequential jobs on the shared test environment: 1. pool-expansion-test (existing) — skipped if the nightly build failed. 2. heal-test — runs after the pool test regardless of its outcome (if: always()): a pool failure makes the run red but does not block the heal regression. Runs the heal script (reset -> install/start 3x4 -> write/outage -> heal -> verify -> reset). * test(heal): address review — camelCase progress, fail-closed, workflow hygiene - Heal progress fields are camelCase in the API (objectsScanned/objectsHealed/ objectsFailed/progressPercentage); read them with a snake_case fallback and distinguish null (absent) progress from zero, logging null as evidence (rustfs/backlog#2035) instead of silently coercing. - Fail closed in step 3: the outage node must actually be inactive after stop, the write target must be reached, and an unobserved outage or incomplete write fails the test instead of warning. - Step 4 waits (bounded) for the cluster to report an active pool after the outage-node restart instead of swallowing the verification error. - Heal start fails fast on 400/403 (deterministic request/auth problems) and only retries transient server errors. - Disable the background scanner (RUSTFS_HEAL_AUTO_HEAL_ENABLE=false) so the explicit heal is the only repair mechanism and the outage is observable. - Workflows: heal and pool share one concurrency group; workflow_run requires an exact successful nightly conclusion; checkout is pinned to the triggering SHA; comma-separated step args are quoted (actionlint SC2054).
This commit is contained in:
@@ -0,0 +1,126 @@
|
|||||||
|
name: RustFS Heal Test
|
||||||
|
|
||||||
|
on:
|
||||||
|
workflow_dispatch:
|
||||||
|
inputs:
|
||||||
|
package_url:
|
||||||
|
description: 'Direct .deb URL (nightly/R2). Defaults to the latest nightly deb.'
|
||||||
|
required: false
|
||||||
|
type: string
|
||||||
|
stop_node_gb:
|
||||||
|
description: 'Stop the outage node when surviving nodes reach N GiB'
|
||||||
|
required: false
|
||||||
|
default: '15'
|
||||||
|
warp_stop_gb:
|
||||||
|
description: 'Stop warp when surviving nodes reach N GiB'
|
||||||
|
required: false
|
||||||
|
default: '40'
|
||||||
|
heal_target_gb:
|
||||||
|
description: 'Outage node must reach N GiB after heal to pass'
|
||||||
|
required: false
|
||||||
|
default: '40'
|
||||||
|
cleanup_before:
|
||||||
|
description: 'Reset the nodes before the test (DESTROYS existing data/config)'
|
||||||
|
type: boolean
|
||||||
|
default: true
|
||||||
|
cleanup_after:
|
||||||
|
description: 'Reset the nodes after the test (DESTROYS test data/config)'
|
||||||
|
type: boolean
|
||||||
|
default: true
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
# Only one test at a time: both this and the pool-expansion workflow mutate
|
||||||
|
# the same test environment, so they share one concurrency group.
|
||||||
|
concurrency:
|
||||||
|
group: rustfs-pool-expansion-test
|
||||||
|
cancel-in-progress: false
|
||||||
|
|
||||||
|
defaults:
|
||||||
|
run:
|
||||||
|
shell: bash
|
||||||
|
|
||||||
|
env:
|
||||||
|
RUSTFS_ACCESS_KEY: ${{ secrets.RUSTFS_ACCESS_KEY }}
|
||||||
|
RUSTFS_SECRET_KEY: ${{ secrets.RUSTFS_SECRET_KEY }}
|
||||||
|
RUSTFS_API_ENDPOINT: ${{ secrets.RUSTFS_API_ENDPOINT || vars.RUSTFS_API_ENDPOINT || vars.RUSTFS_RC_ENDPOINT }}
|
||||||
|
RUSTFS_NODES: ${{ secrets.RUSTFS_NODES || vars.RUSTFS_NODES }}
|
||||||
|
RUSTFS_SSH_USER: ${{ secrets.RUSTFS_SSH_USER || vars.RUSTFS_SSH_USER }}
|
||||||
|
RUSTFS_NIGHTLY_PACKAGE_URL: ${{ vars.RUSTFS_NIGHTLY_PACKAGE_URL || 'https://dl.rustfs.com/artifacts/rustfs/packages/nightly/rustfs-nightly-latest.deb' }}
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
heal-test:
|
||||||
|
runs-on: smoke-testing
|
||||||
|
timeout-minutes: 480
|
||||||
|
steps:
|
||||||
|
- name: Checkout
|
||||||
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- name: Show environment
|
||||||
|
run: |
|
||||||
|
uname -a
|
||||||
|
jq --version
|
||||||
|
openssl version
|
||||||
|
warp --version || true
|
||||||
|
df -h /data | tail -1
|
||||||
|
|
||||||
|
- name: Reset test environment (before)
|
||||||
|
if: ${{ inputs.cleanup_before != 'false' }}
|
||||||
|
run: |
|
||||||
|
chmod +x scripts/test/rustfs_heal_test.sh
|
||||||
|
./scripts/test/rustfs_heal_test.sh --reset -y
|
||||||
|
|
||||||
|
- name: Install RustFS package & start cluster
|
||||||
|
run: |
|
||||||
|
ARGS=(--steps "1,2" -y --endpoint "${{ env.RUSTFS_API_ENDPOINT }}")
|
||||||
|
if [ -n "${{ inputs.package_url }}" ]; then
|
||||||
|
ARGS+=(--package-url "${{ inputs.package_url }}")
|
||||||
|
else
|
||||||
|
ARGS+=(--package-url "${{ env.RUSTFS_NIGHTLY_PACKAGE_URL }}")
|
||||||
|
fi
|
||||||
|
./scripts/test/rustfs_heal_test.sh "${ARGS[@]}"
|
||||||
|
|
||||||
|
- name: Preflight checks
|
||||||
|
run: |
|
||||||
|
ARGS=(--preflight --endpoint "${{ env.RUSTFS_API_ENDPOINT }}")
|
||||||
|
if [ -n "${{ inputs.package_url }}" ]; then
|
||||||
|
ARGS+=(--package-url "${{ inputs.package_url }}")
|
||||||
|
else
|
||||||
|
ARGS+=(--package-url "${{ env.RUSTFS_NIGHTLY_PACKAGE_URL }}")
|
||||||
|
fi
|
||||||
|
./scripts/test/rustfs_heal_test.sh "${ARGS[@]}"
|
||||||
|
|
||||||
|
- name: Run heal test (write -> outage -> heal -> verify)
|
||||||
|
run: |
|
||||||
|
./scripts/test/rustfs_heal_test.sh \
|
||||||
|
--steps "3,4,5,6,7" -y \
|
||||||
|
--endpoint "${{ env.RUSTFS_API_ENDPOINT }}" \
|
||||||
|
--stop-node-gb "${{ inputs.stop_node_gb }}" \
|
||||||
|
--warp-stop-gb "${{ inputs.warp_stop_gb }}" \
|
||||||
|
--heal-target-gb "${{ inputs.heal_target_gb }}" \
|
||||||
|
--log-file /tmp/rustfs-heal-test.log
|
||||||
|
|
||||||
|
- name: Upload test logs
|
||||||
|
if: always()
|
||||||
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
||||||
|
with:
|
||||||
|
name: rustfs-heal-test-${{ github.run_id }}
|
||||||
|
path: |
|
||||||
|
/tmp/rustfs-heal-test.log
|
||||||
|
/tmp/rustfs-warp.*.log
|
||||||
|
if-no-files-found: warn
|
||||||
|
|
||||||
|
- name: Reset test environment (after)
|
||||||
|
if: ${{ always() && inputs.cleanup_after != 'false' }}
|
||||||
|
run: |
|
||||||
|
./scripts/test/rustfs_heal_test.sh --reset -y
|
||||||
|
|
||||||
|
- name: Notify on failure
|
||||||
|
if: failure()
|
||||||
|
run: |
|
||||||
|
echo "RustFS heal test failed"
|
||||||
|
echo "Package source: ${{ inputs.package_url || 'nightly (R2 latest)' }}"
|
||||||
|
echo "See the uploaded log artifact for details."
|
||||||
@@ -30,6 +30,18 @@ on:
|
|||||||
description: 'Run the pool decommission step (3-pool topology only)'
|
description: 'Run the pool decommission step (3-pool topology only)'
|
||||||
type: boolean
|
type: boolean
|
||||||
default: true
|
default: true
|
||||||
|
stop_node_gb:
|
||||||
|
description: 'Heal: stop the outage node when surviving nodes reach N GiB'
|
||||||
|
required: false
|
||||||
|
default: '15'
|
||||||
|
warp_stop_gb:
|
||||||
|
description: 'Heal: stop warp when surviving nodes reach N GiB'
|
||||||
|
required: false
|
||||||
|
default: '40'
|
||||||
|
heal_target_gb:
|
||||||
|
description: 'Heal: outage node must reach N GiB after heal'
|
||||||
|
required: false
|
||||||
|
default: '40'
|
||||||
cleanup_before:
|
cleanup_before:
|
||||||
description: 'Reset the nodes before the test (DESTROYS existing data/config)'
|
description: 'Reset the nodes before the test (DESTROYS existing data/config)'
|
||||||
type: boolean
|
type: boolean
|
||||||
@@ -38,9 +50,10 @@ on:
|
|||||||
description: 'Reset the nodes after the test (DESTROYS test data/config)'
|
description: 'Reset the nodes after the test (DESTROYS test data/config)'
|
||||||
type: boolean
|
type: boolean
|
||||||
default: true
|
default: true
|
||||||
schedule:
|
workflow_run:
|
||||||
# Nightly regression run; remove if you do not want a schedule.
|
# Run after the nightly build completes: pool expansion first, then heal.
|
||||||
- cron: '0 21 * * *'
|
workflows: ["Nightly GNU Build"]
|
||||||
|
types: [completed]
|
||||||
|
|
||||||
permissions:
|
permissions:
|
||||||
contents: read
|
contents: read
|
||||||
@@ -61,19 +74,23 @@ env:
|
|||||||
RUSTFS_API_ENDPOINT: ${{ secrets.RUSTFS_API_ENDPOINT || vars.RUSTFS_API_ENDPOINT || vars.RUSTFS_RC_ENDPOINT }}
|
RUSTFS_API_ENDPOINT: ${{ secrets.RUSTFS_API_ENDPOINT || vars.RUSTFS_API_ENDPOINT || vars.RUSTFS_RC_ENDPOINT }}
|
||||||
RUSTFS_NODES: ${{ secrets.RUSTFS_NODES || vars.RUSTFS_NODES }}
|
RUSTFS_NODES: ${{ secrets.RUSTFS_NODES || vars.RUSTFS_NODES }}
|
||||||
RUSTFS_SSH_USER: ${{ secrets.RUSTFS_SSH_USER || vars.RUSTFS_SSH_USER }}
|
RUSTFS_SSH_USER: ${{ secrets.RUSTFS_SSH_USER || vars.RUSTFS_SSH_USER }}
|
||||||
# Package used by the scheduled run (workflow_dispatch inputs are empty for
|
# Package used by the nightly run (workflow_dispatch inputs are empty for
|
||||||
# schedule events), i.e. the latest nightly deb published by nightly-gnu.yml.
|
# workflow_run events), i.e. the latest nightly deb published by nightly-gnu.yml.
|
||||||
RUSTFS_NIGHTLY_PACKAGE_URL: ${{ vars.RUSTFS_NIGHTLY_PACKAGE_URL || 'https://dl.rustfs.com/artifacts/rustfs/packages/nightly/rustfs-nightly-latest.deb' }}
|
RUSTFS_NIGHTLY_PACKAGE_URL: ${{ vars.RUSTFS_NIGHTLY_PACKAGE_URL || 'https://dl.rustfs.com/artifacts/rustfs/packages/nightly/rustfs-nightly-latest.deb' }}
|
||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
pool-expansion-test:
|
pool-expansion-test:
|
||||||
runs-on: smoke-testing
|
runs-on: smoke-testing
|
||||||
timeout-minutes: 360
|
timeout-minutes: 360
|
||||||
|
# Run on manual dispatch, or when the nightly build completed successfully
|
||||||
|
# (its deb is what the tests install). Skipped when nightly failed.
|
||||||
|
if: ${{ github.event_name == 'workflow_dispatch' || github.event.workflow_run.conclusion == 'success' }}
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout
|
- name: Checkout
|
||||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||||
with:
|
with:
|
||||||
persist-credentials: false
|
persist-credentials: false
|
||||||
|
ref: ${{ github.event.workflow_run.head_sha || github.ref }}
|
||||||
|
|
||||||
- name: Show environment
|
- name: Show environment
|
||||||
run: |
|
run: |
|
||||||
@@ -152,3 +169,75 @@ jobs:
|
|||||||
echo "RustFS pool expansion test failed"
|
echo "RustFS pool expansion test failed"
|
||||||
echo "Package source: ${{ inputs.package_url || inputs.rustfs_version || 'nightly (R2 latest)' }}"
|
echo "Package source: ${{ inputs.package_url || inputs.rustfs_version || 'nightly (R2 latest)' }}"
|
||||||
echo "See the uploaded log artifact for details."
|
echo "See the uploaded log artifact for details."
|
||||||
|
|
||||||
|
# Heal regression runs after the pool test regardless of its outcome: a pool
|
||||||
|
# failure must be reported (it makes the run red) but must not block heal.
|
||||||
|
heal-test:
|
||||||
|
name: Heal test (after pool test)
|
||||||
|
runs-on: smoke-testing
|
||||||
|
timeout-minutes: 480
|
||||||
|
needs: pool-expansion-test
|
||||||
|
if: ${{ always() && (github.event_name == 'workflow_dispatch' || github.event.workflow_run.conclusion == 'success') }}
|
||||||
|
steps:
|
||||||
|
- name: Checkout
|
||||||
|
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
ref: ${{ github.event.workflow_run.head_sha || github.ref }}
|
||||||
|
|
||||||
|
- name: Reset test environment (before)
|
||||||
|
if: ${{ inputs.cleanup_before != 'false' }}
|
||||||
|
run: |
|
||||||
|
chmod +x scripts/test/rustfs_heal_test.sh
|
||||||
|
./scripts/test/rustfs_heal_test.sh --reset -y
|
||||||
|
|
||||||
|
- name: Install RustFS package & start cluster
|
||||||
|
run: |
|
||||||
|
ARGS=(--steps "1,2" -y --endpoint "${{ env.RUSTFS_API_ENDPOINT }}")
|
||||||
|
if [ -n "${{ inputs.package_url }}" ]; then
|
||||||
|
ARGS+=(--package-url "${{ inputs.package_url }}")
|
||||||
|
else
|
||||||
|
ARGS+=(--package-url "${{ env.RUSTFS_NIGHTLY_PACKAGE_URL }}")
|
||||||
|
fi
|
||||||
|
./scripts/test/rustfs_heal_test.sh "${ARGS[@]}"
|
||||||
|
|
||||||
|
- name: Preflight checks
|
||||||
|
run: |
|
||||||
|
ARGS=(--preflight --endpoint "${{ env.RUSTFS_API_ENDPOINT }}")
|
||||||
|
if [ -n "${{ inputs.package_url }}" ]; then
|
||||||
|
ARGS+=(--package-url "${{ inputs.package_url }}")
|
||||||
|
else
|
||||||
|
ARGS+=(--package-url "${{ env.RUSTFS_NIGHTLY_PACKAGE_URL }}")
|
||||||
|
fi
|
||||||
|
./scripts/test/rustfs_heal_test.sh "${ARGS[@]}"
|
||||||
|
|
||||||
|
- name: Run heal test (write -> outage -> heal -> verify)
|
||||||
|
run: |
|
||||||
|
./scripts/test/rustfs_heal_test.sh \
|
||||||
|
--steps 3,4,5,6,7 -y \
|
||||||
|
--endpoint "${{ env.RUSTFS_API_ENDPOINT }}" \
|
||||||
|
--stop-node-gb "${{ inputs.stop_node_gb || '15' }}" \
|
||||||
|
--warp-stop-gb "${{ inputs.warp_stop_gb || '40' }}" \
|
||||||
|
--heal-target-gb "${{ inputs.heal_target_gb || '40' }}" \
|
||||||
|
--log-file /tmp/rustfs-heal-test.log
|
||||||
|
|
||||||
|
- name: Upload test logs
|
||||||
|
if: always()
|
||||||
|
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
|
||||||
|
with:
|
||||||
|
name: rustfs-heal-test-${{ github.run_id }}
|
||||||
|
path: |
|
||||||
|
/tmp/rustfs-heal-test.log
|
||||||
|
/tmp/rustfs-warp.*.log
|
||||||
|
if-no-files-found: warn
|
||||||
|
|
||||||
|
- name: Reset test environment (after)
|
||||||
|
if: ${{ always() && inputs.cleanup_after != 'false' }}
|
||||||
|
run: |
|
||||||
|
./scripts/test/rustfs_heal_test.sh --reset -y
|
||||||
|
|
||||||
|
- name: Notify on failure
|
||||||
|
if: failure()
|
||||||
|
run: |
|
||||||
|
echo "RustFS heal test failed"
|
||||||
|
echo "See the uploaded log artifact for details."
|
||||||
|
|||||||
@@ -0,0 +1,104 @@
|
|||||||
|
# RustFS Heal Test
|
||||||
|
|
||||||
|
Node-outage heal test driven by
|
||||||
|
[`scripts/test/rustfs_heal_test.sh`](rustfs_heal_test.sh), based on the
|
||||||
|
Obsidian note "RustFS Heal 测试步骤". Uses the same 3-node test environment as
|
||||||
|
the pool expansion test (`vm000 vm001 vm002`).
|
||||||
|
|
||||||
|
All status checks talk to the RustFS admin API directly (SigV4-signed,
|
||||||
|
`jq` assertions), no `rc` required.
|
||||||
|
|
||||||
|
## What it does
|
||||||
|
|
||||||
|
1. Downloads the `.deb` package on all nodes (release tag or a direct URL such
|
||||||
|
as the nightly/R2 package).
|
||||||
|
2. Installs it, writes the 3x4 config
|
||||||
|
(`http://rustfs-node{1...3}:9000/data/rustfs{1...4}/mnmd`), starts all
|
||||||
|
three nodes simultaneously, verifies the cluster is up.
|
||||||
|
3. Writes data with `warp` while monitoring disk usage on the surviving nodes
|
||||||
|
(`df -B1G | grep /data/rustfs`):
|
||||||
|
- when both surviving nodes reach `STOP_NODE_AT_GB` (default 15 GiB), stop
|
||||||
|
the outage node (`vm002`, `OUTAGE_NODE_INDEX=2`);
|
||||||
|
- keep writing until both surviving nodes reach `WARP_STOP_AT_GB`
|
||||||
|
(default 40 GiB), then stop warp.
|
||||||
|
4. Restarts the outage node.
|
||||||
|
5. Starts cluster heal: `POST /rustfs/admin/v3/heal/` with body
|
||||||
|
`{"recursive":true}` (retried, returns a `clientToken`).
|
||||||
|
6. Monitors the heal task via `POST /rustfs/admin/v3/heal/?clientToken=<token>`
|
||||||
|
until the summary is a terminal success (`finished`/`completed`),
|
||||||
|
`objects_failed == 0`, **and** the outage node's disk usage reaches
|
||||||
|
`HEAL_TARGET_GB` (default 40 GiB).
|
||||||
|
7. Result analysis: heal stats (scanned/healed/failed), per-node disk usage,
|
||||||
|
pass/fail verdict.
|
||||||
|
|
||||||
|
Success requires **both** the heal API completion (the server's scan/repair
|
||||||
|
verdict) and the outage node's disk reaching the target.
|
||||||
|
|
||||||
|
## Self-hosted runner prerequisites
|
||||||
|
|
||||||
|
- Register the admin host (e.g. `heal`) as a runner with the
|
||||||
|
`smoke-testing` label.
|
||||||
|
- Install `jq`, `openssl`, `curl` and `warp` on the runner. `rc` is **not**
|
||||||
|
required.
|
||||||
|
- The runner user must be able to SSH to `vm000/vm001/vm002` without a
|
||||||
|
password prompt; nodes need passwordless `sudo` for the SSH user and
|
||||||
|
resolvable `rustfs-node*` hostnames.
|
||||||
|
- Admin API credentials need the `admin:server-info`, `admin:heal` and
|
||||||
|
`admin:rebalance` actions.
|
||||||
|
|
||||||
|
## Configuration
|
||||||
|
|
||||||
|
Same repository secrets/variables as the pool expansion workflow:
|
||||||
|
|
||||||
|
| Kind | Name | Purpose |
|
||||||
|
| ------ | --------------------- | ---------------------------------------------- |
|
||||||
|
| Secret | `RUSTFS_ACCESS_KEY` | RustFS access key (default `rustfs@test`) |
|
||||||
|
| Secret | `RUSTFS_SECRET_KEY` | RustFS secret key (default `rustfs@test`) |
|
||||||
|
| Var | `RUSTFS_API_ENDPOINT` | Admin API endpoint, e.g. `http://127.0.0.1:9000` (`RUSTFS_RC_ENDPOINT` fallback) |
|
||||||
|
| Var | `RUSTFS_NODES` | `vm000 vm001 vm002` |
|
||||||
|
| Var | `RUSTFS_SSH_USER` | `azureuser` |
|
||||||
|
| Var | `RUSTFS_NIGHTLY_PACKAGE_URL` | Default nightly deb URL (defaults to the R2 `latest` alias) |
|
||||||
|
|
||||||
|
## Workflow inputs
|
||||||
|
|
||||||
|
| Input | Default | Meaning |
|
||||||
|
| ---------------- | ------- | ----------------------------------------- |
|
||||||
|
| `package_url` | nightly | Direct `.deb` URL; empty = latest nightly |
|
||||||
|
| `stop_node_gb` | `15` | Stop outage node at N GiB on survivors |
|
||||||
|
| `warp_stop_gb` | `40` | Stop warp at N GiB on survivors |
|
||||||
|
| `heal_target_gb` | `40` | Outage node must reach N GiB after heal |
|
||||||
|
| `cleanup_before` | `true` | Reset nodes before the test |
|
||||||
|
| `cleanup_after` | `true` | Reset nodes after the test |
|
||||||
|
|
||||||
|
> ⚠️ `--reset` purges the `rustfs` package and deletes the data directories on
|
||||||
|
> all nodes. Only run against a dedicated test environment.
|
||||||
|
|
||||||
|
## Manual usage
|
||||||
|
|
||||||
|
```bash
|
||||||
|
./scripts/test/rustfs_heal_test.sh --all -y \
|
||||||
|
--package-url https://dl.rustfs.com/artifacts/rustfs/packages/nightly/rustfs-nightly-latest.deb \
|
||||||
|
--endpoint http://127.0.0.1:9000
|
||||||
|
|
||||||
|
./scripts/test/rustfs_heal_test.sh --steps 5,6,7
|
||||||
|
./scripts/test/rustfs_heal_test.sh --reset -y
|
||||||
|
```
|
||||||
|
|
||||||
|
## Known issues
|
||||||
|
|
||||||
|
- Nightly builds gate pool/rebalance activation on a live fleet capability
|
||||||
|
proof (rustfs/backlog#2031); the script retries heal/rebalance starts and
|
||||||
|
prints a hint when the signature appears.
|
||||||
|
- The cluster-level `GET /rustfs/admin/v3/background-heal/status` aggregator
|
||||||
|
returns 501 in the single-pool 3x4 topology (no notification system), so the
|
||||||
|
script monitors the started heal task via its `clientToken` instead.
|
||||||
|
- The heal task may report `progress: null` while running; the script logs
|
||||||
|
this as evidence (rustfs/backlog#2035) rather than coercing it to zero, and
|
||||||
|
reads the canonical camelCase progress fields
|
||||||
|
(`objectsScanned`/`objectsHealed`/`objectsFailed`/`progressPercentage`) with
|
||||||
|
a snake_case fallback.
|
||||||
|
- The server-side per-task heal timeout defaults to 5 minutes; the script
|
||||||
|
writes `RUSTFS_HEAL_TASK_TIMEOUT_SECS=21600` (6h) into the node config so a
|
||||||
|
multi-tens-of-GiB heal can finish. The background scanner is disabled
|
||||||
|
(`RUSTFS_HEAL_AUTO_HEAL_ENABLE=false`) so the explicit heal is the only
|
||||||
|
repair mechanism and the outage effect stays observable.
|
||||||
Executable
+1103
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user