Files
rustfs/docs/operations/internode-grpc-benchmark-runbook.md
T
houseme d887e7e31d test(internode): pin msgpack rollout gates (#5261)
Extend the internode gRPC benchmark driver with P2 request-only, canary, after, and rollback phases. The request-only phase proves that requesting msgpack-only without fleet confirmation keeps compatibility JSON fields, while canary and rollback materialize the operator states needed for mixed-version convergence without claiming real fleet soak evidence.

Verification:

- scripts/test_internode_grpc_ab_bench.sh

- bash -n scripts/run_internode_grpc_ab_bench.sh scripts/test_internode_grpc_ab_bench.sh

- git diff --check

Co-authored-by: heihutu <heihutu@gmail.com>
2026-07-26 03:58:47 +00:00

11 KiB
Raw Blame History

Internode gRPC Optimization — A/B Benchmark Runbook

Reproducible procedure to collect before/after artifacts for each internode gRPC optimization stage (grpc-optimization P0P3). Every stage is env-gated, so "before" and "after" are the same binary with different env — no rebuild between runs.

Live runs need a multi-node cluster (Docker or ≥2 rustfs endpoints), a load tool (warp or s3bench), and a Prometheus scrape of /metrics. They are not runnable in a single-process sandbox. Capture artifacts on a real cluster.

One-click driver

scripts/run_internode_grpc_ab_bench.sh --stage <p0|p1|p2|p3> --phase <before|after|request-only|canary|rollback> [-- <bench args>] wraps the env matrix below: it writes the stage/phase server env to <out-dir>/server-env.sh, then runs the right underlying bench into target/bench/internode-transport/<stage>-<phase>/.

# P1 A/B (restart the cluster with each phase's server-env.sh between the two runs):
scripts/run_internode_grpc_ab_bench.sh --stage p1 --phase before -- --access-key AK --secret-key SK --metrics-url http://node1:9000/metrics
scripts/run_internode_grpc_ab_bench.sh --stage p1 --phase after  -- --access-key AK --secret-key SK --metrics-url http://node1:9000/metrics
# P3 failover A/B (docker four-node):
scripts/run_internode_grpc_ab_bench.sh --stage p3 --phase after
# P2 rollout gates (env preview for request-only rehearsal, canary, and rollback):
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase request-only --dry-run
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase canary --dry-run
scripts/run_internode_grpc_ab_bench.sh --stage p2 --phase rollback --dry-run

RUSTFS_INTERNODE_* are server env: for the load-driven stages (p0/p1/p2) source the emitted server-env.sh on every node and restart rustfs before the run — the driver cannot mutate an already-running server. The P2 canary phase is the exception: source it only on the selected canary node after the release-window counters and fleet-support checks pass, and keep the rest of the fleet on before or request-only while observing fallback/decode-error counters. Use --dry-run to preview the env and command.

Harness

  • Throughput / latency: scripts/run_internode_transport_baseline.sh (drives run_object_batch_bench.sh; writes target/bench/internode-transport-<ts>/). Pass --metrics-url <prometheus> to also capture internode metric deltas.
  • Failover / offline: scripts/run_four_node_cluster_failover_bench.sh (spins up a 4-node compose cluster, kills FAILOVER_NODE, benchmarks; writes target/bench/four-node-failover-<ts>/).

The one-click driver writes each run to target/bench/internode-transport/<stage>-<phase>/ (e.g. p0-before/, p0-after/, p1-before/, p1-after/, p3-before/, p3-after/), each containing the emitted server-env.sh plus the underlying bench artifacts. target/ is gitignored — attach the paired directories to the PR / issue.

Metrics to capture (Prometheus)

Metric Stage signal
rustfs_system_network_internode_operation_duration_ms{operation,backend} control-plane RTT (P0), lock/bulk latency
rustfs_system_network_internode_operation_payload_bytes payload size distribution (P0/P1 sizing)
rustfs_system_network_internode_operation_large_payloads_total large unary RPCs sharing the channel (P1 target)
rustfs_system_network_internode_dial_avg_time_nanos, ..._dial_errors_total connect cost (P3 prewarm)
rustfs_system_network_internode_msgpack_json_fallback_total{direction,message} must be 0 before enabling both msgpack-only gates (P2)
rustfs_system_network_internode_msgpack_json_decode_error_total{direction,message,codec} must be 0 before enabling both msgpack-only gates (P2)
rustfs_cluster_servers_offline_total offline detection correctness (P3 bypass)
lock p99 (lock metrics) P1 head-of-line-blocking win

Per-stage env matrix

Run before with the stage's env at its baseline column, after with the enabled column, everything else at defaults. Roll a restart between runs.

Stage Env before (baseline) after (enabled)
P0 nodelay RUSTFS_INTERNODE_RPC_TCP_NODELAY false true (default)
P0 stream window RUSTFS_INTERNODE_RPC_HTTP2_STREAM_WINDOW_SIZE 0 unset (1 MiB)
P0 conn window RUSTFS_INTERNODE_RPC_HTTP2_CONN_WINDOW_SIZE 0 unset (2 MiB)
P0 msg limit RUSTFS_INTERNODE_RPC_MAX_MESSAGE_SIZE 4194304 unset (100 MiB)
P1 isolation RUSTFS_INTERNODE_CHANNEL_ISOLATION false (default) true
P1 bulk pool RUSTFS_INTERNODE_BULK_CHANNELS 1 24
P2 msgpack-only request RUSTFS_INTERNODE_RPC_MSGPACK_ONLY false (default) true (only after fallback counter = 0 across a window)
P2 fleet confirmation RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED false (default) true (only after mixed-version, fallback-zero, soak, and rollback gates pass)
P3 prewarm RUSTFS_INTERNODE_PREWARM false (default) true
P3 offline bypass RUSTFS_INTERNODE_OFFLINE_BYPASS false (default) true
P3 reprobe / threshold RUSTFS_INTERNODE_OFFLINE_REPROBE_SECS / RUSTFS_INTERNODE_OFFLINE_FAILURE_THRESHOLD defaults 5 / 3

Procedure per stage

  1. Baseline: start the cluster with the stage's env at the before column. Run the relevant bench; save to .../baseline/ (or .../after-P{n-1}/ when chaining stages).
  2. After: restart with the after column; re-run the identical bench; save to .../after-P{n}/.
  3. Diff the object-bench summaries and the metric deltas.
  • P0run_internode_transport_baseline.sh with --sizes 4KiB,1MiB,16MiB,128MiB and --concurrencies 1,16,64. Expect: small-RPC duration_ms (DiskInfo/Ping) down (nodelay), large-metadata (ReadMultiple/BatchReadVersion) throughput up (windows). Functional: a >4 MiB multi-version xl.meta no longer fails out_of_range.
  • P1 — mixed workload (large ReadAll + high-frequency Refresh). Acceptance gate from the design doc: lock p99 down ≥ 20% with RUSTFS_INTERNODE_CHANNEL_ISOLATION=true.
  • P2 — observe msgpack_json_fallback_total and msgpack_json_decode_error_total across a release window; both must stay 0 before flipping both RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true and RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED=true (see the msgpack convergence runbook). Codec allocation via a dhat/heaptrack micro-run.
  • P3 — cold-start: first cross-node op latency should drop ~one connect RTT with prewarm. Failover: run_four_node_cluster_failover_bench.sh, kill a node with RUSTFS_INTERNODE_OFFLINE_BYPASS=true; expect faster failover and a correct rustfs_cluster_servers_offline_total (1 while the node is down, back to 0 after recovery).

Acceptance gates & artifact layout

Each stage's paired run must satisfy an explicit gate before its numbers are accepted. Record the gate verdict (pass/fail + measured delta) in the paired directory's summary and attach it.

Stage Bench Acceptance gate Primary metric(s)
P0 run_internode_transport_baseline.sh small-RPC duration_ms (DiskInfo/Ping) down; large-metadata (ReadMultiple/BatchReadVersion) throughput up; a >4 MiB multi-version xl.meta no longer fails out_of_range (functional). ..._operation_duration_ms{operation}, ..._operation_payload_bytes, object-bench throughput
P1 run_internode_transport_baseline.sh (mixed: large ReadAll + high-frequency Refresh) lock p99 down ≥ 20% with RUSTFS_INTERNODE_CHANNEL_ISOLATION=true vs baseline. lock p99 (lock metrics), ..._operation_large_payloads_total
P3 cold-start run_internode_transport_baseline.sh (fresh cluster, first cross-node op) first cross-node op latency drops ~one connect RTT with RUSTFS_INTERNODE_PREWARM=true. ..._dial_avg_time_nanos, first-op ..._operation_duration_ms
P3 offline dedicated sustained-offline + survivor cross-node access experiment (the standard four-node failover bench is not sensitive to the bypass — quorum holds, recovery_seconds=0) with RUSTFS_INTERNODE_OFFLINE_BYPASS=true, survivor cross-node op latency to the downed peer fast-fails instead of hanging the dial timeout; rustfs_cluster_servers_offline_total = 1 while down, back to 0 after recovery. rustfs_cluster_servers_offline_total, survivor cross-node ..._operation_duration_ms, ..._dial_errors_total

P2 is not a throughput gate. Its acceptance is operational: msgpack_json_fallback_total must read 0 across a full release window before both msgpack-only env gates are enabled (see the msgpack convergence runbook). Do not benchmark P2 as before/after throughput.

Artifact layout per stage (attach both halves + the diff):

target/bench/internode-transport/
  p0-before/  p0-after/     # server-env.sh + object-bench summaries + metric deltas
  p1-before/  p1-after/     # + lock p99 delta (the ≥20% gate)
  p2-before/  p2-request-only/  p2-canary/  p2-after/  p2-rollback/  # env gate artifacts + fallback/decode-error observations
  p3-before/  p3-after/     # cold-start + sustained-offline experiment + offline gauge trace

Bench-host prerequisites (ansible bare-metal)

Captured while running the first real A/B on a 4-node ansible cluster; needed before any live run:

  • RPC secret is mandatory on current main. Internode RPC fails closed: default creds (RUSTFS_SECRET_KEY=rustfsadmin) with no RUSTFS_RPC_SECRET → the node aborts startup immediately with store init aborted: endpoints include remote nodes but RPC authentication secret is not configured (builds before the preflight instead showed No valid auth token and never reached storage_quorum). Set a non-default RUSTFS_RPC_SECRET, identical on every node.
  • systemd start timeout. The install unit is Type=notify; READY only fires after quorum, which on freshly-purged disks exceeds a 30 s TimeoutStartSec → crash loop. Use a drop-in TimeoutStartSec=infinity.
  • Server-side metrics need OTLP. RustFS has no Prometheus pull endpoint (/admin/v3/metrics is NDJSON, not exposition); it only pushes via OTLP. To capture lock/offline/internode metrics, run an otel-collector (OTLP receiver → Prometheus exporter) and set RUSTFS_OBS_ENDPOINT, RUSTFS_OBS_METRICS_EXPORT_ENABLED=true, RUSTFS_OBS_METER_INTERVAL=5. For lock p99 also set RUSTFS_OBJECT_LOCK_DIAG_ENABLE=true (default off).
  • P3 offline method. Standard failover is quorum-insensitive to the bypass. Use a sustained-offline + survivor cross-node access run instead: all nodes up → stop one node (sustained) → warm up to trip offline detection → drive warp on the survivors only (--host excludes the dead node) → compare survivor op p99 and rustfs_cluster_servers_offline_total for RUSTFS_INTERNODE_OFFLINE_BYPASS off/on.

Rollback

Every stage rolls back by setting the env back to its baseline column and restarting. P2 rolls back by unsetting either msgpack-only gate, or by setting both RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=false and RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED=false; no wire format is broken because the JSON fields, _bin fields, and proto field numbers remain additive. Do not remove or reuse JSON proto fields as part of this benchmark stage.