mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-13 08:36:54 +00:00
feat(internode): optimize gRPC transport (#4337)
* feat(internode): P0 gRPC transport tuning, message limits, payload metrics Land the P0 subtask from docs/grpc-optimization: close the client-vs-server transport gaps and add instrumentation to size which unary RPCs need channel isolation in P1. Transport tuning (G3): the client `Endpoint` now disables Nagle and raises the HTTP/2 stream/connection flow-control windows to mirror the server socket, so small lock/health RPCs are not batched and larger metadata responses are not throttled by the 64KiB default window. All env-overridable, 0 opts out. Message-size limits (G1): both `NodeServiceClient` and `NodeServiceServer` set max decode/encode size (default 100MiB) instead of tonic's silent 4MiB cap, so a large multi-version xl.meta or aggregated ReadMultiple no longer fails out_of_range. The server limit is set on `NodeServiceServer` before wrapping in the auth `InterceptedService` (the interceptor type does not expose it). Payload instrumentation (P1 prep): ReadAll/ReadMultiple record a payload-size histogram plus a large-payload counter when a response crosses the configured threshold (default 8MiB), feeding alerting on paths that contend with latency-sensitive control-plane traffic on the shared channel. Threshold-only counter, no per-call hot-path log. Verification: cargo check/test on config, io-metrics, ecstore, rustfs; clippy clean on touched files; make pre-commit green. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(internode): P1 control/bulk gRPC channel isolation (opt-in) Land the P1 subtask from docs/grpc-optimization: physically separate large bytes-carrying unary RPCs from latency-sensitive control-plane RPCs so a big transfer can no longer head-of-line block a lock/health RPC on the shared HTTP/2 connection (G2/G5). Introduce ChannelClass { Control, Bulk } and get_channel_for_class in protos. Control RPCs keep the per-peer connection keyed by the bare address; Bulk RPCs (ReadAll/WriteAll/ReadMultiple/BatchReadVersion, via a new get_bulk_client) are round-robined across a small per-peer bulk pool. Rather than restructuring the global GLOBAL_CONN_MAP (and every consumer), bulk channels are cached under a composite key (addr\0bulk\0idx). The NUL separator cannot appear in a URL, so bulk keys never collide with the control key. This keeps the blast radius small on a consistency-sensitive path. create_new_channel is refactored into build_channel(dial_addr, cache_key) so several physically distinct channels to one peer cache independently while dialing/TLS still use the real address. Gated by RUSTFS_INTERNODE_CHANNEL_ISOLATION (default OFF) so the default build is byte-for-byte the pre-P1 behavior: bulk resolves to the control channel and the switch is a single-env rollback. RUSTFS_INTERNODE_BULK_CHANNELS (default 2, clamped >=1) sizes the pool. On failure, evict_failed_connection drops the whole bulk pool for the peer (round-robin hides which index was used), avoiding half-dead cached channels. Lock RPCs (remote_locker) already use the default Control path, so lock semantics and retry behavior are unchanged. Verification: cargo check/test on config, protos, ecstore, rustfs; new protos tests for bulk key routing and isolation-off passthrough; clippy clean on touched files; make pre-commit green. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(internode): P2 msgpack/JSON codec observability + encode buffer presizing Land the safe, wire-compatible slice of P2 from docs/grpc-optimization: the observability prerequisite for retiring the redundant JSON fields, plus a codec micro-optimization. No proto/wire-format change; JSON is still dual-written. Internode RPCs today dual-encode each metadata value as both msgpack (`*_bin`) and a JSON compatibility string, and decoders prefer `_bin` with a JSON fallback. Before the JSON fields can ever be dropped (a cross-version change), that fallback must be proven unused in production. Add rustfs_system_network_internode_msgpack_json_fallback_total{direction, message}: incremented whenever a decode falls back to the JSON field because the msgpack payload was absent. Wired into both directions — the client decoding peer responses (remote_disk.rs, incl. the list-level read_multiple/batch fallbacks) and the server decoding peer requests (node_service/disk.rs). This counter must read zero across a release window before send paths stop writing JSON and the proto text fields are reserved/removed (the deferred P2-1 steps). Also pre-size the msgpack encode buffers (Vec::with_capacity(512)) on both sides, eliminating the repeated growth reallocations for typical FileInfo payloads with zero added copy. Full thread_local buffer pooling is deferred: it needs either an extra copy (unclear net win) or a send-path buffer-return lifecycle, to be justified by a codec microbenchmark first. Verification: cargo check/test on io-metrics, ecstore, rustfs; new fallback counter smoke test; existing codec decode tests green; clippy clean on touched files; make pre-commit green. Co-Authored-By: heihutu <heihutu@gmail.com> * docs(internode): add msgpack/JSON convergence observation runbook Runbook driving the observation-gated retirement of the redundant JSON compatibility fields on internode gRPC metadata RPCs (grpc-optimization P2-1). Documents the shipped fallback counter (rustfs_system_network_internode_msgpack_json_fallback_total{direction,message}), the PromQL to confirm it reads zero across a release window, a standing alert, and the staged flip/rollback procedure (env-gated msgpack-only send, then proto field removal in N+1). Includes the verified field -> peer-decoder audit: only fields whose peer decodes _bin first may be converged. Notes DeleteVersion.opts (DeleteOptions) is NOT convergence-ready — its server handler is not _bin-first and must gain a decode_msgpack_or_json path first. This gates the send-side change so it cannot empty a JSON field an old peer still needs. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(internode): env-gated msgpack-only send + DeleteVersion _bin support (P2-1) Implements the send-side lever for retiring the redundant JSON compatibility fields on internode gRPC metadata RPCs, plus the missing `_bin` support on the delete path that it depends on (grpc-optimization P2-1). Default-off: the base build is byte-for-byte the prior dual-write behavior. Gated msgpack-only send (RUSTFS_INTERNODE_RPC_MSGPACK_ONLY, default false): - New rustfs_protos::internode_rpc_msgpack_only() reads the flag. - Client (remote_disk.rs) compat_json() and server (node_service/disk.rs) compat_response_json() emit an empty JSON string when the flag is on, so only the msgpack _bin payload is sent. The _bin field is always sent; decoders keep the JSON read fallback. Applied only to fields with a confirmed _bin-first peer decoder (WriteMetadata/UpdateMetadata/RenameData file_info, UpdateMetadata opts, ReadOptions, ReadMultipleReq, BatchReadVersionReq; ReadVersion/ReadXL/RenameData responses and the ReadMultiple/BatchReadVersion response lists). - Only enable after the P2 fallback counter has read zero across a release window (see docs/operations/internode-msgpack-json-convergence-runbook.md). Single-env rollback; no wire-format break. DeleteVersion(s) _bin support (prerequisite): - The DeleteVersion/DeleteVersions protos had NO _bin fields. Add additive (backward-compatible) bytes file_info_bin/opts_bin (DeleteVersion) and repeated bytes versions_bin + bytes opts_bin (DeleteVersions); regenerate the checked-in prost struct. - Client dual-writes them; server decodes them _bin-first with JSON fallback. - These delete fields are kept OUT of the msgpack-only set (always dual-write) until their own fallback counter reads zero across a window with the new decoders fully deployed. DeleteVersion.raw_file_info stays JSON-only (no _bin field yet). Verification: cargo check/test on protos, config, ecstore, rustfs (incl. the six delete request handler tests and a compat_json default-path test); clippy clean on touched files; make pre-commit green. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(internode): P3 cluster peer online/offline health metric Land the safe observability core of P3 (grpc-optimization G6/G8): track each internode peer's reachability and expose the offline count, for parity with MinIO's minio_cluster_servers_offline_total. Pure instrumentation — peer selection and quorum are unchanged. - io-metrics: per-peer PeerHealthState { online, consecutive_failures } registry plus record_peer_reachable/record_peer_unreachable. A peer flips offline after N consecutive failures (dial failures or RPC-triggered evictions) and back online on the next successful dial; the count of offline peers is published to the rustfs_cluster_servers_offline_total gauge. - config: RUSTFS_INTERNODE_OFFLINE_FAILURE_THRESHOLD (default 3, clamped >= 1). - protos: build_channel marks the peer reachable on a successful dial and unreachable on a dial failure; evict_failed_connection feeds the failure signal too. Keyed by the real peer address, so control and bulk channels to one peer share health state. Deferred (documented in docs/grpc-optimization P3): startup prewarm (no clean topology-ready hook yet), the offline fast-bypass in peer routing (consistency- sensitive; must not change quorum), and idempotent-read-only retry. This commit is observability only. Verification: cargo check/test on io-metrics, config, protos (new peer-health state-machine and threshold-clamp tests); clippy clean on touched files; make pre-commit green. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(internode): P3 control-channel prewarm + self-healing offline bypass Add the remaining P3 connection-lifecycle levers (grpc-optimization G6/G8), both env-gated and default-off so the base build is unchanged. Prewarm (RUSTFS_INTERNODE_PREWARM, default off): RemoteDisk::new spawns a best-effort background dial of the peer's control channel, deduped per peer address, moving the connect cost off the first RPC. Failures fall through to the existing lazy connect + recovery monitor. Offline bypass (RUSTFS_INTERNODE_OFFLINE_BYPASS, default off): remote_disk get_client/get_bulk_client fast-fail a peer already marked offline instead of paying the connect timeout, so the erasure layer proceeds on quorum sooner. This does NOT change quorum. It is self-healing: cluster_peer_should_bypass lets one request per RUSTFS_INTERNODE_OFFLINE_REPROBE_SECS (default 5s) through to recover the peer even with no background monitor, and the recovery monitor's own probe path calls the client directly so it is never bypassed. io-metrics gains cluster_peer_is_offline / cluster_peer_should_bypass (with a per-peer re-probe timestamp). Scope: data path only — remote_locker (lock RPCs, most consistency-sensitive) is left dual-writing/unbypassed as a follow-up. Verification: cargo check/test on io-metrics, config, ecstore (new self-healing bypass tests; all 105 rpc tests green); clippy clean on touched files; make pre-commit green. Co-Authored-By: heihutu <heihutu@gmail.com> * docs(internode): add A/B benchmark runbook for gRPC optimization stages Reproducible before/after collection procedure for grpc-optimization P0–P3. Since every stage is env-gated, before/after is the same binary with different env — no rebuild. Documents, per stage: the exact env toggles (baseline vs enabled column), which existing bench script to run (run_internode_transport_baseline.sh / run_four_node_cluster_failover_bench.sh), the Prometheus metrics to capture, and the acceptance gates from the design docs (e.g. lock p99 down >= 20% for P1, msgpack fallback counter = 0 before enabling P2, correct rustfs_cluster_servers_offline_total for P3). Live runs require a multi-node cluster + load tool + Prometheus scrape and cannot be produced in a single-process sandbox; artifacts land under target/bench (gitignored) and attach to the PR. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(internode): P3-2 lock-path offline bypass + P3-3 idempotent read retry Extend the offline bypass to the lock path and add opt-in retries for idempotent reads (grpc-optimization P3-2/P3-3). Both env-gated and default-off/zero. Offline bypass (lock path): factor the bypass decision into a shared pub(crate) internode_offline_bypass_reason(addr) and call it from remote_locker::get_client too, so lock RPCs to an offline peer fast-fail (letting dsync reach quorum sooner) instead of paying the connect timeout. Does not change quorum; the self-healing re-probe keeps peers recoverable. Gated by RUSTFS_INTERNODE_OFFLINE_BYPASS (default off). Idempotent read retry (P3-3): add execute_read_with_retry — a bounded, exponential-backoff retry for read-only/reentrant RPCs on transient network errors — and route disk_info through it. RUSTFS_INTERNODE_IDEMPOTENT_READ_RETRIES defaults to 0 (disabled). Write/lock RPCs are never retried (quorum/idempotency safety, per CLAUDE.md); the wrapper requires an Fn closure so only reads that rebuild their request from borrowed inputs qualify. Deferred: grpc.health.v1 (optional ecosystem-compat only; needs a new tonic-health dep and 3-way hybrid-service wiring — internal needs are met by the existing Ping RPC). Verification: cargo check/test on config, ecstore (105 rpc tests green incl. disk_info now via the retry wrapper); clippy clean on touched files; make pre-commit green. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(scripts): one-click internode gRPC A/B benchmark driver Wrap the per-stage env matrix from the benchmark runbook into scripts/run_internode_grpc_ab_bench.sh: given --stage <p0|p1|p2|p3> and --phase <before|after>, it emits the stage/phase RUSTFS_INTERNODE_* server env to <out-dir>/server-env.sh and runs the right underlying bench (run_internode_transport_baseline.sh for p0/p1/p2, run_four_node_cluster_failover_bench.sh for p3) into a labeled target/bench/internode-transport/<stage>-<phase>/. Passthrough args after `--` reach the underlying bench; --dry-run previews the env + command. The script is explicit that RUSTFS_INTERNODE_* are server env, so for the load-driven stages the operator must restart rustfs with the emitted env before the run; the docker four-node (p3) path exports them for a forwarding compose. shellcheck-clean. Runbook updated with a "One-click driver" section. Co-Authored-By: heihutu <heihutu@gmail.com> * chore(compose): forward RUSTFS_INTERNODE_* into the four-node cluster The four-node local-build compose only forwarded a fixed whitelist of env, so the internode gRPC knobs (grpc-optimization P0-P3) never reached the containers and the A/B bench driver's "after" phase was a no-op. Forward the full RUSTFS_INTERNODE_* set with defaults matching the binary defaults, so leaving them unset is a no-op and the A/B driver can toggle a stage per phase. Co-Authored-By: heihutu <heihutu@gmail.com> * fix(internode): address Copilot review — retry health action + poison-safe peer health Two review nits on #4337: - P3-3 idempotent read retry (remote_disk.rs): execute_read_with_retry ran every attempt through execute_with_timeout_for_op, which hardcodes FailureHealthAction::MarkFailure. So the first transient error could flip the disk faulty and short-circuit the remaining retries, and each attempt over-counted the failure. Route all but the final attempt through execute_with_timeout_for_op_and_health_action with IgnoreFailure; only the last attempt marks faulty/evicts. No default impact (retries default 0). - Peer-health helpers (io-metrics): record_peer_reachable/record_peer_unreachable, cluster_peer_is_offline and cluster_peer_should_bypass early-returned on a poisoned mutex, permanently stalling the offline gauge and bypass state after a single panic. Recover via PoisonError::into_inner(). Verification: cargo check/test on io-metrics + ecstore (105 rpc tests green); clippy clean on touched files; make pre-commit green. Co-Authored-By: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com>
This commit is contained in:
@@ -0,0 +1,97 @@
|
||||
# Internode gRPC Optimization — A/B Benchmark Runbook
|
||||
|
||||
Reproducible procedure to collect **before/after** artifacts for each internode gRPC
|
||||
optimization stage (grpc-optimization P0–P3). Every stage is env-gated, so "before" and
|
||||
"after" are the *same binary* with different env — no rebuild between runs.
|
||||
|
||||
> Live runs need a multi-node cluster (Docker or ≥2 rustfs endpoints), a load tool
|
||||
> (`warp` or `s3bench`), and a Prometheus scrape of `/metrics`. They are not runnable in a
|
||||
> single-process sandbox. Capture artifacts on a real cluster.
|
||||
|
||||
## One-click driver
|
||||
|
||||
`scripts/run_internode_grpc_ab_bench.sh --stage <p0|p1|p2|p3> --phase <before|after> [-- <bench args>]`
|
||||
wraps the env matrix below: it writes the stage/phase **server** env to
|
||||
`<out-dir>/server-env.sh`, then runs the right underlying bench into
|
||||
`target/bench/internode-transport/<stage>-<phase>/`.
|
||||
|
||||
```bash
|
||||
# P1 A/B (restart the cluster with each phase's server-env.sh between the two runs):
|
||||
scripts/run_internode_grpc_ab_bench.sh --stage p1 --phase before -- --access-key AK --secret-key SK --metrics-url http://node1:9000/metrics
|
||||
scripts/run_internode_grpc_ab_bench.sh --stage p1 --phase after -- --access-key AK --secret-key SK --metrics-url http://node1:9000/metrics
|
||||
# P3 failover A/B (docker four-node):
|
||||
scripts/run_internode_grpc_ab_bench.sh --stage p3 --phase after
|
||||
```
|
||||
|
||||
`RUSTFS_INTERNODE_*` are **server** env: for the load-driven stages (p0/p1/p2) source the emitted
|
||||
`server-env.sh` on every node and restart rustfs *before* the run — the driver cannot mutate an
|
||||
already-running server. Use `--dry-run` to preview the env and command.
|
||||
|
||||
## Harness
|
||||
|
||||
- Throughput / latency: `scripts/run_internode_transport_baseline.sh` (drives
|
||||
`run_object_batch_bench.sh`; writes `target/bench/internode-transport-<ts>/`). Pass
|
||||
`--metrics-url <prometheus>` to also capture internode metric deltas.
|
||||
- Failover / offline: `scripts/run_four_node_cluster_failover_bench.sh` (spins up a 4-node
|
||||
compose cluster, kills `FAILOVER_NODE`, benchmarks; writes
|
||||
`target/bench/four-node-failover-<ts>/`).
|
||||
|
||||
Store the paired artifacts as `target/bench/internode-transport/{baseline,after-P0,after-P1,after-P3}/`
|
||||
(the paths the design docs reference). `target/` is gitignored — attach the artifacts to the PR.
|
||||
|
||||
## Metrics to capture (Prometheus)
|
||||
|
||||
| Metric | Stage signal |
|
||||
|---|---|
|
||||
| `rustfs_system_network_internode_operation_duration_ms{operation,backend}` | control-plane RTT (P0), lock/bulk latency |
|
||||
| `rustfs_system_network_internode_operation_payload_bytes` | payload size distribution (P0/P1 sizing) |
|
||||
| `rustfs_system_network_internode_operation_large_payloads_total` | large unary RPCs sharing the channel (P1 target) |
|
||||
| `rustfs_system_network_internode_dial_avg_time_nanos`, `..._dial_errors_total` | connect cost (P3 prewarm) |
|
||||
| `rustfs_system_network_internode_msgpack_json_fallback_total{direction,message}` | must be **0** before enabling msgpack-only (P2) |
|
||||
| `rustfs_cluster_servers_offline_total` | offline detection correctness (P3 bypass) |
|
||||
| lock p99 (lock metrics) | P1 head-of-line-blocking win |
|
||||
|
||||
## Per-stage env matrix
|
||||
|
||||
Run **before** with the stage's env at its baseline column, **after** with the enabled
|
||||
column, everything else at defaults. Roll a restart between runs.
|
||||
|
||||
| Stage | Env | before (baseline) | after (enabled) |
|
||||
|---|---|---|---|
|
||||
| P0 nodelay | `RUSTFS_INTERNODE_RPC_TCP_NODELAY` | `false` | `true` (default) |
|
||||
| P0 stream window | `RUSTFS_INTERNODE_RPC_HTTP2_STREAM_WINDOW_SIZE` | `0` | unset (1 MiB) |
|
||||
| P0 conn window | `RUSTFS_INTERNODE_RPC_HTTP2_CONN_WINDOW_SIZE` | `0` | unset (2 MiB) |
|
||||
| P0 msg limit | `RUSTFS_INTERNODE_RPC_MAX_MESSAGE_SIZE` | `4194304` | unset (100 MiB) |
|
||||
| P1 isolation | `RUSTFS_INTERNODE_CHANNEL_ISOLATION` | `false` (default) | `true` |
|
||||
| P1 bulk pool | `RUSTFS_INTERNODE_BULK_CHANNELS` | `1` | `2`–`4` |
|
||||
| P2 msgpack-only | `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY` | `false` (default) | `true` (only after fallback counter = 0 across a window) |
|
||||
| P3 prewarm | `RUSTFS_INTERNODE_PREWARM` | `false` (default) | `true` |
|
||||
| P3 offline bypass | `RUSTFS_INTERNODE_OFFLINE_BYPASS` | `false` (default) | `true` |
|
||||
| P3 reprobe / threshold | `RUSTFS_INTERNODE_OFFLINE_REPROBE_SECS` / `RUSTFS_INTERNODE_OFFLINE_FAILURE_THRESHOLD` | defaults | `5` / `3` |
|
||||
|
||||
## Procedure per stage
|
||||
|
||||
1. **Baseline**: start the cluster with the stage's env at the *before* column. Run the
|
||||
relevant bench; save to `.../baseline/` (or `.../after-P{n-1}/` when chaining stages).
|
||||
2. **After**: restart with the *after* column; re-run the identical bench; save to
|
||||
`.../after-P{n}/`.
|
||||
3. Diff the object-bench summaries and the metric deltas.
|
||||
|
||||
- **P0** — `run_internode_transport_baseline.sh` with `--sizes 4KiB,1MiB,16MiB,128MiB` and
|
||||
`--concurrencies 1,16,64`. Expect: small-RPC `duration_ms` (DiskInfo/Ping) down (nodelay),
|
||||
large-metadata (ReadMultiple/BatchReadVersion) throughput up (windows). Functional: a
|
||||
`>4 MiB` multi-version `xl.meta` no longer fails `out_of_range`.
|
||||
- **P1** — mixed workload (large `ReadAll` + high-frequency `Refresh`). Acceptance gate from
|
||||
the design doc: **lock p99 down ≥ 20%** with `RUSTFS_INTERNODE_CHANNEL_ISOLATION=true`.
|
||||
- **P2** — observe `msgpack_json_fallback_total` across a release window; it must stay **0**
|
||||
before flipping `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true` (see the msgpack convergence
|
||||
runbook). Codec allocation via a `dhat`/`heaptrack` micro-run.
|
||||
- **P3** — cold-start: first cross-node op latency should drop ~one connect RTT with prewarm.
|
||||
Failover: `run_four_node_cluster_failover_bench.sh`, kill a node with
|
||||
`RUSTFS_INTERNODE_OFFLINE_BYPASS=true`; expect faster failover and a correct
|
||||
`rustfs_cluster_servers_offline_total` (1 while the node is down, back to 0 after recovery).
|
||||
|
||||
## Rollback
|
||||
|
||||
Every stage is a single-env rollback (set the env back to its baseline column and restart).
|
||||
No wire-format is broken except P2 stage 2 (proto field removal), which is a separate release.
|
||||
@@ -0,0 +1,145 @@
|
||||
# Internode msgpack/JSON Convergence Runbook
|
||||
|
||||
Operational runbook for retiring the redundant JSON compatibility fields on internode
|
||||
gRPC metadata RPCs (grpc-optimization **P2-1**). This is a **cross-version** change: it
|
||||
proceeds strictly by observation-gated stages, never in one step.
|
||||
|
||||
## Background
|
||||
|
||||
Internode RPCs dual-encode each metadata value as **both**:
|
||||
|
||||
- a msgpack binary field (`*_bin`, e.g. `file_info_bin`), and
|
||||
- a JSON compatibility string (e.g. `file_info`).
|
||||
|
||||
Decoders prefer the `_bin` payload and fall back to the JSON string only when `_bin` is
|
||||
empty (`decode_msgpack_or_json`). The dual-write costs bandwidth and CPU. Before the JSON
|
||||
fields can be dropped, the fallback branch must be proven **unused** in production —
|
||||
otherwise a rolling upgrade with mixed node versions could read an emptied field.
|
||||
|
||||
## The observation metric (already shipped)
|
||||
|
||||
```
|
||||
rustfs_system_network_internode_msgpack_json_fallback_total{direction, message}
|
||||
```
|
||||
|
||||
Incremented whenever a decode falls back to the JSON field because the msgpack payload was
|
||||
absent.
|
||||
|
||||
- `direction="request"` — a server decoding a peer's request (`node_service/disk.rs`).
|
||||
- `direction="response"` — a client decoding a peer's response (`cluster/rpc/remote_disk.rs`),
|
||||
including the list-level `ReadMultiple` / `BatchReadVersion` fallbacks.
|
||||
- `message` — the value name, e.g. `FileInfo`, `RawFileInfo`, `ReadMultipleResp`.
|
||||
|
||||
## Stage 0 — Observe (current stage)
|
||||
|
||||
Ship the current release (which contains the counter) and let it run for **at least one
|
||||
full release window** across the whole fleet. The counter must stay at **zero**.
|
||||
|
||||
Confirm zero across the observation window (adjust `[30d]` to the window length):
|
||||
|
||||
```promql
|
||||
sum by (direction, message) (
|
||||
increase(rustfs_system_network_internode_msgpack_json_fallback_total[30d])
|
||||
)
|
||||
```
|
||||
|
||||
Every series must be `0`. A non-zero value means some peer is still emitting an empty
|
||||
`_bin` (an old node, or a message whose sender does not fill `_bin`) — investigate the
|
||||
`{direction, message}` label before proceeding.
|
||||
|
||||
Standing alert (keep enabled through all stages):
|
||||
|
||||
```yaml
|
||||
- alert: InternodeMsgpackJsonFallback
|
||||
expr: sum by (direction, message) (increase(rustfs_system_network_internode_msgpack_json_fallback_total[15m])) > 0
|
||||
for: 5m
|
||||
labels: { severity: warning }
|
||||
annotations:
|
||||
summary: "Internode RPC fell back to JSON decode ({{ $labels.direction }}/{{ $labels.message }})"
|
||||
description: "A peer sent an empty msgpack _bin payload. Do NOT advance msgpack-only convergence while this fires."
|
||||
```
|
||||
|
||||
## Field → peer-decoder audit
|
||||
|
||||
The send-side change (Stage 1) may only empty a JSON field whose **peer decodes `_bin`
|
||||
first**. The following mapping is verified against the current code.
|
||||
|
||||
### Convergence-ready (peer decodes `_bin` first)
|
||||
|
||||
| Direction | Message / field | Peer decoder |
|
||||
|---|---|---|
|
||||
| request | `WriteMetadata.file_info` | `FileInfo` (node_service/disk.rs) |
|
||||
| request | `UpdateMetadata.file_info` | `FileInfo` |
|
||||
| request | `UpdateMetadata.opts` | `UpdateMetadataOpts` |
|
||||
| request | `RenameData.file_info` | `FileInfo` |
|
||||
| request | `ReadMultiple.read_multiple_req` | `ReadMultipleReq` |
|
||||
| request | `BatchReadVersion.batch_read_version_req` | `BatchReadVersionReq` |
|
||||
| request | `Read*.opts` | `ReadOptions` |
|
||||
| response | `ReadVersion.file_info` | `FileInfo` (cluster/rpc/remote_disk.rs) |
|
||||
| response | `ReadXL.raw_file_info` | `RawFileInfo` |
|
||||
| response | `RenameData.rename_data_resp` | `RenameDataResp` |
|
||||
| response | `ReadMultiple` resp list | per-item + list fallback |
|
||||
| response | `BatchReadVersion` resp list | per-item + list fallback |
|
||||
|
||||
### `_bin` support added THIS release — converge after their own window
|
||||
|
||||
The `DeleteVersion`/`DeleteVersions` protos had **no `_bin` fields**. They gained additive
|
||||
`*_bin` fields plus bin-first server decoders in this release, and the client now dual-writes
|
||||
them. They are **kept out** of the msgpack-only set (always dual-write) until their own
|
||||
fallback counter has read zero across a window with the new decoders fully deployed.
|
||||
|
||||
| Direction | Message / field | Status |
|
||||
|---|---|---|
|
||||
| request | `DeleteVersion.file_info` (`FileInfo`) | `_bin` added; dual-write; converge after window |
|
||||
| request | `DeleteVersion.opts` (`DeleteOptions`) | `_bin` added; dual-write; converge after window |
|
||||
| request | `DeleteVersions.versions` (`FileInfoVersions`) | `_bin` added; dual-write; converge after window |
|
||||
| request | `DeleteVersions.opts` (`DeleteOptions`) | `_bin` added; dual-write; converge after window |
|
||||
|
||||
> Any `*_bin` proto field not in the tables above must be mapped to a confirmed `_bin`-first
|
||||
> peer decoder before it is added to the convergence set.
|
||||
|
||||
### Still JSON-only (no `_bin` field)
|
||||
|
||||
| Direction | Message / field | Note |
|
||||
|---|---|---|
|
||||
| response | `DeleteVersion.raw_file_info` | proto has no `_bin`; needs an additive proto field before it can converge. |
|
||||
|
||||
## Stage 1 — Stop writing JSON (env-gated, after Stage 0 reads zero)
|
||||
|
||||
The send-side lever is **implemented** and gated by the default-off env flag
|
||||
`RUSTFS_INTERNODE_RPC_MSGPACK_ONLY`. When enabled, the convergence-ready fields above send
|
||||
only `_bin` and leave the JSON string empty; the `_bin` payload is always sent and decoders
|
||||
keep the JSON read fallback unchanged. The delete fields are excluded (dual-write) per the
|
||||
section above.
|
||||
|
||||
Only enable it **after** Stage 0 has read zero for a full window across the fleet:
|
||||
|
||||
1. Ship with the flag **off** (no behavior change).
|
||||
2. Enable it on one node (`RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true`, restart) and watch the
|
||||
fallback counter for a soak period. If it stays zero, enable fleet-wide.
|
||||
3. **Rollback:** set `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=false` (or unset) and restart. No
|
||||
wire-format was broken in this stage, so rollback is immediate and safe.
|
||||
|
||||
## Stage 2 — Remove the proto JSON fields (next release, N+1)
|
||||
|
||||
Only after Stage 1 has been stable with the flag on for a full window and the counter is
|
||||
still zero.
|
||||
|
||||
1. Mark the retired text fields `reserved` in `crates/protos/src/node.proto` (never reuse
|
||||
the field numbers) and delete the JSON read-fallback branches; codec becomes msgpack-only.
|
||||
2. This is a hard wire-format change — it requires the mixed-version upgrade rehearsal
|
||||
(four-node scripts) to pass, and it cannot be rolled back by env alone.
|
||||
|
||||
## Rollback matrix
|
||||
|
||||
| Stage | Wire-format broken? | Rollback |
|
||||
|---|---|---|
|
||||
| 0 Observe | no | n/a (metric only) |
|
||||
| 1 msgpack-only send | no | `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=false` + restart |
|
||||
| 2 remove fields | yes | redeploy prior release; field numbers stay `reserved` |
|
||||
|
||||
## Related
|
||||
|
||||
- Codec + counter implementation: commit `feat(internode): P2 msgpack/JSON codec observability + encode buffer presizing`.
|
||||
- Decoders: `decode_msgpack_or_json` in `crates/ecstore/src/cluster/rpc/remote_disk.rs` (client)
|
||||
and `rustfs/src/storage/rpc/node_service/disk.rs` (server).
|
||||
Reference in New Issue
Block a user