mirror of
https://github.com/Studio-Saelix/sencho.git
synced 2026-08-09 02:12:59 +00:00
fix(mesh): route peer→central traffic over the existing forward WS (#1094)
* fix(mesh): route peer→central traffic over the existing forward WS The reverse mesh callback path (`/api/mesh/proxy-tunnel-from-peer`) needed SENCHO_PRIMARY_URL on central plus a publicly reachable origin from the peer's perspective. In a typical homelab where central sits behind NAT, peer→central dispatch silently failed at the dialer's short-circuit and the headline "call any service on any node by hostname" worked one way only. The forward WS at `/api/mesh/proxy-tunnel` is already bidirectional end to end. Make the bridge a persistent control-plane primitive: dial every mesh-enabled proxy peer at startup, reconcile every 60 s, never idle-close. Peer→central traffic multiplexes over the same WS via `tcp_open_reverse`. Removed: - `meshProxyTunnelFromPeer.ts` WS handler and dispatch - `MeshCentralRegistry`, `PeerToCentralMeshSessionDialer` - `mesh_handshake` first-frame state machine in `meshProxyTunnel.ts` - `maybeSendBootstrap`, `buildHandshakeFrame` in the dialer - `mesh_proxy_callback_bootstrap` capability and `maybeWarnUnsetPrimaryUrl` - `mesh_centrals` table (drop migration; greenfield, no users) - `PilotTunnelManager.replaceOrRegisterProxyBridge` (dead after handler removal) - twelve associated unit/integration tests plus the peer-recovery branch in `MeshService.openCrossNode` Added: - `MeshService.proactiveBridgeFanout` selects every mesh-enabled proxy peer (no longer gated on `mesh_stacks` rows) - `startBridgeReconcileLoop` runs the fanout every 60 s (override via `SENCHO_MESH_RECONCILE_INTERVAL_MS`) - `MeshProxyTunnelDialer` default idle TTL is now `0` and exposes `isDialing(nodeId)` for the status surface - `MeshNodeStatus.reverseCallbackStatus` discriminator (`connected | connecting | unavailable | not_applicable`) surfaced via `/api/mesh/status` and rendered as a pill in the Routing tab - `openCrossNode` error message distinguishes "no proxy target" from "waiting for central to dial the reverse bridge" - New tests: `mesh-service-proxy-tunnel-reconcile`, `mesh-status-reverse-callback`, `mesh-proxy-tunnel-dialer-no-idle-close` SENCHO_PRIMARY_URL is no longer required for any mesh function. * fix(mesh): rewrite proxy-tunnel reconcile test contents The previous commit renamed the file but the rewritten test bodies stayed unstaged on top of the rename. This commit lands the actual rewrite: the fanout assertion now requires every mesh-enabled proxy peer to be dialed, not just those with `mesh_stacks` rows, and adds a reconcile-tick repeated-call test.
This commit is contained in:
@@ -71,19 +71,16 @@ services:
|
||||
|
||||
Sencho's static IP on the network is `<network address> + 2` (so `10.42.0.2` for the example above). Operators can configure each node independently; the override generator pushes alias hostnames that resolve to whichever IP the deploying node uses locally.
|
||||
|
||||
## Bidirectional secure callback
|
||||
## Bidirectional traffic
|
||||
|
||||
Proxy-mode mesh nodes can re-establish their secure mesh callback to central
|
||||
when traffic starts from that node. Central is the hub; the callback is signed
|
||||
with a `mesh_tunnel`-scoped credential bound to the node's API token.
|
||||
|
||||
<Note>
|
||||
**Prerequisite:** set `SENCHO_PRIMARY_URL` on central to its canonical public
|
||||
origin. This value is used as the audience for the mesh callback credential
|
||||
and as the callback URL on the node side. Without it, central infers an origin
|
||||
from inbound request headers; the inference works in standard installs but is
|
||||
not the intended deploy shape for fleets where reverse-proxy headers may differ.
|
||||
</Note>
|
||||
Mesh traffic flows in both directions over a single persistent WebSocket
|
||||
between central and each proxy-mode peer. Central dials the bridge to every
|
||||
mesh-enabled peer on startup and re-dials any dropped bridge on its
|
||||
reconcile tick (every 60 seconds by default). When a container on the peer
|
||||
needs to reach a service on central, the request multiplexes over that same
|
||||
bridge: no extra inbound listener on central, no public URL required on
|
||||
the central side. This is the path that makes "call a service on any node
|
||||
by hostname" work in NAT'd homelab topologies out of the box.
|
||||
|
||||
## Test upstream
|
||||
|
||||
|
||||
Reference in New Issue
Block a user