mirror of
https://github.com/taylanbakircioglu/haproxy-openmanager.git
synced 2026-09-20 17:43:29 +00:00
main
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0ee227363e |
fix(agent): adoption cannot hide a node; Config Import on fresh installs (v1.11.1)
Six fixes to the agent's discovery path and one long-standing parity gap, found
by auditing it in loops against a real fleet.
DISCOVERY REPORTS THAT NEVER REACHED THE SERVER
The agent posts an unmanaged keepalived.conf for adoption and caches the hash of
what it sent, so the file - which carries the VRRP password - is re-posted only
when it changes. Delivery was judged by curl's exit code, and curl without -f
exits 0 on 5xx too, so a report the server REJECTED was recorded as delivered.
Since a hand-maintained config does not change on its own, that node dropped out
of "Unmanaged keepalived detected" permanently; the only cure was deleting a
cache file on the node by hand.
- the report is cached only on a 2xx;
- GET /agents/{name}/keepalived-config now reports whether the server actually
holds a discovery for that agent, and the cache may only suppress while it
says yes - which is what lets nodes stuck from earlier releases recover on
their own, with nobody touching them;
- the flag is parsed with has() + tostring, not `// empty`: jq's alternative
operator returns the alternative for **false** as well as null, so the naive
form could not tell "no record" from "older backend" and the recovery would
have been completely inert;
- a 400/413/422 records the refusal so identical bytes are not re-posted
forever - 4xx and 5xx agent calls are never sampled out of the request log,
so an unattended loop would write a row carrying the whole config every
cycle - while 401 and 404 keep retrying, because here they mean a token
rotation or an agent row briefly absent, not a bad payload;
- the CLEAR path had the same exit-code defect, where it left a stale row
offering a managed node for adoption with nothing to ever retry it.
CONFIG IMPORT WAS A NO-OP ON FRESHLY INSTALLED AGENTS
check_config_requests uploads a node's live haproxy.cfg on request. It was
defined in the installer body and in the self-upgrade daemon, but not in the
heredoc a fresh install writes, and its call site is guarded by `type` - so on
such a node the operator asked for a config and nothing arrived, with no error
anywhere. Any agent that had self-upgraded at least once already had it, which
is why it went unnoticed. The self-upgrade definition is copied verbatim
(verified line-for-line). A freshly installed agent now polls that endpoint once
per cycle exactly as every upgraded agent already does; no node running today
changes behaviour.
DETERMINISTIC CONFIG PATH
A pool may hold several clusters and the join that resolves keepalived_config_path
was unordered, so the path handed to an agent could differ between polls whenever
two clusters disagreed - the agent would inspect a file that is not there and the
node would never appear, intermittently. A customised path now wins over the
shipped default, then the lowest cluster id. Verified against a real PostgreSQL
over seven arrangements: with one cluster per pool, or when every cluster carries
the default, the value is byte-identical to before.
Verified end to end on a production fleet and, for each decision, against the
real _kp_discover block rather than a paraphrase.
Backend suite: 1674 passed, 152 skipped. bash -n passes on the whole file and on
the fresh-install body in isolation. The keepalived path is logic-identical
across both daemon copies, now pinned by a test.
|
||
|
|
1d4e4286af |
fix(agent): keep acknowledging once converged, so a lost report self-heals (v1.10.14)
The deploy report is the server's only evidence that a member node applied its keepalived.conf, and it was sent on the write path alone. Once the rendered config was on disk the agent took the idempotency early return every cycle and never reported again, so a single lost report - a backend restart, a 5xx, a network blip - left the VIP reading SYNCING with an empty "Last ack" forever while the node was demonstrably running the right config. Nothing would ever reconcile the two; the only escape was to change the rendered config so the agent wrote it again, which means touching a live VIP to fix a display problem. The agent now re-asserts its state on that path too: one request per node per poll cycle (~2.5 min), nothing written, keepalived not reloaded. This gap dates from the original HA/VIP work rather than this release series; it only became visible when acknowledgements were dropped for an unrelated reason. A test pins that both daemon copies report BEFORE the early return, since placing it after would silently restore the old behaviour. Verified end to end on a real HA pair: discovery, instance-based adoption of both nodes, PENDING, Apply, agent pull, the validation gate, the hash-pinned takeover, the acknowledgement, and retirement of the one-shot authorisation. |
||
|
|
0eb587dfa8 |
fix: keepalived validation gate and deploy acknowledgements (v1.10.13)
Two agent-side fixes found while taking the VIP adoption flow through a real
HA pair, released together.
1. A VALID CONFIG WAS REJECTED BY ITS OWN WARNING (v1.10.12)
Before writing a rendered keepalived.conf the agent validates it with
keepalived -t and, on failure, keeps the running config and does not restart
keepalived. That fail-safe is right, but it treated ANY non-zero exit as
invalid - and keepalived's config-test exit code does not separate fatal from
benign. Measured on 2.2.8:
clean config ..................... 0
auth_pass longer than 8 chars .... 5 "Truncating auth_pass to 8 characters"
missing '}' ...................... 5 "There are 1 missing '}'s"
unknown keyword .................. 5 "Unknown keyword '...'"
script without script_security ... 6 "SECURITY VIOLATION ..."
Exit 5 covers both a harmless truncation and a broken file, so a VRRP password
over eight characters was enough to block every apply - including on a node
whose own running config emits the same warning and had been serving the VIP
for weeks. Accepting exit 5 would have accepted broken configs, so the gate now
judges the OUTPUT: known-benign messages are dropped and anything remaining
still fails. It fails CLOSED - an unrecognised message, or a non-zero exit with
no readable output at all, is fatal - and the filter is an allowlist, never a
denylist. The agent also reports what keepalived said, in its log and in the
status the HA/VIP page shows; discarding it left a correct refusal that nobody
could act on.
2. EVERY DEPLOY ACKNOWLEDGEMENT WAS DROPPED
The takeover-retirement clause on POST /agents/{name}/keepalived-status reused
one placeholder for both the assignment `last_deploy_hash=$n` and the
comparison inside its CASE. PostgreSQL types a placeholder per USE, so it came
out as text in one and character varying in the other and asyncpg rejected the
statement with AmbiguousParameterError. The whole UPDATE never ran, so no
member recorded an acknowledgement: VIPs sat at SYNCING with an empty Last ack
while the nodes were verifiably running the config, and teardown acks were lost
the same way. The hash now has its own placeholder, compared only against the
column.
Verified against real keepalived and a real PostgreSQL rather than by
inspection, including on busybox and bash 3.2, and a test asserts every $n in
those statements is bound exactly once.
|
||
|
|
a87994e06a |
fix(vip): refuse adoption that strands a node or normalises a peer (v1.10.9)
Three findings from a second pass over the adoption flow, all of the same class: something real leaving the set silently. 1. STRANDING. _collect_instance_participants can only match a node it can READ, that is ENABLED, and that is in the SAME pool. Each of those is a door a genuine member of the VRRP group leaves through without a word, and the nodes that remain are rewritten while it keeps serving the same address from an unmanaged config. Found on a live pool: one node of a pair had an unclosed vrrp_instance block, so it parsed to nothing while its partner parsed cleanly. Rather than guard each door, ask the question directly: does any reported keepalived.conf mention THIS virtual address without being one of the nodes we are about to adopt? Refuses naming the node and the reason. Scoped on the address so an unrelated file elsewhere cannot block every adoption, and excluding nodes already under management (a standing VIP, or our ownership marker) because those are not stranded. 2. SILENT NORMALISATION. prefix_length, unicast/multicast mode, HAProxy tracking and the VRRP password are stored ONCE on the VIP and re-rendered onto EVERY member, so whichever node was clicked imposed its settings on the others. prefix_length is the sharpest: the design refuses to GUESS a netmask for a live VIP, and copying one node's netmask onto another is that same change wearing a different hat. All four must now agree, with both values named in the refusal. The VRRP secret is compared by decrypting each node's token - Fernet is non-deterministic, so ciphertexts cannot be compared - and a token that will not decrypt is an error rather than an assumed match. 3. THE TAKEOVER AUTHORISATION WAS NOT ONE-SHOT. takeover_expected_hash is the permission to overwrite a keepalived.conf that lacks our ownership marker. It was written at adoption and never cleared, so it stayed valid for that file content indefinitely: restoring the pre-adoption file would have been overwritten again with no fresh human approval. It is now retired when a member acknowledges our rendered config, gated on the acked hash matching applied_config_hash so a partial or failed deploy never drops it and leaves the VIP unable to converge. The panel applies the stranding rule too, so the Adopt button is disabled with the reason instead of letting the operator click into a 422. No schema change, no agent change, no API-shape break. Backend suite: 1366 passed, 152 skipped. Frontend build clean. |
||
|
|
7d95c737f0 |
fix(vip): adopt the whole VRRP instance, not one node (v1.10.8)
Four defects found while tracing the adoption flow end to end after v1.10.4
reached a live HA pair.
B3/B4 (one root, one fix). Adoption took only the node whose row was clicked:
- adopting the BACKUP alone produced a VIP that apply always rejects, because
apply requires exactly one MASTER member;
- adopting the MASTER alone left the peer unmanaged, and adopting it
afterwards hit the VRID-collision guard with 409, so a pair could never be
completed from the panel;
- on a UNICAST instance the single-member render dropped the unicast block
entirely (render_keepalived_conf emits it only when peer_ips is non-empty),
so keepalived fell back to multicast on the adopted node while its peer
stayed unicast. They stop seeing each other and BOTH claim the VIP.
Adoption now resolves the whole instance via _collect_instance_participants,
keyed on (virtual_router_id, virtual address) - the same key keepalived uses to
group nodes. Each participant becomes a member with the role, priority and
interface its own file declares, and its own one-shot takeover hash, so the
per-node overwrite guard is unchanged. It refuses, naming the reason, when the
group has other than one MASTER, when advert_int differs across nodes, when a
declared unicast peer is not among the nodes being adopted, or when a node is
already in a live VIP - the one-active-VIP-per-agent rule that create/update
enforce via _validate_members_against_pool and adoption never called.
B1. The Apply Management "View Change" regex matched vip-(create|update|delete)
only, so an adopt version fell through to the generic HAProxy diff and rendered
the cluster's whole haproxy.cfg as removed. Display-only, but alarming. A test
now asserts every action _stage_vip_version can stage is in that alternation.
B2. Rejecting an adoption hid the node from the panel permanently:
vip_discoveries.adopted_vip_id is write-once, a VIP is only ever soft-deleted so
the column's ON DELETE SET NULL never fires, and the agent does not re-report a
file whose hash has not changed. Rather than clearing the column on each path,
adoptability is derived from whether the linked VIP is still active, which
self-heals reject, undo-reject and approved teardown alike.
The panel now lists one row per instance instead of per node, and the adopt
dialog names every node that will be taken over. Blockers are aggregated across
all of them, matching what the endpoint checks.
No schema change, no agent change, no API-shape break: /api/vip/discoveries
gains a derived field and /api/vip/adopt keeps its request body.
Backend suite: 1359 passed, 152 skipped. Frontend build clean (no new lint
warnings in VIPManagement.js).
|