Six fixes to the agent's discovery path and one long-standing parity gap, found
by auditing it in loops against a real fleet.
DISCOVERY REPORTS THAT NEVER REACHED THE SERVER
The agent posts an unmanaged keepalived.conf for adoption and caches the hash of
what it sent, so the file - which carries the VRRP password - is re-posted only
when it changes. Delivery was judged by curl's exit code, and curl without -f
exits 0 on 5xx too, so a report the server REJECTED was recorded as delivered.
Since a hand-maintained config does not change on its own, that node dropped out
of "Unmanaged keepalived detected" permanently; the only cure was deleting a
cache file on the node by hand.
- the report is cached only on a 2xx;
- GET /agents/{name}/keepalived-config now reports whether the server actually
holds a discovery for that agent, and the cache may only suppress while it
says yes - which is what lets nodes stuck from earlier releases recover on
their own, with nobody touching them;
- the flag is parsed with has() + tostring, not `// empty`: jq's alternative
operator returns the alternative for **false** as well as null, so the naive
form could not tell "no record" from "older backend" and the recovery would
have been completely inert;
- a 400/413/422 records the refusal so identical bytes are not re-posted
forever - 4xx and 5xx agent calls are never sampled out of the request log,
so an unattended loop would write a row carrying the whole config every
cycle - while 401 and 404 keep retrying, because here they mean a token
rotation or an agent row briefly absent, not a bad payload;
- the CLEAR path had the same exit-code defect, where it left a stale row
offering a managed node for adoption with nothing to ever retry it.
CONFIG IMPORT WAS A NO-OP ON FRESHLY INSTALLED AGENTS
check_config_requests uploads a node's live haproxy.cfg on request. It was
defined in the installer body and in the self-upgrade daemon, but not in the
heredoc a fresh install writes, and its call site is guarded by `type` - so on
such a node the operator asked for a config and nothing arrived, with no error
anywhere. Any agent that had self-upgraded at least once already had it, which
is why it went unnoticed. The self-upgrade definition is copied verbatim
(verified line-for-line). A freshly installed agent now polls that endpoint once
per cycle exactly as every upgraded agent already does; no node running today
changes behaviour.
DETERMINISTIC CONFIG PATH
A pool may hold several clusters and the join that resolves keepalived_config_path
was unordered, so the path handed to an agent could differ between polls whenever
two clusters disagreed - the agent would inspect a file that is not there and the
node would never appear, intermittently. A customised path now wins over the
shipped default, then the lowest cluster id. Verified against a real PostgreSQL
over seven arrangements: with one cluster per pool, or when every cluster carries
the default, the value is byte-identical to before.
Verified end to end on a production fleet and, for each decision, against the
real _kp_discover block rather than a paraphrase.
Backend suite: 1674 passed, 152 skipped. bash -n passes on the whole file and on
the fresh-install body in isolation. The keepalived path is logic-identical
across both daemon copies, now pinned by a test.
47 KiB
Upgrade Notes — v1.11.1 (adoption cannot hide a node; Config Import on fresh installs)
Agent-script change, no schema change. No SCHEMA_VERSION bump, so the built-in roles are
not re-seeded. After deploying, sync the Linux agent script from Agent Management and let
the agents upgrade, or none of this reaches the nodes.
- A node that never appeared under Unmanaged keepalived detected now recovers by itself.
The agent caches the hash of its last discovery report and skips re-posting while it matches.
Delivery was judged by
curl's exit code, which is 0 for 5xx too, so a rejected report was cached as delivered and a hand-maintained config — which never changes on its own — kept that node hidden. The report is now cached only on a 2xx, and the config endpoint reports whether the server actually holds a discovery for that agent, so the cache can only suppress while the server agrees. No access to the nodes is needed; affected nodes reappear within one poll cycle (~2.5 min) after the agents pick up the new script. - Permanent refusals do not loop. A 400, 413 or 422 means the payload itself is unacceptable,
so the refusal is recorded and the same bytes are not re-posted; fixing the file releases the
brake, because it is keyed to the content hash. 401 and 404 keep retrying — in this system
they mean a token rotation or an agent row briefly absent while it re-registers, and braking on
them would have re-created the very failure above. This matters beyond noise: 4xx and 5xx agent
calls are never sampled out of the request log, so a loop would write a row carrying the whole
keepalived.confevery cycle on every affected node. - The clear path had the same defect. When a node becomes managed the agent tells the server to drop the discovery; that too was judged by the exit code, so a rejected clear left a stale row offering a managed node for adoption, with nothing to ever retry it.
- Config Import now works on a freshly installed agent.
check_config_requestsuploads a node's livehaproxy.cfgwhen you ask for it. It was defined in the installer and in the self-upgrade daemon but not in the body a fresh install writes, and its call site is guarded bytype, so on such a node the feature was a silent no-op: the request was made and nothing ever arrived. Any agent that had self-upgraded at least once already had it, which is why it went unnoticed. A freshly installed agent now polls the pending-requests endpoint once per cycle, exactly as every upgraded agent already does — no node running today changes behaviour. - The keepalived.conf path is resolved deterministically. A pool may hold more than one cluster and the join was unordered, so the path handed to an agent could differ between polls whenever two clusters disagreed on it — the agent would inspect a file that is not there and the node would never appear, intermittently. A customised path now wins over the shipped default, then the lowest cluster id. With one cluster per pool, or when every cluster carries the default, the value is byte-identical to before.
Rollback: safe. No schema or data change; reverting restores the previous behaviour, in which a rejected discovery report is never retried and Config Import is absent on fresh installs.
Upgrade Notes — v1.11.0 (Unified request/response log)
Adds one new table and bumps SCHEMA_VERSION 11 → 12. The migration runs automatically on the
first backend start. No agent impact, no rendered-config change, no change to any existing API
shape or response body.
12, not 11. The feature was developed against v1.10.3, where
SCHEMA_VERSIONwas still 10, and proposed 11. v1.10.4 took 11 in the meantime. Shipping it as 11 would have been silently inert:run_all_migrations()returns early onapplied_version >= SCHEMA_VERSION, so every database already at 11 would have skipped the entire sequence and received neitherrequest_logsnor therequestlog.*permissions, while a fresh install would have received both. If you are coming from a pre-release build that recorded 11 for this feature, no action is needed: the bump to 12 re-runs the (idempotent) sequence and creates whatever is missing.
-
New table
request_logs(BIGSERIAL primary key, 9 indexes). Created empty and starts filling immediately. The shipped defaults keep 7 days of successful requests, 30 days of failed ones, and at most 500 000 rows — whichever limit is reached first. Change any of it in Settings → Request Log. -
Successful agent polls are NOT logged (
requestlog.capture_agent_success, default off). Failed agent calls always are. This is what keeps the table's size a function of operator activity rather than of how many nodes you run. Measured at 2 424 bytes/row on PostgreSQL 15 against the real schema and all nine indexes, with each agent issuing ~9 792 logged calls a day:fleet if successful polls were logged with the shipped default 20 nodes 196k rows/day, 453 MB/day, row cap in 61 h ~21k rows/day, 48 MB/day, cap in 24 d 200 nodes 2.0M rows/day, 4.5 GB/day, row cap in 6 h ~30k rows/day, 69 MB/day, cap in 17 d 500 nodes 4.9M rows/day, 11 GB/day, row cap in 2.5 h ~44k rows/day, 103 MB/day, cap in 11 d The row cap always holds, but it holds by deleting — so without this default the configured "7 days of successes, 30 days of failures" silently becomes a few hours of both, and the forensic record the feature exists for is evicted by polling noise. Turn it on temporarily when debugging a specific node, then turn it back off.
-
Runtime cost is measured, not estimated. The middleware adds 27.7 µs (p50) per request and 1.4 µs on an excluded path; redaction runs on the writer task, off the request path, at 18.8 µs per row. At 500 nodes that is 0.096 % of one core in total.
-
New permissions
requestlog.readandrequestlog.manage. Both are granted tosuper_adminandsecurity_admin;operatorgetsrequestlog.readonly;viewergets neither, because captured request/response bodies are a broader disclosure surface than the read-only configuration views a viewer is meant to have. Custom roles can be granted either from Users → Roles. -
⚠️ Built-in roles are re-seeded to their defaults. This is the pre-existing behaviour of every
SCHEMA_VERSIONbump, not something new in this release, but it bites here because this release bumps: the version gate re-runs the whole sequence andupdate_system_roles_to_enterprise_rbac()issues an unconditionalUPDATE roles SET … permissions = <defaults> WHERE name = …for the four built-in roles (super_admin,operator,security_admin,viewer). Any customisation you made to a built-in role is reverted. Roles you created yourself are untouched (the update matches on name). To preserve customisation, export withGET /api/rolesbefore upgrading and re-apply withPUT /api/roles/{id}, or move the customisation into a custom role. -
Bodies are captured, redacted and capped at 8 KB. Passwords, tokens, API keys, private-key PEMs, JWT-shaped values,
Authorization/Cookieheaders, ACME JWS payloads and DNS-provider credentials are never stored. Headers use an allowlist — anything not on it is dropped rather than saved. Review Settings → Request Log before enabling body capture in a regulated environment;capture_bodiescan be turned off while still recording who called what, with what result. -
Excluded by default: health checks, the API docs, the ACME HTTP-01 challenge endpoint (it returns
key_authorization), the agent heartbeat (the highest-volume POST in the system), static assets, and the log viewer's own endpoints. The list is editable, except the log viewer itself, which is a hard floor so the page cannot end up logging you reading it. -
Disk growth is the main operational consideration. On a busy install, in order of bluntness: leave
capture_agent_successoff (the single biggest lever on a large fleet), lowersample_rate(errors are always kept at 100 %), turn offcapture_get, turn offcapture_bodies, or shortensuccess_retention_days. -
operatorcan see agent rows. A caller holding onlyrequestlog.readsees their own inbound requests plus the fleet's, and nothing belonging to another user. Agent rows are included because an apply fails on the node and the node reports it over its own API key — scoping to own-rows only would have hidden the diagnosis from exactly the role the grant exists for. Anonymous traffic (failed logins and the usernames they carry, unauthenticated probes) is not agent traffic and remains visible only torequestlog.manage. -
New environment variables, all optional:
REQUEST_LOG_ENABLED(defaulttrue),REQUEST_LOG_QUEUE_MAX(2000 rows),REQUEST_LOG_QUEUE_MAX_BYTES(64 MiB),REQUEST_LOG_BATCH_SIZE(100),REQUEST_LOG_FLUSH_MS(500). See.env.template.REQUEST_LOG_QUEUE_MAX_BYTESis a hard memory ceiling per worker: the row count alone does not bound memory, becausemax_body_bytesis operator-editable up to 256 KB and a row can carry it twice — at the ceiling the default 2 000-row queue would hold ~1 GiB, which is the whole pod limit. Whichever limit binds first stops the queue. If you raiseREQUEST_LOG_QUEUE_MAX, raise this with it. -
GET /api/request-logs/statssink counters are per worker, labelled as such in the response. WithUVICORN_WORKERS > 1each process keeps its own queue and its own drop counter. -
Clearing the exclude-path list does not mean "log everything." An empty list falls back to the shipped defaults, so the log viewer and the raw-body agent heartbeat stay excluded; the UI now says so and shows what was actually applied.
-
To disable entirely: set
REQUEST_LOG_ENABLED=falsein the backend environment and restart — the middleware is then not registered at all and costs nothing, not even a settings lookup. Theenabledtoggle in Settings is the no-restart equivalent (it takes effect immediately). -
Default admin password is not reset by this bump; user seeding is guarded by an existence check, not an upsert.
-
Rollback: downgrade the backend image freely.
request_logsis purely additive and is simply ignored by v1.10.x. Drop the table manually if you want the space back:DROP TABLE IF EXISTS request_logs;
Upgrade Notes — v1.10.14 (a converged node keeps acknowledging)
Agent-script change, no schema change. No SCHEMA_VERSION bump. After deploying, sync the
Linux agent script from Agent Management and let the agents upgrade, or the fix does not
reach the nodes.
- Symptom: a VIP shows
SYNCING (0/n)with an empty Last ack even though every member node has the renderedkeepalived.confon disk, keepalived is running and the VIP is held. - Cause: the deploy report was sent only when the agent actually wrote the config. Once the node matched, it took the idempotency early return every cycle and never reported again, so any report lost in transit was never retried and the server's view stayed stale permanently.
- Fix: the agent re-asserts its state on the idempotent path as well. One request per node per poll cycle (~2.5 min); nothing is written and keepalived is not reloaded.
- Recovery is automatic. A VIP stuck at SYNCING converges on the first poll after the agents pick up the new script. No action on the nodes, no re-apply, no edit to force a rewrite.
- This is not new in 1.10.12. The gap dates from the original HA/VIP work; it only became visible when acknowledgements were dropped for an unrelated reason.
Rollback: safe. Reverting restores the previous behaviour, in which a lost acknowledgement is never recovered.
Upgrade Notes — v1.10.13 (deploy acknowledgements were dropped)
Backend only, no schema change. No SCHEMA_VERSION bump, no agent change. If you deployed
v1.10.12, deploy this one too.
- Regression in v1.10.12. The takeover-retirement clause added to the keepalived status
endpoint reused a query placeholder for both an assignment and a comparison. PostgreSQL types a
placeholder per use, so it was deduced as
textin one place andcharacter varyingin the other, and asyncpg refused the statement outright. - Symptom: a VIP stayed at
SYNCING (0/n)with an empty Last ack even though the agent log showedapplied config for VIP <id>on every member. Teardown acknowledgements were lost the same way, so a deletion never showed as complete. - Nothing was damaged. The failure was on the write of the acknowledgement, not on the node. Configs were deployed correctly throughout; only the reporting was lost. Existing VIPs converge on the next poll once this is deployed, with no action on the nodes.
- Verified against a real PostgreSQL, not by inspection: both statements execute, a matching
hash retires the takeover authorisation, a non-matching hash and a NULL
applied_config_hashleave it in place, and the acknowledgement is recorded in every case.
Rollback: do not roll back to v1.10.12; roll back to v1.10.11 instead, which predates the clause entirely.
Upgrade Notes — v1.10.12 (valid config rejected by its own warning)
Agent-script change, no schema change. No SCHEMA_VERSION bump. After deploying, sync the
Linux agent script from Agent Management and let the agents upgrade, or the fix does not
reach the nodes.
- Symptom: applying a VIP left it stuck at
SYNCING, the node kept its previous config and the agent logged onlyconfig validation failed (keepalived -t). - Cause: the agent treated any non-zero exit from
keepalived -tas invalid. keepalived's config-test exit code does not separate fatal from benign: on 2.2.8 a clean config exits 0, whileTruncating auth_pass to 8 charactersexits 5 and so do a missing}and an unknown keyword. A VRRP password longer than eight characters was enough to block every apply, even on a node whose own running config emits the same warning. - Fix: the gate judges the output instead. Known-benign messages are dropped and anything
left still fails, so it fails closed. Verified against real keepalived: the truncation warning
passes; a missing brace, an unknown keyword and a
SECURITY VIOLATIONare refused. - Also: the agent now reports what keepalived actually said, in its log and in the status the HA/VIP page shows. The refusal was correct but unactionable without reproducing it by hand.
- The fail-safe itself is unchanged: a config that genuinely fails validation is never written and keepalived is never restarted.
Rollback: safe. Reverting restores the stricter gate, which rejects valid configs whose password exceeds eight characters.
Upgrade Notes — v1.10.11 (Adoptable tag names the real blocker)
Frontend only, no schema change. No SCHEMA_VERSION bump, no API change, no agent change.
- The Adoptable tag and the disabled Adopt button were derived separately, so they could
name different problems. A pair blocked by a peer whose config could not be parsed showed
MASTER missing, because the unreadable node's
state MASTERhad not been counted — true, but it sent the operator to the wrong node. Both now come from one ordered decision. - New label blocked by peer for a group held up by a node that references the same address but cannot be taken over with it. Two MASTERs is now distinct from none.
- Display only. The endpoint's checks and refusals are unchanged.
Rollback: safe; purely presentational.
Upgrade Notes — v1.10.10 (Adoption blockers listed once per instance)
Frontend only, no schema change. No SCHEMA_VERSION bump, no API change, no agent change.
- The adoption panel merged every member's blocker list, so a two-node pair showed each shared
problem twice. The two files report different line numbers for the same directive, so exact
de-duplication did not collapse them. Blockers are now merged on the message with the leading
line N:ignored. - Display only. The endpoint already evaluated the combined set across all nodes, and what it accepts or refuses is unchanged.
Rollback: safe; purely presentational.
Upgrade Notes — v1.10.9 (Adoption refuses to strand a node)
Backend + frontend, no schema change. No SCHEMA_VERSION bump, so the built-in roles are
not re-seeded. No agent change.
- Adoption will not leave a node behind. v1.10.8 resolved the whole VRRP instance, but only
from nodes it could parse, that were enabled and that were in the same pool. Anything else fell
out of the set silently while its peers were rewritten. Adoption now refuses if any reported
keepalived.confmentions the virtual address and is not among the nodes being taken over, and says which node and why. - The nodes must agree on the shared fields.
prefix_length, unicast/multicast mode, HAProxy tracking and the VRRP password live on the VIP and are re-rendered onto every member, so one node's value used to be imposed on the rest. A disagreement is now refused with both values shown. - The takeover authorisation is now retired on acknowledgement. It is the permission to
overwrite a
keepalived.confthat does not carry our marker. It was never cleared, so it stayed valid for that exact file content indefinitely; restoring the pre-adoption file would have been overwritten again without fresh approval. It is now dropped once the member acks our rendered config, gated on the acked hash matchingapplied_config_hashso a failed deploy cannot strand the VIP. - Nothing to do on upgrade. Existing adopted VIPs keep working; their authorisation is retired on the next successful acknowledgement.
Rollback: safe. No schema or data migration; reverting restores the previous (more permissive) adoption checks.
Upgrade Notes — v1.10.8 (VIP adoption takes the whole VRRP instance)
Backend + frontend, no schema change. No SCHEMA_VERSION bump, so the built-in roles are
not re-seeded. No agent impact: nothing about what the agent reports or how it takes a
config over changes.
- Adoption is now per VRRP instance, not per node. Every node in the pool reporting the same
virtual_router_idand virtual address becomes a member of one VIP, each with the role, priority and interface its ownkeepalived.confdeclares, and each with its own one-shot takeover hash. The panel lists one row per instance. - Why this mattered: single-node adoption could not produce a working pair. The BACKUP alone failed apply, the MASTER alone left the peer unmanaged and the peer could not then be adopted (VRID collision). On a unicast instance it was worse than inconvenient: the render drops the unicast block when there are no peers, so the adopted node fell back to multicast while its peer stayed unicast and both could hold the address.
- New refusals, each with the reason in the message: the group does not have exactly one
MASTER; the nodes disagree on
advert_int; a declared unicast peer is not among the nodes being adopted; a node is already a member of a live VIP. - Apply Management "View Change" now renders the adopt diff correctly. It did not recognise
the
adoptaction and fell through to the generic HAProxy diff, which compared the stagedkeepalived.confagainst the cluster's previoushaproxy.cfgand showed the whole HAProxy config as removed. Alarming, but display-only — nothing was ever applied from that view. - Rejecting an adoption is recoverable again. It used to hide the node from the panel
permanently. Nothing clears
vip_discoveries.adopted_vip_id, a VIP is only soft-deleted so the column'sON DELETE SET NULLnever fires, and the agent does not re-report a file whose hash has not changed. Adoptability is now derived from whether the linked VIP is still active. - If you adopted a VIP on 1.10.4-1.10.7, check it before applying: it may have only one member. Add the peer from the VIP's edit form, or reject the pending adoption and adopt again — the node reappears in the panel under this release.
Rollback: safe. No schema or data change; reverting restores the previous single-node adoption behaviour.
Upgrade Notes — v1.10.7 (HA / VIP follows the selected cluster)
Backend + frontend, no schema change. No SCHEMA_VERSION bump, so the built-in roles are
not re-seeded. No agent impact.
- The HA / VIP page ignored the cluster picker. Both the VIP table and the Unmanaged
keepalived detected panel queried the whole fleet, so on an install with more than one
cluster the lists never changed when the selection did. Both now pass
cluster_id, mapped to the cluster's pool the same wayGET /api/vip?cluster_id=already worked for Apply Management. - Behaviour change worth knowing: the VIP table is now scoped to the selected cluster. It
used to show every VIP in the fleet. If you relied on the fleet-wide view, the API still
supports it —
GET /api/vipandGET /api/vip/discoverieswithoutcluster_idreturn everything, unchanged. - API compatibility:
cluster_idis optional on both endpoints. Existing integrations that do not send it behave exactly as before.
Rollback: safe. The change is a query parameter plus the page that sends it; reverting restores the fleet-wide lists and touches no data.
Upgrade Notes — v1.10.6 (VIP adoption panel was unreachable)
One backend fix, no schema change. No SCHEMA_VERSION bump, so the built-in roles are
not re-seeded. No API-shape change, no frontend change and zero agent impact.
- v1.10.4's adoption panel never appeared.
GET /discoverieswas declared afterGET /{vip_id}inrouters/vip.py. FastAPI matches routes in declaration order, so the discovery list was routed into the get-one-VIP handler, which declaresvip_id: intand answered 422 before the real handler ran. The HA/VIP page treats any non-OK response as "nothing to show", so the feature was invisible with no error in any log. - Nothing was lost. The agent side always worked: discoveries were reported and stored in
vip_discoveries. Deploy this backend and the rows appear immediately — no agent upgrade, no re-sync of the agent script, no re-report needed. - If you are upgrading straight from 1.10.3 or earlier, follow the v1.10.4 notes below as
well: that release does bump
SCHEMA_VERSION(10 → 11), which re-seeds the four built-in roles, and its agent script has to reach the nodes before discovery starts. - Regression guard. A static source scan now fails the build if any literal API path in any router is declared after a parameterised route that would swallow it. The whole router tree is clean as of this release.
Rollback: safe and immediate. The change is a route declaration order plus a test; reverting to 1.10.5 restores the previous (broken-panel) behaviour and touches no data.
Upgrade Notes — v1.10.5 (HTTP-01 challenge backend on split deployments)
Bug fixes, no schema change. No SCHEMA_VERSION bump, so the built-in roles are not
re-seeded. No API-shape change and zero agent impact.
- HTTP-01 could fail silently when HAProxy runs on different hosts than the management stack.
The rendered config wrote
server _acme_mgmt <mgmt>:8080from a value that defaults to loopback — and HAProxy resolves that address on the HAProxy node, so it pointed at the wrong box. Every check still reported success. The per-clusteracme_backend_urlnow has a UI field (Cluster Management), changing it actually mints a config version, and the value is validated at the write boundary. - A config-generation failure could be pushed to agents as the cluster's whole
haproxy.cfg. The generator reported failure by returning# Error ...instead of raising, and the apply path hashed that comment and stored it as an APPLIED version. Both persisting call sites now refuse with 422 and leave the running config in force. This is worth knowing even if you never touch ACME, since any exception in the generator could trigger it. frontends.modeis nullable and was interpolated raw, emitting a literalmode Nonethat HAProxy rejects — which fails the whole cluster config, not just that frontend. Normalised now.- Cluster creation ignored the ACME fields: a cluster created with ACME switched on came back switched off, with no error.
docker-compose.ymlhardcodedPUBLIC_URL/MANAGEMENT_BASE_URL, so a value in your.envor host environment was silently ignored. They are interpolated now, with the previous literals as defaults, so behaviour is unchanged unless you actually set them.- Diagnostics stop over-reporting health. The port-80 check now reads the body, so a reverse proxy answering 200 with a web page is no longer counted as a working challenge endpoint. Every new condition is a warning, never a failure — the Site Wizard blocks submit on a failing check, so a new failing condition would have locked installs on upgrade day.
- Rollback: downgrade freely. No schema or data change.
Upgrade Notes — v1.10.4 (Adopt an existing keepalived VIP)
Additive, but this release DOES bump the schema — read the role warning below. Nothing on any node changes until you adopt a VIP and apply it.
- Schema:
SCHEMA_VERSIONbumps to11, so on first start the (idempotent) migration sequence re-runs once and adds one new table (vip_discoveries) plus two additive columns (vip_instances.adopted_at,vip_members.takeover_expected_hash). No existing table is altered, no existing row changes, and the admin password is not reset. - ⚠️ Built-in roles are re-seeded to their defaults — the pre-existing behaviour of every
SCHEMA_VERSIONbump. If you customisedsuper_admin/operator/security_admin/viewer, re-apply those changes after upgrading. (The three previous releases did not bump the version, so this is the first re-seed since v1.9.0.) No new permission strings are introduced: discovery and adoption are governed by the existingvip.read/vip.create. - ⚠️ The Linux agent script changed, and discovery does not start until nodes run it. The
fallback latest Linux agent version moves
2.0.0→2.1.0, so nodes will pull the new script through the normal agent-upgrade path. The addition is read-only: the agent reads thekeepalived.confit does not own and reports it, rate-limited to once per content change. It writes nothing new to the node. Until a node has upgraded, it simply never appears under Unmanaged keepalived detected. - Nothing is taken over implicitly. The agent still refuses to overwrite a
keepalived.confthat lacks OpenManager's ownership marker. Adoption authorises exactly one takeover of exactly the file that was analysed, pinned to its md5: if the file changes between adoption and Apply, the agent refuses again and reportsexternally_managedrather than clobbering your edit. Re-adopt to pick up the current file. - Adoption can refuse, on purpose. It replaces the file with OpenManager's render, so anything
the renderer cannot reproduce would be destroyed. Those directives are listed as blockers —
notify_*failover hooks,vrrp_sync_group, LVSvirtual_serversections, a second address in one instance, a customtrack_script, extraglobal_defs. You can accept that loss explicitly with a tick, but a value that is unknown rather than lost (an absentvirtual_router_id, or an address with no prefix length) cannot be waived — the VRID is fatal to guess and the prefix has to be supplied, because picking a netmask for a live VIP would change its routing. - Multi-node VIPs need every node. Adoption covers the node that reported. Its unicast peers
hold their own
keepalived.conf, so adopt or add them as members before applying — otherwise the render has no peers. The UI says so after a successful adopt. - Secrets: the reported config may contain the VRRP
auth_pass. It is split at ingest — the password is Fernet-encrypted into its own column (same key path asvip_instances,VIP_ENCRYPTION_KEYfalling back to a key derived fromSECRET_KEY) and the stored copy of the file has it masked, so nothing readable through the API, the UI preview or a DB dump carries it in cleartext. - Rollback: downgrading to 1.10.3 leaves
vip_discoveriesas an unused table and the two new columns unread; managed VIPs keep working. One caveat: a VIP adopted on 1.10.4 but not yet applied loses its takeover authorisation on downgrade, so the node's original config stays in place and the VIP sits PENDING — harmless, but re-adopt after upgrading again. Agents already on script 2.1.0 keep reporting discoveries to an endpoint that no longer exists; the report fails quietly and nothing on the node is affected.
Upgrade Notes — v1.10.3 (Multi-account ACME wizard fix)
Frontend only. Nothing to do on upgrade. No schema, no SCHEMA_VERSION bump, no API change, no
environment variable, zero agent impact. Installations with a single ACME account behave exactly as
before.
- What was broken: with more than one ACME account registered, the Request ACME
Certificate wizard did not honour the account you selected. Choosing an HTTP-01 account still
submitted a DNS-01 request, which the API rejected with
The selected ACME account has no DNS provider configured for DNS-01.The Review step also named the default account rather than the chosen one, so the mismatch was invisible before submitting. - Default account: the wizard previously previewed the oldest valid account while the
backend uses the newest (
ORDER BY created_at DESC). If you never picked an account explicitly and have several, requests were already going to the newest one — only the preview was wrong. The wizard now previews that same account, marks it(default), and sendsaccount_idexplicitly so the two can no longer diverge. - Wildcard guard: the client-side "wildcard requires a DNS-01 account" block silently stopped applying on the Review step. Requests were still rejected by the backend, so nothing incorrect was ever issued — you now get the warning before submitting instead of an error after.
- No action needed on existing certificates or orders. Nothing about issuance, renewal or the stored accounts changes; only how the wizard resolves which account a new request uses.
- Rollback: downgrade freely. This release changes frontend behaviour only.
Upgrade Notes — v1.10.2 (Dark mode fixes on Apply Management)
Frontend only. Nothing to do on upgrade. No schema, no SCHEMA_VERSION bump, no API change,
no environment variable, zero agent impact. Light mode is byte-identical: every colour swapped in
this release resolves, under the default algorithm, to exactly the literal it replaced
(colorWarningBg → #fffbe6, colorSuccessBg → #f6ffed, colorErrorBg → #fff2f0, …), so
only dark mode changes.
- Apply Management panels were painted with light-mode colour literals, so in dark mode the "Pending Changes" box rendered as a cream panel with light text on it. Measured contrast was 1.03:1 — effectively invisible. It is now 11.50:1. The same class of bug affected the diff rows in View Change (2.21:1 and 2.99:1, now 5.49:1 and 4.01:1), the ACME/pending version panels, the VIP pending-delete row and the agent-error recommendation box.
- Static confirm dialogs came up white in dark mode. In Ant Design 5 the static
Modal.confirm/message/notificationAPIs render into their own detached root and never see the app'sConfigProvider, so they always used the light algorithm. This release registersConfigProvider.config({ holderRender })once at the app root, which fixes every static dialog in the app (12 components use them), not just Apply Management. - Rollback: downgrade freely. This release changes rendering only.
Upgrade Notes — v1.10.1 (CSR private key encrypted at rest)
Backward compatible. Nothing to do on upgrade, and nothing changes for existing clusters, agents or certificates:
- Schema: no
SCHEMA_VERSIONbump and no migration. The Fernet token replaces the PEM inside the existingssl_csrs.private_key_pemTEXT column. As in v1.10.0, this means the four built-in roles are not re-seeded, so any customization ofsuper_admin/operator/security_admin/viewersurvives. - Existing pending CSRs keep working. Rows written before this release hold a raw PEM and are still read transparently, so a CSR that is already out for signature can be imported normally after the upgrade. There is no data migration and no downtime step. Those rows stay plaintext until they are imported (which NULLs the key) — if you want everything encrypted immediately, delete and re-create any long-pending CSRs.
- Scope: this covers the PENDING CSR key only. It is the one key in the system that sits idle
for the whole signing window and is never transmitted.
ssl_certificates.private_key_contentand the ACME order keys are unchanged, because agents must receive those in plaintext on every poll. - Optional env:
CSR_ENCRYPTION_KEY(see.env.template). If unset, the key is derived fromSECRET_KEYvia HKDF with its own info string, so it is independent of the VIP, MFA and DNS provider keys.- ⚠️ Rotating
SECRET_KEYwhileCSR_ENCRYPTION_KEYis unset makes pending CSR keys unrecoverable. Import then fails with an explicit "delete this CSR and create a new one" error rather than a misleading key-mismatch. Set an explicitCSR_ENCRYPTION_KEYif you rotateSECRET_KEY. Certificates already imported are unaffected — their key lives on the certificate row.
- ⚠️ Rotating
- API / UI / agents: unchanged. No CSR endpoint ever returned the private key before or now, and nothing about the CSR tab changes.
- Rollback: the application downgrades cleanly — 1.10.0 starts normally against the same
database and every other feature is unaffected. The one casualty is a CSR created on 1.10.1
and still pending: 1.10.0 has no decrypt step, so it hands the Fernet token straight to the
key-pairing check. Measured on a real downgrade, the import then fails with
HTTP 500 — Could not verify the certificate/key pair: key parse failed (encrypted?); it does not silently pair the wrong key, and it does not corrupt anything. Import or delete CSRs created on 1.10.1 before downgrading. Certificates already imported are unaffected, since their key lives on the certificate row, and CSRs created before 1.10.1 are plaintext and still work.
Upgrade Notes — v1.10.0 (GoDaddy DNS-01 provider)
Backward compatible & additive. Nothing changes unless you select GoDaddy as an ACME account's DNS provider:
- Schema: no
SCHEMA_VERSIONbump. The GoDaddy credentials (API Key + Secret) are stored as two keys inside the existing encryptedletsencrypt_account_dns_credentials.credentials_encryptedblob — no new table, no new column, no migration. - ✅ Built-in roles are NOT re-seeded. The re-seed warning in the v1.9.0 notes below is
triggered by a
SCHEMA_VERSIONbump. This release does not bump it, so any customization you made tosuper_admin/operator/security_admin/viewersurvives untouched. - Permissions / API shape: unchanged.
GET /api/letsencrypt/dns-providerssimply returns one extra entry in itsprovidersarray; every request and response shape is identical, and the credential form is rendered from that schema, so there is no frontend behaviour change either. - Environment: no new variable. GoDaddy credentials use the same Fernet-at-rest path as
Cloudflare (
DNS_PROVIDER_ENCRYPTION_KEY, falling back to a key derived fromSECRET_KEY). - Agents: zero agent changes. DNS-01 is invisible to agents; an issued certificate follows the normal PENDING → Apply Management → agent pull pipeline exactly as before.
- Using it: the API Key must be a Production key from
developer.godaddy.com/keys(the first key that dashboard issues is an OTE/test key and is rejected), the zone must be in the same GoDaddy account, and that account needs at least one registered domain before GoDaddy permits DNS API access. A Personal Access Token also works — paste it as the Key and leave the Secret blank. Credentials are checked against the GoDaddy API before they are stored, so an invalid, OTE or ineligible key fails at save time. Note the check is a read: a Personal Access Token that hasdomains.domain:readbut notdomains.dns:updatesaves successfully and only fails at the first publish, with a 403 in the order timeline. - Rollback: don't select GoDaddy. Existing Manual and Cloudflare accounts and all HTTP-01
issuance are untouched. Downgrading after adopting GoDaddy is not a no-op: on 1.9.0
godaddyis not a known provider, so any account still set to it degrades to the manual-confirm path (in-flight DNS-01 orders wait for a confirmation nobody can give and expire after 48h, and renewals stop), and the cleanup sweep marks published TXT records cleaned without removing them. Before downgrading, switch affected accounts back to Manual or Cloudflare and let the reconcile sweep remove outstanding_acme-challengerecords first. The stored credential row itself is inert — an encrypted blob for an unknown provider.
Upgrade Notes — v1.9.0 (CSR creation)
Backward compatible & additive. Upgrading to v1.9.0 changes nothing for existing clusters/agents until you create a CSR:
- Schema:
SCHEMA_VERSIONbumps to10, so on first start the (idempotent) migration sequence re-runs once and adds one new table (ssl_csrs) plus its indexes. No existing table is altered, existing rows are untouched, and the admin password is not reset (the default-user seeding is guarded by an existence check, not an upsert). No new permission strings are introduced — all CSR endpoints are governed by the existingssl.create/ssl.read/ssl.deletepermissions. - ⚠️ Built-in roles are re-seeded to their defaults (pre-existing behaviour of
every
SCHEMA_VERSIONbump — verified in a v1.8.10 → v1.9.0 upgrade drill). Because the version gate re-runs the whole sequence,update_system_roles_to_enterprise_rbac()issues an unconditionalUPDATE roles SET … permissions = <defaults> WHERE name = …for the four built-in roles (super_admin,operator,security_admin,viewer). Any customization you made to a built-in role is reverted. In the drill, anoperatorrole that had been narrowed by removingapply.executeandconfig.bulk_importcame back with both restored (57 → 59 permissions).- Roles you created yourself are NOT affected — the re-seed matches on the four built-in names only.
- This is not new in v1.9.0: it happens on every release that bumps
SCHEMA_VERSION(v1.7.0, v1.8.0, v1.8.8 …). It is documented as intentional atbackend/database/migrations.py(the "BUMP THIS … OR seeded/role data" note) — the migration is treated as the authority on built-in-role contents. - If you have hardened a built-in role, do this: export it before upgrading
(
GET /api/roles), then re-apply your changes after the first start (PUT /api/roles/{id}) — or, preferably, move your customization into a purpose-made custom role, which survives every upgrade.
- Key storage: CSR private keys are stored in the database like every other key
in the system (
ssl_certificates.private_key_contentand the ACME order keys). The key is never returned by any CSR API endpoint, and after a successful import the CSR row's key copy is set to NULL (the key then lives only on the certificate row). - Agents: zero agent changes. Agents never read the new table; a CSR becomes visible to agents only after its signed certificate is imported and applied via Apply Management (the standard PENDING pipeline).
- Rollback: simply don't use the CSR tab. The
ssl_csrstable is inert when empty; downgrading the application leaves it as an ignored extra table.
Upgrade Notes — v1.7.0 (HA / VIP Keepalived management, Issue #27)
Backward compatible & opt-in. Upgrading to v1.7.0 changes nothing for existing clusters/agents until you create a VIP:
- Schema:
SCHEMA_VERSIONbumps to3, so on first start the (idempotent) migration sequence re-runs once and adds two new tables (vip_instances,vip_members) plus an additivevip_instances.applied_snapshotcolumn (enables rejecting a pending VIP change and restoring the previous applied state). No existing table is altered. Existing rows and the admin password are not reset (default users are create-if-missing). The only data effect is that the four built-in system roles (super_admin/operator/security_admin/viewer) are re-seeded to their canonical permission sets plus the newvip.*permissions — this is the long-standing behavior of the role seeder; custom roles are untouched. - Agents: the agent script gains an opt-in keepalived deploy that is a no-op on
any node without an applied VIP, and it never overwrites a hand-managed
/etc/keepalived/keepalived.conf(it reports "externally managed" instead). - Scope: VRRP VIPs target bare-metal / VMware / on-prem L2 networks. On AWS/Azure/GCP the cloud fabric doesn't honor VRRP/gratuitous-ARP; the UI surfaces this. Ensure host firewalls permit VRRP (IP protocol 112).
- Optional env:
VIP_ENCRYPTION_KEY(see.env.template) — if unset, the VRRP secret encryption key is derived fromSECRET_KEY(like MFA).
No rollback steps are required to disable the feature: simply don't create VIPs (or delete them — agents tear down their managed keepalived on the next poll).
Agent Upgrade Guide - Dashboard Stats Fix
Problem
Dashboard showing 0 metrics after agent auto-upgrade due to missing environment variables (SOCAT_BIN, STATS_SOCKET_PATH) when agent restarts in daemon mode.
Solution
Implemented lazy initialization in both get_haproxy_stats_csv() and get_server_statuses() functions. These functions now initialize their dependencies on first call, making them completely independent of global variable initialization.
Deployment Steps
Step 1: Wait for Pipeline ⏳
# Pipeline is currently running after git push
# Check status: https://[your-azure-devops]/pipeline
# Wait for deployment to complete (~3-5 minutes)
Step 2: Verify Backend Deployment ✅
# Check backend logs for successful deployment
kubectl logs -f deployment/backend -n haproxy-manager | head -20
# Expected: New pod started with latest code
Step 3: Update Agent Version in UI 🔄
Option A: Automatic Script Sync (Recommended)
# Get admin token
TOKEN=$(curl -k -s -X POST "https://haproxy-manager.example.com/api/auth/login" \
-H "Content-Type: application/json" \
-d '{"username": "admin", "password": "admin123"}' | jq -r '.access_token')
# Sync scripts from files to database (creates version 1.0.0)
curl -k -X POST "https://haproxy-manager.example.com/api/agents/sync-scripts-from-files" \
-H "Authorization: Bearer $TOKEN" | jq .
# Expected: {"status": "success", "synced": ["linux", "macos"]}
Option B: Manual UI Update
- Go to Agent Management page
- Click Settings → Agent Versions
- For Linux platform:
- Click Edit Script
- Version: Keep current or increment (e.g., 1.0.3)
- Changelog: Add "Fixed stats collection after upgrade"
- Click Save (this syncs file content to database)
- Repeat for macOS platform
Step 4: Upgrade Agents 🚀
Option A: UI (Single Agent)
- Go to Agent Management page
- Select agent (e.g.,
demo-agent) - Click Upgrade button
- Wait 30 seconds for agent to restart
Option B: API (Batch Upgrade)
# Get all agents
curl -k -s -X GET "https://haproxy-manager.example.com/api/agents" \
-H "Authorization: Bearer $TOKEN" | jq '.agents[] | {id, name, version}'
# Upgrade specific agent (replace {agent_id})
curl -k -X POST "https://haproxy-manager.example.com/api/agents/{agent_id}/upgrade" \
-H "Authorization: Bearer $TOKEN" | jq .
Step 5: Verify Stats Collection 📊
Backend Logs (30 seconds after upgrade):
kubectl logs -f deployment/backend -n haproxy-manager | grep -E "haproxy_stats_csv|demo-agent"
# Expected logs:
# ✅ "Has haproxy_stats_csv: True"
# ✅ "CSV preview: IyBweG..." (base64 data)
# ✅ "STATS: Parsed 15 rows for cluster demo-cluster1"
Agent Logs (on agent server):
sudo tail -f /var/log/haproxy-agent/agent.log | grep STATS
# Expected logs:
# ✅ "STATS: Initialized socat: /usr/bin/socat"
# ✅ "STATS: Using default socket path: /var/run/haproxy/admin.sock"
# ✅ "STATS: Fetched CSV: 2121 bytes, 9 lines, base64: 2828 chars"
Dashboard UI:
- Open Dashboard page
- Refresh page (F5)
- Check metrics:
- ✅ Frontend/Backend filters populated
- ✅ Overview metrics showing real data (not 0)
- ✅ Charts showing data points
- ✅ "Waiting for Agent Data" warning gone
Troubleshooting
Issue: Backend still shows "Has haproxy_stats_csv: False"
Cause: Agent hasn't upgraded yet or using old script version
Solution:
# Check agent version on agent server
grep "AGENT_VERSION" /usr/local/bin/haproxy-agent | head -1
# Force agent restart
sudo systemctl restart haproxy-agent
# Check logs immediately
sudo tail -20 /var/log/haproxy-agent/agent.log
Issue: "STATS: socat not available"
Cause: socat not installed
Solution:
# Install socat
sudo yum install -y socat # RHEL/CentOS
sudo apt install -y socat # Debian/Ubuntu
# Restart agent
sudo systemctl restart haproxy-agent
Issue: "STATS: Socket not found: /var/run/haproxy/admin.sock"
Cause: HAProxy stats socket not configured
Solution:
# Check HAProxy config for stats socket
grep "stats socket" /etc/haproxy/haproxy.cfg
# Add if missing (in global section):
# stats socket /var/run/haproxy/admin.sock mode 666 level admin
# Reload HAProxy
sudo systemctl reload haproxy
Verification Checklist
- Pipeline completed successfully
- Backend pod restarted with new code
- Agent scripts synced to database (version visible in UI)
- Agents upgraded to new version
- Backend logs show "Has haproxy_stats_csv: True"
- Agent logs show STATS messages
- Dashboard showing real metrics
- Charts populated with data
- No "Waiting for Agent Data" warning
Next Upgrades
This fix is permanent. Future agent upgrades will NOT break stats collection because:
- ✅ Functions are self-contained with lazy initialization
- ✅ No dependency on global variable initialization order
- ✅ Works in any restart scenario (systemd, daemon mode, upgrade)
- ✅ Backward compatible with existing agents
No manual intervention needed for future upgrades! 🎉