mirror of
https://github.com/taylanbakircioglu/haproxy-openmanager.git
synced 2026-09-16 15:45:11 +00:00
v1.10.11
276 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9e64002f1c |
fix(vip): make the Adoptable tag name the real blocker (v1.10.11)
The tag and the disabled Adopt button were computed by two separate ladders and could disagree. Seen on a live pair: one node's keepalived.conf had an unbalanced brace, so it was excluded from the instance; its partner was then tagged "MASTER missing" - technically true, because the unreadable node's state MASTER had not been counted - while the actual reason (the peer cannot be taken over, so adopting would strand it) sat only in the button's tooltip. The label pointed the operator at the wrong node. groupState now makes ONE ordered decision and returns the label, its colour and the reason together, so the tag can never describe a different condition than the one disabling the button. A group held up by a node that references the same address but cannot be adopted with it reads "blocked by peer"; two MASTERs is distinguished from none. Display only: the endpoint's checks and refusals are untouched. Backend suite: 1366 passed, 152 skipped. Frontend build clean.v1.10.11 |
||
|
|
4f24d5bdd9 |
fix(vip): list adoption blockers once per instance (v1.10.10)
Since the panel groups a VRRP instance into one row, the blocker lists of all its members are merged - and every node reports the SAME problems about the SAME shared config. A two-node pair therefore showed each issue twice, in the Adoptable tooltip and in the adopt dialog. Plain de-duplication does not collapse them because the two files report different line numbers for the same directive. mergeBlockers keys on the message with a leading "line N:" stripped and keeps the first occurrence, so each distinct problem appears once while the text the operator reads still carries a line reference. Display only: the endpoint already evaluated the combined set across every node, and what it accepts or refuses is unchanged. Backend suite: 1366 passed, 152 skipped. Frontend build clean. |
||
|
|
a87994e06a |
fix(vip): refuse adoption that strands a node or normalises a peer (v1.10.9)
Three findings from a second pass over the adoption flow, all of the same class: something real leaving the set silently. 1. STRANDING. _collect_instance_participants can only match a node it can READ, that is ENABLED, and that is in the SAME pool. Each of those is a door a genuine member of the VRRP group leaves through without a word, and the nodes that remain are rewritten while it keeps serving the same address from an unmanaged config. Found on a live pool: one node of a pair had an unclosed vrrp_instance block, so it parsed to nothing while its partner parsed cleanly. Rather than guard each door, ask the question directly: does any reported keepalived.conf mention THIS virtual address without being one of the nodes we are about to adopt? Refuses naming the node and the reason. Scoped on the address so an unrelated file elsewhere cannot block every adoption, and excluding nodes already under management (a standing VIP, or our ownership marker) because those are not stranded. 2. SILENT NORMALISATION. prefix_length, unicast/multicast mode, HAProxy tracking and the VRRP password are stored ONCE on the VIP and re-rendered onto EVERY member, so whichever node was clicked imposed its settings on the others. prefix_length is the sharpest: the design refuses to GUESS a netmask for a live VIP, and copying one node's netmask onto another is that same change wearing a different hat. All four must now agree, with both values named in the refusal. The VRRP secret is compared by decrypting each node's token - Fernet is non-deterministic, so ciphertexts cannot be compared - and a token that will not decrypt is an error rather than an assumed match. 3. THE TAKEOVER AUTHORISATION WAS NOT ONE-SHOT. takeover_expected_hash is the permission to overwrite a keepalived.conf that lacks our ownership marker. It was written at adoption and never cleared, so it stayed valid for that file content indefinitely: restoring the pre-adoption file would have been overwritten again with no fresh human approval. It is now retired when a member acknowledges our rendered config, gated on the acked hash matching applied_config_hash so a partial or failed deploy never drops it and leaves the VIP unable to converge. The panel applies the stranding rule too, so the Adopt button is disabled with the reason instead of letting the operator click into a 422. No schema change, no agent change, no API-shape break. Backend suite: 1366 passed, 152 skipped. Frontend build clean. |
||
|
|
7d95c737f0 |
fix(vip): adopt the whole VRRP instance, not one node (v1.10.8)
Four defects found while tracing the adoption flow end to end after v1.10.4
reached a live HA pair.
B3/B4 (one root, one fix). Adoption took only the node whose row was clicked:
- adopting the BACKUP alone produced a VIP that apply always rejects, because
apply requires exactly one MASTER member;
- adopting the MASTER alone left the peer unmanaged, and adopting it
afterwards hit the VRID-collision guard with 409, so a pair could never be
completed from the panel;
- on a UNICAST instance the single-member render dropped the unicast block
entirely (render_keepalived_conf emits it only when peer_ips is non-empty),
so keepalived fell back to multicast on the adopted node while its peer
stayed unicast. They stop seeing each other and BOTH claim the VIP.
Adoption now resolves the whole instance via _collect_instance_participants,
keyed on (virtual_router_id, virtual address) - the same key keepalived uses to
group nodes. Each participant becomes a member with the role, priority and
interface its own file declares, and its own one-shot takeover hash, so the
per-node overwrite guard is unchanged. It refuses, naming the reason, when the
group has other than one MASTER, when advert_int differs across nodes, when a
declared unicast peer is not among the nodes being adopted, or when a node is
already in a live VIP - the one-active-VIP-per-agent rule that create/update
enforce via _validate_members_against_pool and adoption never called.
B1. The Apply Management "View Change" regex matched vip-(create|update|delete)
only, so an adopt version fell through to the generic HAProxy diff and rendered
the cluster's whole haproxy.cfg as removed. Display-only, but alarming. A test
now asserts every action _stage_vip_version can stage is in that alternation.
B2. Rejecting an adoption hid the node from the panel permanently:
vip_discoveries.adopted_vip_id is write-once, a VIP is only ever soft-deleted so
the column's ON DELETE SET NULL never fires, and the agent does not re-report a
file whose hash has not changed. Rather than clearing the column on each path,
adoptability is derived from whether the linked VIP is still active, which
self-heals reject, undo-reject and approved teardown alike.
The panel now lists one row per instance instead of per node, and the adopt
dialog names every node that will be taken over. Blockers are aggregated across
all of them, matching what the endpoint checks.
No schema change, no agent change, no API-shape break: /api/vip/discoveries
gains a derived field and /api/vip/adopt keeps its request body.
Backend suite: 1359 passed, 152 skipped. Frontend build clean (no new lint
warnings in VIPManagement.js).
|
||
|
|
eda7f36c93 |
fix(vip): scope the HA/VIP page to the selected cluster (v1.10.7)
The page ignored the cluster picker in the header. Both the VIP table and the v1.10.4 "Unmanaged keepalived detected" panel queried the whole fleet, so on an install with more than one cluster the lists showed every cluster's nodes at once and did not change when the selection did - the panel looked stuck on one cluster's keepalived. - GET /api/vip/discoveries takes the same optional cluster_id the VIP list already took, mapped to the cluster's pool via haproxy_clusters exactly like list_vips does. Omitting it still returns the whole fleet, so no existing caller changes behaviour. - VIPManagement reads selectedCluster from ClusterContext (it only took the cluster list before) and sends cluster_id on both fetches. The fetch callbacks depend on the scope, so switching cluster refetches instead of showing stale rows. - tests/test_vip_discoveries_scope.py pins both defects this endpoint has had: that the route reaches its own handler rather than being parsed as a vip_id (asserting != 422 specifically, since the repo's generic endpoint-auth tests accept 422 alongside 401/403 and therefore could not catch it), and that the cluster filter resolves cluster -> pool and short-circuits when absent. Behaviour change worth calling out: the VIP table is now scoped to the selected cluster where it was fleet-wide before. The API still serves the fleet-wide view to any caller that omits cluster_id. Backend suite: 1345 passed, 152 skipped. Frontend production build clean. |
||
|
|
11e5bf57d9 |
fix(vip): make the adoption panel reachable again (v1.10.6)
Follow-up to #58. The "Unmanaged keepalived detected" panel it shipped never appeared on any deployment. GET /discoveries was declared after GET /{vip_id} in routers/vip.py, and FastAPI matches routes in declaration order, so every request for the discovery list was routed into get_vip, which takes vip_id: int and rejected "discoveries" with 422 before list_vip_discoveries ever ran. The failure was completely silent. The agents reported their configs correctly, the rows landed in vip_discoveries, and the HA/VIP page treats any non-OK response as "nothing to show" - so the feature was invisible with no error in any log. Confirmed against a real fleet: two discovery rows present in the database, one with a parsed candidate, and an empty panel. - move list_vip_discoveries above the /{vip_id} routes, with a comment stating the ordering requirement - add tests/test_router_path_shadowing.py: a static source scan that fails if any literal path in any router is declared after a parameterised route that would swallow it. The whole router tree is clean; the detector is itself tested against the pre-fix ordering so the guard cannot pass vacuously. POST /adopt is unaffected: no POST /{vip_id} route exists. Also carries the release metadata for #58 (v1.10.4) and #60 (v1.10.5), which are published together with this fix rather than as separate artifacts. No schema, API-shape, frontend or agent change. Data reported under the earlier code is not lost - existing rows show up as soon as this backend is deployed, with no agent action needed. |
||
|
|
a4c74f2a52 |
Merge pull request #60 from mustafaulukaya/fix/acme-http01-split-deployment
HTTP-01 challenge backend on split deployments |
||
|
|
a36dd87a74 |
fix(agent): port VIP adoption into the in-script daemon so self-upgraded agents get it
Found during an impact analysis of the agent-script change in PR #58. linux_install.sh contains TWO daemon implementations and which one runs depends on how the agent reached its current state: - Fresh install: the heredoc at lines 923-2646 is written to /usr/local/bin/haproxy-agent and systemd runs that file. - Self-upgrade: perform_agent_upgrade copies the downloaded INSTALLER script over that same path (`cp "$temp_script" "$current_script"` with current_script=/usr/local/bin/haproxy-agent). systemd then runs the installer with `daemon` + SKIP_TO_DAEMON=true, which takes the separate in-script daemon that lives after the heredoc. PR #58 added _kp_discover and the one-shot takeover only to the heredoc copy, so both were absent from the path that self-upgraded agents actually run — and self-upgrade is exactly the path the release notes tell operators to rely on ("nodes will pull the new script through the normal agent-upgrade path"). The feature would have worked on a freshly installed node and been silently inert on every upgraded one. The file's own banner warns about this ("check_agent_upgrade() - Multiple locations ... TIP: Search for function name to find all occurrences!"), and send_heartbeat / fetch_and_deploy_keepalived_config / get_haproxy_stats_csv are already maintained as parallel copies for the same reason. Verified empirically rather than by reading: the script was instrumented and run in a container exactly as systemd invokes it after an upgrade (SKIP_TO_DAEMON=true, `bash linux_install.sh daemon`), inspecting the live `declare -f fetch_and_deploy_keepalived_config`. before: DISCOVERY_YOK TAKEOVER_YOK after: DISCOVERY_VAR TAKEOVER_VAR ENDPOINT_VAR Both blocks are ported verbatim from the heredoc copy with indentation adjusted; no logic changed, so the guard semantics are identical in both paths — takeover still requires allow_takeover AND a non-empty expected hash AND a matching on-disk md5, and anything else falls through to the existing "externally managed — refusing to overwrite" branch. `bash -n` passes. Backend suite unchanged at 1263 passed. |
||
|
|
2f125da043 |
Merge pull request #58 from mustafaulukaya/feature/keepalived-vip-adoption
Adopt an existing keepalived VIP (Issue #27 follow-up) |
||
|
|
9e6e4dd03b |
fix(acme): repair the diagnostics contract the GET probe broke
Running the suite properly — the previous rounds could not, pytest was not installed locally — surfaced four failures, all in the code this branch touches. main is clean at 1410 passed, so these were mine. Two were real defects, not stale expectations: check_port80 turned an unreadable body into a hard failure. The GET probe read the body outside any guard, so a connection reset mid-response, or a server that hangs after headers, made a healthy 404 fail. The body is EVIDENCE, not a precondition: when it cannot be read the check now falls back to the status-only semantics it has always had, and records body_class 'unread'. The stricter rule applies only when there is something to judge. check_routing raised KeyError on a row without `acme_enabled`. The column gates a warning, so a row shape lacking it should not take down the whole diagnostic. The rest were expectations that had to change, because the contract did: - The probe is a GET now, so the test doubles needed a body. `_FakeHEADResp` became `_FakeGETResp` with headers and a readable content stream. - "A port-80 frontend row exists" no longer means ok. That assertion is exactly the bug: it describes what the database wants, while the nodes run whatever was last applied — which is how this check reported success throughout an incident where the live config had no usable challenge route. New coverage for the branches that had none, including the case that started all of this: a proxy that has lost its /.well-known/acme-challenge/ location serves its SPA with HTTP 200, which the old `status in (200, 404)` rule accepted as healthy. Also the tcp-only cluster, which must warn rather than fail — SiteWizard blocks submit on any failing check, so failing there would lock those installs on upgrade day. Verified against real dependencies rather than a stub harness: the app boots with the challenge route registered; the field validator fires through pydantic on both cluster models, normalising whitespace and rejecting scheme-less and loopback values; `model_fields_set` really does distinguish an explicit null from an omitted key, which is what makes clearing the field work; the settings validator rejects both raw and jsonb-encoded bad values; and a legacy scheme-less value is skipped so the next source renders. Suite: 1477 passed on the branch, 1410 on main, 0 failed on either. |
||
|
|
47cc79dcf7 |
fix(acme): address review findings on the challenge-backend hardening
Five confirmed findings from an adversarial review of the branch, four of them regressions introduced by it. Fall through to the next source when a stored URL cannot be resolved. Stopping at the first non-empty candidate emitted a backend section with no `server` line: the section exists so `haproxy -c` passes and Apply succeeds, then every challenge request 503s from an empty backend with nothing to show for it. Scheme-less values are common — the settings field was free text until this branch — so this was reachable on real installs. Selection moved into `select_acme_backend_source()` so it is testable and the skipped candidates are logged rather than silently dropped. Report a challenge backend with no server line. `extract_acme_backend_target` returns None for that section, and the loopback filter skipped falsy targets, so the case above would have been reported as "challenge route present in applied config" — the new check confirming the very state it exists to catch. Do not narrow the row set feeding the routing check's `fail` branch. Adding a mode filter to the WHERE clause turned a tcp-only port-80 cluster from "ok" into "fail", and the site wizard blocks submit on any failing check, so those installs would have been locked on upgrade day. Mode is now examined in Python and only downgrades to `warn`, using an expression that is character-for- character the renderer's normalisation. Match the agent's config selector. The applied-config lookup omitted `is_active = TRUE`, so it could read a superseded row and report on a config the nodes never received. Extraction now happens in SQL rather than pulling whole configs — these run to hundreds of KB. Select `acme_backend_url` when loading the existing cluster. It was absent, so the entity snapshot recorded old_values as NULL unconditionally and rejecting the pending version wiped the operator's per-cluster URL back to the global loopback default — re-creating the exact failure this branch removes. Also carry `acme_enabled` and `acme_backend_url` through cluster creation. The create model declared neither and the INSERT wrote neither, so a cluster created with ACME switched on came back switched off with no error shown. Refuted and deliberately not changed: settings PUT re-validating a stored loopback value (it validates only what is submitted), an apply-path connection leak (the 422 propagates to a handler that closes it), and the modal discarding backend rejection reasons (the envelope matches). |
||
|
|
bb774141d4 |
fix(acme): make the challenge backend fixable from the panel
Correcting a wrong ACME challenge backend was impossible without a shell, and
even with one the correction did not reach the nodes.
The mint gate only fired when `acme_enabled` flipped. `acme_backend_url` was
written to the DB and minted nothing, so Apply answered "No pending changes to
apply" and the nodes kept the old address forever. It is now decided by
comparing the rendered `server _acme_mgmt` line against the active version —
the one line that answers "would the nodes talk to a different address?".
Comparing whole configs would flag every unrelated pending edit.
The field had no UI at all. Added to the cluster form with validation that
mirrors the backend rules, and keyed on `model_fields_set` so clearing it
reverts to the global setting — with a plain `is not None` test an empty box is
indistinguishable from "not submitted", so a value could never be removed.
Validation is asymmetric on purpose (utils/acme_backend_url):
- at the write boundary, reject what cannot express a reachable target —
including the two silent traps: a scheme-less value became `localhost`, and
an out-of-range port raised inside the generator and destroyed the config
- at render time, never reject. The shipped defaults are themselves loopback,
so refusing to render would make every acme_enabled cluster unappliable,
including for changes unrelated to ACME. Problems are logged and surfaced.
The port-less default stays 8080 rather than moving to HTTP's 80: the bundled
compose publishes nginx on 8080, so installs relying on it work today and the
first sign of breaking them would be the unattended renewal loop months later.
The omission is warned about instead.
RFC1918 is allowed and is usually the right answer here, and no DNS resolution
is performed — both deliberate departures from utils/ssrf_guard, whose policy
is the opposite of what this address needs. What the management host can
resolve says nothing about what the HAProxy node can reach.
Diagnostics stop reporting success on a dead path:
- check_port80 uses GET instead of HEAD and classifies the body. A proxy that
has lost its /.well-known/acme-challenge/ location serves its SPA with HTTP
200, which `status in (200, 404)` accepted as healthy. Warnings also surface
when other domains pass, which previously hid the most diagnostic outcome.
- check_routing filters `mode`, joins `acme_enabled` and reads the APPLIED
config instead of counting database rows, and reports a loopback target.
- every new condition is `warn`, never `fail`: the site wizard blocks submit on
any fail, so a new failing condition would lock every install on upgrade day.
Also: normalise `frontends.mode` once per frontend. It is nullable, and the
raw value was interpolated into `mode {}`, emitting a literal `mode None` that
HAProxy rejects — taking down the whole cluster config. The ACME gate and the
backend-mode check now read the same normalised value.
And stop hardcoding PUBLIC_URL / MANAGEMENT_BASE_URL in docker-compose, which
silently ignored the operator's .env and made the wrong default load-bearing.
|
||
|
|
acfd32dd63 |
fix(acme): stop shipping config-generation failures as applied config
`generate_haproxy_config_for_cluster` reports failure by RETURNING a one-line
comment ("# Error generating configuration: ...") rather than raising. Nothing
in the backend checked for it, so the apply path hashed that comment, stored it
as an APPLIED config version and pushed it to every agent — silently replacing
a cluster's entire haproxy.cfg.
Any exception inside the generator triggers this. An out-of-range port in
`acme_backend_url` is enough: `urlparse('http://h:99999').port` raises
ValueError, the outer `except` swallows it, and the cluster loses its config.
- add `is_config_generation_error()` next to the sentinel definitions so call
sites stop matching the string by hand
- guard the apply path: refuse with 422 and leave the running config in force
- guard the ACME-toggle PENDING mint, and let HTTPException through the local
`except Exception`, which would otherwise report success while no pending
version exists
Also make the challenge backend diagnosable without shell access, since this
block previously emitted no log line at all:
- log the rendered `host:port` and WHICH source chose it (cluster override,
system setting, or the MANAGEMENT_BASE_URL fallback) under the greppable
`ACME-BACKEND` keyword
- warn when the rendered address is loopback: HAProxy resolves it on the node,
not on the management host, so it can only work on an all-in-one install —
and it is exactly what the shipped defaults produce
- log peer, X-Forwarded-For and Host on the challenge endpoint, which answers
"did the request arrive at all?" — the question that separates a wrong
backend address from a wrong response body
Refs the HTTP-01 investigation: a split deployment rendered
`server _acme_mgmt <mgmt>:8080` against a port with no listener, and every
existing check reported success.
|
||
|
|
78af849fdc |
docs(v1.10.4): document VIP adoption and its upgrade caveats
Release note covering why the page was empty, why the heartbeat could not drive adoption, and how the blocker gate decides what may and may not be waived. The upgrade notes lead with the two things an operator has to act on rather than burying them. This release bumps SCHEMA_VERSION, which re-seeds the four built-in roles - the first re-seed since v1.9.0, because the three releases in between did not bump it - so role customizations have to be re-applied. And the Linux agent script changed, so discovery does not start until nodes pull it; until then a node simply never appears in the list. Also states the parts that are easy to get wrong: nothing is taken over implicitly, adoption can refuse on purpose and why, a multi-node VIP needs every peer adopted before applying, where the VRRP password lives, and what actually happens on a downgrade (the table goes unread, an adopted-but-unapplied VIP loses its takeover authorisation and the node keeps its original config). |
||
|
|
709817fec3 | chore(version): bump to 1.10.4 - adopt an existing keepalived VIP | ||
|
|
822c441d34 |
test(vip): cover the adoption gate and the VRRP password masking
The gate decides whether an operator's working keepalived.conf gets replaced, so its rule is pinned directly: a supplied prefix resolves only the prefix blocker, accepting data loss resolves only the loss blocker, and setting both still cannot wave through an impossibility like an unknown VRID or an unsupported auth_type. The gate matches on substrings of the blocker prose the UI displays, which means a reworded message would silently stop being waivable. One test therefore feeds real parser output through it in both directions rather than hand-written strings, so the message and the rule are checked together. Also asserts that masking leaves no trace of a password containing spaces while keeping the rest of the config readable, using the router's own regex so the test breaks if it is ever loosened. |
||
|
|
1ca811e211 |
feat(ui): surface unmanaged keepalived on HA/VIP with an Adopt flow
The page came up empty on a fleet that already runs keepalived, with nothing to explain why. It now lists the nodes whose keepalived.conf the agent found and deliberately left alone, in a section separate from managed VIPs so the distinction is visible: OpenManager is not managing these. Each discovered vrrp_instance shows the address, VRID, and this node's own role, priority and interface, plus whether it can be adopted. Adopt opens a modal that states what will happen rather than just asking for confirmation: which directives would be deleted on takeover (with an explicit tick to accept that, disabled otherwise), which values were assumed from keepalived's documented defaults rather than read from the file, and the config itself with the VRRP password masked. Blockers are split the same way the backend splits them, from the same marker strings, so the button cannot offer an adoption the API would reject: a loss is waivable with the tick, a missing prefix is resolvable by supplying it, and an impossibility disables Adopt outright with the reason shown. After a successful adopt the peers of the adopted instance are named, because those nodes hold their own keepalived.conf and the VIP is not a complete VRRP group until they are members too. |
||
|
|
164841219a |
feat(agent): report an unmanaged keepalived.conf and honour a one-shot takeover
Two additions to the Linux agent, both inside the existing keepalived converge function so no new poll or timer is introduced. Discovery is strictly read-only: when the node has a keepalived.conf without OpenManager's ownership marker, the agent posts it so an existing VIP can be adopted from the UI. Nothing is written to the node. It is rate-limited by content - the md5 of the last report is cached next to the config, so a file that may carry the VRRP password is posted only when it actually changes rather than every cycle. Once we own the file there is nothing left to adopt, so the record is cleared exactly once. The content is JSON-encoded with `jq -Rs` so newlines survive verbatim and the hash the server pins the takeover to is the hash of what is really on disk. The ownership guard now has exactly one exception, and it does not weaken it. Previously any file without the marker was refused, which is what protects a hand-maintained setup - and also what would block adoption forever. The server authorises a single takeover of a specific file by sending the md5 the operator adopted from, and the agent overwrites only when the on-disk hash still matches. If the file changed in between, the agent refuses again and reports why, so an edit made after adoption wins over the stale adoption instead of being destroyed. The fallback latest Linux agent version moves 2.0.0 -> 2.1.0 so nodes pull the new script through the normal upgrade path. Discovery simply does not happen on a node that has not upgraded yet. |
||
|
|
d92a7e9660 |
feat(vip): adopt a discovered keepalived instance into a managed VIP
GET /api/vip/discoveries lists what the agents found; POST /api/vip/adopt turns
one vrrp_instance into a managed VIP using the values from the node's own file
instead of retyping them. The VIP is created PENDING like any other, so nothing
reaches the node until it is applied from Apply Management.
Adoption replaces the operator's file with our render, so the gate is the
feature. Blockers fall into three kinds and only two are resolvable:
- a LOSS ("our renderer cannot reproduce this, so adopting would delete it")
can be accepted explicitly - that is an informed choice about a notify hook
or an LVS section;
- an UNKNOWN prefix length can be supplied, because picking a netmask for a
live VIP would change its routing;
- anything else is an IMPOSSIBILITY, not a loss: an absent virtual_router_id,
a fractional advert_int, an unsupported auth_type. No flag waves those
through.
That rule now lives in one place, remaining_blockers(), so the endpoint and the
UI cannot drift apart - and it is unit-testable, which matters because getting
it wrong destroys a working config.
Adoption keeps the VRRP identity it found: unlike create_vip, which allocates
the next free VRID, a VRID already used in the pool is a hard 409. Silently
renumbering would put the adopted node in a different VRRP domain from the
peers that still run the original config.
The member row records the reporting node's own role, priority and interface,
and carries the one-shot takeover hash. The response returns the instance's
unicast peers, because those nodes hold their own keepalived.conf and have to
be adopted or added as members before the render describes a complete group.
|
||
|
|
7dfd31832a |
feat(vip): store the keepalived.conf an agent finds on a node
Ingest side of adopting an existing VIP. A node reports the keepalived.conf it found and does NOT own to a new endpoint, and the finding is kept in a new vip_discoveries table (one row per agent, since the file is per-node). Reporting is read-only on the node. The heartbeat cannot carry this: it has the VIP address and a best-effort MASTER/BACKUP, while rendering a node's config needs eleven fields, so the file itself has to be read and parsed server-side - parsing keepalived's block syntax in bash is not something to attempt on a production load balancer. Secrets are split at ingest. The reported content may contain the VRRP auth_pass, so the password is Fernet-encrypted into its own column through the same key path as vip_instances, and the stored copy of the file has it masked. Nothing readable through the API, the UI preview or a database dump carries it in cleartext, and the parse result is never logged. A file that does not parse records its error rather than failing the agent's poll loop, and a report of "the file is gone" clears the row so the UI stops offering a stale candidate. The keepalived-config delivery gains allow_takeover and takeover_expected_hash. The agent refuses to overwrite a keepalived.conf without our ownership marker, which is the guard that protects a hand-maintained setup - and is exactly the guard adoption has to pass. Rather than weaken it, an adopted VIP authorises exactly ONE takeover of exactly the file that was analysed by pinning its md5, so a config edited between adoption and Apply is still refused instead of being silently overwritten. SCHEMA_VERSION 10 -> 11 for the new table and two additive columns. Nothing existing is altered, but the bump re-seeds the four built-in roles, which the upgrade notes call out. |
||
|
|
8ac567dfe0 |
feat(vip): parse an existing keepalived.conf so a VIP can be adopted
Groundwork for adopting a hand-maintained keepalived setup into HA/VIP management. The page is empty today because the flow is one-way: VIPs are declared in OpenManager and pushed to the node, and nothing reads what is already there. The heartbeat cannot drive adoption. It carries two keepalived facts - keepalive_state (MASTER/BACKUP, best-effort from logs) and keepalive_ip (the first address grepped out of virtual_ipaddress) - while render_keepalived_conf needs eleven: virtual_router_id, auth_pass, interface, priority, prefix_length, advert_int, unicast peers, track_haproxy, role, address and name. Guessing the rest is not a cosmetic risk: a wrong VRID puts the nodes in separate VRRP domains and a wrong auth_pass makes them reject each other, and either way both nodes claim the VIP. So the config itself has to be read. Extracting the fields is the easy half. Adoption REPLACES the operator's file with our render, so anything their file contains that the renderer cannot reproduce would be destroyed on takeover - a notify_master failover hook, an LVS virtual_server section, a sync group, a second address in one instance, a custom track_script. The parser therefore also returns everything it could not model, and build_adoption_candidate turns each entry into a blocker with the file's own line number. Values that are unknowable rather than unreproducible block too: a missing virtual_router_id, and a missing prefix length, because our renderer always writes an explicit prefix and picking one would silently change a live VIP's netmask. keepalived's own documented defaults (state BACKUP, priority 100, advert_int 1) are applied but reported in `defaulted`, so the UI can say which values were assumed rather than read. Handles the layout variation real files have: nested braces, `#` and `!` comments, blocks opened and closed on one line, quoted script paths containing spaces, and several vrrp_instance blocks in one file. A parse result carries auth_pass in cleartext, since that is the only way to re-render an identical config, so it must never be logged - noted on every function that returns one. Tests pin each blocker and the layout variants, and include the invariant that keeps the parser honest: a config the renderer itself produced must parse back with zero blockers, so adding a directive to render_keepalived_conf without teaching the parser fails the suite instead of making OpenManager's own output look unadoptable. Verified by mutation - eight deliberate weakenings of the safety checks are each caught by at least one test. No endpoint, no schema change and no agent change yet; nothing calls this. |
||
|
|
02667bbda4 |
test(acme): make the wizard regression suite pass under the default jest timeout
Follow-up to #57. The contributed tests render the whole ACMEAutomation tree (antd Steps + Form + Select) and drive it through all three wizard steps, which takes 4-9 seconds per test. Under jest's default 5s per-test limit two of them failed, so `npm test` did not pass as shipped: ✕ an explicitly picked HTTP-01 account survives the step change and is what gets submitted -> Exceeded timeout of 5000 ms ✕ the wildcard guard still applies on the Review step, where Submit lives -> Exceeded timeout of 5000 ms The PR's reported 5/5 holds only when the runner is invoked with an explicit --testTimeout. Setting it in the file instead means the suite passes however it is invoked, which matters because the frontend image build runs `npm run build` and never the tests, so nothing else would have caught this. Verified with the default runner (no flags) after the change: 5/5 pass. Test-only. No production code touched.v1.10.3 |
||
|
|
c97df53da8 |
Merge pull request #57 from appouse/feature/multiple-account-acme-automation
Multi-account ACME: the certificate wizard honours the selected account (v1.10.3). Verified before merge. The contributed regression tests were run against the PRE-FIX component to confirm they actually catch the bug: 4 failed / 1 passed, reproducing the reported symptom exactly (Expected "http-01" / Received "dns-01", Review rendering the default account's email instead of the picked one, and the wildcard warning absent). With the fix: 5/5 pass. The three root causes were each confirmed against the tree — rc-field-form's useWatch honours options.preserve (es/useWatch.js:66), the backend default is ORDER BY created_at DESC (routers/letsencrypt.py:505) while the list is served ORDER BY id (:213), and the null-dns_provider normalization matches the backend's own at :473. Frontend production build succeeds; backend suite unchanged at 1243 passed. |
||
|
|
be01ddd119 |
test(acme): drive the certificate wizard to pin multi-account account resolution
Renders the real component and walks it through all three wizard steps, because the bug these cover was invisible to any unit test: it only appeared once the wizard advanced PAST the step that owns the account Select, since Form.useWatch reports only currently-rendered fields. Assertions describe behaviour rather than markup - the primary evidence is the POST body (account_id paired with challenge_type), compared against a fixture whose DNS-01 account is deliberately both the lower id and the older account, which is the exact shape that made the UI default and the backend default disagree. Verified by running the suite against the pre-fix component: the picked-account test reports challenge_type "dns-01" where "http-01" is expected, the Review test shows "Active (dns@example.com) / Challenge Method: DNS-01 / DNS provider: godaddy" for a chosen HTTP-01 account - the reported symptom reproduced - and the wildcard-guard test finds Submit enabled. Four fail, one passes: the DNS-01 selection path, kept as a positive control because it worked before (the default the wizard fell back to happened to be the DNS-01 account) and must keep working after. Adds src/setupTests.js with the ResizeObserver and matchMedia polyfills jsdom lacks and Ant Design 5 needs before any Select can open. |
||
|
|
bbd8359f50 |
docs(v1.10.3): document the multi-account ACME wizard fix
Release note covering the three compounding faults, and upgrade notes stating that this is frontend-only with nothing to do on upgrade. Two points are called out for operators rather than glossed: installations that never picked an account explicitly were already using the newest valid account, so only the preview was wrong; and the wildcard guard that stopped applying on Review was a lost warning, not a correctness hole, since the backend still rejected those requests. |
||
|
|
dd7b7822bf | chore(version): bump to 1.10.3 - multi-account ACME wizard fix | ||
|
|
c4139bb11a |
fix(acme): honour the selected ACME account in the certificate wizard
With more than one account registered, picking an HTTP-01 account in Request ACME Certificate still submitted a DNS-01 request, which the API rejected with "The selected ACME account has no DNS provider configured for DNS-01." Three faults compounded: 1. Form.useWatch reports only fields that are currently rendered. The account Select lives on the Configuration step, so the moment the wizard advanced to Review the watch read undefined and the wizard fell back to the default account - even though the value was still in the form store. Both wizard watches now pass preserve: true. The same fault silently disabled the wildcard guard on Review, the one step where Submit lives. 2. The UI and the backend disagreed on which account is the default. The backend takes the newest valid account (ORDER BY created_at DESC); the UI took the first valid entry of a list ordered by id, i.e. the oldest - the opposite account whenever the two differ. The wizard now resolves the same account and sends account_id explicitly, so there is no guess left to disagree about. 3. account_id was read from the form store while challenge_type came from the reverted account object. Both are now derived from one resolved account, so the pair can no longer describe two different accounts. Also: the Review step showed the default account's address instead of the chosen one, and Submit stayed enabled when the resolved account was deactivated (the fallback can land on a non-valid account, and the Select lists deactivated accounts). Single-account installations are unaffected. |
||
|
|
dbb9189f16 |
fix(ui): make Apply Management readable in dark mode (v1.10.2)
Three dark-mode defects reported on the Apply Management page, all the same
class of bug: light-mode colour literals hardcoded where theme tokens belong.
1. The "Pending Changes" box was painted background #fffbe6 with border
#ffe58f. In dark mode the text on top is light, so the version name,
timestamp and "View Change" link sat on a cream panel and were unreadable.
Measured contrast was 1.03:1; it is now 11.50:1 (secondary text 1.03:1 ->
7.02:1).
2. The added/removed rows in the View Change diff used #f6ffed/#52c41a and
#fff2f0/#ff4d4f, which stayed near-white inside the otherwise dark diff
panel. Now 5.49:1 (added) and 4.01:1 (removed), from 2.21:1 and 2.99:1.
3. The "Apply All Configuration Changes" confirm dialog came up white. This one
is not a colour literal: in Ant Design 5 the STATIC Modal.confirm / message /
notification APIs render into their own detached root and never see the app's
ConfigProvider, so they always fall back to the light algorithm. Registering
ConfigProvider.config({ holderRender }) once at the app root wraps that
detached root in the same ConfigProvider. Verified against the installed antd
5.29.3 source rather than assumed: config-provider/index.js sets
globalHolderRender, and modal/confirm.js wraps the dialog with it. This fixes
EVERY static dialog in the application — 12 components call Modal.confirm —
not just this page.
While in the file, six more instances of the same bug were fixed: the error
Alert border, the VIP pending-delete row, two ACME/pending version panels, the
applied-version panel and the agent-error recommendation box.
Light mode is byte-identical. Each token resolves under the default algorithm
to exactly the literal it replaced (colorWarningBg -> #fffbe6, colorSuccessBg ->
#f6ffed, colorErrorBg -> #fff2f0, colorInfoBg, colorErrorBorder, ...), so this
release can only change dark mode. Contrast was measured by resolving the real
design tokens under both algorithms and computing WCAG ratios, not by eye.
Note the diff rows were already low-contrast in LIGHT mode (2.21:1 and 2.99:1)
and remain so; that is the design system's own success/error pair and changing
it would alter the established light-mode appearance, so it is left alone.
holderRender is registered in an effect rather than during render, since
ConfigProvider.config() mutates antd module state; effects still run long
before a user can click anything that opens a static dialog.
Frontend only: no schema, no SCHEMA_VERSION bump, no API change, no environment
variable, zero agent impact. Backend suite unchanged at 1243 passed.
v1.10.2
|
||
|
|
eee0a4716a |
feat(ssl): encrypt the pending CSR private key at rest (v1.10.1, closes #53)
Closes the follow-up filed during the v1.9.0 CSR review. The private key of a
PENDING CSR is now Fernet-encrypted in the database instead of being stored as
a raw PEM.
Why this key specifically: it is the one key in the system that sits idle. It
is generated at CSR creation, waits for an external CA to sign the request
(days to weeks), and is destroyed the moment the signed certificate is
imported. It is never transmitted to an agent and never leaves the server.
ssl_certificates.private_key_content and the ACME order keys are deliberately
NOT covered, because agents must receive those in plaintext on every poll, so
encrypting them at rest buys nothing without an end-to-end redesign.
Implementation follows the pattern already used for the VRRP secret, TOTP
secrets and DNS provider credentials: a new utils/csr_key_crypto.py with its
own CSR_ENCRYPTION_KEY env var and its own HKDF info string
("csr-private-key-v1"), so rotating one secret class never affects another.
No schema change and deliberately NO SCHEMA_VERSION bump: the Fernet token
replaces the PEM inside the existing ssl_csrs.private_key_pem TEXT column. A
bump would re-run the migration sequence and re-seed the four built-in roles to
their defaults, which is a needless side effect for a storage-format change.
Backward compatible with no data migration. Rows written before this release
hold a raw PEM and are still read unchanged; the discriminator is exact rather
than a heuristic, since a Fernet token is base64url and can never contain the
"-----BEGIN" marker. Legacy rows drain naturally because a CSR's key copy is
NULLed on import.
A key that cannot be decrypted (SECRET_KEY rotated while CSR_ENCRYPTION_KEY was
unset) now fails with an explicit "delete this CSR and create a new one" error.
Previously that situation would have surfaced as the far more confusing
"certificate does not match this CSR's private key".
Also documents all four per-purpose encryption keys in .env.template. Only
VIP_ENCRYPTION_KEY was listed; MFA_ENCRYPTION_KEY and
DNS_PROVIDER_ENCRYPTION_KEY had been missing since v1.6.0 and v1.8.0.
Verified before release, on a corporate pre-production environment and locally:
- Full backend suite 1234 -> 1243 passed (+9 new tests), 0 failed.
- Against a real Postgres: a CSR created through the API stores a Fernet token
with no PEM header in the column, and imports successfully.
- Full 1.10.0 -> 1.10.1 -> 1.10.0 drill on one database volume. The upgrade
logs "Schema already at version 10 (>= 10); skipping migration run", so no
migration executes and the built-in roles are not re-seeded. A CSR created on
1.10.0 with a plaintext key imports successfully after the upgrade, which is
the backward-compatibility guarantee proven against a real row rather than a
mock.
- rsa-2048, rsa-4096 and ecdsa-p384 all round-trip through create, encrypt,
decrypt and import.
- Key derivation is stable across processes: two independent containers sharing
SECRET_KEY decrypt each other's tokens (required for UVICORN_WORKERS > 1 and
multi-replica deployments), while a different SECRET_KEY yields None rather
than a wrong key or an exception.
- Downgrade behaviour was measured, not assumed: 1.10.0 cannot parse the token
and fails with HTTP 500 "key parse failed (encrypted?)" rather than pairing a
wrong key. The rollback note states the measured behaviour.
- No CSR endpoint returns the key in any form: list and detail responses
contain neither a PEM nor a Fernet token.
Not changed here, from the issue's "worth folding in" list: the create rate
limit is not a concurrency guard, create_csr holds a pooled connection across
RSA key generation, detail=str(e) echoes internal error text (a repo-wide
convention), and is_global skips cluster validation in both routers/ssl.py and
routers/csr.py. None are storage concerns and each is a separate change.
v1.10.1
|
||
|
|
3c8832330a |
Merge pull request #54 from appouse/feature/godaddy-dns-provider
GoDaddy DNS-01 provider for ACME (v1.10.0). Validated on a corporate pre-production environment before merge: backend suite 1221 to 1234 passed (+13, exactly the new GoDaddy tests) with 0 failures; a wire-level harness against a fake GoDaddy API confirmed the apex+wildcard pair coexists, removing the last value uses DELETE rather than PUT [], an unreadable read fails closed with no write attempted, and across every scenario not one request reached a zone-wide endpoint (SPF, DKIM and DMARC survived untouched). The path-guard premise was measured directly: with yarl 1.24.5 a '.' segment normalizes onto the zone-wide TXT endpoint and '..' onto the whole-zone endpoint, so the guard in _rrset_path is load-bearing. No schema, environment, frontend or agent change. Closes #55v1.10.0 |
||
|
|
81ab674072 |
docs(v1.10.0): document the GoDaddy DNS provider and upgrade notes
README: add GoDaddy to the two feature bullets and to the DNS-01 provider catalog, spelling out that the API Key must be a Production key (the first key the developer dashboard issues is an OTE/test key and is rejected), that the zone must be in the same account, that the account needs a registered domain before GoDaddy permits DNS API access, and that a Personal Access Token works with the Secret left blank. Note that publishing is automatic for GoDaddy as well as Cloudflare, and add the release-notes entry. UPGRADE_GUIDE: new section stating there is no SCHEMA_VERSION bump, so the built-in-role re-seed warning from v1.9.0 does not apply, and no new environment variable, API-shape or agent change. Two limits are stated explicitly rather than glossed: the credential check is a read, so a token with read but not write scope saves successfully and only fails at the first publish; and downgrading after adopting GoDaddy is not a no-op, because an unknown provider name degrades DNS-01 orders to the manual-confirm path and leaves published TXT records marked cleaned without being removed. |
||
|
|
6bf6d016f5 | chore(version): bump to 1.10.0 - GoDaddy DNS-01 provider | ||
|
|
0a0226c758 |
test(dns): cover GoDaddy relative-name derivation, RRset merge and request handling
Twelve tests, no network and no database, in the existing pure-logic style. The merge helpers are covered directly (additive add, idempotent republish, tombstone filtering, remove-one-of-many, remove-the-last-value signalling DELETE), but helper math alone would stay green if the write path stopped using it, so add_txt_record and remove_txt_record are also driven against a recording stub: the assertions pin that a sibling value survives a publish, that an already-published value issues no write, that an unreadable read raises instead of replacing the set, that removing the last value emits DELETE and never an empty PUT, and that no call is ever aimed at a zone-wide path. _request is exercised through a fake response for the cases that only appear against the real API: an empty 204 body must not raise, a 3xx must not read as success (redirects are not followed), a transport failure mid-read must not be mistaken for an empty body, and each error status must produce a message naming what the operator has to fix. Also covers the auth header in both forms, that the sanitizer strips credentials from composed error text, that the module does not log at all, the credential-field schema against the upsert validator's own key and length rules, and the two-key encryption round trip. Verified by mutation: nine deliberate breakages of the provider - single-value PUT, empty PUT instead of DELETE, coercing an unreadable read to empty, treating 3xx as success, following redirects, swallowing transport errors, dropping the dot-segment guard, lowering the TTL below the API floor, and removing sanitization - are each caught by at least one test. |
||
|
|
8e534ef170 |
feat(dns): add GoDaddy DNS-01 provider (API Key+Secret / PAT, additive RRset writes)
Registers a third pluggable DNS provider for ACME DNS-01 alongside Manual and Cloudflare. Credentials are an API Key + Secret pair; leaving the Secret blank sends the Key as a Personal Access Token (Bearer), which is the migration path as GoDaddy retires the sso-key scheme. GoDaddy's Domains API v1 has no per-value TXT write: PUT on a record set replaces every value at that name. A certificate covering example.com and *.example.com publishes two different TXT values at the same _acme-challenge.example.com, so add/remove are read-modify-write - read the current set, merge, put the whole list back - with empty-data tombstone rows filtered out (they are rejected on echo) and DELETE used for the last value, since PUT with an empty array is rejected. The zone-wide sibling endpoints (.../records/TXT and .../records) would wipe SPF/DKIM/DMARC and the whole zone respectively, so the record path is built in one place that refuses an empty or dot segment. An unreadable record-set read fails closed rather than being treated as an empty set, because the PUT that follows would otherwise destroy the coexisting values. Zone lookup walks name suffixes probing the records API rather than the domain listing, so zones delegated to GoDaddy nameservers resolve and accounts that are rejected from the domain-details endpoint still work. Credential and eligibility failures during the walk surface instead of being reported as "no managed domain". Provider errors are sanitized at the single point where GoDaddy-supplied text enters a message, since those strings are persisted to order events and shown in the UI. No new dependency, no schema change, no frontend change - the credential form is rendered from the provider schema. |
||
|
|
71786200dd |
docs(v1.9.0): correct upgrade notes on built-in role re-seed + backfill 1.8.8-1.8.10 release notes
Found during a v1.8.10 to v1.9.0 upgrade drill on a populated database
(schema v9 to v10) before releasing. Documentation only, no code change.
1. The v1.9.0 upgrade notes said "custom roles need no changes", which reads
as "role data is untouched". It is not: because the SCHEMA_VERSION bump
re-runs the whole idempotent sequence, update_system_roles_to_enterprise_rbac()
issues an unconditional UPDATE roles SET ... permissions = <defaults> for
the four BUILT-IN roles. In the drill an `operator` role that had been
narrowed by removing apply.execute and config.bulk_import came back with
both restored (57 to 59 permissions). Operator-created roles are NOT
affected; the re-seed matches the four built-in names only.
This is pre-existing behaviour of every SCHEMA_VERSION bump and is
documented as intentional in migrations.py, so it is not introduced by the
CSR feature. The v1.7.0 upgrade notes carried this caveat and it was not
carried forward. Restored, with an export and re-apply procedure.
Also clarified why the admin password is safe: the default-user seeding is
guarded by an existence check ("safer than ON CONFLICT"), not an upsert,
so an operator-changed password survives.
2. README release notes jumped from v1.8.7 straight to v1.9.0 because
v1.8.8, v1.8.9 and v1.8.10 were never backfilled. Added all three.
v1.9.0
|
||
|
|
33e3e8ef9d |
Merge pull request #50 from mustafaulukaya/feature/csr-creation
CSR creation (v1.9.0): in-app key + CSR generation and signed-certificate import. Validated on a corporate pre-production environment before merge: backend suite 1145 to 1221 passed (+76), config generator output byte-identical to 1.8.10 against the same database, real haproxy 2.8 -c accepts a CSR-issued certificate, API surface additive only (+5 endpoints), populated v9 to v10 upgrade drill preserved all data, rollback to 1.8.10 starts cleanly, zero diff in the agent scripts. Closes #49 |
||
|
|
69e12f7459 |
chore(version): bump to 1.9.0 - CSR creation
- backend/version.json + frontend package version to 1.9.0 - README: feature list entry, CSR workflow section, SSL CSR API reference, v1.9.0 release notes - UPGRADE_GUIDE: v1.9.0 section (additive ssl_csrs table, SCHEMA_VERSION 9 -> 10, no RBAC changes, zero agent impact, rollback note) |
||
|
|
af07d72514 |
feat(ui): add CSR tab to the SSL Certificates page
New CSRManagement component as a third tab (deep-linkable via ?tab=csr): - Create modal: name (path-traversal-safe client rules mirroring the server), CN with wildcard support, SAN tag input, key algorithm select, optional subject fields in a collapse panel. On success the view modal opens immediately with the CSR PEM. - View modal: subject/SAN summary, read-only CSR PEM with copy and a Download .csr button (Blob download). - Import modal: paste signed certificate + optional chain, usage type, global/cluster scope with cluster multi-select, and an optional certificate-name override for collisions that appeared after CSR creation; SAN-drift warnings surface in a warning dialog. - Duplicate action pre-fills the create modal (forceRender so the form accepts values before first open); delete confirm spells out that a pending CSR key is destroyed permanently. - SSL certificate table now renders a distinct CSR source tag next to the existing Manual / Auto (ACME) tags. |
||
|
|
a6166d11b9 |
feat(ssl): add CSR generation and signed-certificate import (backend)
New /api/ssl/csrs endpoint group: generate a private key + CSR server-side
(RSA 2048/4096, ECDSA P-256/P-384; full subject + DNS SANs with wildcard
support), list/detail/delete CSRs, and import the CA-signed certificate.
- New ssl_csrs table (SCHEMA_VERSION 9 -> 10, additive + idempotent); the
migration re-raises on failure so a failed run is retried instead of being
stamped as applied.
- Import verifies the certificate against the stored key as a hard gate
(match=None is treated as an integrity error, not a lenient pass), rejects
malformed and expired certificates with 400, warns on SAN drift, and
creates a normal ssl_certificates row (source=csr, cluster_id=NULL,
last_config_status=PENDING) so it flows through the standard
Apply Management -> agent pull pipeline.
- Concurrency: FOR UPDATE row lock serialises double-import and
delete-during-import; a partial unique index reserves pending CSR names;
soft-deleted same-name certs are reactivated preserving the row id.
- Security: no CSR endpoint ever returns the private key (explicit column
lists, enforced by a static test); the key copy on the CSR row is NULLed
after import; ssl.create/read/delete permissions enforced on every
endpoint incl. reads; per-user rate limit on key generation, which runs
in a worker thread; csr_id and cluster_ids are int32-guarded.
- ssl_service: extract _prepare_cert_fields from create_cert_row (behaviour
unchanged, extraction tests untouched) and add stage_ssl_config_versions
reusing the exact ssl-{id}-create-{ts} version-name scheme.
- Tests: crypto round-trip for all four algorithms, model validation,
import-flow unit tests, endpoint auth/permission pinning, migration and
key-non-exposure static assertions.
|
||
|
|
882d25bb68 |
Merge pull request #44 from taylanbakircioglu/chore/bump-1.8.10-github
chore(version): bump to 1.8.10 (security release)v1.8.10 |
||
|
|
0ebf6583ea |
chore(version): bump to 1.8.10 — security release (RCE/auth/SSRF advisories fixed)
Single-source version bump (backend/version.json) with frontend/package.json and package-lock kept in sync (test_version_consistency). Marks the release that ships the GHSA-7rhv / GHSA-3p5c / GHSA-3vh4 fixes. |
||
|
|
6be19f0bb5 |
Merge pull request #43 from taylanbakircioglu/fix/security-advisories-github
fix(security): remediate RCE, missing-auth and SSRF advisories (backend-only) |
||
|
|
9e5185c458 |
fix(security): post-review hardening — agent-inventory regression, coverage gaps, SSRF newNonce
Follow-up to the RCE/missing-auth/SSRF remediation, from a thorough multi-lens
review (3 agents + a black-box audit of all 201 routes). Backend-only; no
agent-script changes.
Regression fix (introduced by the previous commit):
- GET /api/agents was made JWT-only, but deployed agents call it WITH X-API-Key
(not a JWT) to read their applied_config_version and avoid re-applying config on
restart. It now accepts EITHER a valid operator JWT OR a valid agent X-API-Key,
so agents no longer get 401 (which caused a spurious HAProxy reload every restart).
Completeness (GHSA-3p5c siblings the first pass missed — same data class, now JWT):
- dashboard.py: GET /api/haproxy-cluster-pools/{id}/agents (full agent inventory —
a direct anonymous bypass of the GET /api/agents lockdown), /api/pools,
/api/haproxy-cluster-pools, /api/dashboard/stats, /api/dashboard/overview
(auth was optional -> leaked stats/names/health/alerts anonymously),
/api/haproxy/stats.
- waf.py: GET /api/waf/rules. health.py: GET /api/health/errors.
- agent.py: GET /api/agents/generate-uninstall-script/{platform} (agent-management
endpoint; was anonymous) now requires JWT or agent key, like generate-install-script.
- config.py: POST /api/config/{validate,optimize,templates/{id}/generate} were
optional-auth (logging only) and run a HAProxy validator on caller input; now
require a JWT. (bulk-create, parse-bulk, diff and configuration/request were
already mandatory-auth — verified.)
All newly-gated endpoints are frontend-only (axios sends the JWT) or unused;
agents never call them.
SSRF (GHSA-3vh4) gap:
- acme_service._get_nonce fetched directory['newNonce'] (from the attacker-
influenceable directory JSON) with a bare session, http allowed, dual-stack, and
BEFORE the guarded _signed_request POST. Now guarded (assert_public_url +
safe_connector + no redirects + timeout), matching the other ACME sinks.
Correctness:
- Three agent webhooks (config-applied, config-validation-failed, config-sync)
swallowed their auth 401 into a 200 error body via a bare `except Exception`.
Added `except HTTPException: raise` so the 401/403 propagates.
Audit result (live black-box, all 201 routes probed unauthenticated): no data
leak and no unauthenticated mutation anywhere; every sensitive route returns
401/403 (a pre-existing group of read handlers wraps the 401 into a 500 via a
broad except — no data is exposed; left as-is, documented as cosmetic).
Verified: full pytest tests/ (1145 passed, 0 failed; +16 regression tests) + live
localtest stack smoke — agent-key GET /api/agents=200, anonymous=401, all newly
gated endpoints reject anonymous and admit JWT, the 3 webhooks return 401.
|
||
|
|
520b69a1c6 |
fix(security): remediate RCE, missing-auth and SSRF advisories (backend-only, no agent changes)
Addresses three reported advisories, all verified against the code. Fixes are
entirely server-side — deployed agents already send a valid X-API-Key on every
call, so enforcing it does not require any agent-script change or upgrade.
GHSA-7rhv-c5pc-69r8 (CRITICAL RCE — agent script-template poisoning):
- POST/GET /api/agents/script-templates/{platform} now require the agents.version
permission (was authentication-only), matching POST /versions. Blocks a viewer
JWT from overwriting the root install/upgrade script.
GHSA-3p5c-m5m4-mjpx (missing authentication):
- Agent data-plane endpoints now REQUIRE a valid X-API-Key (was optional/skipped
when the header was absent), checked before any DB access: config,
ssl-certificates (private keys!), upgrade-status, heartbeat (by-name and the
previously auth-less by-id), configuration pending-requests. Removes keyless
heartbeat spoofing and keyless rogue-agent auto-registration.
- Operator/UI endpoints now require a JWT: GET /api/agents, the entire
/api/dashboard-stats router, /api/health/{deep,agents,clusters}, and
/api/ssl/certificates/{id}/config-versions. The simple /api/health liveness
probe stays public. Adds shared auth_middleware.require_authenticated_user.
GHSA-3vh4-gvxx-wm2p (SSRF via ACME directory_url):
- New utils/ssrf_guard.py (https-only + public-IP-only, IPv4-pinned, no redirects),
applied to settings test-connection, acme_service.get_directory and
_signed_request, and validated at Let's Encrypt account creation. The
test-connection response no longer reflects arbitrary upstream JSON keys
(information-disclosure oracle) — only fixed ACME field names.
Verified: full pytest tests/ (1128 passed, 0 failed) + live localtest stack smoke
(valid JWT/key paths return 200/404 as expected; anonymous requests 401; SSRF to
metadata/private/loopback refused). No changes to backend/utils/agent_scripts/*.
|
||
|
|
56107fa86f |
Merge pull request #42 from taylanbakircioglu/fix/security-deps-round2
fix(deps): patch websocket-driver (CRITICAL) + resolve remaining postcss (#12/#2) |
||
|
|
f86a4331e8 |
fix(deps): patch websocket-driver (CRITICAL #52) and resolve remaining postcss (#12)
Follow-up to the consolidated security bump: - websocket-driver -> 0.7.5 (CRITICAL, message corruption; dev/build tooling) - resolve-url-loader -> 5.0.0 (pulls postcss ^8), resolving the last postcss<8.5.10 instance (#12) and #2 at the source. Project uses no SASS, so resolve-url-loader v4->v5 is inert at build time. Verified: full production docker build succeeds on node 18; postcss now resolves to a single 8.5.10 across the tree; deferred dev-only deps unchanged. |
||
|
|
1c47e246ec |
Merge pull request #39 from taylanbakircioglu/fix/security-deps
fix(deps): patch frontend security advisories (11 Dependabot alerts, incl. all 4 HIGH) |
||
|
|
d914f2398b |
fix(deps): patch frontend security advisories (11 Dependabot alerts, incl. all 4 HIGH)
Bump react-router-dom to ^6.30.4 (runtime open-redirect fix, CVE-2026-40181) and add scoped npm overrides to patch dev/build-toolchain transitive deps: ws (7.5.11 / wds-scoped 8.21.0), form-data (4.0.6 / jsdom-scoped 3.0.5), js-yaml (3.15.0 / eslint-scoped 4.2.0), http-proxy-middleware 2.0.10, launch-editor 2.14.1, postcss 8.5.10 (resolve-url-loader kept at 7.0.39), @babel/core 7.29.6. Deferred (breaking major / node20, not in production bundle): webpack-dev-server, serialize-javascript, uuid, @tootallnate/once, resolve-url-loader's postcss. Verified: plain `npm install` (no --legacy-peer-deps) + full production docker build succeed on node 18; only package.json + regenerated package-lock.json change. |
||
|
|
9c1f3c811b |
chore(redis): upgrade Redis image from 7-alpine to 8.8.0-alpine
Infra-only change: image tag bump in docker-compose and k8s manifest. Backend redis-py client (redis>=5.0.0, resolves to 8.x) verified compatible against Redis 8.8.0 for all commands in use (get/setex/incr/expire/delete/ping). No application code or version change. |
||
|
|
c79391cd13 |
feat(acl): accept HAProxy -f pattern-file references with advisory warnings (v1.8.9, Issue #38)
The manual Frontend editor, wizard and visual ACL builder hard-rejected the ACL `-f <file>` flag while bulk import accepted it. Worse, a frontend imported with an `-f` ACL could not be edited at all (422) until the ACL was dropped. The original guard predated the fail-safe apply flow: the agent runs `haproxy -c` before every reload, so a missing pattern file is rejected safely and the previous config keeps running. Pattern files are operator-managed host files — the same policy adopted for SPOE filter configs in v1.8.8. - models: remove the 5 `-f` hard rejects (frontend acl/redirect/use_backend validators + wizard string/dict-redirect guards); `$(`/backtick and X!X contradiction guards unchanged - routers/frontend: `_pattern_file_warnings` helper; non-blocking warning on create + update responses listing referenced pattern files (empty when no rule uses `-f` — zero noise) - routers/config: bulk-import preview advisory listing pattern files per frontend (cluster config-dir aware, next to the SPOE advisories) - React: remove the FrontendManagement submit gate and SiteWizard step gate; ACLRuleBuilder renders informational notes instead of errors and re-adds `-f (pattern file on host)` to the flag dropdown; create path now renders server warnings like update - tests: 4 reject-pins inverted to accept-pins; new test_acl_pattern_file_allow.py (accept/guards-kept/zero-noise/advisory); full suite green (1094 passed)v1.8.9 |