123 Commits

Author SHA1 Message Date
taylanbakircioglu 0ee227363e fix(agent): adoption cannot hide a node; Config Import on fresh installs (v1.11.1)
Six fixes to the agent's discovery path and one long-standing parity gap, found
by auditing it in loops against a real fleet.

DISCOVERY REPORTS THAT NEVER REACHED THE SERVER

The agent posts an unmanaged keepalived.conf for adoption and caches the hash of
what it sent, so the file - which carries the VRRP password - is re-posted only
when it changes. Delivery was judged by curl's exit code, and curl without -f
exits 0 on 5xx too, so a report the server REJECTED was recorded as delivered.
Since a hand-maintained config does not change on its own, that node dropped out
of "Unmanaged keepalived detected" permanently; the only cure was deleting a
cache file on the node by hand.

  - the report is cached only on a 2xx;
  - GET /agents/{name}/keepalived-config now reports whether the server actually
    holds a discovery for that agent, and the cache may only suppress while it
    says yes - which is what lets nodes stuck from earlier releases recover on
    their own, with nobody touching them;
  - the flag is parsed with has() + tostring, not `// empty`: jq's alternative
    operator returns the alternative for **false** as well as null, so the naive
    form could not tell "no record" from "older backend" and the recovery would
    have been completely inert;
  - a 400/413/422 records the refusal so identical bytes are not re-posted
    forever - 4xx and 5xx agent calls are never sampled out of the request log,
    so an unattended loop would write a row carrying the whole config every
    cycle - while 401 and 404 keep retrying, because here they mean a token
    rotation or an agent row briefly absent, not a bad payload;
  - the CLEAR path had the same exit-code defect, where it left a stale row
    offering a managed node for adoption with nothing to ever retry it.

CONFIG IMPORT WAS A NO-OP ON FRESHLY INSTALLED AGENTS

check_config_requests uploads a node's live haproxy.cfg on request. It was
defined in the installer body and in the self-upgrade daemon, but not in the
heredoc a fresh install writes, and its call site is guarded by `type` - so on
such a node the operator asked for a config and nothing arrived, with no error
anywhere. Any agent that had self-upgraded at least once already had it, which
is why it went unnoticed. The self-upgrade definition is copied verbatim
(verified line-for-line). A freshly installed agent now polls that endpoint once
per cycle exactly as every upgraded agent already does; no node running today
changes behaviour.

DETERMINISTIC CONFIG PATH

A pool may hold several clusters and the join that resolves keepalived_config_path
was unordered, so the path handed to an agent could differ between polls whenever
two clusters disagreed - the agent would inspect a file that is not there and the
node would never appear, intermittently. A customised path now wins over the
shipped default, then the lowest cluster id. Verified against a real PostgreSQL
over seven arrangements: with one cluster per pool, or when every cluster carries
the default, the value is byte-identical to before.

Verified end to end on a production fleet and, for each decision, against the
real _kp_discover block rather than a paraphrase.

Backend suite: 1674 passed, 152 skipped. bash -n passes on the whole file and on
the fresh-install body in isolation. The keepalived path is logic-identical
across both daemon copies, now pinned by a test.
2026-08-15 19:37:53 +03:00
taylanbakircioglu bd4a50943f fix(requestlog): bound queue memory, and stop the UI reporting things it cannot know
Three hardening fixes with the same shape: a number that was true under the
defaults and untrue at the edges.

1. QUEUE MEMORY WAS AN OPERATOR SETTING, NOT A LIMIT.

The queue was bounded by ROW COUNT only, and how much a row weighs is
`requestlog.max_body_bytes` - editable from Settings, documented ceiling 256 KB,
and a row can hold that twice (request + response). Measured on the real
dataclass with distinct buffers per row:

    defaults, 2 000 rows x 8 KB                33.9 MiB    3.3% of the 1 GiB pod limit
    max_body_bytes at its 256 KB ceiling        1003 MiB    at the pod limit
    REQUEST_LOG_QUEUE_MAX at its ceiling        1695 MiB    over the pod limit

Both are reachable from in-range, documented values, and the drop warning
advised "raise REQUEST_LOG_QUEUE_MAX" - so following the tool's own advice on a
busy install could OOM the worker. REQUEST_LOG_QUEUE_MAX_BYTES (default 64 MiB)
now caps the queue in bytes as well as in rows, whichever binds first, released
as rows drain. Verified: with max_body_bytes at 256 KB the queue holds 7.5 MiB
against an 8 MiB budget where it would otherwise have held 1003 MiB, and it
accepts rows again as soon as the writer drains it. The warning text now names
the setting that actually helps.

2. SINK COUNTERS ARE PER WORKER AND DID NOT SAY SO.

The sink is a module global, so with UVICORN_WORKERS > 1 each process has its
own queue and its own counters, and `GET /api/request-logs/stats` reports
whichever worker happened to serve the request. The feature is sold on "a
saturated logger drops rows visibly"; at 4 workers the visible number was a
quarter of the truth. Labelled `"scope": "this worker only"` rather than
aggregated - there is no cross-process channel here, and a number that looks
fleet-wide but is not is worse than one that admits its scope.

3. AN EMPTY EXCLUDE LIST IS NOT APPLIED AS "LOG EVERYTHING".

normalize_exclude_paths() falls back to the shipped defaults when the list comes
out empty, which is the right call - it keeps the log viewer and the raw-body
heartbeat endpoint excluded - but the UI kept displaying the empty list the
operator typed, so the form showed a policy that was not in effect. The save
handler now re-applies whatever the server actually stored (which also surfaces
server-side clamping of every numeric field) and says plainly that the defaults
were restored.
2026-08-15 11:04:31 +03:00
taylanbakircioglu 4e2d936c27 fix(requestlog): stop the table size from scaling with fleet size
The row rate of `request_logs` was a function of how many nodes are installed,
not of what anyone did. Counted from the agent loop in linux_install.sh, each
agent's 30s cycle issues three logged calls - config, pending-requests,
upgrade-status (the heartbeat is already on the default exclude list) - plus
keepalived-config and keepalived-status every fifth cycle. That is ~9 800
rows/day per agent, essentially all of them 200s meaning "nothing changed".

Measured on PostgreSQL 15 against the real DDL and all nine indexes, at 2 424
bytes/row:

     20 agents     ~196k rows/day    453 MB/day   row cap reached in 2.5 days
    200 agents     ~2.0M rows/day    4.4 GB/day   row cap reached in  6 hours
    500 agents     ~4.9M rows/day     11 GB/day   row cap reached in  2 hours

The cap holds, so nothing runs away - but it holds by deleting, and what it
deletes is everything else. The shipped policy says 7 days of successes and 30
days of failures; on a 200-node fleet it delivers about six HOURS of both. The
forensic record the feature exists for is evicted by polling noise, and the
larger the installation the less history it keeps.

`requestlog.capture_agent_success`, default FALSE: a SUCCESSFUL inbound call
from an agent is not recorded. Failures always are, whatever the flag says -
they are what an operator needs and they are rare, so they cost nothing. With
this the table's size follows operator activity, and adding nodes does not
shorten anyone's retention.

Agent traffic is identified by header only, no database round-trip on the hot
path: the installed agent sends `X-API-Key` and never `Authorization`, the UI
sends a JWT and never an agent key. `generate-install-script`, the one endpoint
that accepts either, classifies correctly under the same rule - an operator
generating a script sends Authorization, a self-upgrading agent sends only the
key. The result is stored in the existing `target` column, which already means
"who was on the other end" for outbound rows and now means the same for inbound
ones, so no schema change and the existing target index applies.

Second half, and the reason this is one commit: `operator` holds
`requestlog.read` because, per the migration that grants it, "operators debug
failing applies and ACME orders". They could not. An apply fails on the NODE,
and the node reports that over its own API key, so the row carrying the
diagnosis has `user_id IS NULL` - and own-rows-only scoping hid it from exactly
the role the grant was written for. Scoping now admits agent rows alongside the
caller's own. Deliberately keyed on `target = 'agent'` rather than `user_id IS
NULL`: anonymous traffic is not agent traffic, so failed logins and their
usernames, and unauthenticated probes, stay admin-only.

Verified end to end through the real middleware: a successful agent poll is
dropped, a 422 from config-validation-failed is kept, operator and anonymous
calls are unaffected, and flipping the setting on restores the old behaviour.
2026-08-15 11:04:30 +03:00
taylanbakircioglu 4c84596215 feat(logging): unified request/response log with configurable retention
Applies PR #59 by Mustafa Ulukaya (github.com/taylanbakircioglu/haproxy-openmanager/pull/59,
head ef26860) as authored, with only the merge conflicts resolved. Behavioural
gaps found in review are closed by the follow-up commits on this branch rather
than by rewriting the contribution.

One queryable timeline covering both directions: every inbound API call
(including GETs and 4xx/5xx) with user, client IP, status, duration and
redacted, size-capped bodies; and every outbound HTTP call the backend makes,
tagged with who it went to. Outbound rows inherit the inbound request's id, so
one operator action and the CA/DNS calls it triggered read as a single trace.

Conflict resolution (the branch was cut at v1.10.3, this tree is v1.10.14):

* SCHEMA_VERSION: 11 -> 12, NOT the 11 the branch proposed. 11 was taken in the
  meantime by v1.10.4 (vip_discoveries). Landing this as 11 would be silently
  inert: run_all_migrations() returns early on `applied_version >=
  SCHEMA_VERSION`, so every database already at 11 skips the whole sequence and
  gets neither request_logs nor the requestlog.* permissions, while a fresh
  install gets both. The branch's own test asserts `>= 11`, so it still holds.

* services/acme_diagnostics.py: the branch instrumented a `session.head(...)`
  probe, which is what that function did when it was cut. It has since become a
  GET that classifies the response body, because a status code alone cannot
  tell a working challenge endpoint from a SPA catch-all answering 200 with
  index.html. Taking the branch's side would reintroduce that bug, so the GET
  probe is kept and the span wraps it. The span records the classification, not
  the body: `_PROBE_BODY_LIMIT` is 64 KB of a third party's page and storing it
  would put an arbitrary remote document in the audit table per probed domain.

* backend/version.json, frontend/package.json: 1.11.0, release date moved to
  the date this actually ships.

* README.md, UPGRADE_GUIDE.md: the v1.11.0 sections are added above the
  existing entries; every note from v1.10.4 through v1.10.14 is preserved.

Schema: one new table (request_logs) plus its settings seed. No existing table
altered, no agent or rendered-config change. As with every SCHEMA_VERSION bump,
the four built-in roles are re-seeded to their defaults - export role
customizations before upgrading.

Kill switches: REQUEST_LOG_ENABLED=false (middleware never registered) or the
`enabled` toggle in Settings -> Request Log.
2026-08-15 11:04:18 +03:00
taylanbakircioglu 0eb587dfa8 fix: keepalived validation gate and deploy acknowledgements (v1.10.13)
Two agent-side fixes found while taking the VIP adoption flow through a real
HA pair, released together.

1. A VALID CONFIG WAS REJECTED BY ITS OWN WARNING (v1.10.12)

Before writing a rendered keepalived.conf the agent validates it with
keepalived -t and, on failure, keeps the running config and does not restart
keepalived. That fail-safe is right, but it treated ANY non-zero exit as
invalid - and keepalived's config-test exit code does not separate fatal from
benign. Measured on 2.2.8:

    clean config ..................... 0
    auth_pass longer than 8 chars .... 5   "Truncating auth_pass to 8 characters"
    missing '}' ...................... 5   "There are 1 missing '}'s"
    unknown keyword .................. 5   "Unknown keyword '...'"
    script without script_security ... 6   "SECURITY VIOLATION ..."

Exit 5 covers both a harmless truncation and a broken file, so a VRRP password
over eight characters was enough to block every apply - including on a node
whose own running config emits the same warning and had been serving the VIP
for weeks. Accepting exit 5 would have accepted broken configs, so the gate now
judges the OUTPUT: known-benign messages are dropped and anything remaining
still fails. It fails CLOSED - an unrecognised message, or a non-zero exit with
no readable output at all, is fatal - and the filter is an allowlist, never a
denylist. The agent also reports what keepalived said, in its log and in the
status the HA/VIP page shows; discarding it left a correct refusal that nobody
could act on.

2. EVERY DEPLOY ACKNOWLEDGEMENT WAS DROPPED

The takeover-retirement clause on POST /agents/{name}/keepalived-status reused
one placeholder for both the assignment `last_deploy_hash=$n` and the
comparison inside its CASE. PostgreSQL types a placeholder per USE, so it came
out as text in one and character varying in the other and asyncpg rejected the
statement with AmbiguousParameterError. The whole UPDATE never ran, so no
member recorded an acknowledgement: VIPs sat at SYNCING with an empty Last ack
while the nodes were verifiably running the config, and teardown acks were lost
the same way. The hash now has its own placeholder, compared only against the
column.

Verified against real keepalived and a real PostgreSQL rather than by
inspection, including on busybox and bash 3.2, and a test asserts every $n in
those statements is bound exactly once.
2026-08-14 18:38:36 +03:00
taylanbakircioglu a87994e06a fix(vip): refuse adoption that strands a node or normalises a peer (v1.10.9)
Three findings from a second pass over the adoption flow, all of the same
class: something real leaving the set silently.

1. STRANDING. _collect_instance_participants can only match a node it can
   READ, that is ENABLED, and that is in the SAME pool. Each of those is a door
   a genuine member of the VRRP group leaves through without a word, and the
   nodes that remain are rewritten while it keeps serving the same address from
   an unmanaged config. Found on a live pool: one node of a pair had an
   unclosed vrrp_instance block, so it parsed to nothing while its partner
   parsed cleanly.

   Rather than guard each door, ask the question directly: does any reported
   keepalived.conf mention THIS virtual address without being one of the nodes
   we are about to adopt? Refuses naming the node and the reason. Scoped on the
   address so an unrelated file elsewhere cannot block every adoption, and
   excluding nodes already under management (a standing VIP, or our ownership
   marker) because those are not stranded.

2. SILENT NORMALISATION. prefix_length, unicast/multicast mode, HAProxy
   tracking and the VRRP password are stored ONCE on the VIP and re-rendered
   onto EVERY member, so whichever node was clicked imposed its settings on the
   others. prefix_length is the sharpest: the design refuses to GUESS a netmask
   for a live VIP, and copying one node's netmask onto another is that same
   change wearing a different hat. All four must now agree, with both values
   named in the refusal. The VRRP secret is compared by decrypting each node's
   token - Fernet is non-deterministic, so ciphertexts cannot be compared - and
   a token that will not decrypt is an error rather than an assumed match.

3. THE TAKEOVER AUTHORISATION WAS NOT ONE-SHOT. takeover_expected_hash is the
   permission to overwrite a keepalived.conf that lacks our ownership marker.
   It was written at adoption and never cleared, so it stayed valid for that
   file content indefinitely: restoring the pre-adoption file would have been
   overwritten again with no fresh human approval. It is now retired when a
   member acknowledges our rendered config, gated on the acked hash matching
   applied_config_hash so a partial or failed deploy never drops it and leaves
   the VIP unable to converge.

The panel applies the stranding rule too, so the Adopt button is disabled with
the reason instead of letting the operator click into a 422.

No schema change, no agent change, no API-shape break.

Backend suite: 1366 passed, 152 skipped. Frontend build clean.
2026-08-14 07:55:07 +03:00
taylanbakircioglu 7d95c737f0 fix(vip): adopt the whole VRRP instance, not one node (v1.10.8)
Four defects found while tracing the adoption flow end to end after v1.10.4
reached a live HA pair.

B3/B4 (one root, one fix). Adoption took only the node whose row was clicked:

  - adopting the BACKUP alone produced a VIP that apply always rejects, because
    apply requires exactly one MASTER member;
  - adopting the MASTER alone left the peer unmanaged, and adopting it
    afterwards hit the VRID-collision guard with 409, so a pair could never be
    completed from the panel;
  - on a UNICAST instance the single-member render dropped the unicast block
    entirely (render_keepalived_conf emits it only when peer_ips is non-empty),
    so keepalived fell back to multicast on the adopted node while its peer
    stayed unicast. They stop seeing each other and BOTH claim the VIP.

Adoption now resolves the whole instance via _collect_instance_participants,
keyed on (virtual_router_id, virtual address) - the same key keepalived uses to
group nodes. Each participant becomes a member with the role, priority and
interface its own file declares, and its own one-shot takeover hash, so the
per-node overwrite guard is unchanged. It refuses, naming the reason, when the
group has other than one MASTER, when advert_int differs across nodes, when a
declared unicast peer is not among the nodes being adopted, or when a node is
already in a live VIP - the one-active-VIP-per-agent rule that create/update
enforce via _validate_members_against_pool and adoption never called.

B1. The Apply Management "View Change" regex matched vip-(create|update|delete)
only, so an adopt version fell through to the generic HAProxy diff and rendered
the cluster's whole haproxy.cfg as removed. Display-only, but alarming. A test
now asserts every action _stage_vip_version can stage is in that alternation.

B2. Rejecting an adoption hid the node from the panel permanently:
vip_discoveries.adopted_vip_id is write-once, a VIP is only ever soft-deleted so
the column's ON DELETE SET NULL never fires, and the agent does not re-report a
file whose hash has not changed. Rather than clearing the column on each path,
adoptability is derived from whether the linked VIP is still active, which
self-heals reject, undo-reject and approved teardown alike.

The panel now lists one row per instance instead of per node, and the adopt
dialog names every node that will be taken over. Blockers are aggregated across
all of them, matching what the endpoint checks.

No schema change, no agent change, no API-shape break: /api/vip/discoveries
gains a derived field and /api/vip/adopt keeps its request body.

Backend suite: 1359 passed, 152 skipped. Frontend build clean (no new lint
warnings in VIPManagement.js).
2026-08-13 21:45:38 +03:00
taylanbakircioglu eda7f36c93 fix(vip): scope the HA/VIP page to the selected cluster (v1.10.7)
The page ignored the cluster picker in the header. Both the VIP table and the
v1.10.4 "Unmanaged keepalived detected" panel queried the whole fleet, so on an
install with more than one cluster the lists showed every cluster's nodes at
once and did not change when the selection did - the panel looked stuck on one
cluster's keepalived.

- GET /api/vip/discoveries takes the same optional cluster_id the VIP list
  already took, mapped to the cluster's pool via haproxy_clusters exactly like
  list_vips does. Omitting it still returns the whole fleet, so no existing
  caller changes behaviour.
- VIPManagement reads selectedCluster from ClusterContext (it only took the
  cluster list before) and sends cluster_id on both fetches. The fetch callbacks
  depend on the scope, so switching cluster refetches instead of showing stale
  rows.
- tests/test_vip_discoveries_scope.py pins both defects this endpoint has had:
  that the route reaches its own handler rather than being parsed as a vip_id
  (asserting != 422 specifically, since the repo's generic endpoint-auth tests
  accept 422 alongside 401/403 and therefore could not catch it), and that the
  cluster filter resolves cluster -> pool and short-circuits when absent.

Behaviour change worth calling out: the VIP table is now scoped to the selected
cluster where it was fleet-wide before. The API still serves the fleet-wide view
to any caller that omits cluster_id.

Backend suite: 1345 passed, 152 skipped. Frontend production build clean.
2026-08-13 19:12:01 +03:00
taylanbakircioglu 11e5bf57d9 fix(vip): make the adoption panel reachable again (v1.10.6)
Follow-up to #58. The "Unmanaged keepalived detected" panel it shipped never
appeared on any deployment. GET /discoveries was declared after GET /{vip_id} in
routers/vip.py, and FastAPI matches routes in declaration order, so every request
for the discovery list was routed into get_vip, which takes vip_id: int and
rejected "discoveries" with 422 before list_vip_discoveries ever ran.

The failure was completely silent. The agents reported their configs correctly,
the rows landed in vip_discoveries, and the HA/VIP page treats any non-OK
response as "nothing to show" - so the feature was invisible with no error in
any log. Confirmed against a real fleet: two discovery rows present in the
database, one with a parsed candidate, and an empty panel.

- move list_vip_discoveries above the /{vip_id} routes, with a comment stating
  the ordering requirement
- add tests/test_router_path_shadowing.py: a static source scan that fails if
  any literal path in any router is declared after a parameterised route that
  would swallow it. The whole router tree is clean; the detector is itself
  tested against the pre-fix ordering so the guard cannot pass vacuously.

POST /adopt is unaffected: no POST /{vip_id} route exists.

Also carries the release metadata for #58 (v1.10.4) and #60 (v1.10.5), which
are published together with this fix rather than as separate artifacts.

No schema, API-shape, frontend or agent change. Data reported under the earlier
code is not lost - existing rows show up as soon as this backend is deployed,
with no agent action needed.
2026-08-13 18:55:54 +03:00
taylanbakircioglu a4c74f2a52 Merge pull request #60 from mustafaulukaya/fix/acme-http01-split-deployment
HTTP-01 challenge backend on split deployments
2026-08-13 18:53:39 +03:00
Mustafa ULUKAYA 47cc79dcf7 fix(acme): address review findings on the challenge-backend hardening
Five confirmed findings from an adversarial review of the branch, four of them
regressions introduced by it.

Fall through to the next source when a stored URL cannot be resolved.
Stopping at the first non-empty candidate emitted a backend section with no
`server` line: the section exists so `haproxy -c` passes and Apply succeeds,
then every challenge request 503s from an empty backend with nothing to show
for it. Scheme-less values are common — the settings field was free text until
this branch — so this was reachable on real installs. Selection moved into
`select_acme_backend_source()` so it is testable and the skipped candidates are
logged rather than silently dropped.

Report a challenge backend with no server line. `extract_acme_backend_target`
returns None for that section, and the loopback filter skipped falsy targets,
so the case above would have been reported as "challenge route present in
applied config" — the new check confirming the very state it exists to catch.

Do not narrow the row set feeding the routing check's `fail` branch. Adding a
mode filter to the WHERE clause turned a tcp-only port-80 cluster from "ok"
into "fail", and the site wizard blocks submit on any failing check, so those
installs would have been locked on upgrade day. Mode is now examined in Python
and only downgrades to `warn`, using an expression that is character-for-
character the renderer's normalisation.

Match the agent's config selector. The applied-config lookup omitted
`is_active = TRUE`, so it could read a superseded row and report on a config
the nodes never received. Extraction now happens in SQL rather than pulling
whole configs — these run to hundreds of KB.

Select `acme_backend_url` when loading the existing cluster. It was absent, so
the entity snapshot recorded old_values as NULL unconditionally and rejecting
the pending version wiped the operator's per-cluster URL back to the global
loopback default — re-creating the exact failure this branch removes.

Also carry `acme_enabled` and `acme_backend_url` through cluster creation. The
create model declared neither and the INSERT wrote neither, so a cluster
created with ACME switched on came back switched off with no error shown.

Refuted and deliberately not changed: settings PUT re-validating a stored
loopback value (it validates only what is submitted), an apply-path connection
leak (the 422 propagates to a handler that closes it), and the modal discarding
backend rejection reasons (the envelope matches).
2026-08-11 19:13:07 +03:00
Mustafa ULUKAYA bb774141d4 fix(acme): make the challenge backend fixable from the panel
Correcting a wrong ACME challenge backend was impossible without a shell, and
even with one the correction did not reach the nodes.

The mint gate only fired when `acme_enabled` flipped. `acme_backend_url` was
written to the DB and minted nothing, so Apply answered "No pending changes to
apply" and the nodes kept the old address forever. It is now decided by
comparing the rendered `server _acme_mgmt` line against the active version —
the one line that answers "would the nodes talk to a different address?".
Comparing whole configs would flag every unrelated pending edit.

The field had no UI at all. Added to the cluster form with validation that
mirrors the backend rules, and keyed on `model_fields_set` so clearing it
reverts to the global setting — with a plain `is not None` test an empty box is
indistinguishable from "not submitted", so a value could never be removed.

Validation is asymmetric on purpose (utils/acme_backend_url):

- at the write boundary, reject what cannot express a reachable target —
  including the two silent traps: a scheme-less value became `localhost`, and
  an out-of-range port raised inside the generator and destroyed the config
- at render time, never reject. The shipped defaults are themselves loopback,
  so refusing to render would make every acme_enabled cluster unappliable,
  including for changes unrelated to ACME. Problems are logged and surfaced.

The port-less default stays 8080 rather than moving to HTTP's 80: the bundled
compose publishes nginx on 8080, so installs relying on it work today and the
first sign of breaking them would be the unattended renewal loop months later.
The omission is warned about instead.

RFC1918 is allowed and is usually the right answer here, and no DNS resolution
is performed — both deliberate departures from utils/ssrf_guard, whose policy
is the opposite of what this address needs. What the management host can
resolve says nothing about what the HAProxy node can reach.

Diagnostics stop reporting success on a dead path:

- check_port80 uses GET instead of HEAD and classifies the body. A proxy that
  has lost its /.well-known/acme-challenge/ location serves its SPA with HTTP
  200, which `status in (200, 404)` accepted as healthy. Warnings also surface
  when other domains pass, which previously hid the most diagnostic outcome.
- check_routing filters `mode`, joins `acme_enabled` and reads the APPLIED
  config instead of counting database rows, and reports a loopback target.
- every new condition is `warn`, never `fail`: the site wizard blocks submit on
  any fail, so a new failing condition would lock every install on upgrade day.

Also: normalise `frontends.mode` once per frontend. It is nullable, and the
raw value was interpolated into `mode {}`, emitting a literal `mode None` that
HAProxy rejects — taking down the whole cluster config. The ACME gate and the
backend-mode check now read the same normalised value.

And stop hardcoding PUBLIC_URL / MANAGEMENT_BASE_URL in docker-compose, which
silently ignored the operator's .env and made the wrong default load-bearing.
2026-08-11 17:15:55 +03:00
Mustafa ULUKAYA acfd32dd63 fix(acme): stop shipping config-generation failures as applied config
`generate_haproxy_config_for_cluster` reports failure by RETURNING a one-line
comment ("# Error generating configuration: ...") rather than raising. Nothing
in the backend checked for it, so the apply path hashed that comment, stored it
as an APPLIED config version and pushed it to every agent — silently replacing
a cluster's entire haproxy.cfg.

Any exception inside the generator triggers this. An out-of-range port in
`acme_backend_url` is enough: `urlparse('http://h:99999').port` raises
ValueError, the outer `except` swallows it, and the cluster loses its config.

- add `is_config_generation_error()` next to the sentinel definitions so call
  sites stop matching the string by hand
- guard the apply path: refuse with 422 and leave the running config in force
- guard the ACME-toggle PENDING mint, and let HTTPException through the local
  `except Exception`, which would otherwise report success while no pending
  version exists

Also make the challenge backend diagnosable without shell access, since this
block previously emitted no log line at all:

- log the rendered `host:port` and WHICH source chose it (cluster override,
  system setting, or the MANAGEMENT_BASE_URL fallback) under the greppable
  `ACME-BACKEND` keyword
- warn when the rendered address is loopback: HAProxy resolves it on the node,
  not on the management host, so it can only work on an all-in-one install —
  and it is exactly what the shipped defaults produce
- log peer, X-Forwarded-For and Host on the challenge endpoint, which answers
  "did the request arrive at all?" — the question that separates a wrong
  backend address from a wrong response body

Refs the HTTP-01 investigation: a split deployment rendered
`server _acme_mgmt <mgmt>:8080` against a port with no listener, and every
existing check reported success.
2026-08-11 16:58:00 +03:00
mustafa.ulukaya ef26860df9 feat(logging): unified request/response log with configurable retention (v1.11.0)
Until now the only record of what happened was `user_activity_logs`, which
stores non-GET 2xx operations with no bodies. When something failed you could
see that a counter went up, never what was sent or what came back.

This adds one queryable timeline covering both directions:

- inbound: every API call, including GETs and including 4xx/5xx, with the
  user, client IP, status, duration and — redacted, size-capped — the request
  and response bodies.
- outbound: every HTTP call the backend makes, tagged with who it went to
  (ACME/Let's Encrypt, Cloudflare, GoDaddy, HAProxy stats, agents, the ACME
  diagnostics probe).

Outbound rows inherit the inbound request's id, so one operator action and the
CA/DNS calls it triggered read as a single trace: opening a failed "Request
Certificate" shows the exact POST /acme/new-order and the CA's 429 underneath.

Implementation notes:

- Capture is a pure-ASGI middleware that TEES the request and response streams
  rather than draining them. `await request.body()` inside a BaseHTTPMiddleware
  would consume the receive channel and break the raw-body agent heartbeat
  handler. Registered last so it is outermost: it then sees the final
  client-visible response and seeds correlation_id_context before the error
  handler reads it.
- Rows are written by a batching background writer with a bounded queue, so the
  request path never awaits the database and a saturated logger drops rows
  visibly (surfaced on the page) instead of blocking. Redaction runs on the
  writer, off the request coroutine.
- Secrets never land: headers are an allowlist with Authorization/Cookie kept
  only as a presence marker; body keys and value shapes are redacted
  (passwords, tokens, api_token, API keys, private-key PEMs, JWTs); the ACME
  JWS request body is never stored, because a stored protected+signature pair
  is a replayable credential — a summary is logged instead; DNS-provider errors
  record only the exception type; the ACME HTTP-01 challenge endpoint is
  excluded so key_authorization is never captured.
- Retention is operator-configurable in Settings -> Request Log: separate day
  counts for successful and failed rows (7 / 30) plus a hard row cap (500k),
  whichever is reached first. Pruned in batches under a Postgres advisory lock,
  with the day counts bound as parameters, never interpolated.
- New permissions requestlog.read / requestlog.manage. super_admin and
  security_admin get both, operator gets read, viewer gets neither.

Schema: one new table (request_logs) plus its settings seed, SCHEMA_VERSION
10 -> 11, auto-migrated. No existing table altered, no agent or rendered-config
change. Kill switches: REQUEST_LOG_ENABLED=false (middleware never registered)
or the `enabled` toggle in Settings.

Tests: 245 new (7 backend files + 1 frontend), full suite 1655 backend +
17 frontend passing.
2026-08-11 02:36:03 +03:00
mustafa.ulukaya d92a7e9660 feat(vip): adopt a discovered keepalived instance into a managed VIP
GET /api/vip/discoveries lists what the agents found; POST /api/vip/adopt turns
one vrrp_instance into a managed VIP using the values from the node's own file
instead of retyping them. The VIP is created PENDING like any other, so nothing
reaches the node until it is applied from Apply Management.

Adoption replaces the operator's file with our render, so the gate is the
feature. Blockers fall into three kinds and only two are resolvable:

  - a LOSS ("our renderer cannot reproduce this, so adopting would delete it")
    can be accepted explicitly - that is an informed choice about a notify hook
    or an LVS section;
  - an UNKNOWN prefix length can be supplied, because picking a netmask for a
    live VIP would change its routing;
  - anything else is an IMPOSSIBILITY, not a loss: an absent virtual_router_id,
    a fractional advert_int, an unsupported auth_type. No flag waves those
    through.

That rule now lives in one place, remaining_blockers(), so the endpoint and the
UI cannot drift apart - and it is unit-testable, which matters because getting
it wrong destroys a working config.

Adoption keeps the VRRP identity it found: unlike create_vip, which allocates
the next free VRID, a VRID already used in the pool is a hard 409. Silently
renumbering would put the adopted node in a different VRRP domain from the
peers that still run the original config.

The member row records the reporting node's own role, priority and interface,
and carries the one-shot takeover hash. The response returns the instance's
unicast peers, because those nodes hold their own keepalived.conf and have to
be adopted or added as members before the render describes a complete group.
2026-08-11 01:35:59 +03:00
mustafa.ulukaya 7dfd31832a feat(vip): store the keepalived.conf an agent finds on a node
Ingest side of adopting an existing VIP. A node reports the keepalived.conf it
found and does NOT own to a new endpoint, and the finding is kept in a new
vip_discoveries table (one row per agent, since the file is per-node).

Reporting is read-only on the node. The heartbeat cannot carry this: it has the
VIP address and a best-effort MASTER/BACKUP, while rendering a node's config
needs eleven fields, so the file itself has to be read and parsed server-side -
parsing keepalived's block syntax in bash is not something to attempt on a
production load balancer.

Secrets are split at ingest. The reported content may contain the VRRP
auth_pass, so the password is Fernet-encrypted into its own column through the
same key path as vip_instances, and the stored copy of the file has it masked.
Nothing readable through the API, the UI preview or a database dump carries it
in cleartext, and the parse result is never logged. A file that does not parse
records its error rather than failing the agent's poll loop, and a report of
"the file is gone" clears the row so the UI stops offering a stale candidate.

The keepalived-config delivery gains allow_takeover and
takeover_expected_hash. The agent refuses to overwrite a keepalived.conf
without our ownership marker, which is the guard that protects a
hand-maintained setup - and is exactly the guard adoption has to pass. Rather
than weaken it, an adopted VIP authorises exactly ONE takeover of exactly the
file that was analysed by pinning its md5, so a config edited between adoption
and Apply is still refused instead of being silently overwritten.

SCHEMA_VERSION 10 -> 11 for the new table and two additive columns. Nothing
existing is altered, but the bump re-seeds the four built-in roles, which the
upgrade notes call out.
2026-08-11 01:35:59 +03:00
mustafa.ulukaya a6166d11b9 feat(ssl): add CSR generation and signed-certificate import (backend)
New /api/ssl/csrs endpoint group: generate a private key + CSR server-side
(RSA 2048/4096, ECDSA P-256/P-384; full subject + DNS SANs with wildcard
support), list/detail/delete CSRs, and import the CA-signed certificate.

- New ssl_csrs table (SCHEMA_VERSION 9 -> 10, additive + idempotent); the
  migration re-raises on failure so a failed run is retried instead of being
  stamped as applied.
- Import verifies the certificate against the stored key as a hard gate
  (match=None is treated as an integrity error, not a lenient pass), rejects
  malformed and expired certificates with 400, warns on SAN drift, and
  creates a normal ssl_certificates row (source=csr, cluster_id=NULL,
  last_config_status=PENDING) so it flows through the standard
  Apply Management -> agent pull pipeline.
- Concurrency: FOR UPDATE row lock serialises double-import and
  delete-during-import; a partial unique index reserves pending CSR names;
  soft-deleted same-name certs are reactivated preserving the row id.
- Security: no CSR endpoint ever returns the private key (explicit column
  lists, enforced by a static test); the key copy on the CSR row is NULLed
  after import; ssl.create/read/delete permissions enforced on every
  endpoint incl. reads; per-user rate limit on key generation, which runs
  in a worker thread; csr_id and cluster_ids are int32-guarded.
- ssl_service: extract _prepare_cert_fields from create_cert_row (behaviour
  unchanged, extraction tests untouched) and add stage_ssl_config_versions
  reusing the exact ssl-{id}-create-{ts} version-name scheme.
- Tests: crypto round-trip for all four algorithms, model validation,
  import-flow unit tests, endpoint auth/permission pinning, migration and
  key-non-exposure static assertions.
2026-08-04 21:23:57 +03:00
taylanbakircioglu 9e5185c458 fix(security): post-review hardening — agent-inventory regression, coverage gaps, SSRF newNonce
Follow-up to the RCE/missing-auth/SSRF remediation, from a thorough multi-lens
review (3 agents + a black-box audit of all 201 routes). Backend-only; no
agent-script changes.

Regression fix (introduced by the previous commit):
- GET /api/agents was made JWT-only, but deployed agents call it WITH X-API-Key
  (not a JWT) to read their applied_config_version and avoid re-applying config on
  restart. It now accepts EITHER a valid operator JWT OR a valid agent X-API-Key,
  so agents no longer get 401 (which caused a spurious HAProxy reload every restart).

Completeness (GHSA-3p5c siblings the first pass missed — same data class, now JWT):
- dashboard.py: GET /api/haproxy-cluster-pools/{id}/agents (full agent inventory —
  a direct anonymous bypass of the GET /api/agents lockdown), /api/pools,
  /api/haproxy-cluster-pools, /api/dashboard/stats, /api/dashboard/overview
  (auth was optional -> leaked stats/names/health/alerts anonymously),
  /api/haproxy/stats.
- waf.py: GET /api/waf/rules. health.py: GET /api/health/errors.
- agent.py: GET /api/agents/generate-uninstall-script/{platform} (agent-management
  endpoint; was anonymous) now requires JWT or agent key, like generate-install-script.
- config.py: POST /api/config/{validate,optimize,templates/{id}/generate} were
  optional-auth (logging only) and run a HAProxy validator on caller input; now
  require a JWT. (bulk-create, parse-bulk, diff and configuration/request were
  already mandatory-auth — verified.)
  All newly-gated endpoints are frontend-only (axios sends the JWT) or unused;
  agents never call them.

SSRF (GHSA-3vh4) gap:
- acme_service._get_nonce fetched directory['newNonce'] (from the attacker-
  influenceable directory JSON) with a bare session, http allowed, dual-stack, and
  BEFORE the guarded _signed_request POST. Now guarded (assert_public_url +
  safe_connector + no redirects + timeout), matching the other ACME sinks.

Correctness:
- Three agent webhooks (config-applied, config-validation-failed, config-sync)
  swallowed their auth 401 into a 200 error body via a bare `except Exception`.
  Added `except HTTPException: raise` so the 401/403 propagates.

Audit result (live black-box, all 201 routes probed unauthenticated): no data
leak and no unauthenticated mutation anywhere; every sensitive route returns
401/403 (a pre-existing group of read handlers wraps the 401 into a 500 via a
broad except — no data is exposed; left as-is, documented as cosmetic).

Verified: full pytest tests/ (1145 passed, 0 failed; +16 regression tests) + live
localtest stack smoke — agent-key GET /api/agents=200, anonymous=401, all newly
gated endpoints reject anonymous and admit JWT, the 3 webhooks return 401.
2026-07-20 13:43:32 +03:00
taylanbakircioglu 520b69a1c6 fix(security): remediate RCE, missing-auth and SSRF advisories (backend-only, no agent changes)
Addresses three reported advisories, all verified against the code. Fixes are
entirely server-side — deployed agents already send a valid X-API-Key on every
call, so enforcing it does not require any agent-script change or upgrade.

GHSA-7rhv-c5pc-69r8 (CRITICAL RCE — agent script-template poisoning):
- POST/GET /api/agents/script-templates/{platform} now require the agents.version
  permission (was authentication-only), matching POST /versions. Blocks a viewer
  JWT from overwriting the root install/upgrade script.

GHSA-3p5c-m5m4-mjpx (missing authentication):
- Agent data-plane endpoints now REQUIRE a valid X-API-Key (was optional/skipped
  when the header was absent), checked before any DB access: config,
  ssl-certificates (private keys!), upgrade-status, heartbeat (by-name and the
  previously auth-less by-id), configuration pending-requests. Removes keyless
  heartbeat spoofing and keyless rogue-agent auto-registration.
- Operator/UI endpoints now require a JWT: GET /api/agents, the entire
  /api/dashboard-stats router, /api/health/{deep,agents,clusters}, and
  /api/ssl/certificates/{id}/config-versions. The simple /api/health liveness
  probe stays public. Adds shared auth_middleware.require_authenticated_user.

GHSA-3vh4-gvxx-wm2p (SSRF via ACME directory_url):
- New utils/ssrf_guard.py (https-only + public-IP-only, IPv4-pinned, no redirects),
  applied to settings test-connection, acme_service.get_directory and
  _signed_request, and validated at Let's Encrypt account creation. The
  test-connection response no longer reflects arbitrary upstream JSON keys
  (information-disclosure oracle) — only fixed ACME field names.

Verified: full pytest tests/ (1128 passed, 0 failed) + live localtest stack smoke
(valid JWT/key paths return 200/404 as expected; anonymous requests 401; SSRF to
metadata/private/loopback refused). No changes to backend/utils/agent_scripts/*.
2026-07-20 12:45:28 +03:00
taylanbakircioglu c79391cd13 feat(acl): accept HAProxy -f pattern-file references with advisory warnings (v1.8.9, Issue #38)
The manual Frontend editor, wizard and visual ACL builder hard-rejected the ACL
`-f <file>` flag while bulk import accepted it. Worse, a frontend imported with
an `-f` ACL could not be edited at all (422) until the ACL was dropped.

The original guard predated the fail-safe apply flow: the agent runs `haproxy -c`
before every reload, so a missing pattern file is rejected safely and the previous
config keeps running. Pattern files are operator-managed host files — the same
policy adopted for SPOE filter configs in v1.8.8.

- models: remove the 5 `-f` hard rejects (frontend acl/redirect/use_backend
  validators + wizard string/dict-redirect guards); `$(`/backtick and X!X
  contradiction guards unchanged
- routers/frontend: `_pattern_file_warnings` helper; non-blocking warning on
  create + update responses listing referenced pattern files (empty when no
  rule uses `-f` — zero noise)
- routers/config: bulk-import preview advisory listing pattern files per
  frontend (cluster config-dir aware, next to the SPOE advisories)
- React: remove the FrontendManagement submit gate and SiteWizard step gate;
  ACLRuleBuilder renders informational notes instead of errors and re-adds
  `-f (pattern file on host)` to the flag dropdown; create path now renders
  server warnings like update
- tests: 4 reject-pins inverted to accept-pins; new test_acl_pattern_file_allow.py
  (accept/guards-kept/zero-noise/advisory); full suite green (1094 passed)
2026-07-14 00:27:11 +03:00
taylanbakircioglu 9e2ea04777 feat(haproxy): preserve SPOE filter + frontend log-format on import/edit (v1.8.8, Issue #38)
Bulk import / manual edit silently dropped `filter spoe engine ...` (Coraza WAF)
and frontend `log-format` because the parser recognised only a fixed directive
set. The regenerated config then missed the SPOE engine, so HAProxy failed with
"unable to find SPOE engine 'coraza' used by the send-spoe-group".

- parser: capture `filter` + `log-format`/`log-format-sd` into new ParsedFrontend fields
- db: additive nullable `log_format` + `filters` TEXT columns on frontends (SCHEMA_VERSION 8->9)
- generator: new `filter` bucket flushed before http-request rules so `filter` precedes
  `send-spoe-group`; `log-format` kept in prelude
- bulk import: preview dict, change-detection, persist (create + merge-update); cluster-aware
  SPOE pre-flight advisories (missing-filter + host-prerequisite) surfaced in the UI
- manual CRUD: full round-trip (get/create/update) incl. React form fields (no null-wipe)
- reject/rollback: restore the new columns; restore path + wizard helper kept in parity
- backend `option spop-check` recognised (suppresses spurious warning for coraza-spoa)
- tests: test_spoe_filter_import.py; full suite green (1079 passed)
2026-07-10 18:34:36 +03:00
taylanbakircioglu 23257b02cf perf(api): opt-in uvicorn workers + heartbeat query consolidation (v1.8.6, Issue #35)
A user running the API on a 2-core/4GB host reported slow-feeling API
responses (Issue #35 follow-up). Review of the hot paths found no
pathological defect; the dominant factors are the single uvicorn worker
(one core serves all requests) and the constant agent-poll baseline
(4 requests per agent every 30s). Two zero-risk improvements:

- backend/Dockerfile: CMD now honors UVICORN_WORKERS, falling back to
  WEB_CONCURRENCY and then 1. Flagless uvicorn natively honors
  WEB_CONCURRENCY, so the fallback keeps any deployment that relied on it
  byte-for-byte compatible; with the final default of 1 worker uvicorn runs
  in-process exactly as before. >1 enables the multiprocess supervisor so
  multi-core hosts can use all cores. Background tasks are already
  multi-replica safe (FOR UPDATE SKIP LOCKED / advisory locks), as exercised
  by the k8s HPA deployment (2-10 replicas). `exec` keeps uvicorn as PID 1
  (clean SIGTERM, verified ~1s docker stop with 2 workers).
  docker-compose.yml passes UVICORN_WORKERS through as empty-when-unset so a
  user-set WEB_CONCURRENCY is never overridden; .env.template documents it.

- routers/agent.py heartbeat (by-name endpoint): the agent's
  status/version/upgrade_status were read with three separate single-column
  SELECTs against the same row; now one SELECT. Identical values and None
  semantics (single consistent snapshot instead of three reads); saves two
  round-trips per heartbeat per agent every 30s. The legacy by-id heartbeat
  endpoint is untouched; the heartbeat API contract is unchanged for agents
  of every version.

- README: new "Performance Tuning" section (worker/replica scaling, and how
  to use the X-Response-Time header plus "Slow request detected" logs to
  pinpoint slow endpoints).

Verification: full backend suite in docker green (1063 passed, 151 skipped;
also re-run by the runtime image build); worker-count expansion matrix
(unset->1, UVICORN_WORKERS=2->2, WEB_CONCURRENCY=3->3, both->UVICORN_WORKERS,
empty->fallback) all correct; default run confirmed single-process with
uvicorn as PID 1 and healthy API; UVICORN_WORKERS=2 confirmed parent + 2
workers, healthy API, clean shutdown; live heartbeats verified for
register + existing-agent paths AND degraded agents (no stats socket /
haproxy stopped / garbage stats CSV / unknown backend in server_statuses):
all return 200, agent row updates correctly, zero backend errors.
No schema, API, or agent changes.
2026-07-06 02:27:39 +03:00
taylanbakircioglu 27fbe48c4b fix(agent): tolerate empty system_info in heartbeat JSON (v1.8.3)
A self-hosted agent could fail every heartbeat with HTTP 400
`Invalid JSON: Expecting property name enclosed in double quotes` when the
system-info block it collects came back empty on an unusual host. The agent
builds the heartbeat JSON as text, so an empty `$system_info` collapsed the
`$system_info,` line to a bare comma and broke the payload.

- Agent script (linux + macos, kept in sync): guard the fragment-form
  register_agent and send_heartbeat builders so an empty system_info falls back
  to a valid key and can never emit a bare comma. Uses the most portable bash
  glob test (no POSIX class / pattern substitution; verified on bash 3.2-5.2 and
  on Ubuntu/Debian/Rocky/Alpine/Amazon Linux). True no-op for healthy agents.
- Backend heartbeat endpoint: parse the body as-is first and only run the
  malformed-JSON repair when parsing fails, so a valid heartbeat from any agent
  version is byte-for-byte untouched. The repair (now a testable helper) recovers
  a leading or doubled comma (the empty-system_info artifact) in addition to the
  existing empty-value / trailing-comma fixes.

No agent version bump; self-upgrade and daemon mode are unaffected. Healthy
agents of every version behave identically. Full backend suite green.

Addresses #31.
2026-06-25 14:42:34 +03:00
taylanbakircioglu 70ebc02e09 fix(acme): Cloudflare token, ZeroSSL EAB, Apply Management (v1.8.1)
Follow-up fixes for the DNS-01 feature reported on #35:

- Cloudflare: sanitize the API token (strip surrounding quotes + any non
  token68 chars) so a pasted token with quotes/spaces no longer fails with
  "Invalid request headers"; verify-on-save shows a precise hint when it
  cleaned the input. Covers the automated orchestrator path too.
- ZeroSSL/Google EAB: enter the EAB Key ID and HMAC Key per-account in the
  Register Account dialog (falls back to the global Settings value when blank);
  base64-validate the HMAC key; humanize the externalAccountRequired failure;
  and preserve the deliberate 409/422 instead of downgrading them to 400.
- Apply Management: cluster ACME enable/disable changes now show under a
  dedicated "ACME Challenge Routing" section, are counted in the Apply/Reject
  dialogs, and Apply/Reject All process them (previously "Rejected 0 HA/VIP
  change(s)") - consistent with every other entity. Reject rolls acme_enabled
  back to the original via ORDER BY created_at ASC over the snapshot chain.
- getErrorMsg surfaces field-level validation messages.

Backward compatible (additive / strict superset; HTTP-01 unchanged).

Addresses #35.
2026-06-24 20:37:05 +03:00
taylanbakircioglu c492b26bb1 feat(acme): DNS-01 challenge support with pluggable DNS providers (v1.8.0)
Add ACME DNS-01 (TXT-record) validation alongside the existing HTTP-01,
for internal/isolated clusters with no public port 80 and for wildcard
certificates. Opt-in via a global kill-switch (default off); HTTP-01 is
byte-for-byte unchanged, with zero agent or rendered-config changes.

- Pluggable DNS provider interface (Manual + Cloudflare). Per-account
  credentials are Fernet-encrypted at rest, verified on save, and never
  returned by the API or written to logs/events/error_detail.
- Non-blocking per-cycle orchestrator: publish (CAS) -> propagation grace
  (across cycles, no in-loop sleep) -> respond -> finalize/download, with a
  bounded fresh-order retry chain (1 original + 3 retries) on propagation lag.
- Manual flow: user publishes the TXT record and confirms; manual DNS-01
  cannot auto-renew unattended (auto-renew forced off and surfaced in the UI).
- Migration v8: additive, idempotent columns on letsencrypt_accounts/orders
  and acme_challenges, plus a new letsencrypt_account_dns_credentials table.
- Challenge-type-aware diagnostics (port80/routing/DNS checks skipped for
  DNS-01) and a DNS-01 event timeline in the order detail.
- Frontend: DNS-01 account + credentials management, cert wizard adaptation,
  order-detail TXT records + verify, orders/renewal Method columns, and a
  Settings kill-switch. README, release notes, and API docs updated.

Implements #35.
2026-06-24 02:24:33 +03:00
taylanbakircioglu a1192e602d feat: HA / VIP (Keepalived) management from the UI (#27)
Manage highly-available virtual IPs backed by Keepalived (VRRP) directly from the OpenManager
UI — no more SSHing into nodes to install/configure Keepalived by hand. Builds on the agent
pull-architecture: define the VIP centrally, click Apply, and the agents converge.

Highlights:
- New "HA / VIP" tab: create a virtual IP, pick a per-node interface, select which pool nodes
  participate (MASTER/BACKUP roles + priorities); live MASTER/BACKUP/FAULT per node.
- On Apply, agents install & configure Keepalived (unicast VRRP, cloud-safe default) across the
  major distros (Debian/Ubuntu, RHEL/CentOS/Alma/Rocky, Fedora, SUSE/openSUSE, Alpine) with a
  HAProxy health-check, so the VIP fails over automatically when HAProxy drops.
- Single-node (a managed floating IP without failover) and multi-node VRRP failover both work.
- VIP changes ride the standard Apply Management flow with the standard "View Change" diff.
- Approval-gated deletion (safety): deleting a running VIP is staged for approval and the VIP
  keeps running, untouched, until you approve it — an agent never tears a VIP down without an
  explicit human approval. Per-VIP Diagnostics view; opt-in package uninstall (only on nodes
  where OpenManager installed it). A node already running a hand-managed Keepalived is detected
  and never overwritten ("externally managed").
- Fully opt-in and backward compatible: nodes/clusters without a VIP are unaffected. Adds
  vip_instances + vip_members tables (idempotent SCHEMA_VERSION bump; existing data and
  passwords unaffected) and a `vip` RBAC permission group.
- Also includes a HAProxy config-generator robustness fix: auto-inject a stick-table when a
  frontend uses a stick counter (track-sc / sc_*_rate) but declares none.

On-prem / L2 (VRRP) scope; the UI notes the cloud caveat.
2026-06-07 01:52:20 +03:00
taylanbakircioglu b34d7cf811 fix: reactivate disabled backend servers from the UI + v1.6.3 (Issue #24)
A backend server toggled OFF (is_active=false) vanished from the UI with no way
to reactivate it: GET /api/backends honored include_inactive for backends but the
server sub-queries hardcoded 'AND is_active = TRUE'.

- get_backends: server sub-queries now honor include_inactive (default callers
  unchanged); added last_config_status to the server payload so the UI can tell a
  DISABLED server (re-enableable) from a DELETION (pending delete).
- toggle_server: persists an entity snapshot so an Apply-Management Reject rolls
  back is_active (previously left the server stuck disabled).
- BackendServers.js: requests include_inactive, shows disabled servers with the
  ON/OFF switch + an 'Inactive' tag, hides only DELETION-pending servers, and
  keeps soft-deleted BACKENDS hidden (so include_inactive doesn't resurface them).
- Config generation unchanged: disabled servers stay '# DISABLED:' comments and
  convert back to live lines when re-enabled.

Startup migration hardening (multi-replica / rolling-deploy safety): create_essential_tables
fails fast on lock contention and retries; run_all_migrations is serialized by a
session advisory lock and gated by a schema_migrations version marker, so an
already-current schema is skipped instead of issuing lock-heavy DDL that a serving
replica's traffic could block at startup. Idempotent and fail-open.

Version reported consistently across all layers (version.json, backend fallback,
frontend package) -> 1.6.3.
2026-06-02 02:42:48 +03:00
taylanbakircioglu f3d4fb11bb fix(agent): accept agent token (X-API-Key) on cluster read endpoints (Issue #22)
Assigning a HAProxy agent failed with '401: Authorization header missing'
on GET /api/clusters. A cluster-read hardening had made GET /api/clusters and
GET /api/clusters/{id} accept only a user JWT in the Authorization header;
agents authenticate with their agent token in the X-API-Key header, so the
token was never read.

Both endpoints now accept either a user JWT (Authorization) or an agent token
(X-API-Key via validate_agent_api_key), mirroring the existing dual-auth on
POST /api/agents/generate-install-script. Anonymous access is still rejected,
so the original hardening is preserved. The auth guard is placed before the
try block so the failure surfaces as a clean 401 (not the 500-wrapped-401 in
the report). Agent install scripts now consistently send the token via
X-API-Key (pre-flight cluster check on linux/macos, and macOS get_cluster_paths
which previously used the wrong Authorization: Bearer header).

Also normalizes the platform in the uninstall-script generator so macOS agents
(which report platform 'darwin') no longer get a 400 from
GET /api/agents/generate-uninstall-script/darwin.

version 1.6.0 -> 1.6.2.
2026-05-30 19:09:26 +03:00
taylanbakircioglu bd6a31cb0d feat: v1.6.0 — Multi-Factor Authentication (Issue #18)
Adds opt-in TOTP-based Multi-Factor Authentication that is fully
backwards compatible with existing logins. Operators choose to enable
MFA per account; nothing changes for users who do not opt in.

Highlights
==========

* RFC 6238 TOTP (6 digits, 30s period, SHA1) with ±30s skew tolerance,
  compatible with Microsoft / Google Authenticator, Authy, Duo, 1Password.
* Per-step replay protection (`mfa_last_used_totp_step`) so a captured
  code cannot be reused inside the same window.
* Fernet-encrypted TOTP secrets at rest, key resolution via
  `MFA_ENCRYPTION_KEY` env (HKDF-derived from `SECRET_KEY` as fallback).
* 10 single-use, bcrypt-hashed backup codes per user, formatted
  `XXXX-YYYY` from a confusion-free alphabet (no 0/O/1/I/L).
* Two-step login flow: `POST /api/auth/login` returns `mfa_required`
  + `mfa_token`, then `POST /api/auth/login/mfa-verify` accepts a TOTP
  code OR a backup code. JWT is minted only after MFA succeeds.
* Self-service: users enable / disable MFA from their own row in the
  Users page; admins reset (single user or bulk) but never enable on
  behalf of someone else (matches AWS IAM / GitHub / Google Workspace).
* Bulk emergency reset CLI: `scripts/admin-mfa-reset-all.sh`.

Security hardening
==================

* Atomic transactions with `SELECT … FOR UPDATE` on `mfa_pending_logins`
  and `users` rows so concurrent verify / enroll calls cannot race.
* `/api/mfa/enroll/start` refuses re-enrollment when MFA is already on
  (prevents silent secret rotation via a stolen JWT).
* Pydantic `ValidationError` messages are sanitized before reaching the
  audit log so request bodies (TOTP / backup codes in flight) never
  appear in plaintext.
* Slowapi rate limits are per-USER, not per-IP, with a trusted-proxy
  XFF strategy so a single ingress address cannot exhaust the bucket
  for thousands of operators (`MFA_TRUSTED_PROXY_CIDRS`,
  `MFA_RATE_LIMIT_*` env-overridable).
* Login query now scopes to `is_active = TRUE` so a soft-deleted row
  with the same username can no longer occlude the active user
  (also closes a small account-enumeration side channel).

Database
========

Additive migrations (idempotent `ADD COLUMN IF NOT EXISTS`,
`CREATE TABLE IF NOT EXISTS`):

  - users: mfa_enabled, mfa_method, mfa_secret_encrypted,
    mfa_enrolled_at, mfa_last_used_at, mfa_last_used_totp_step
  - mfa_backup_codes (user_id ON DELETE CASCADE)
  - mfa_pending_logins (user_id ON DELETE CASCADE, challenge_token,
    attempts, expires_at)
  - mfa_pending_enrollments (user_id ON DELETE CASCADE)

Frontend
========

* Login page becomes a 3-phase state machine
  (credentials → MFA → submitting); legacy single-step login is
  preserved for users who haven't enrolled.
* New MFAEnrollModal (3-step wizard: QR + secret → verify → backup
  codes) using `qrcode.react`.
* Users page shows MFA column + per-row enable/disable/reset actions.
  Admins viewing other users with MFA off see a non-actionable info
  icon explaining that only the user themselves can enable MFA.

Deployment
==========

* `MFA_ENCRYPTION_KEY` is added to `k8s/manifests/03-secrets.yaml` as
  a placeholder; `SECRET_KEY` is also placeholder-ized so both are
  injected by the existing pipeline pattern (sed-replace + apply).
* No new build-time env vars are required for the frontend. The SPA
  uses `window.location.host` for `/api/*` and is routed by the
  existing nginx ingress configuration.
* `frontend/.dockerignore` ensures host `.env*` files cannot bleed
  into the production bundle.

Tests
=====

* New unit suites:
  - `test_mfa_service.py` (TOTP, encryption, backup codes)
  - `test_mfa_backwards_compat.py` (regression — non-MFA flow unchanged)
  - `test_mfa_rate_limits.py` (env override + dataclass immutability)
  - `test_mfa_rate_limit_key.py` (JWT key, trusted-proxy XFF, fallbacks)
* All existing 1000+ unit tests continue to pass.

Documentation
=============

* README MFA section (overview, day-to-day operations, emergency
  reset CLI, env variables, rate-limit tuning).
* `scripts/README.md` documents the bulk reset script.

Issue: #18
2026-05-19 04:35:16 +03:00
taylanbakircioglu d7208528f7 fix: v1.5.2 — ACME Diagnostics Panel Hardening + AGPL-3.0 relicense (Bulgu #94/#95/#96)
A focused hardening pass on the v1.5.0 ACME Diagnostic Panel
surface, exercised against a live production deployment (Round-25
+ Round-26 audits) and supplemented by an AGPL-3.0 relicense.

------------------------------------------------------------------
LICENSE — Relicense to AGPL-3.0-or-later
------------------------------------------------------------------
Effective v1.5.2 the project is licensed under the **GNU Affero
General Public License v3.0 (or later)**. v1.5.0 and v1.5.1
remain under the prior MIT terms.

The relicense is consistent with the project's intent as a
community-operated HAProxy management surface: forks that run
HAProxy OpenManager as a network service for third parties are
now required to publish their modifications under the same
license (AGPL §13). Day-to-day single-tenant deployments,
internal corporate use, and ordinary forks-for-fixes are
unaffected.

Changes:
  * LICENSE replaced with full AGPL-3.0 text.
  * README "## License" section rewritten with the AGPL summary
    + the network-service obligation.
  * frontend/package.json gains `"license": "AGPL-3.0-or-later"`.

------------------------------------------------------------------
BULGU #94 / #95 — Diagnostic Panel Must Never Opaque-500
------------------------------------------------------------------
Live exercise of the v1.5.0 Diagnostic Panel against a deployed
build surfaced two opaque-500 paths. The panel exists to make
ACME failures legible; producing an opaque HTTP 500 defeats the
entire feature. Fix shape: every endpoint now returns either a
canonical 4xx (auth / not-found / rate-limit) or an HTTP 200
"structured failure envelope" that the React UI knows how to
render — never a 500 for an in-suite failure.

Affected paths:

POST /api/letsencrypt/orders/{order_id}/diagnostics
  Pre-fix: a UndefinedColumnError or DB-connectivity failure
  inside `run_checks` bubbled out of the bare try/finally and
  surfaced as a generic 500 with no operator-actionable detail.
  Post-fix: setup-stage and run-stage failures are caught
  separately and converted to a `status: diagnostics_unavailable`
  envelope carrying `error_stage`, `error_type`, `error_message`,
  and a `correlation_id` that the operator can grep in the
  backend log. Individual checks are wrapped in `_safe_check`
  so one broken check (e.g. DNS lookup timeout) never crashes
  the suite — the failing check shows up as `status: "fail"`
  with its message, the others still run.

GET /api/letsencrypt/orders/{order_id}/events
  Pre-fix: the SQL `SELECT … status FROM user_activity_logs`
  referenced a column that did not exist in the canonical
  migration; every diagnostic-panel open against an order with
  any user-activity-log correlation got an `UndefinedColumnError`
  500. Post-fix: the endpoint now introspects
  `information_schema.columns` and projects only the columns
  actually present. Partial failures (one source dies, the
  other works) are reported via `meta.errors[]` rather than
  collapsing the whole timeline.

POST /api/letsencrypt/orders/{order_id}/diagnostics/{check_id}/rerun
  Same structured-envelope contract as the full-suite POST,
  scoped to a single check row.

Frontend (`frontend/src/components/ACMEAutomation.js`):
  * Distinct `diagRunError` / `diagEventsError` / `diagMeta`
    states so the modal can render the cause inline (Antd Alert)
    instead of a silent dropdown.
  * Event-log auto-tail polling backs off after 3 consecutive
    failures so the Network tab does not get spammed with 500s
    every 5s.
  * Correlation IDs visible in every error banner.

------------------------------------------------------------------
BULGU #96 — Clean 404 for Out-Of-Range order_id
------------------------------------------------------------------
A live exercise of the post-#94 diagnostic panel against the
deployed build surfaced one remaining contract gap. A path-
param `order_id` outside the Postgres int4 range
(e.g. > 2_147_483_647) caused `_load_order` to raise
`asyncpg.exceptions.DataError: invalid input for query
argument $1: 2147483648 (value out of int32 range)`. Round-25
correctly surfaced this in a `diagnostics_unavailable`
envelope — but that envelope leaked SQL implementation detail
("query argument $1", "int32 range", DataError class name)
into the operator-facing response body.

Semantically an out-of-range integer can never reference a
real order — it's just "not found". `_load_order` now catches
`asyncpg.exceptions.DataError` and re-raises a canonical
`HTTPException(404, "Order {id} not found")`. Because all
three endpoints re-raise `HTTPException` from their outer
try/except (the Round-25 envelope only fires for non-
HTTPException crashes), the canonical 404 path now wins
end-to-end across /diagnostics, /events, and /rerun.

------------------------------------------------------------------
TEST / LINT / LIVE VERIFICATION
------------------------------------------------------------------
  * Backend pytest 1104/1104 (the +20 vs v1.5.1's 1084 are the
    Round-25 and #96 contract pins; see
    test_acme_diagnostics_router_round25.py).
  * Live prod-canary verification: every endpoint return shape
    confirmed against the deployed build — int4 overflow returns
    clean 404 with no SQL leak, normal paths return Round-25
    envelopes, HTTP method matrix returns 405 on wrong verbs,
    no auth returns 401, invalid `check_id` returns 400, and
    `meta.correlation_id` is present on every diagnostic
    response.

------------------------------------------------------------------
COMPATIBILITY
------------------------------------------------------------------
  * No breaking API contract changes: `status` field on the
    diagnostic response can now be `"diagnostics_unavailable"`
    in addition to the existing pass-through of the
    underlying order status (`pending` / `valid` / `invalid` /
    `cancelled` / …) — older UIs that only switch on the
    existing values render the `diagnostics_unavailable`
    case as "unknown status" rather than crashing.
  * Frontend handles the new envelope shape AND the legacy
    HTTP 4xx/5xx paths.
2026-05-14 00:07:00 +03:00
taylanbakircioglu 2e7db4d99f fix: v1.5.1 — Round-23 + Round-24 audit follow-ups (Bulgu #83#93)
A live-deployment audit pass over the v1.5.0 Site Wizard + ACME
Diagnostic Panel surface. Two adversarial review rounds (R23, R24)
each capped by an end-to-end smoke test against a multi-cluster
staging deployment.

Bulgu #83 — Frontend Management page warned about stale data
without a clear retry CTA. The toast now carries an in-place
"Reload" action and the page-level Empty state surfaces the same
recovery affordance, so operators never get stuck on a stale-data
view without an obvious way out.

Bulgu #84 — ACME diagnostics ran with the wrong "last_heartbeat"
column reference against the agents table. Aligned the SELECT
with the actual schema column (`last_seen`); pinned by an idempotent
regression test in `test_acme_diagnostics.py`.

Bulgu #85 — ACME order error_detail rendering could leak the raw
asyncpg/SQL exception class name when humanize_error_detail
encountered an unhandled CA response shape. Added a backwards-
compatible fallback branch that emits an "ACME error (raw)" panel
without exposing parse_error class name to the user.

Bulgu #86 — Multi-cluster apply with concurrent rejects could
leave wizard_staged orders dangling without their parent draft.
Pinned via reject_order_with_cluster_orphan test.

Bulgu #87 — Frontend Management page list virtualization
mis-keyed during a re-sort + stale-row replace race; fixed by
keying rows on `id + version` so React reconciler does not reuse
DOM for a logically different row.

Bulgu #88 — Site Wizard "Cancel" mid-flow now surfaces an
unsaved-draft prompt with explicit Save / Discard buttons (and
the same prompt on browser tab close), so the operator never
loses 5 steps of input to an accidental ESC.

Bulgu #89 — Existing-cert SSL mode showed an empty dropdown when
the cluster had >100 certs because the listing endpoint
default-limited results. Endpoint now exposes pagination AND
the wizard switches to client-side filtering above 50 rows.

Bulgu #90 — ACME pre-check on the wizard preview path did NOT
re-validate the account against `letsencrypt_accounts` if the
operator stepped Back/Forward between SSL and Review. Added a
debounced re-validation on Review entry.

Bulgu #93 — Site Wizard hsts_enabled toggle in HTTPS frontend
was idempotent-by-name (the generated `http-response set-header
Strict-Transport-Security` line could duplicate across a Save +
Apply cycle). The renderer now upserts the header in place.

Cumulative outcome: backend pytest 1084/1084, frontend lint
clean, and a 6-hour live-deployment smoke session against staging
with no regressions reported.
2026-05-14 00:06:02 +03:00
taylanbakircioglu 02b1cb2bca feat: v1.5.0 — Site Wizard (Issue #14) + ACME Diagnostic Panel (Issue #13)
Closes #13, Closes #14.

This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0
shipped the ACME stability & enterprise audit (Issues #10/#11/#12).
v1.5.0 builds on that foundation with two co-equal headline features
plus a 22-round audit campaign hardening the prior configuration
surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0
lands in v1.5.2).

------------------------------------------------------------------
HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13)
------------------------------------------------------------------
A live pre-flight + post-failure diagnostic surface for every ACME
order, reachable from the ACME Automation page. The panel exists
to make ACME failures legible to operators who do NOT have shell
access to the API host.

Endpoints (`backend/routers/acme_diagnostics.py`):
  POST /api/letsencrypt/orders/{order_id}/diagnostics
       Run the full 5-check suite (DNS / port-80 / routing /
       account / agents) and humanize the order's `error_detail`
       (>=11 RFC-8555 problem types, backwards compatible with
       legacy plain-string failures).
  POST /api/letsencrypt/orders/{order_id}/diagnostics/
                                {check_id}/rerun
       Re-run a single check in place — used by the "Re-run"
       button on every row of the modal's pre-flight table.
  GET  /api/letsencrypt/orders/{order_id}/events
       Merged event timeline combining the typed
       `acme_order_events` rows with correlated
       `user_activity_logs` entries (resource_type =
       'letsencrypt_order' AND resource_id = order_id). The
       diagnostic modal auto-tails this timeline every 5 seconds
       while open.

Service-level checks (`backend/services/acme_diagnostics.py`):
  * DNS resolution via stdlib socket.gethostbyname_ex through
    run_in_executor (intentionally avoiding an aiodns runtime
    dep for v1.5.0).
  * Port-80 HEAD probe, target locked to the order's domains,
    success on HTTP 200 OR 404, warns on egress timeout
    (corp egress policies routinely blackhole outbound 80 —
    fail-hard would be too noisy).
  * SSRF guard: probe refuses non-public IPs and surfaces the
    skip in the diagnostic result; IPv4-mapped IPv6 normalisation
    closes the `::ffff:169.254.169.254` cloud-metadata vector.
  * HAProxy routing presence check: matches the order's
    cluster_ids to a port-80 HTTP frontend.
  * ACME account validity check against `letsencrypt_accounts`.
  * Agent presence check (>=1 active agent in target cluster).
  * Every sub-check wrapped in a wall-clock timeout to bound
    impact on the API event loop.

RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate
limit on both run and rerun, backed by the (user_id, action,
created_at DESC) composite index.

Frontend (`frontend/src/components/ACMEAutomation.js`):
  * "Diagnose" button on every order row + the existing
    "stuck order" warning row.
  * Modal with two tabs:
    - Pre-flight Checks (Antd Table with status pills + Re-run
      buttons + humanized error banner)
    - Event Log (Antd Timeline with auto-tail polling, scroll-
      to-bottom, pause-on-hover)
  * Correlation IDs surfaced in error banners and individual
    check fail details for backend-log lookup.

------------------------------------------------------------------
HEADLINE FEATURE B — Site Setup Wizard (Issue #14)
------------------------------------------------------------------
A single guided flow that creates a Backend + Servers + HTTP
Frontend (and optional HTTPS Frontend) in one atomic transaction.

Endpoints (`backend/routers/site_wizard.py`):
  POST /api/site-wizard/preview     — diff-preview the changeset
  POST /api/site-wizard/create      — atomic execute
  POST /api/site-wizard/reject      — clean rollback (including
                                       any wizard_staged ACME
                                       orders)
  GET  /api/site-wizard/drafts      — draft persistence
  PUT  /api/site-wizard/drafts/{id} — save/update
  DELETE /api/site-wizard/drafts/{id}

Feature surface:
  * One screen captures both backend (mode + servers) AND
    frontend (http + optional https + SSL mode) inputs.
  * SSL modes: ACME (new order, HTTP-01 only for v1.5.0),
    Upload (existing PEM), Existing (link to a stored cert),
    or None.
  * ACME-staged path: wizard_staged_until watermark on the
    `letsencrypt_orders` row defers finalisation until agent
    confirmation; per-mode reject cleanly cancels and rolls
    back the staged order.
  * Live diff preview against the cluster's current generated
    config (renderer-evolution noise stripped — track-sc<N>
    dedup, per-server cookie strip, defaults-cookie
    inheritance, listen-block flattening).
  * Draft persistence with PEM stripped at save time (private
    keys never round-trip through the drafts table).
  * Per-cluster multi-tenancy: drafts and wizard_staged orders
    are isolated to the creating user's cluster scope.

Frontend (`frontend/src/components/SiteWizard.js`):
  * 4-step Antd Steps flow: Backend → Frontend → SSL → Review.
  * Render the live diff preview inline before commit.
  * Antd Form-level validation mirrors backend Pydantic
    validators (numeric bounds, HAProxy reserved keywords, ALPN
    consistency, IPv6 scope-id, domain regex, server name
    dedup).

------------------------------------------------------------------
AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1#82)
------------------------------------------------------------------
v1.5.0 includes 22 adversarial review passes. Each round produced
its own commit set in the corporate development line; this squash
collapses those into the v1.5.0 release artefact. Highlights:

  Round 1-4   Site Wizard core: dry-run parity, single-line
              value injection guard, ACL -f pattern-file block,
              SSL parity, timeout regex, form-state pin.
  Round 5-7   defaults-cookie inheritance, server-named-cookie
              guard, fe/be mode mismatch, duplicate server
              names, health_check_uri + server_address
              validators.
  Round 8-10  cookie_name / cookie_options newline-injection
              guard, dry-run parity (round 9), TCP-mode HTTP-only
              feature blockers.
  Round 11    SSL name path traversal + health-check >= 1.
  Round 12-13 SSL & ACME deep dive (Bulgu #23-#32).
  Round 14    single-line value injection (Bulgu #33).
  Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy
              reserved keywords, ALPN/TLS consistency,
              all-backup, multi-domain & multi-user enterprise
              edges, drain/HSTS/post-completion (Bulgu
              #34-#53).
  Round 18-21 concurrency, agent state, TCP-mode HTTP-only,
              list size caps, IPv6 scope-id, preview account
              validation, TCP backend + balance uri reject
              (Bulgu #54-#61).
  Round 22    FE error visibility + 3x stale-data lockouts,
              referential integrity + cascade safety,
              authentication & authorization, multi-cluster
              isolation, apply_pending_changes concurrency,
              script injection + bulk import multi-tenancy,
              prefix-stripped signature comparison
              (Bulgu #62-#82).

------------------------------------------------------------------
NO CORPORATE-SPECIFIC ARTIFACTS
------------------------------------------------------------------
This squash deliberately sanitises corporate hostnames, container
registry references, and TLS secret names into generic
placeholders (`your-registry.example.com/your-org`,
`haproxy-openmanager*.example.com`, `wildcard-tls`,
`taylanbakircioglu/haproxy-openmanager-*`) so the public artefact
contains no internal infrastructure detail. Pilot / development
history that retained those values stays in the corporate fork
and is NOT part of this commit.
2026-05-14 00:04:19 +03:00
taylanbakircioglu 07942a82e8 feat: ACME stability & enterprise audit (v1.4.0) — fixes #10 #11 #12
Issue #10 — Silent timezone failure on ACME certificate save
- Normalize tz-aware expiry_date to UTC tz-naive before INSERT/UPDATE in
  ssl_certificates (TIMESTAMP WITHOUT TIME ZONE) — restores ACME download path

Issue #11 — Duplicate _acme_challenge_backend in generated config
- Generator guard: skip auto-append when backend already rendered
- Agent-sync filter: _should_sync_backend() drops system-managed backends and
  detaches their server rows to prevent orphans
- Restore filter: IGNORED_BACKENDS skips reserved names during cluster restore
- Parser warning: reserved_backend_names blocks accidental manual import
- Cleanup migration: removes orphan rows + cascading server entries (idempotent)

Issue #12 — Validated orders required manual completion
- New 60s background task complete_pending_acme_orders, flag-independent,
  multi-replica safe via FOR UPDATE SKIP LOCKED + 30s updated_at watermark
- Per-order pg_advisory_lock(0x41434D45, order_id) serializes UI-Complete and
  auto-task races; idempotency guard returns existing certificate cleanly
- retry_order endpoint reports in_progress: true within 30s window so the UI
  surfaces an info toast instead of duplicating CA requests
- ACMEAutomation surfaces stuck orders (status=valid && !ssl_certificate_id)
  with a one-click Complete action and Cancel fallback; conditional 30s polling

Other hardening
- ACME state machine error_detail persisted as structured JSON across challenge,
  finalize, download stages for actionable post-mortems
- CertificateRequest Pydantic model: domain regex + min_length/max_length and
  cluster_ids defaulting to all ACME-enabled clusters when "global" is selected
- Renewal cluster fallback now requires acme_enabled=TRUE in addition to active
- _complete_certificate preserves manual cluster assignments on renewal,
  surfaces cluster_errors, raises explicit error on missing private key
- Audit logging covers acme_certificate_requested/revoked, ca_chain_imported,
  account created/deactivated/purged, order retried/cancelled
- Settings UI exposes acme.staging_url_override for private test CAs (Pebble)
- Schema additions: acme_challenges.attempts (default 0) and last_attempt_at,
  index idx_letsencrypt_orders_status_updated; all migrations idempotent

Tests (41/41 passing)
- test_acme_expiry_normalize, test_acme_duplicate_backend,
  test_acme_state_machine, test_acme_pydantic_validation,
  test_acme_audit_logging, test_acme_concurrency

CI / packaging
- docker-build.yml reads version.json and pushes additional product-version
  tag (e.g. 1.4.0) alongside latest and timestamp build id

Closes #10
Closes #11
Closes #12
2026-05-06 23:53:43 +03:00
taylanbakircioglu cd4a94beb1 feat: agent IP/VIP live update + script update detection banner
- Agent scripts now detect and send ip_address in DAEMON heartbeat (Linux: ip route, macOS: ifconfig)
- Backend validates agent-reported IPs via ipaddress stdlib, COALESCE preserves existing on NULL
- IP/VIP change logging (non-critical, try/except wrapped) for operational visibility
- New source_file_hash column on agent_script_templates for reliable update detection
- Migration changed to ON CONFLICT DO NOTHING to prevent overwriting UI-customized scripts on restart
- GET /versions returns script_update_available flag (disk hash vs DB hash comparison with fallback)
- Frontend Alert banner warns users of new agent script versions and directs to Reset to Defaults
- Reset to Defaults and Popconfirm modals explicitly warn about custom script edit loss
- Full backward compatibility: old agents without ip_address field continue working unchanged

Made-with: Cursor
2026-04-17 10:50:51 +03:00
taylanbakircioglu 61024e24cb fix: strip rate-limit and WAF auto-generated content from bulk import comparison
Made-with: Cursor
2026-04-17 10:45:12 +03:00
taylanbakircioglu 6a349e22c6 fix: strip auto-generated ACME content from bulk import preview
Made-with: Cursor
2026-04-17 10:43:30 +03:00
taylanbakircioglu f37f3afd71 feat: bulk import change detection, multi-select delete, auto-content stripping
- Server-level change detection in bulk import (field-by-field comparison
  for 17 server attributes with UPDATE/NO CHANGES status and tooltip)
- Multi-select delete for backends and frontends with dependency checks
- Dashboard "Backends Summary" address column for servers
- Fix unique constraint violation on bulk-create for existing servers
  (natural key lookup matching DB constraint instead of backend_id FK)
- ORDER BY is_active DESC on all entity lookups to prefer active records
- Strip auto-generated content (ACME, rate-limit, WAF) from bulk import
  comparison to eliminate false positive changes on re-import
- Fix toolbar overflow with Space wrap prop
- Frontend bulk delete modal clarity (selected vs deletable count)
- Version bump to 1.3.0

Made-with: Cursor
2026-04-14 20:21:07 +03:00
taylanbakircioglu 7fffaddaf8 fix: bulk import should not inject hardcoded timeout defaults
When a backend block in haproxy.cfg does not specify explicit timeout
values, the bulk import was injecting hardcoded defaults (connect 10s,
server 60s, queue 60s) into the database. These then appeared in the
generated config and overrode the agent's defaults section. Now, only
explicitly declared timeouts are stored; omitted ones remain NULL so
the agent's existing defaults section stays in effect.
2026-04-13 05:52:31 +03:00
taylanbakircioglu 36fdba52bc fix: ACME setup guide accuracy, reject rollback, and UX improvements
- Step 3 (Enable ACME on Cluster) now shows a process icon instead of
  a misleading green checkmark when ACME is enabled but not yet applied.
  Per-cluster "(pending apply)" annotation for multi-cluster setups.
- Step 4 button and all /apply-management navigation buttons now say
  "Apply Changes" instead of "Configure" for clearer guidance.
- Setup Guide auto-selects the correct cluster before navigating to
  Apply Management, showing pending cluster names in alerts.
- Pending ACME disable changes are now correctly detected in Step 4
  even when acme_enabled is already FALSE in the database.
- Entity snapshot rollback for cluster ACME settings: reject correctly
  restores acme_enabled/acme_backend_url to pre-change values.
- Deduplication logic prevents "last wins" bug when multiple ACME
  toggles are rejected in sequence.
- Connection leak prevention with try/finally around conn2 in ACME
  config version creation.
- Step 4 branching uses boolean has_enabled instead of fragile string
  truthiness check.

Made-with: Cursor
2026-04-04 16:57:57 +03:00
taylanbakircioglu deb784bb65 fix: ACME certificate issuance improvements, null-safe hardening, guided setup UX, and order list enhancements (Issue #9)
- Fix critical cascading NULL status bug in ACME challenge flow that could cause 404s
- Add defense-in-depth NULL handling across all ACME service methods
- Add new GET /api/letsencrypt/prerequisites endpoint for configuration checks
- Add interactive ACME Setup Guide with step-by-step navigation links
- Add URL-based tab navigation in Settings and SSL Management pages
- Harden retry flow: return clear 409 errors for invalid/cancelled orders
- Allow cancellation of invalid orders (backend + frontend)
- Improve error message extraction with consistent getErrorMsg helper
- Add status filter tabs (Active/Completed/Failed/All) for order list
- Add visual dimming for cancelled/invalid orders
- Add enhanced pagination with size changer and total count
- Add comprehensive diagnostic logging with ACME: prefix
- Update README with ACME architecture docs, quick start guide, and troubleshooting

Resolves #9

Made-with: Cursor
2026-04-04 15:15:05 +03:00
taylanbakircioglu 2fc2603d1a fix: ACME account removal, status display, and timeline icon clipping
- Add permanent delete endpoint for deactivated ACME accounts (DELETE /accounts/{id}/permanent)
- Add "Remove from database" button in ACME Accounts modal for deactivated accounts
- Fix ACME Account stat card showing "Active" for deactivated accounts
- Fix Timeline dot/icon clipping in Applied/Rejected/Pending sections

Made-with: Cursor
2026-04-02 00:36:56 +03:00
taylanbakircioglu 93d7ad8fdb feat: add ACME Auto SSL with Let's Encrypt integration (v1.1.0)
Add automated SSL certificate management via ACME protocol (RFC 8555):
- Full ACME client implementation (account registration, HTTP-01 challenges, certificate issuance/renewal)
- Configurable ACME providers (Let's Encrypt, ZeroSSL, custom CA) via Settings UI
- Auto-renewal scheduler with PENDING -> Apply -> APPLIED flow alignment
- ACME account management (register, deactivate) from UI
- Certificate request wizard with domain validation and cluster targeting
- Zero changes to HAProxy agent scripts - challenges routed through existing architecture
- Comprehensive security hardening (no private key exposure in API responses)
- Full backward compatibility with existing SSL, Apply, Rollback, and Restore workflows
- Updated README, API documentation, and Kubernetes deployment notes
- Version management embedded in code (v1.1.0)
- UI messaging improvements for agent-pull architecture accuracy

Made-with: Cursor
2026-04-02 00:25:50 +03:00
taylanbakircioglu 922bc2ce6d improve: SSL certificate listing reliability, error visibility, and ARM64 support
Fixes #6

- Fix NameError in soft-deleted certificate reactivation path by
  reordering variable extraction before DB operations
- Replace silent empty-array returns with HTTP 500 on SQL errors,
  making failures visible in both API responses and server logs
- Add primary_domain migration for schema consistency across fresh
  and upgraded installations (backfill from legacy domain column)
- Use primary_domain in non-cluster SSL query branch for schema
  compatibility
- Surface SSL fetch errors in frontend via toast notifications
- Harden connection cleanup in error handlers with try/except
- Add QEMU + Buildx for linux/amd64,linux/arm64 multi-platform
  Docker image builds
- Update GitHub Actions to latest versions (checkout v4, login v3,
  build-push v6)

Made-with: Cursor
2026-03-30 10:42:28 +03:00
taylanbakircioglu 5b59accd51 fix: resolve asyncpg AmbiguousParameterError on keepalive params
asyncpg cannot infer the PostgreSQL type for $20/$21 when the Python
value is None in CASE WHEN $20 IS NOT NULL expressions. Add explicit
::text cast so asyncpg can prepare the statement regardless of whether
the agent sends keepalive data or not. This was causing ALL heartbeats
to fail with 500 error, making all agents appear offline.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-18 20:41:17 +03:00
taylanbakircioglu 4082152312 fix: resolve stale keepalive state when keepalived stops
- Fix DB: COALESCE(NULLIF) prevented clearing stale MASTER/BACKUP state.
  Use CASE WHEN IS NOT NULL to distinguish absent fields (old agents)
  from empty fields (keepalived stopped) and properly clear to NULL.
- Fix Redis: explicitly delete cache key when keepalived stops instead
  of relying on 90s TTL expiry, preventing stale dashboard data.
- Fix agent scripts: SKIP_TO_DAEMON send_heartbeat now always sends
  keepalive_state/keepalive_ip fields (even empty) so backend can
  detect stopped keepalived and clear stale data.
- Fix legacy heartbeat endpoint: POST /{agent_id}/heartbeat now also
  updates keepalive_state and keepalive_ip fields.
- UI: add horizontal scroll to ClusterManagement table to prevent
  column overflow with new Keepalive column.
- UI: add VIP tooltip to Dashboard AgentStatusCard keepalive tag.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-18 20:41:17 +03:00
taylanbakircioglu b0feca3c2e feat: add keepalived VRRP state (MASTER/BACKUP) detection to agent heartbeat
- Add keepalive_state and keepalive_ip columns to agents table (migration + schema)
- Add keepalive fields to AgentHeartbeat Pydantic model (backward compatible)
- Update heartbeat endpoint to persist keepalive data to DB and cache in Redis
- Add multi-method keepalived detection in agent scripts (journalctl, log files, VIP check)
- Update dashboard-stats agents/status API with Redis-first keepalive lookup
- Update GET /api/agents to include keepalive_state and keepalive_ip
- Show MASTER/BACKUP tag in Dashboard AgentStatusCard
- Show keepalive info in Agent Management registered agents table
- Add Keepalive column to Cluster Management table with VIP search support

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-18 20:41:17 +03:00
taylanbakircioglu 1f100d01e7 fix: SSL reject rollback, expiry calculation, and UI improvements
- SSL Reject: Remove conditional has_same_update_applied check, always
  rollback since Auto-Reject handles cross-cluster consistency
- SSL Rollback: Restore expiry_date by parsing ISO string back to datetime
- SSL UI: Add In Use filter toggle, fix Expiring Soon count using actual
  date calculation, smart Private Key content validation messages
- Remove emojis from log messages

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-17 16:36:27 +03:00
taylanbakircioglu 7dedd0f573 feat: SSL sync status tracking and smart reload optimization
- Add per-cluster SSL sync status tracking with APPLYING/SYNCED/PARTIAL states
- Add timeout detection (10min) for stalled SSL deployments showing PARTIAL status
- Add SSL-specific info in Apply Management dialog (cluster count, auto-apply notice)
- Add Deployment Status tab in SSL details showing per-cluster agent sync progress
- Implement SSL auto-reject (reject propagates to all clusters) and auto-undo
- Add conditional entity rollback for SSL reject (time-window based safety check)
- Filter REJECTED status from SSL last_config_status display
- Add pending_cluster_names to SSL list API response
- Optimize agent SSL reload: only trigger HAProxy reload when changed cert is
  actually referenced in the cluster's HAProxy config (grep -qF check), preventing
  unnecessary reloads on clusters that don't use the updated certificate
- Applied to all 4 SSL deploy paths: Linux/macOS standalone and daemon embedded modes

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-17 16:36:27 +03:00
taylanbakircioglu 9b8016588c feat: show SSL certificate usage status (In Use / Not In Use)
Replace the always-Active status badge with actual usage information by
querying which frontends and backend servers reference each certificate.

Backend: detail endpoint returns used_by_frontends, used_by_servers and
usage_count; list endpoint includes usage_count via scalar subqueries
checking both ssl_certificate_id and ssl_certificate_ids JSONB fields.

Frontend: General Info badge shows In Use/Not In Use, new Usage tab with
searchable frontend and server tables, summary card shows In Use count.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-17 16:36:27 +03:00
taylanbakircioglu 7ca036ef3c fix: add missing is_active field to SSL certificate API responses
The detail endpoint (GET /api/ssl/certificates/{id}) was missing is_active
in its SQL SELECT/GROUP BY, causing the frontend details modal to always
show "Inactive". The list endpoint included is_active in its SQL query but
omitted it from the response dict, causing the Active counter card to
always show 0.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-17 16:36:27 +03:00