ZeroSSL/Google account registration failed with
`malformed: The Replay Nonce could not be base64url-decoded`: the ACME client
(a process-wide singleton) kept a single anti-replay nonce shared across
certificate authorities, so a nonce issued by one CA could be sent to another,
and the auto-retry only covered `badNonce`.
- Scope the nonce per CA (self._nonce_by_dir keyed by directory_url): a nonce
from one CA is never sent to another; account registration always uses a fresh
nonce from the target CA.
- Broaden the 400 retry to also recover from the nonce-malformed rejection.
- Fix _b64url_decode padding (used for the EAB HMAC key).
Backend-only; HTTP-01 and Let's Encrypt are unaffected.
Addresses #35.
Follow-up fixes for the DNS-01 feature reported on #35:
- Cloudflare: sanitize the API token (strip surrounding quotes + any non
token68 chars) so a pasted token with quotes/spaces no longer fails with
"Invalid request headers"; verify-on-save shows a precise hint when it
cleaned the input. Covers the automated orchestrator path too.
- ZeroSSL/Google EAB: enter the EAB Key ID and HMAC Key per-account in the
Register Account dialog (falls back to the global Settings value when blank);
base64-validate the HMAC key; humanize the externalAccountRequired failure;
and preserve the deliberate 409/422 instead of downgrading them to 400.
- Apply Management: cluster ACME enable/disable changes now show under a
dedicated "ACME Challenge Routing" section, are counted in the Apply/Reject
dialogs, and Apply/Reject All process them (previously "Rejected 0 HA/VIP
change(s)") - consistent with every other entity. Reject rolls acme_enabled
back to the original via ORDER BY created_at ASC over the snapshot chain.
- getErrorMsg surfaces field-level validation messages.
Backward compatible (additive / strict superset; HTTP-01 unchanged).
Addresses #35.
Add ACME DNS-01 (TXT-record) validation alongside the existing HTTP-01,
for internal/isolated clusters with no public port 80 and for wildcard
certificates. Opt-in via a global kill-switch (default off); HTTP-01 is
byte-for-byte unchanged, with zero agent or rendered-config changes.
- Pluggable DNS provider interface (Manual + Cloudflare). Per-account
credentials are Fernet-encrypted at rest, verified on save, and never
returned by the API or written to logs/events/error_detail.
- Non-blocking per-cycle orchestrator: publish (CAS) -> propagation grace
(across cycles, no in-loop sleep) -> respond -> finalize/download, with a
bounded fresh-order retry chain (1 original + 3 retries) on propagation lag.
- Manual flow: user publishes the TXT record and confirms; manual DNS-01
cannot auto-renew unattended (auto-renew forced off and surfaced in the UI).
- Migration v8: additive, idempotent columns on letsencrypt_accounts/orders
and acme_challenges, plus a new letsencrypt_account_dns_credentials table.
- Challenge-type-aware diagnostics (port80/routing/DNS checks skipped for
DNS-01) and a DNS-01 event timeline in the order detail.
- Frontend: DNS-01 account + credentials management, cert wizard adaptation,
order-detail TXT records + verify, orders/renewal Method columns, and a
Settings kill-switch. README, release notes, and API docs updated.
Implements #35.
Two friction points surfaced in issue #31: (1) the default-login note said
admin/admin but the seeded password is admin123; (2) a reporter logged in at
:3000 (the raw static frontend, no /api behind it) instead of :8080 (nginx,
which serves the UI and proxies /api). Fix the credentials note and add a
clear pointer that :8080 is the single entry point, with the Quick Reference
table annotated accordingly.
Patches the critical shell-quote advisory (quote() does not escape
newlines in object .op values, CVSS 8.1). shell-quote is a dev/build-only
transitive dependency (react-dev-utils / launch-editor) and is not present
in the production image or the browser bundle, so there is no runtime
exposure; this clears the alert and the dev-time risk. Lockfile-only change
verified to install with shell-quote resolving to 1.8.4.
The build workflow tagged the Docker images with the product version but
never created the matching git tag, so the repo Tags/Releases drifted
behind (stuck at the last manual tag, v1.6.0) while Docker Hub had 1.7.8.
Add a step that, after the images are pushed, creates a Release (and its
tag) for the current version.json when one does not already exist, and
grant the job contents:write so it can do so.
Manage highly-available virtual IPs backed by Keepalived (VRRP) directly from the OpenManager
UI — no more SSHing into nodes to install/configure Keepalived by hand. Builds on the agent
pull-architecture: define the VIP centrally, click Apply, and the agents converge.
Highlights:
- New "HA / VIP" tab: create a virtual IP, pick a per-node interface, select which pool nodes
participate (MASTER/BACKUP roles + priorities); live MASTER/BACKUP/FAULT per node.
- On Apply, agents install & configure Keepalived (unicast VRRP, cloud-safe default) across the
major distros (Debian/Ubuntu, RHEL/CentOS/Alma/Rocky, Fedora, SUSE/openSUSE, Alpine) with a
HAProxy health-check, so the VIP fails over automatically when HAProxy drops.
- Single-node (a managed floating IP without failover) and multi-node VRRP failover both work.
- VIP changes ride the standard Apply Management flow with the standard "View Change" diff.
- Approval-gated deletion (safety): deleting a running VIP is staged for approval and the VIP
keeps running, untouched, until you approve it — an agent never tears a VIP down without an
explicit human approval. Per-VIP Diagnostics view; opt-in package uninstall (only on nodes
where OpenManager installed it). A node already running a hand-managed Keepalived is detected
and never overwritten ("externally managed").
- Fully opt-in and backward compatible: nodes/clusters without a VIP are unaffected. Adds
vip_instances + vip_members tables (idempotent SCHEMA_VERSION bump; existing data and
passwords unaffected) and a `vip` RBAC permission group.
- Also includes a HAProxy config-generator robustness fix: auto-inject a stick-table when a
frontend uses a stick counter (track-sc / sc_*_rate) but declares none.
On-prem / L2 (VRRP) scope; the UI notes the cloud caveat.
Supersedes the v1.6.4 stable pin (1.30.2-alpine) with the mainline
patched release. No config, schema, or behavior changes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The bundled nginx reverse proxy was flagged for the nginx 'poolslip' advisory
(affected: mainline <=1.31.0; fixed: stable 1.30.2+ / mainline 1.31.1+). The
config-level mitigation (named capture groups instead of $1/$2 in rewrite) does
not apply — the product's nginx config (nginx/nginx.conf and the k8s configmap)
has no rewrite capture-group directives, only prefix locations + proxy_pass. So
the fix is the version: pin nginx:alpine -> nginx:1.30.2-alpine in
docker-compose.yml and k8s/manifests/10-nginx.yaml.
No config/schema/behavior change (nginx only reverse-proxies). Version bumped to
1.6.4 across all layers. Verified in Docker: nginx -v=1.30.2; nginx -t OK on both
the compose and production configmap configs; all proxied routes work through
nginx; no nginx errors.
A backend server toggled OFF (is_active=false) vanished from the UI with no way
to reactivate it: GET /api/backends honored include_inactive for backends but the
server sub-queries hardcoded 'AND is_active = TRUE'.
- get_backends: server sub-queries now honor include_inactive (default callers
unchanged); added last_config_status to the server payload so the UI can tell a
DISABLED server (re-enableable) from a DELETION (pending delete).
- toggle_server: persists an entity snapshot so an Apply-Management Reject rolls
back is_active (previously left the server stuck disabled).
- BackendServers.js: requests include_inactive, shows disabled servers with the
ON/OFF switch + an 'Inactive' tag, hides only DELETION-pending servers, and
keeps soft-deleted BACKENDS hidden (so include_inactive doesn't resurface them).
- Config generation unchanged: disabled servers stay '# DISABLED:' comments and
convert back to live lines when re-enabled.
Startup migration hardening (multi-replica / rolling-deploy safety): create_essential_tables
fails fast on lock contention and retries; run_all_migrations is serialized by a
session advisory lock and gated by a schema_migrations version marker, so an
already-current schema is skipped instead of issuing lock-heavy DDL that a serving
replica's traffic could block at startup. Idempotent and fail-open.
Version reported consistently across all layers (version.json, backend fallback,
frontend package) -> 1.6.3.
Assigning a HAProxy agent failed with '401: Authorization header missing'
on GET /api/clusters. A cluster-read hardening had made GET /api/clusters and
GET /api/clusters/{id} accept only a user JWT in the Authorization header;
agents authenticate with their agent token in the X-API-Key header, so the
token was never read.
Both endpoints now accept either a user JWT (Authorization) or an agent token
(X-API-Key via validate_agent_api_key), mirroring the existing dual-auth on
POST /api/agents/generate-install-script. Anonymous access is still rejected,
so the original hardening is preserved. The auth guard is placed before the
try block so the failure surfaces as a clean 401 (not the 500-wrapped-401 in
the report). Agent install scripts now consistently send the token via
X-API-Key (pre-flight cluster check on linux/macos, and macOS get_cluster_paths
which previously used the wrong Authorization: Bearer header).
Also normalizes the platform in the uninstall-script generator so macOS agents
(which report platform 'darwin') no longer get a 400 from
GET /api/agents/generate-uninstall-script/darwin.
version 1.6.0 -> 1.6.2.
Adds opt-in TOTP-based Multi-Factor Authentication that is fully
backwards compatible with existing logins. Operators choose to enable
MFA per account; nothing changes for users who do not opt in.
Highlights
==========
* RFC 6238 TOTP (6 digits, 30s period, SHA1) with ±30s skew tolerance,
compatible with Microsoft / Google Authenticator, Authy, Duo, 1Password.
* Per-step replay protection (`mfa_last_used_totp_step`) so a captured
code cannot be reused inside the same window.
* Fernet-encrypted TOTP secrets at rest, key resolution via
`MFA_ENCRYPTION_KEY` env (HKDF-derived from `SECRET_KEY` as fallback).
* 10 single-use, bcrypt-hashed backup codes per user, formatted
`XXXX-YYYY` from a confusion-free alphabet (no 0/O/1/I/L).
* Two-step login flow: `POST /api/auth/login` returns `mfa_required`
+ `mfa_token`, then `POST /api/auth/login/mfa-verify` accepts a TOTP
code OR a backup code. JWT is minted only after MFA succeeds.
* Self-service: users enable / disable MFA from their own row in the
Users page; admins reset (single user or bulk) but never enable on
behalf of someone else (matches AWS IAM / GitHub / Google Workspace).
* Bulk emergency reset CLI: `scripts/admin-mfa-reset-all.sh`.
Security hardening
==================
* Atomic transactions with `SELECT … FOR UPDATE` on `mfa_pending_logins`
and `users` rows so concurrent verify / enroll calls cannot race.
* `/api/mfa/enroll/start` refuses re-enrollment when MFA is already on
(prevents silent secret rotation via a stolen JWT).
* Pydantic `ValidationError` messages are sanitized before reaching the
audit log so request bodies (TOTP / backup codes in flight) never
appear in plaintext.
* Slowapi rate limits are per-USER, not per-IP, with a trusted-proxy
XFF strategy so a single ingress address cannot exhaust the bucket
for thousands of operators (`MFA_TRUSTED_PROXY_CIDRS`,
`MFA_RATE_LIMIT_*` env-overridable).
* Login query now scopes to `is_active = TRUE` so a soft-deleted row
with the same username can no longer occlude the active user
(also closes a small account-enumeration side channel).
Database
========
Additive migrations (idempotent `ADD COLUMN IF NOT EXISTS`,
`CREATE TABLE IF NOT EXISTS`):
- users: mfa_enabled, mfa_method, mfa_secret_encrypted,
mfa_enrolled_at, mfa_last_used_at, mfa_last_used_totp_step
- mfa_backup_codes (user_id ON DELETE CASCADE)
- mfa_pending_logins (user_id ON DELETE CASCADE, challenge_token,
attempts, expires_at)
- mfa_pending_enrollments (user_id ON DELETE CASCADE)
Frontend
========
* Login page becomes a 3-phase state machine
(credentials → MFA → submitting); legacy single-step login is
preserved for users who haven't enrolled.
* New MFAEnrollModal (3-step wizard: QR + secret → verify → backup
codes) using `qrcode.react`.
* Users page shows MFA column + per-row enable/disable/reset actions.
Admins viewing other users with MFA off see a non-actionable info
icon explaining that only the user themselves can enable MFA.
Deployment
==========
* `MFA_ENCRYPTION_KEY` is added to `k8s/manifests/03-secrets.yaml` as
a placeholder; `SECRET_KEY` is also placeholder-ized so both are
injected by the existing pipeline pattern (sed-replace + apply).
* No new build-time env vars are required for the frontend. The SPA
uses `window.location.host` for `/api/*` and is routed by the
existing nginx ingress configuration.
* `frontend/.dockerignore` ensures host `.env*` files cannot bleed
into the production bundle.
Tests
=====
* New unit suites:
- `test_mfa_service.py` (TOTP, encryption, backup codes)
- `test_mfa_backwards_compat.py` (regression — non-MFA flow unchanged)
- `test_mfa_rate_limits.py` (env override + dataclass immutability)
- `test_mfa_rate_limit_key.py` (JWT key, trusted-proxy XFF, fallbacks)
* All existing 1000+ unit tests continue to pass.
Documentation
=============
* README MFA section (overview, day-to-day operations, emergency
reset CLI, env variables, rate-limit tuning).
* `scripts/README.md` documents the bulk reset script.
Issue: #18
The backend Docker image is built with `context: ./backend`, so the
repo-root `version.json` is outside the build context and never
reaches the container. Backend `main.py` falls back to a compile-
time `_version_info` constant when `/app/version.json` is missing.
In practice this produced a real production drift: a successful
redeploy of the v1.5.2 tree silently reported `"v1.5.0"` in
`/api/version` for a window of releases because the constant in
main.py had not been bumped in lockstep with `version.json`, and
the canonical file was never available to read inside the
container.
Fix is workflow-only:
* New "stage version.json into backend build context" step
(between `read product version` and `set up qemu`) that runs
`cp version.json backend/version.json` so the next
`docker buildx build` includes it.
* `.gitignore` entry for `backend/version.json` keeps `git status`
clean for developers (the canonical file remains at repo root;
`backend/version.json` is a transient CI artefact).
Backend reading logic is unchanged: the loop in `main.py` first
tries `/app/version.json`, then falls back to the constant.
Post-fix, the first path WILL find the file and produce the
correct response; the constant becomes a pure defensive fallback
(rather than the production hot path it accidentally became).
No code or test changes needed: existing tests assert against the
`_version_info` dict regardless of whether it was populated from
JSON or the fallback constant.
A focused hardening pass on the v1.5.0 ACME Diagnostic Panel
surface, exercised against a live production deployment (Round-25
+ Round-26 audits) and supplemented by an AGPL-3.0 relicense.
------------------------------------------------------------------
LICENSE — Relicense to AGPL-3.0-or-later
------------------------------------------------------------------
Effective v1.5.2 the project is licensed under the **GNU Affero
General Public License v3.0 (or later)**. v1.5.0 and v1.5.1
remain under the prior MIT terms.
The relicense is consistent with the project's intent as a
community-operated HAProxy management surface: forks that run
HAProxy OpenManager as a network service for third parties are
now required to publish their modifications under the same
license (AGPL §13). Day-to-day single-tenant deployments,
internal corporate use, and ordinary forks-for-fixes are
unaffected.
Changes:
* LICENSE replaced with full AGPL-3.0 text.
* README "## License" section rewritten with the AGPL summary
+ the network-service obligation.
* frontend/package.json gains `"license": "AGPL-3.0-or-later"`.
------------------------------------------------------------------
BULGU #94 / #95 — Diagnostic Panel Must Never Opaque-500
------------------------------------------------------------------
Live exercise of the v1.5.0 Diagnostic Panel against a deployed
build surfaced two opaque-500 paths. The panel exists to make
ACME failures legible; producing an opaque HTTP 500 defeats the
entire feature. Fix shape: every endpoint now returns either a
canonical 4xx (auth / not-found / rate-limit) or an HTTP 200
"structured failure envelope" that the React UI knows how to
render — never a 500 for an in-suite failure.
Affected paths:
POST /api/letsencrypt/orders/{order_id}/diagnostics
Pre-fix: a UndefinedColumnError or DB-connectivity failure
inside `run_checks` bubbled out of the bare try/finally and
surfaced as a generic 500 with no operator-actionable detail.
Post-fix: setup-stage and run-stage failures are caught
separately and converted to a `status: diagnostics_unavailable`
envelope carrying `error_stage`, `error_type`, `error_message`,
and a `correlation_id` that the operator can grep in the
backend log. Individual checks are wrapped in `_safe_check`
so one broken check (e.g. DNS lookup timeout) never crashes
the suite — the failing check shows up as `status: "fail"`
with its message, the others still run.
GET /api/letsencrypt/orders/{order_id}/events
Pre-fix: the SQL `SELECT … status FROM user_activity_logs`
referenced a column that did not exist in the canonical
migration; every diagnostic-panel open against an order with
any user-activity-log correlation got an `UndefinedColumnError`
500. Post-fix: the endpoint now introspects
`information_schema.columns` and projects only the columns
actually present. Partial failures (one source dies, the
other works) are reported via `meta.errors[]` rather than
collapsing the whole timeline.
POST /api/letsencrypt/orders/{order_id}/diagnostics/{check_id}/rerun
Same structured-envelope contract as the full-suite POST,
scoped to a single check row.
Frontend (`frontend/src/components/ACMEAutomation.js`):
* Distinct `diagRunError` / `diagEventsError` / `diagMeta`
states so the modal can render the cause inline (Antd Alert)
instead of a silent dropdown.
* Event-log auto-tail polling backs off after 3 consecutive
failures so the Network tab does not get spammed with 500s
every 5s.
* Correlation IDs visible in every error banner.
------------------------------------------------------------------
BULGU #96 — Clean 404 for Out-Of-Range order_id
------------------------------------------------------------------
A live exercise of the post-#94 diagnostic panel against the
deployed build surfaced one remaining contract gap. A path-
param `order_id` outside the Postgres int4 range
(e.g. > 2_147_483_647) caused `_load_order` to raise
`asyncpg.exceptions.DataError: invalid input for query
argument $1: 2147483648 (value out of int32 range)`. Round-25
correctly surfaced this in a `diagnostics_unavailable`
envelope — but that envelope leaked SQL implementation detail
("query argument $1", "int32 range", DataError class name)
into the operator-facing response body.
Semantically an out-of-range integer can never reference a
real order — it's just "not found". `_load_order` now catches
`asyncpg.exceptions.DataError` and re-raises a canonical
`HTTPException(404, "Order {id} not found")`. Because all
three endpoints re-raise `HTTPException` from their outer
try/except (the Round-25 envelope only fires for non-
HTTPException crashes), the canonical 404 path now wins
end-to-end across /diagnostics, /events, and /rerun.
------------------------------------------------------------------
TEST / LINT / LIVE VERIFICATION
------------------------------------------------------------------
* Backend pytest 1104/1104 (the +20 vs v1.5.1's 1084 are the
Round-25 and #96 contract pins; see
test_acme_diagnostics_router_round25.py).
* Live prod-canary verification: every endpoint return shape
confirmed against the deployed build — int4 overflow returns
clean 404 with no SQL leak, normal paths return Round-25
envelopes, HTTP method matrix returns 405 on wrong verbs,
no auth returns 401, invalid `check_id` returns 400, and
`meta.correlation_id` is present on every diagnostic
response.
------------------------------------------------------------------
COMPATIBILITY
------------------------------------------------------------------
* No breaking API contract changes: `status` field on the
diagnostic response can now be `"diagnostics_unavailable"`
in addition to the existing pass-through of the
underlying order status (`pending` / `valid` / `invalid` /
`cancelled` / …) — older UIs that only switch on the
existing values render the `diagnostics_unavailable`
case as "unknown status" rather than crashing.
* Frontend handles the new envelope shape AND the legacy
HTTP 4xx/5xx paths.
A live-deployment audit pass over the v1.5.0 Site Wizard + ACME
Diagnostic Panel surface. Two adversarial review rounds (R23, R24)
each capped by an end-to-end smoke test against a multi-cluster
staging deployment.
Bulgu #83 — Frontend Management page warned about stale data
without a clear retry CTA. The toast now carries an in-place
"Reload" action and the page-level Empty state surfaces the same
recovery affordance, so operators never get stuck on a stale-data
view without an obvious way out.
Bulgu #84 — ACME diagnostics ran with the wrong "last_heartbeat"
column reference against the agents table. Aligned the SELECT
with the actual schema column (`last_seen`); pinned by an idempotent
regression test in `test_acme_diagnostics.py`.
Bulgu #85 — ACME order error_detail rendering could leak the raw
asyncpg/SQL exception class name when humanize_error_detail
encountered an unhandled CA response shape. Added a backwards-
compatible fallback branch that emits an "ACME error (raw)" panel
without exposing parse_error class name to the user.
Bulgu #86 — Multi-cluster apply with concurrent rejects could
leave wizard_staged orders dangling without their parent draft.
Pinned via reject_order_with_cluster_orphan test.
Bulgu #87 — Frontend Management page list virtualization
mis-keyed during a re-sort + stale-row replace race; fixed by
keying rows on `id + version` so React reconciler does not reuse
DOM for a logically different row.
Bulgu #88 — Site Wizard "Cancel" mid-flow now surfaces an
unsaved-draft prompt with explicit Save / Discard buttons (and
the same prompt on browser tab close), so the operator never
loses 5 steps of input to an accidental ESC.
Bulgu #89 — Existing-cert SSL mode showed an empty dropdown when
the cluster had >100 certs because the listing endpoint
default-limited results. Endpoint now exposes pagination AND
the wizard switches to client-side filtering above 50 rows.
Bulgu #90 — ACME pre-check on the wizard preview path did NOT
re-validate the account against `letsencrypt_accounts` if the
operator stepped Back/Forward between SSL and Review. Added a
debounced re-validation on Review entry.
Bulgu #93 — Site Wizard hsts_enabled toggle in HTTPS frontend
was idempotent-by-name (the generated `http-response set-header
Strict-Transport-Security` line could duplicate across a Save +
Apply cycle). The renderer now upserts the header in place.
Cumulative outcome: backend pytest 1084/1084, frontend lint
clean, and a 6-hour live-deployment smoke session against staging
with no regressions reported.
Closes#13, Closes#14.
This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0
shipped the ACME stability & enterprise audit (Issues #10/#11/#12).
v1.5.0 builds on that foundation with two co-equal headline features
plus a 22-round audit campaign hardening the prior configuration
surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0
lands in v1.5.2).
------------------------------------------------------------------
HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13)
------------------------------------------------------------------
A live pre-flight + post-failure diagnostic surface for every ACME
order, reachable from the ACME Automation page. The panel exists
to make ACME failures legible to operators who do NOT have shell
access to the API host.
Endpoints (`backend/routers/acme_diagnostics.py`):
POST /api/letsencrypt/orders/{order_id}/diagnostics
Run the full 5-check suite (DNS / port-80 / routing /
account / agents) and humanize the order's `error_detail`
(>=11 RFC-8555 problem types, backwards compatible with
legacy plain-string failures).
POST /api/letsencrypt/orders/{order_id}/diagnostics/
{check_id}/rerun
Re-run a single check in place — used by the "Re-run"
button on every row of the modal's pre-flight table.
GET /api/letsencrypt/orders/{order_id}/events
Merged event timeline combining the typed
`acme_order_events` rows with correlated
`user_activity_logs` entries (resource_type =
'letsencrypt_order' AND resource_id = order_id). The
diagnostic modal auto-tails this timeline every 5 seconds
while open.
Service-level checks (`backend/services/acme_diagnostics.py`):
* DNS resolution via stdlib socket.gethostbyname_ex through
run_in_executor (intentionally avoiding an aiodns runtime
dep for v1.5.0).
* Port-80 HEAD probe, target locked to the order's domains,
success on HTTP 200 OR 404, warns on egress timeout
(corp egress policies routinely blackhole outbound 80 —
fail-hard would be too noisy).
* SSRF guard: probe refuses non-public IPs and surfaces the
skip in the diagnostic result; IPv4-mapped IPv6 normalisation
closes the `::ffff:169.254.169.254` cloud-metadata vector.
* HAProxy routing presence check: matches the order's
cluster_ids to a port-80 HTTP frontend.
* ACME account validity check against `letsencrypt_accounts`.
* Agent presence check (>=1 active agent in target cluster).
* Every sub-check wrapped in a wall-clock timeout to bound
impact on the API event loop.
RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate
limit on both run and rerun, backed by the (user_id, action,
created_at DESC) composite index.
Frontend (`frontend/src/components/ACMEAutomation.js`):
* "Diagnose" button on every order row + the existing
"stuck order" warning row.
* Modal with two tabs:
- Pre-flight Checks (Antd Table with status pills + Re-run
buttons + humanized error banner)
- Event Log (Antd Timeline with auto-tail polling, scroll-
to-bottom, pause-on-hover)
* Correlation IDs surfaced in error banners and individual
check fail details for backend-log lookup.
------------------------------------------------------------------
HEADLINE FEATURE B — Site Setup Wizard (Issue #14)
------------------------------------------------------------------
A single guided flow that creates a Backend + Servers + HTTP
Frontend (and optional HTTPS Frontend) in one atomic transaction.
Endpoints (`backend/routers/site_wizard.py`):
POST /api/site-wizard/preview — diff-preview the changeset
POST /api/site-wizard/create — atomic execute
POST /api/site-wizard/reject — clean rollback (including
any wizard_staged ACME
orders)
GET /api/site-wizard/drafts — draft persistence
PUT /api/site-wizard/drafts/{id} — save/update
DELETE /api/site-wizard/drafts/{id}
Feature surface:
* One screen captures both backend (mode + servers) AND
frontend (http + optional https + SSL mode) inputs.
* SSL modes: ACME (new order, HTTP-01 only for v1.5.0),
Upload (existing PEM), Existing (link to a stored cert),
or None.
* ACME-staged path: wizard_staged_until watermark on the
`letsencrypt_orders` row defers finalisation until agent
confirmation; per-mode reject cleanly cancels and rolls
back the staged order.
* Live diff preview against the cluster's current generated
config (renderer-evolution noise stripped — track-sc<N>
dedup, per-server cookie strip, defaults-cookie
inheritance, listen-block flattening).
* Draft persistence with PEM stripped at save time (private
keys never round-trip through the drafts table).
* Per-cluster multi-tenancy: drafts and wizard_staged orders
are isolated to the creating user's cluster scope.
Frontend (`frontend/src/components/SiteWizard.js`):
* 4-step Antd Steps flow: Backend → Frontend → SSL → Review.
* Render the live diff preview inline before commit.
* Antd Form-level validation mirrors backend Pydantic
validators (numeric bounds, HAProxy reserved keywords, ALPN
consistency, IPv6 scope-id, domain regex, server name
dedup).
------------------------------------------------------------------
AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1 → #82)
------------------------------------------------------------------
v1.5.0 includes 22 adversarial review passes. Each round produced
its own commit set in the corporate development line; this squash
collapses those into the v1.5.0 release artefact. Highlights:
Round 1-4 Site Wizard core: dry-run parity, single-line
value injection guard, ACL -f pattern-file block,
SSL parity, timeout regex, form-state pin.
Round 5-7 defaults-cookie inheritance, server-named-cookie
guard, fe/be mode mismatch, duplicate server
names, health_check_uri + server_address
validators.
Round 8-10 cookie_name / cookie_options newline-injection
guard, dry-run parity (round 9), TCP-mode HTTP-only
feature blockers.
Round 11 SSL name path traversal + health-check >= 1.
Round 12-13 SSL & ACME deep dive (Bulgu #23-#32).
Round 14 single-line value injection (Bulgu #33).
Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy
reserved keywords, ALPN/TLS consistency,
all-backup, multi-domain & multi-user enterprise
edges, drain/HSTS/post-completion (Bulgu
#34-#53).
Round 18-21 concurrency, agent state, TCP-mode HTTP-only,
list size caps, IPv6 scope-id, preview account
validation, TCP backend + balance uri reject
(Bulgu #54-#61).
Round 22 FE error visibility + 3x stale-data lockouts,
referential integrity + cascade safety,
authentication & authorization, multi-cluster
isolation, apply_pending_changes concurrency,
script injection + bulk import multi-tenancy,
prefix-stripped signature comparison
(Bulgu #62-#82).
------------------------------------------------------------------
NO CORPORATE-SPECIFIC ARTIFACTS
------------------------------------------------------------------
This squash deliberately sanitises corporate hostnames, container
registry references, and TLS secret names into generic
placeholders (`your-registry.example.com/your-org`,
`haproxy-openmanager*.example.com`, `wildcard-tls`,
`taylanbakircioglu/haproxy-openmanager-*`) so the public artefact
contains no internal infrastructure detail. Pilot / development
history that retained those values stays in the corporate fork
and is NOT part of this commit.
- Agent scripts now detect and send ip_address in DAEMON heartbeat (Linux: ip route, macOS: ifconfig)
- Backend validates agent-reported IPs via ipaddress stdlib, COALESCE preserves existing on NULL
- IP/VIP change logging (non-critical, try/except wrapped) for operational visibility
- New source_file_hash column on agent_script_templates for reliable update detection
- Migration changed to ON CONFLICT DO NOTHING to prevent overwriting UI-customized scripts on restart
- GET /versions returns script_update_available flag (disk hash vs DB hash comparison with fallback)
- Frontend Alert banner warns users of new agent script versions and directs to Reset to Defaults
- Reset to Defaults and Popconfirm modals explicitly warn about custom script edit loss
- Full backward compatibility: old agents without ip_address field continue working unchanged
Made-with: Cursor
- Server-level change detection in bulk import (field-by-field comparison
for 17 server attributes with UPDATE/NO CHANGES status and tooltip)
- Multi-select delete for backends and frontends with dependency checks
- Dashboard "Backends Summary" address column for servers
- Fix unique constraint violation on bulk-create for existing servers
(natural key lookup matching DB constraint instead of backend_id FK)
- ORDER BY is_active DESC on all entity lookups to prefer active records
- Strip auto-generated content (ACME, rate-limit, WAF) from bulk import
comparison to eliminate false positive changes on re-import
- Fix toolbar overflow with Space wrap prop
- Frontend bulk delete modal clarity (selected vs deletable count)
- Version bump to 1.3.0
Made-with: Cursor
Bulk import accepted dots in frontend/backend/server names but UI and
backend validators rejected them with ^[a-zA-Z0-9_-]+$. After import,
entities with dots could not be edited. HAProxy itself allows dots in
section names, so the regex is expanded to ^[a-zA-Z0-9_.-]+$ across
all 12 validation points (5 React form rules, 1 ACL char-strip,
3 Pydantic validators, 1 WAF validator, 2 config-validator warnings).
When a backend block in haproxy.cfg does not specify explicit timeout
values, the bulk import was injecting hardcoded defaults (connect 10s,
server 60s, queue 60s) into the database. These then appeared in the
generated config and overrode the agent's defaults section. Now, only
explicitly declared timeouts are stored; omitted ones remain NULL so
the agent's existing defaults section stays in effect.
- Fix frontend port conflict validation to consider bind_address+port
combination instead of port-only. HAProxy allows same port on different
bind addresses (e.g., bind 10.0.0.1:443 vs bind 10.0.0.2:443). This
was blocking frontend edit/save in multi-VIP environments.
- Add form field dependency so port re-validates when bind_address changes.
- Fix API URL construction using window.location.host instead of hostname
to preserve non-standard ports (e.g., :8080), preventing CORS errors
in BulkConfigImport.
- Make CORS_ORIGINS configurable via environment variable.
- Step 3 (Enable ACME on Cluster) now shows a process icon instead of
a misleading green checkmark when ACME is enabled but not yet applied.
Per-cluster "(pending apply)" annotation for multi-cluster setups.
- Step 4 button and all /apply-management navigation buttons now say
"Apply Changes" instead of "Configure" for clearer guidance.
- Setup Guide auto-selects the correct cluster before navigating to
Apply Management, showing pending cluster names in alerts.
- Pending ACME disable changes are now correctly detected in Step 4
even when acme_enabled is already FALSE in the database.
- Entity snapshot rollback for cluster ACME settings: reject correctly
restores acme_enabled/acme_backend_url to pre-change values.
- Deduplication logic prevents "last wins" bug when multiple ACME
toggles are rejected in sequence.
- Connection leak prevention with try/finally around conn2 in ACME
config version creation.
- Step 4 branching uses boolean has_enabled instead of fragile string
truthiness check.
Made-with: Cursor
- Fix critical cascading NULL status bug in ACME challenge flow that could cause 404s
- Add defense-in-depth NULL handling across all ACME service methods
- Add new GET /api/letsencrypt/prerequisites endpoint for configuration checks
- Add interactive ACME Setup Guide with step-by-step navigation links
- Add URL-based tab navigation in Settings and SSL Management pages
- Harden retry flow: return clear 409 errors for invalid/cancelled orders
- Allow cancellation of invalid orders (backend + frontend)
- Improve error message extraction with consistent getErrorMsg helper
- Add status filter tabs (Active/Completed/Failed/All) for order list
- Add visual dimming for cancelled/invalid orders
- Add enhanced pagination with size changer and total count
- Add comprehensive diagnostic logging with ACME: prefix
- Update README with ACME architecture docs, quick start guide, and troubleshooting
Resolves#9
Made-with: Cursor
- Full dark mode support across all pages with lightbulb toggle in header
- Theme preference persisted in localStorage across sessions
- Ant Design 5 token-based theming (40+ components updated)
- Recharts dark mode: axes, grids, tooltips adapt to theme
- Login page redesigned with product-consistent blue-gray palette
- Overscroll bounce background matches dark theme
- All Servers search with multi-field filtering
- Bulk Config Import UI streamlined with collapsible guidelines
- ConfigProvider moved above AppContent for correct token resolution
Made-with: Cursor
Agent startup left HAPROXY_CONFIG_PATH empty when the /api/clusters jq
select() returned no output (e.g. transient API failure or cluster_id
mismatch). This caused "No existing HAProxy config found at: " errors
and partial config merge failures for newly added agents.
Three-layer fix:
1. Load HAProxy paths from config.json as baseline before get_cluster_paths()
2. Use local variables in get_cluster_paths() - only override globals when
API returns non-empty values (defensive against jq select() empty output)
3. Dynamically update paths from /api/agents/{name}/config response in
check_config_updates() - allows cluster path changes without reinstall
Applied to both linux_install.sh and macos_install.sh.
Made-with: Cursor
Root cause: config_status enum was created with only PENDING and APPLIED
values. The REJECTED value was never added due to a silent duplicate_object
exception in create_essential_tables(). This caused SSL certificate listing
to crash with "invalid input value for enum config_status: REJECTED" on
fresh installations.
Also adds scrollable containers to Apply Management page to prevent
Agent Sync Status card from being pushed off-screen.
Closes#7
Made-with: Cursor
Fixes#6
- Fix NameError in soft-deleted certificate reactivation path by
reordering variable extraction before DB operations
- Replace silent empty-array returns with HTTP 500 on SQL errors,
making failures visible in both API responses and server logs
- Add primary_domain migration for schema consistency across fresh
and upgraded installations (backfill from legacy domain column)
- Use primary_domain in non-cluster SSL query branch for schema
compatibility
- Surface SSL fetch errors in frontend via toast notifications
- Harden connection cleanup in error handlers with try/except
- Add QEMU + Buildx for linux/amd64,linux/arm64 multi-platform
Docker image builds
- Update GitHub Actions to latest versions (checkout v4, login v3,
build-push v6)
Made-with: Cursor
When user changes redirect type from location/prefix to scheme, the
target field (e.g. "https://example.com") stays as-is but the UI shows
a blank Select. If saved without re-selecting, invalid config would be
sent. Now auto-sets target to "https" when switching to scheme type.
Made-with: Cursor
- Enforce mutual exclusivity for -m flags (only one match method at a time)
- Show error status on value field when -f flag is used without absolute
file path, with tooltip explaining the requirement
- Replace free-text input with Select dropdown for redirect scheme type,
restricting to valid values (http/https) only
- Fix flag serialization order (-i → -m → -f) to prevent HAProxy parse errors
Made-with: Cursor
Replace plain TextArea inputs with an interactive card-based visual builder
for ACL rules, backend routing rules, and redirect rules. Includes:
- Structured ACL definition cards with match type, flags, and value fields
- Backend routing cards with operator (if/unless) and ACL condition selector
- Redirect rule cards with type, target, code, and condition fields
- Visual ↔ Raw mode toggle for each section
- HAProxy config preview panel
- Safe flag serialization order (-i → -m → -f) to prevent parse errors
- Guard against empty/incomplete rules in serializers
- Context-aware flag placeholder hints per match type category
Made-with: Cursor
Added standalone installation as a third deployment option alongside
Docker and Kubernetes. The react-scripts 5.0.1 upgrade in commit d4b58d6
resolved the Docker build failure reported in #4.
Made-with: Cursor
Comprehensive guide for installing HAProxy OpenManager on a standalone
server, VM, or LXC container without Docker or Kubernetes. Covers
PostgreSQL, Redis, backend, frontend, nginx reverse proxy setup,
systemd services, and firewall configuration.
Made-with: Cursor
- Backend/frontend now pull pre-built images from Docker Hub by default
- Build sections commented out (can be uncommented for local builds)
- Removed dev volume mount (./backend:/app) for production use
- Backend healthcheck changed from curl to python urllib (curl not in slim image)
- Added nginx /health endpoint for compose healthcheck
Made-with: Cursor