mirror of
https://github.com/taylanbakircioglu/haproxy-openmanager.git
synced 2026-09-16 23:55:13 +00:00
main
36 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0ee227363e |
fix(agent): adoption cannot hide a node; Config Import on fresh installs (v1.11.1)
Six fixes to the agent's discovery path and one long-standing parity gap, found
by auditing it in loops against a real fleet.
DISCOVERY REPORTS THAT NEVER REACHED THE SERVER
The agent posts an unmanaged keepalived.conf for adoption and caches the hash of
what it sent, so the file - which carries the VRRP password - is re-posted only
when it changes. Delivery was judged by curl's exit code, and curl without -f
exits 0 on 5xx too, so a report the server REJECTED was recorded as delivered.
Since a hand-maintained config does not change on its own, that node dropped out
of "Unmanaged keepalived detected" permanently; the only cure was deleting a
cache file on the node by hand.
- the report is cached only on a 2xx;
- GET /agents/{name}/keepalived-config now reports whether the server actually
holds a discovery for that agent, and the cache may only suppress while it
says yes - which is what lets nodes stuck from earlier releases recover on
their own, with nobody touching them;
- the flag is parsed with has() + tostring, not `// empty`: jq's alternative
operator returns the alternative for **false** as well as null, so the naive
form could not tell "no record" from "older backend" and the recovery would
have been completely inert;
- a 400/413/422 records the refusal so identical bytes are not re-posted
forever - 4xx and 5xx agent calls are never sampled out of the request log,
so an unattended loop would write a row carrying the whole config every
cycle - while 401 and 404 keep retrying, because here they mean a token
rotation or an agent row briefly absent, not a bad payload;
- the CLEAR path had the same exit-code defect, where it left a stale row
offering a managed node for adoption with nothing to ever retry it.
CONFIG IMPORT WAS A NO-OP ON FRESHLY INSTALLED AGENTS
check_config_requests uploads a node's live haproxy.cfg on request. It was
defined in the installer body and in the self-upgrade daemon, but not in the
heredoc a fresh install writes, and its call site is guarded by `type` - so on
such a node the operator asked for a config and nothing arrived, with no error
anywhere. Any agent that had self-upgraded at least once already had it, which
is why it went unnoticed. The self-upgrade definition is copied verbatim
(verified line-for-line). A freshly installed agent now polls that endpoint once
per cycle exactly as every upgraded agent already does; no node running today
changes behaviour.
DETERMINISTIC CONFIG PATH
A pool may hold several clusters and the join that resolves keepalived_config_path
was unordered, so the path handed to an agent could differ between polls whenever
two clusters disagreed - the agent would inspect a file that is not there and the
node would never appear, intermittently. A customised path now wins over the
shipped default, then the lowest cluster id. Verified against a real PostgreSQL
over seven arrangements: with one cluster per pool, or when every cluster carries
the default, the value is byte-identical to before.
Verified end to end on a production fleet and, for each decision, against the
real _kp_discover block rather than a paraphrase.
Backend suite: 1674 passed, 152 skipped. bash -n passes on the whole file and on
the fresh-install body in isolation. The keepalived path is logic-identical
across both daemon copies, now pinned by a test.
|
||
|
|
ef26860df9 |
feat(logging): unified request/response log with configurable retention (v1.11.0)
Until now the only record of what happened was `user_activity_logs`, which stores non-GET 2xx operations with no bodies. When something failed you could see that a counter went up, never what was sent or what came back. This adds one queryable timeline covering both directions: - inbound: every API call, including GETs and including 4xx/5xx, with the user, client IP, status, duration and — redacted, size-capped — the request and response bodies. - outbound: every HTTP call the backend makes, tagged with who it went to (ACME/Let's Encrypt, Cloudflare, GoDaddy, HAProxy stats, agents, the ACME diagnostics probe). Outbound rows inherit the inbound request's id, so one operator action and the CA/DNS calls it triggered read as a single trace: opening a failed "Request Certificate" shows the exact POST /acme/new-order and the CA's 429 underneath. Implementation notes: - Capture is a pure-ASGI middleware that TEES the request and response streams rather than draining them. `await request.body()` inside a BaseHTTPMiddleware would consume the receive channel and break the raw-body agent heartbeat handler. Registered last so it is outermost: it then sees the final client-visible response and seeds correlation_id_context before the error handler reads it. - Rows are written by a batching background writer with a bounded queue, so the request path never awaits the database and a saturated logger drops rows visibly (surfaced on the page) instead of blocking. Redaction runs on the writer, off the request coroutine. - Secrets never land: headers are an allowlist with Authorization/Cookie kept only as a presence marker; body keys and value shapes are redacted (passwords, tokens, api_token, API keys, private-key PEMs, JWTs); the ACME JWS request body is never stored, because a stored protected+signature pair is a replayable credential — a summary is logged instead; DNS-provider errors record only the exception type; the ACME HTTP-01 challenge endpoint is excluded so key_authorization is never captured. - Retention is operator-configurable in Settings -> Request Log: separate day counts for successful and failed rows (7 / 30) plus a hard row cap (500k), whichever is reached first. Pruned in batches under a Postgres advisory lock, with the day counts bound as parameters, never interpolated. - New permissions requestlog.read / requestlog.manage. super_admin and security_admin get both, operator gets read, viewer gets neither. Schema: one new table (request_logs) plus its settings seed, SCHEMA_VERSION 10 -> 11, auto-migrated. No existing table altered, no agent or rendered-config change. Kill switches: REQUEST_LOG_ENABLED=false (middleware never registered) or the `enabled` toggle in Settings. Tests: 245 new (7 backend files + 1 frontend), full suite 1655 backend + 17 frontend passing. |
||
|
|
dd7b7822bf | chore(version): bump to 1.10.3 - multi-account ACME wizard fix | ||
|
|
dbb9189f16 |
fix(ui): make Apply Management readable in dark mode (v1.10.2)
Three dark-mode defects reported on the Apply Management page, all the same
class of bug: light-mode colour literals hardcoded where theme tokens belong.
1. The "Pending Changes" box was painted background #fffbe6 with border
#ffe58f. In dark mode the text on top is light, so the version name,
timestamp and "View Change" link sat on a cream panel and were unreadable.
Measured contrast was 1.03:1; it is now 11.50:1 (secondary text 1.03:1 ->
7.02:1).
2. The added/removed rows in the View Change diff used #f6ffed/#52c41a and
#fff2f0/#ff4d4f, which stayed near-white inside the otherwise dark diff
panel. Now 5.49:1 (added) and 4.01:1 (removed), from 2.21:1 and 2.99:1.
3. The "Apply All Configuration Changes" confirm dialog came up white. This one
is not a colour literal: in Ant Design 5 the STATIC Modal.confirm / message /
notification APIs render into their own detached root and never see the app's
ConfigProvider, so they always fall back to the light algorithm. Registering
ConfigProvider.config({ holderRender }) once at the app root wraps that
detached root in the same ConfigProvider. Verified against the installed antd
5.29.3 source rather than assumed: config-provider/index.js sets
globalHolderRender, and modal/confirm.js wraps the dialog with it. This fixes
EVERY static dialog in the application — 12 components call Modal.confirm —
not just this page.
While in the file, six more instances of the same bug were fixed: the error
Alert border, the VIP pending-delete row, two ACME/pending version panels, the
applied-version panel and the agent-error recommendation box.
Light mode is byte-identical. Each token resolves under the default algorithm
to exactly the literal it replaced (colorWarningBg -> #fffbe6, colorSuccessBg ->
#f6ffed, colorErrorBg -> #fff2f0, colorInfoBg, colorErrorBorder, ...), so this
release can only change dark mode. Contrast was measured by resolving the real
design tokens under both algorithms and computing WCAG ratios, not by eye.
Note the diff rows were already low-contrast in LIGHT mode (2.21:1 and 2.99:1)
and remain so; that is the design system's own success/error pair and changing
it would alter the established light-mode appearance, so it is left alone.
holderRender is registered in an effect rather than during render, since
ConfigProvider.config() mutates antd module state; effects still run long
before a user can click anything that opens a static dialog.
Frontend only: no schema, no SCHEMA_VERSION bump, no API change, no environment
variable, zero agent impact. Backend suite unchanged at 1243 passed.
|
||
|
|
eee0a4716a |
feat(ssl): encrypt the pending CSR private key at rest (v1.10.1, closes #53)
Closes the follow-up filed during the v1.9.0 CSR review. The private key of a
PENDING CSR is now Fernet-encrypted in the database instead of being stored as
a raw PEM.
Why this key specifically: it is the one key in the system that sits idle. It
is generated at CSR creation, waits for an external CA to sign the request
(days to weeks), and is destroyed the moment the signed certificate is
imported. It is never transmitted to an agent and never leaves the server.
ssl_certificates.private_key_content and the ACME order keys are deliberately
NOT covered, because agents must receive those in plaintext on every poll, so
encrypting them at rest buys nothing without an end-to-end redesign.
Implementation follows the pattern already used for the VRRP secret, TOTP
secrets and DNS provider credentials: a new utils/csr_key_crypto.py with its
own CSR_ENCRYPTION_KEY env var and its own HKDF info string
("csr-private-key-v1"), so rotating one secret class never affects another.
No schema change and deliberately NO SCHEMA_VERSION bump: the Fernet token
replaces the PEM inside the existing ssl_csrs.private_key_pem TEXT column. A
bump would re-run the migration sequence and re-seed the four built-in roles to
their defaults, which is a needless side effect for a storage-format change.
Backward compatible with no data migration. Rows written before this release
hold a raw PEM and are still read unchanged; the discriminator is exact rather
than a heuristic, since a Fernet token is base64url and can never contain the
"-----BEGIN" marker. Legacy rows drain naturally because a CSR's key copy is
NULLed on import.
A key that cannot be decrypted (SECRET_KEY rotated while CSR_ENCRYPTION_KEY was
unset) now fails with an explicit "delete this CSR and create a new one" error.
Previously that situation would have surfaced as the far more confusing
"certificate does not match this CSR's private key".
Also documents all four per-purpose encryption keys in .env.template. Only
VIP_ENCRYPTION_KEY was listed; MFA_ENCRYPTION_KEY and
DNS_PROVIDER_ENCRYPTION_KEY had been missing since v1.6.0 and v1.8.0.
Verified before release, on a corporate pre-production environment and locally:
- Full backend suite 1234 -> 1243 passed (+9 new tests), 0 failed.
- Against a real Postgres: a CSR created through the API stores a Fernet token
with no PEM header in the column, and imports successfully.
- Full 1.10.0 -> 1.10.1 -> 1.10.0 drill on one database volume. The upgrade
logs "Schema already at version 10 (>= 10); skipping migration run", so no
migration executes and the built-in roles are not re-seeded. A CSR created on
1.10.0 with a plaintext key imports successfully after the upgrade, which is
the backward-compatibility guarantee proven against a real row rather than a
mock.
- rsa-2048, rsa-4096 and ecdsa-p384 all round-trip through create, encrypt,
decrypt and import.
- Key derivation is stable across processes: two independent containers sharing
SECRET_KEY decrypt each other's tokens (required for UVICORN_WORKERS > 1 and
multi-replica deployments), while a different SECRET_KEY yields None rather
than a wrong key or an exception.
- Downgrade behaviour was measured, not assumed: 1.10.0 cannot parse the token
and fails with HTTP 500 "key parse failed (encrypted?)" rather than pairing a
wrong key. The rollback note states the measured behaviour.
- No CSR endpoint returns the key in any form: list and detail responses
contain neither a PEM nor a Fernet token.
Not changed here, from the issue's "worth folding in" list: the create rate
limit is not a concurrency guard, create_csr holds a pooled connection across
RSA key generation, detail=str(e) echoes internal error text (a repo-wide
convention), and is_global skips cluster validation in both routers/ssl.py and
routers/csr.py. None are storage concerns and each is a separate change.
|
||
|
|
6bf6d016f5 | chore(version): bump to 1.10.0 - GoDaddy DNS-01 provider | ||
|
|
69e12f7459 |
chore(version): bump to 1.9.0 - CSR creation
- backend/version.json + frontend package version to 1.9.0 - README: feature list entry, CSR workflow section, SSL CSR API reference, v1.9.0 release notes - UPGRADE_GUIDE: v1.9.0 section (additive ssl_csrs table, SCHEMA_VERSION 9 -> 10, no RBAC changes, zero agent impact, rollback note) |
||
|
|
0ebf6583ea |
chore(version): bump to 1.8.10 — security release (RCE/auth/SSRF advisories fixed)
Single-source version bump (backend/version.json) with frontend/package.json and package-lock kept in sync (test_version_consistency). Marks the release that ships the GHSA-7rhv / GHSA-3p5c / GHSA-3vh4 fixes. |
||
|
|
f86a4331e8 |
fix(deps): patch websocket-driver (CRITICAL #52) and resolve remaining postcss (#12)
Follow-up to the consolidated security bump: - websocket-driver -> 0.7.5 (CRITICAL, message corruption; dev/build tooling) - resolve-url-loader -> 5.0.0 (pulls postcss ^8), resolving the last postcss<8.5.10 instance (#12) and #2 at the source. Project uses no SASS, so resolve-url-loader v4->v5 is inert at build time. Verified: full production docker build succeeds on node 18; postcss now resolves to a single 8.5.10 across the tree; deferred dev-only deps unchanged. |
||
|
|
d914f2398b |
fix(deps): patch frontend security advisories (11 Dependabot alerts, incl. all 4 HIGH)
Bump react-router-dom to ^6.30.4 (runtime open-redirect fix, CVE-2026-40181) and add scoped npm overrides to patch dev/build-toolchain transitive deps: ws (7.5.11 / wds-scoped 8.21.0), form-data (4.0.6 / jsdom-scoped 3.0.5), js-yaml (3.15.0 / eslint-scoped 4.2.0), http-proxy-middleware 2.0.10, launch-editor 2.14.1, postcss 8.5.10 (resolve-url-loader kept at 7.0.39), @babel/core 7.29.6. Deferred (breaking major / node20, not in production bundle): webpack-dev-server, serialize-javascript, uuid, @tootallnate/once, resolve-url-loader's postcss. Verified: plain `npm install` (no --legacy-peer-deps) + full production docker build succeed on node 18; only package.json + regenerated package-lock.json change. |
||
|
|
c79391cd13 |
feat(acl): accept HAProxy -f pattern-file references with advisory warnings (v1.8.9, Issue #38)
The manual Frontend editor, wizard and visual ACL builder hard-rejected the ACL `-f <file>` flag while bulk import accepted it. Worse, a frontend imported with an `-f` ACL could not be edited at all (422) until the ACL was dropped. The original guard predated the fail-safe apply flow: the agent runs `haproxy -c` before every reload, so a missing pattern file is rejected safely and the previous config keeps running. Pattern files are operator-managed host files — the same policy adopted for SPOE filter configs in v1.8.8. - models: remove the 5 `-f` hard rejects (frontend acl/redirect/use_backend validators + wizard string/dict-redirect guards); `$(`/backtick and X!X contradiction guards unchanged - routers/frontend: `_pattern_file_warnings` helper; non-blocking warning on create + update responses listing referenced pattern files (empty when no rule uses `-f` — zero noise) - routers/config: bulk-import preview advisory listing pattern files per frontend (cluster config-dir aware, next to the SPOE advisories) - React: remove the FrontendManagement submit gate and SiteWizard step gate; ACLRuleBuilder renders informational notes instead of errors and re-adds `-f (pattern file on host)` to the flag dropdown; create path now renders server warnings like update - tests: 4 reject-pins inverted to accept-pins; new test_acl_pattern_file_allow.py (accept/guards-kept/zero-noise/advisory); full suite green (1094 passed) |
||
|
|
9e2ea04777 |
feat(haproxy): preserve SPOE filter + frontend log-format on import/edit (v1.8.8, Issue #38)
Bulk import / manual edit silently dropped `filter spoe engine ...` (Coraza WAF) and frontend `log-format` because the parser recognised only a fixed directive set. The regenerated config then missed the SPOE engine, so HAProxy failed with "unable to find SPOE engine 'coraza' used by the send-spoe-group". - parser: capture `filter` + `log-format`/`log-format-sd` into new ParsedFrontend fields - db: additive nullable `log_format` + `filters` TEXT columns on frontends (SCHEMA_VERSION 8->9) - generator: new `filter` bucket flushed before http-request rules so `filter` precedes `send-spoe-group`; `log-format` kept in prelude - bulk import: preview dict, change-detection, persist (create + merge-update); cluster-aware SPOE pre-flight advisories (missing-filter + host-prerequisite) surfaced in the UI - manual CRUD: full round-trip (get/create/update) incl. React form fields (no null-wipe) - reject/rollback: restore the new columns; restore path + wizard helper kept in parity - backend `option spop-check` recognised (suppresses spurious warning for coraza-spoa) - tests: test_spoe_filter_import.py; full suite green (1079 passed) |
||
|
|
97b2452bd2 |
fix(version): single-source version reporting to stop UI version drift (v1.8.7)
The version shown in the UI (backend-sourced via /api/version) could lag the
real release. The canonical version lived at repo-root version.json, but the
backend image is built with context ./backend, so that file did not reach the
container unless a pipeline staged it; the backend then fell back to a hardcoded
constant in main.py that had to be bumped by hand and had drifted (it reported
1.8.4 after 1.8.5 and 1.8.6 shipped).
- version.json moves to backend/version.json (single source of truth), now
committed and co-located with main.py, so COPY . . bakes it into every image
directly - correct version in every deployment, no pipeline staging required.
- backend/main.py reads the co-located file; its in-code fallback is no longer a
real version ("unknown") so it can never silently drift again.
- docker-build.yml reads backend/version.json and drops the staging step;
docker-compose.localtest drops the stale root mount.
- backend/tests/test_version_consistency.py fails the build if main.py hardcodes
a real version, if the canonical file is missing/invalid, or if
frontend/package.json drifts from it.
Bumped to 1.8.7. No functional, schema, API, or agent change. Full suite green.
|
||
|
|
23257b02cf |
perf(api): opt-in uvicorn workers + heartbeat query consolidation (v1.8.6, Issue #35)
A user running the API on a 2-core/4GB host reported slow-feeling API responses (Issue #35 follow-up). Review of the hot paths found no pathological defect; the dominant factors are the single uvicorn worker (one core serves all requests) and the constant agent-poll baseline (4 requests per agent every 30s). Two zero-risk improvements: - backend/Dockerfile: CMD now honors UVICORN_WORKERS, falling back to WEB_CONCURRENCY and then 1. Flagless uvicorn natively honors WEB_CONCURRENCY, so the fallback keeps any deployment that relied on it byte-for-byte compatible; with the final default of 1 worker uvicorn runs in-process exactly as before. >1 enables the multiprocess supervisor so multi-core hosts can use all cores. Background tasks are already multi-replica safe (FOR UPDATE SKIP LOCKED / advisory locks), as exercised by the k8s HPA deployment (2-10 replicas). `exec` keeps uvicorn as PID 1 (clean SIGTERM, verified ~1s docker stop with 2 workers). docker-compose.yml passes UVICORN_WORKERS through as empty-when-unset so a user-set WEB_CONCURRENCY is never overridden; .env.template documents it. - routers/agent.py heartbeat (by-name endpoint): the agent's status/version/upgrade_status were read with three separate single-column SELECTs against the same row; now one SELECT. Identical values and None semantics (single consistent snapshot instead of three reads); saves two round-trips per heartbeat per agent every 30s. The legacy by-id heartbeat endpoint is untouched; the heartbeat API contract is unchanged for agents of every version. - README: new "Performance Tuning" section (worker/replica scaling, and how to use the X-Response-Time header plus "Slow request detected" logs to pinpoint slow endpoints). Verification: full backend suite in docker green (1063 passed, 151 skipped; also re-run by the runtime image build); worker-count expansion matrix (unset->1, UVICORN_WORKERS=2->2, WEB_CONCURRENCY=3->3, both->UVICORN_WORKERS, empty->fallback) all correct; default run confirmed single-process with uvicorn as PID 1 and healthy API; UVICORN_WORKERS=2 confirmed parent + 2 workers, healthy API, clean shutdown; live heartbeats verified for register + existing-agent paths AND degraded agents (no stats socket / haproxy stopped / garbage stats CSV / unknown backend in server_statuses): all return 200, agent row updates correctly, zero backend errors. No schema, API, or agent changes. |
||
|
|
c8d144ca9d |
fix(acme): remove stray paren breaking ACME completion task (v1.8.5, Issue #35)
The background order-completion task (complete_pending_acme_orders, 60s
cycle) failed on EVERY cycle since v1.8.0 with:
[ACME-COMPLETE] Error in completion task: syntax error at or near ")"
Root cause: the bounded DNS-01 retry OR-arm added to the atomic order-claim
query in v1.8.0 (13b65d9) carried one extra closing parenthesis, making the
whole SELECT invalid PostgreSQL. The claim is the task's first statement, so
the generic except swallowed it each minute and NO background ACME work ever
ran on v1.8.0-v1.8.4:
- orders were never claimed for finalize -> download -> save (http-01 too);
- advance_dns01_order never ran, so the DNS-01 TXT record was never
published - DNS-01 with an automated provider (e.g. Cloudflare) could
never validate (exactly the report in Issue #35);
- wizard-staged orders never left wizard_staged (same try block);
- retry_invalid_dns01 / reconcile_dns01_cleanup never executed;
- hourly-created renewal orders could never complete in the background.
Fix: drop the stray ')' (one line). Query semantics are unchanged.
Why the suite missed it: the unit tests mock asyncpg, so raw SQL never
reaches a real parser. Added a regression test that AST-scans the ACME
modules' SQL string literals (comments/quoted literals stripped) and fails
on unbalanced parentheses - it is red on the pre-fix tree and would have
caught the v1.8.0 regression at commit time. Scanned all six ACME modules:
this query was the only unbalanced SQL.
Verification: full backend suite in docker green (1062 passed, 151
skipped); the fixed query EXPLAINs cleanly on postgres:15; live localtest
run shows zero completion-task errors and a seeded pending order is claimed
("[ACME-COMPLETE] Claimed 1 order(s)"). Backend-only, no schema/API/agent
changes; fully backward compatible.
Reported-by: @tkkost (GitHub Issue #35)
|
||
|
|
64d42663cd |
fix(agent): stop installer self-kill in pre-installation cleanup (v1.8.4)
The Linux/macOS agent installer could abort during "pre-installation cleanup"
(terminal showed `Killing processes matching: haproxy-agent` then `Killed`,
returning to the prompt) when the install script's own command line contained
"haproxy-agent". The cleanup killed processes via `pgrep -f "$pattern"` starting
with the bare string "haproxy-agent", which also matched the running installer's
own command line and a sudo/PAM ancestor that the $$/$PPID self-exclusion did not
cover, so the installer terminated itself before installing.
- linux_install.sh / macos_install.sh: the cleanup kill loop now targets ONLY
the installed agent - "$INSTALL_DIR/haproxy-agent" (the daemon binary path) and
the agent service/label ("haproxy-agent.service" / "com.haproxy.agent") - never
the bare "haproxy-agent" substring. Neither pattern can match the installer's
own command line. The redundant bare pattern is dropped (the service is stopped
separately, and the binary-path pattern still catches a running daemon).
- frontend (AgentManagement.js): name the downloaded scripts
install-agent-<platform>.sh / uninstall-agent-<platform>.sh (matching the
backend's suggested filename) - defense in depth so this cannot resurface.
Installer-only change. The running agent and its privilege model are unchanged
(it runs as root for HAProxy reload, config writes, keepalived, and self-upgrade).
The cleanup runs only on a full interactive install (gated by SKIP_TO_DAEMON), so
daemon mode, self-upgrade, and config/version apply are unaffected. Both agent
scripts kept in sync. Scripts parse on bash 4.2-5.2; full backend suite green.
Addresses #31.
|
||
|
|
27fbe48c4b |
fix(agent): tolerate empty system_info in heartbeat JSON (v1.8.3)
A self-hosted agent could fail every heartbeat with HTTP 400 `Invalid JSON: Expecting property name enclosed in double quotes` when the system-info block it collects came back empty on an unusual host. The agent builds the heartbeat JSON as text, so an empty `$system_info` collapsed the `$system_info,` line to a bare comma and broke the payload. - Agent script (linux + macos, kept in sync): guard the fragment-form register_agent and send_heartbeat builders so an empty system_info falls back to a valid key and can never emit a bare comma. Uses the most portable bash glob test (no POSIX class / pattern substitution; verified on bash 3.2-5.2 and on Ubuntu/Debian/Rocky/Alpine/Amazon Linux). True no-op for healthy agents. - Backend heartbeat endpoint: parse the body as-is first and only run the malformed-JSON repair when parsing fails, so a valid heartbeat from any agent version is byte-for-byte untouched. The repair (now a testable helper) recovers a leading or doubled comma (the empty-system_info artifact) in addition to the existing empty-value / trailing-comma fixes. No agent version bump; self-upgrade and daemon mode are unaffected. Healthy agents of every version behave identically. Full backend suite green. Addresses #31. |
||
|
|
e86e86a53c |
fix(acme): scope ACME nonce per CA - fixes ZeroSSL registration (v1.8.2)
ZeroSSL/Google account registration failed with `malformed: The Replay Nonce could not be base64url-decoded`: the ACME client (a process-wide singleton) kept a single anti-replay nonce shared across certificate authorities, so a nonce issued by one CA could be sent to another, and the auto-retry only covered `badNonce`. - Scope the nonce per CA (self._nonce_by_dir keyed by directory_url): a nonce from one CA is never sent to another; account registration always uses a fresh nonce from the target CA. - Broaden the 400 retry to also recover from the nonce-malformed rejection. - Fix _b64url_decode padding (used for the EAB HMAC key). Backend-only; HTTP-01 and Let's Encrypt are unaffected. Addresses #35. |
||
|
|
70ebc02e09 |
fix(acme): Cloudflare token, ZeroSSL EAB, Apply Management (v1.8.1)
Follow-up fixes for the DNS-01 feature reported on #35: - Cloudflare: sanitize the API token (strip surrounding quotes + any non token68 chars) so a pasted token with quotes/spaces no longer fails with "Invalid request headers"; verify-on-save shows a precise hint when it cleaned the input. Covers the automated orchestrator path too. - ZeroSSL/Google EAB: enter the EAB Key ID and HMAC Key per-account in the Register Account dialog (falls back to the global Settings value when blank); base64-validate the HMAC key; humanize the externalAccountRequired failure; and preserve the deliberate 409/422 instead of downgrading them to 400. - Apply Management: cluster ACME enable/disable changes now show under a dedicated "ACME Challenge Routing" section, are counted in the Apply/Reject dialogs, and Apply/Reject All process them (previously "Rejected 0 HA/VIP change(s)") - consistent with every other entity. Reject rolls acme_enabled back to the original via ORDER BY created_at ASC over the snapshot chain. - getErrorMsg surfaces field-level validation messages. Backward compatible (additive / strict superset; HTTP-01 unchanged). Addresses #35. |
||
|
|
c492b26bb1 |
feat(acme): DNS-01 challenge support with pluggable DNS providers (v1.8.0)
Add ACME DNS-01 (TXT-record) validation alongside the existing HTTP-01, for internal/isolated clusters with no public port 80 and for wildcard certificates. Opt-in via a global kill-switch (default off); HTTP-01 is byte-for-byte unchanged, with zero agent or rendered-config changes. - Pluggable DNS provider interface (Manual + Cloudflare). Per-account credentials are Fernet-encrypted at rest, verified on save, and never returned by the API or written to logs/events/error_detail. - Non-blocking per-cycle orchestrator: publish (CAS) -> propagation grace (across cycles, no in-loop sleep) -> respond -> finalize/download, with a bounded fresh-order retry chain (1 original + 3 retries) on propagation lag. - Manual flow: user publishes the TXT record and confirms; manual DNS-01 cannot auto-renew unattended (auto-renew forced off and surfaced in the UI). - Migration v8: additive, idempotent columns on letsencrypt_accounts/orders and acme_challenges, plus a new letsencrypt_account_dns_credentials table. - Challenge-type-aware diagnostics (port80/routing/DNS checks skipped for DNS-01) and a DNS-01 event timeline in the order detail. - Frontend: DNS-01 account + credentials management, cert wizard adaptation, order-detail TXT records + verify, orders/renewal Method columns, and a Settings kill-switch. README, release notes, and API docs updated. Implements #35. |
||
|
|
a1192e602d |
feat: HA / VIP (Keepalived) management from the UI (#27)
Manage highly-available virtual IPs backed by Keepalived (VRRP) directly from the OpenManager
UI — no more SSHing into nodes to install/configure Keepalived by hand. Builds on the agent
pull-architecture: define the VIP centrally, click Apply, and the agents converge.
Highlights:
- New "HA / VIP" tab: create a virtual IP, pick a per-node interface, select which pool nodes
participate (MASTER/BACKUP roles + priorities); live MASTER/BACKUP/FAULT per node.
- On Apply, agents install & configure Keepalived (unicast VRRP, cloud-safe default) across the
major distros (Debian/Ubuntu, RHEL/CentOS/Alma/Rocky, Fedora, SUSE/openSUSE, Alpine) with a
HAProxy health-check, so the VIP fails over automatically when HAProxy drops.
- Single-node (a managed floating IP without failover) and multi-node VRRP failover both work.
- VIP changes ride the standard Apply Management flow with the standard "View Change" diff.
- Approval-gated deletion (safety): deleting a running VIP is staged for approval and the VIP
keeps running, untouched, until you approve it — an agent never tears a VIP down without an
explicit human approval. Per-VIP Diagnostics view; opt-in package uninstall (only on nodes
where OpenManager installed it). A node already running a hand-managed Keepalived is detected
and never overwritten ("externally managed").
- Fully opt-in and backward compatible: nodes/clusters without a VIP are unaffected. Adds
vip_instances + vip_members tables (idempotent SCHEMA_VERSION bump; existing data and
passwords unaffected) and a `vip` RBAC permission group.
- Also includes a HAProxy config-generator robustness fix: auto-inject a stick-table when a
frontend uses a stick counter (track-sc / sc_*_rate) but declares none.
On-prem / L2 (VRRP) scope; the UI notes the cloud caveat.
|
||
|
|
9d7a718671 |
fix(security): re-pin nginx to 1.31.1-alpine for the poolslip advisory (v1.6.5)
Supersedes the v1.6.4 stable pin (1.30.2-alpine) with the mainline patched release. No config, schema, or behavior changes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
62b1599354 |
fix(security): pin nginx to 1.30.2-alpine for the poolslip advisory (v1.6.4)
The bundled nginx reverse proxy was flagged for the nginx 'poolslip' advisory (affected: mainline <=1.31.0; fixed: stable 1.30.2+ / mainline 1.31.1+). The config-level mitigation (named capture groups instead of $1/$2 in rewrite) does not apply — the product's nginx config (nginx/nginx.conf and the k8s configmap) has no rewrite capture-group directives, only prefix locations + proxy_pass. So the fix is the version: pin nginx:alpine -> nginx:1.30.2-alpine in docker-compose.yml and k8s/manifests/10-nginx.yaml. No config/schema/behavior change (nginx only reverse-proxies). Version bumped to 1.6.4 across all layers. Verified in Docker: nginx -v=1.30.2; nginx -t OK on both the compose and production configmap configs; all proxied routes work through nginx; no nginx errors. |
||
|
|
b34d7cf811 |
fix: reactivate disabled backend servers from the UI + v1.6.3 (Issue #24)
A backend server toggled OFF (is_active=false) vanished from the UI with no way to reactivate it: GET /api/backends honored include_inactive for backends but the server sub-queries hardcoded 'AND is_active = TRUE'. - get_backends: server sub-queries now honor include_inactive (default callers unchanged); added last_config_status to the server payload so the UI can tell a DISABLED server (re-enableable) from a DELETION (pending delete). - toggle_server: persists an entity snapshot so an Apply-Management Reject rolls back is_active (previously left the server stuck disabled). - BackendServers.js: requests include_inactive, shows disabled servers with the ON/OFF switch + an 'Inactive' tag, hides only DELETION-pending servers, and keeps soft-deleted BACKENDS hidden (so include_inactive doesn't resurface them). - Config generation unchanged: disabled servers stay '# DISABLED:' comments and convert back to live lines when re-enabled. Startup migration hardening (multi-replica / rolling-deploy safety): create_essential_tables fails fast on lock contention and retries; run_all_migrations is serialized by a session advisory lock and gated by a schema_migrations version marker, so an already-current schema is skipped instead of issuing lock-heavy DDL that a serving replica's traffic could block at startup. Idempotent and fail-open. Version reported consistently across all layers (version.json, backend fallback, frontend package) -> 1.6.3. |
||
|
|
d9ef86f548 |
chore(deps): bump axios from 1.15.2 to 1.16.0 in /frontend (#23)
Bumps [axios](https://github.com/axios/axios) from 1.15.2 to 1.16.0. - [Release notes](https://github.com/axios/axios/releases) - [Changelog](https://github.com/axios/axios/blob/v1.x/CHANGELOG.md) - [Commits](https://github.com/axios/axios/compare/v1.15.2...v1.16.0) --- updated-dependencies: - dependency-name: axios dependency-version: 1.16.0 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
f3d4fb11bb |
fix(agent): accept agent token (X-API-Key) on cluster read endpoints (Issue #22)
Assigning a HAProxy agent failed with '401: Authorization header missing'
on GET /api/clusters. A cluster-read hardening had made GET /api/clusters and
GET /api/clusters/{id} accept only a user JWT in the Authorization header;
agents authenticate with their agent token in the X-API-Key header, so the
token was never read.
Both endpoints now accept either a user JWT (Authorization) or an agent token
(X-API-Key via validate_agent_api_key), mirroring the existing dual-auth on
POST /api/agents/generate-install-script. Anonymous access is still rejected,
so the original hardening is preserved. The auth guard is placed before the
try block so the failure surfaces as a clean 401 (not the 500-wrapped-401 in
the report). Agent install scripts now consistently send the token via
X-API-Key (pre-flight cluster check on linux/macos, and macOS get_cluster_paths
which previously used the wrong Authorization: Bearer header).
Also normalizes the platform in the uninstall-script generator so macOS agents
(which report platform 'darwin') no longer get a 400 from
GET /api/agents/generate-uninstall-script/darwin.
version 1.6.0 -> 1.6.2.
|
||
|
|
f2b3df517f |
chore(deps): bump axios from 1.14.0 to 1.15.2 in /frontend (#17)
Bumps [axios](https://github.com/axios/axios) from 1.14.0 to 1.15.2. - [Release notes](https://github.com/axios/axios/releases) - [Changelog](https://github.com/axios/axios/blob/v1.x/CHANGELOG.md) - [Commits](https://github.com/axios/axios/compare/v1.14.0...v1.15.2) --- updated-dependencies: - dependency-name: axios dependency-version: 1.15.2 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
bd6a31cb0d |
feat: v1.6.0 — Multi-Factor Authentication (Issue #18)
Adds opt-in TOTP-based Multi-Factor Authentication that is fully
backwards compatible with existing logins. Operators choose to enable
MFA per account; nothing changes for users who do not opt in.
Highlights
==========
* RFC 6238 TOTP (6 digits, 30s period, SHA1) with ±30s skew tolerance,
compatible with Microsoft / Google Authenticator, Authy, Duo, 1Password.
* Per-step replay protection (`mfa_last_used_totp_step`) so a captured
code cannot be reused inside the same window.
* Fernet-encrypted TOTP secrets at rest, key resolution via
`MFA_ENCRYPTION_KEY` env (HKDF-derived from `SECRET_KEY` as fallback).
* 10 single-use, bcrypt-hashed backup codes per user, formatted
`XXXX-YYYY` from a confusion-free alphabet (no 0/O/1/I/L).
* Two-step login flow: `POST /api/auth/login` returns `mfa_required`
+ `mfa_token`, then `POST /api/auth/login/mfa-verify` accepts a TOTP
code OR a backup code. JWT is minted only after MFA succeeds.
* Self-service: users enable / disable MFA from their own row in the
Users page; admins reset (single user or bulk) but never enable on
behalf of someone else (matches AWS IAM / GitHub / Google Workspace).
* Bulk emergency reset CLI: `scripts/admin-mfa-reset-all.sh`.
Security hardening
==================
* Atomic transactions with `SELECT … FOR UPDATE` on `mfa_pending_logins`
and `users` rows so concurrent verify / enroll calls cannot race.
* `/api/mfa/enroll/start` refuses re-enrollment when MFA is already on
(prevents silent secret rotation via a stolen JWT).
* Pydantic `ValidationError` messages are sanitized before reaching the
audit log so request bodies (TOTP / backup codes in flight) never
appear in plaintext.
* Slowapi rate limits are per-USER, not per-IP, with a trusted-proxy
XFF strategy so a single ingress address cannot exhaust the bucket
for thousands of operators (`MFA_TRUSTED_PROXY_CIDRS`,
`MFA_RATE_LIMIT_*` env-overridable).
* Login query now scopes to `is_active = TRUE` so a soft-deleted row
with the same username can no longer occlude the active user
(also closes a small account-enumeration side channel).
Database
========
Additive migrations (idempotent `ADD COLUMN IF NOT EXISTS`,
`CREATE TABLE IF NOT EXISTS`):
- users: mfa_enabled, mfa_method, mfa_secret_encrypted,
mfa_enrolled_at, mfa_last_used_at, mfa_last_used_totp_step
- mfa_backup_codes (user_id ON DELETE CASCADE)
- mfa_pending_logins (user_id ON DELETE CASCADE, challenge_token,
attempts, expires_at)
- mfa_pending_enrollments (user_id ON DELETE CASCADE)
Frontend
========
* Login page becomes a 3-phase state machine
(credentials → MFA → submitting); legacy single-step login is
preserved for users who haven't enrolled.
* New MFAEnrollModal (3-step wizard: QR + secret → verify → backup
codes) using `qrcode.react`.
* Users page shows MFA column + per-row enable/disable/reset actions.
Admins viewing other users with MFA off see a non-actionable info
icon explaining that only the user themselves can enable MFA.
Deployment
==========
* `MFA_ENCRYPTION_KEY` is added to `k8s/manifests/03-secrets.yaml` as
a placeholder; `SECRET_KEY` is also placeholder-ized so both are
injected by the existing pipeline pattern (sed-replace + apply).
* No new build-time env vars are required for the frontend. The SPA
uses `window.location.host` for `/api/*` and is routed by the
existing nginx ingress configuration.
* `frontend/.dockerignore` ensures host `.env*` files cannot bleed
into the production bundle.
Tests
=====
* New unit suites:
- `test_mfa_service.py` (TOTP, encryption, backup codes)
- `test_mfa_backwards_compat.py` (regression — non-MFA flow unchanged)
- `test_mfa_rate_limits.py` (env override + dataclass immutability)
- `test_mfa_rate_limit_key.py` (JWT key, trusted-proxy XFF, fallbacks)
* All existing 1000+ unit tests continue to pass.
Documentation
=============
* README MFA section (overview, day-to-day operations, emergency
reset CLI, env variables, rate-limit tuning).
* `scripts/README.md` documents the bulk reset script.
Issue: #18
|
||
|
|
d7208528f7 |
fix: v1.5.2 — ACME Diagnostics Panel Hardening + AGPL-3.0 relicense (Bulgu #94/#95/#96)
A focused hardening pass on the v1.5.0 ACME Diagnostic Panel
surface, exercised against a live production deployment (Round-25
+ Round-26 audits) and supplemented by an AGPL-3.0 relicense.
------------------------------------------------------------------
LICENSE — Relicense to AGPL-3.0-or-later
------------------------------------------------------------------
Effective v1.5.2 the project is licensed under the **GNU Affero
General Public License v3.0 (or later)**. v1.5.0 and v1.5.1
remain under the prior MIT terms.
The relicense is consistent with the project's intent as a
community-operated HAProxy management surface: forks that run
HAProxy OpenManager as a network service for third parties are
now required to publish their modifications under the same
license (AGPL §13). Day-to-day single-tenant deployments,
internal corporate use, and ordinary forks-for-fixes are
unaffected.
Changes:
* LICENSE replaced with full AGPL-3.0 text.
* README "## License" section rewritten with the AGPL summary
+ the network-service obligation.
* frontend/package.json gains `"license": "AGPL-3.0-or-later"`.
------------------------------------------------------------------
BULGU #94 / #95 — Diagnostic Panel Must Never Opaque-500
------------------------------------------------------------------
Live exercise of the v1.5.0 Diagnostic Panel against a deployed
build surfaced two opaque-500 paths. The panel exists to make
ACME failures legible; producing an opaque HTTP 500 defeats the
entire feature. Fix shape: every endpoint now returns either a
canonical 4xx (auth / not-found / rate-limit) or an HTTP 200
"structured failure envelope" that the React UI knows how to
render — never a 500 for an in-suite failure.
Affected paths:
POST /api/letsencrypt/orders/{order_id}/diagnostics
Pre-fix: a UndefinedColumnError or DB-connectivity failure
inside `run_checks` bubbled out of the bare try/finally and
surfaced as a generic 500 with no operator-actionable detail.
Post-fix: setup-stage and run-stage failures are caught
separately and converted to a `status: diagnostics_unavailable`
envelope carrying `error_stage`, `error_type`, `error_message`,
and a `correlation_id` that the operator can grep in the
backend log. Individual checks are wrapped in `_safe_check`
so one broken check (e.g. DNS lookup timeout) never crashes
the suite — the failing check shows up as `status: "fail"`
with its message, the others still run.
GET /api/letsencrypt/orders/{order_id}/events
Pre-fix: the SQL `SELECT … status FROM user_activity_logs`
referenced a column that did not exist in the canonical
migration; every diagnostic-panel open against an order with
any user-activity-log correlation got an `UndefinedColumnError`
500. Post-fix: the endpoint now introspects
`information_schema.columns` and projects only the columns
actually present. Partial failures (one source dies, the
other works) are reported via `meta.errors[]` rather than
collapsing the whole timeline.
POST /api/letsencrypt/orders/{order_id}/diagnostics/{check_id}/rerun
Same structured-envelope contract as the full-suite POST,
scoped to a single check row.
Frontend (`frontend/src/components/ACMEAutomation.js`):
* Distinct `diagRunError` / `diagEventsError` / `diagMeta`
states so the modal can render the cause inline (Antd Alert)
instead of a silent dropdown.
* Event-log auto-tail polling backs off after 3 consecutive
failures so the Network tab does not get spammed with 500s
every 5s.
* Correlation IDs visible in every error banner.
------------------------------------------------------------------
BULGU #96 — Clean 404 for Out-Of-Range order_id
------------------------------------------------------------------
A live exercise of the post-#94 diagnostic panel against the
deployed build surfaced one remaining contract gap. A path-
param `order_id` outside the Postgres int4 range
(e.g. > 2_147_483_647) caused `_load_order` to raise
`asyncpg.exceptions.DataError: invalid input for query
argument $1: 2147483648 (value out of int32 range)`. Round-25
correctly surfaced this in a `diagnostics_unavailable`
envelope — but that envelope leaked SQL implementation detail
("query argument $1", "int32 range", DataError class name)
into the operator-facing response body.
Semantically an out-of-range integer can never reference a
real order — it's just "not found". `_load_order` now catches
`asyncpg.exceptions.DataError` and re-raises a canonical
`HTTPException(404, "Order {id} not found")`. Because all
three endpoints re-raise `HTTPException` from their outer
try/except (the Round-25 envelope only fires for non-
HTTPException crashes), the canonical 404 path now wins
end-to-end across /diagnostics, /events, and /rerun.
------------------------------------------------------------------
TEST / LINT / LIVE VERIFICATION
------------------------------------------------------------------
* Backend pytest 1104/1104 (the +20 vs v1.5.1's 1084 are the
Round-25 and #96 contract pins; see
test_acme_diagnostics_router_round25.py).
* Live prod-canary verification: every endpoint return shape
confirmed against the deployed build — int4 overflow returns
clean 404 with no SQL leak, normal paths return Round-25
envelopes, HTTP method matrix returns 405 on wrong verbs,
no auth returns 401, invalid `check_id` returns 400, and
`meta.correlation_id` is present on every diagnostic
response.
------------------------------------------------------------------
COMPATIBILITY
------------------------------------------------------------------
* No breaking API contract changes: `status` field on the
diagnostic response can now be `"diagnostics_unavailable"`
in addition to the existing pass-through of the
underlying order status (`pending` / `valid` / `invalid` /
`cancelled` / …) — older UIs that only switch on the
existing values render the `diagnostics_unavailable`
case as "unknown status" rather than crashing.
* Frontend handles the new envelope shape AND the legacy
HTTP 4xx/5xx paths.
|
||
|
|
2e7db4d99f |
fix: v1.5.1 — Round-23 + Round-24 audit follow-ups (Bulgu #83 → #93)
A live-deployment audit pass over the v1.5.0 Site Wizard + ACME Diagnostic Panel surface. Two adversarial review rounds (R23, R24) each capped by an end-to-end smoke test against a multi-cluster staging deployment. Bulgu #83 — Frontend Management page warned about stale data without a clear retry CTA. The toast now carries an in-place "Reload" action and the page-level Empty state surfaces the same recovery affordance, so operators never get stuck on a stale-data view without an obvious way out. Bulgu #84 — ACME diagnostics ran with the wrong "last_heartbeat" column reference against the agents table. Aligned the SELECT with the actual schema column (`last_seen`); pinned by an idempotent regression test in `test_acme_diagnostics.py`. Bulgu #85 — ACME order error_detail rendering could leak the raw asyncpg/SQL exception class name when humanize_error_detail encountered an unhandled CA response shape. Added a backwards- compatible fallback branch that emits an "ACME error (raw)" panel without exposing parse_error class name to the user. Bulgu #86 — Multi-cluster apply with concurrent rejects could leave wizard_staged orders dangling without their parent draft. Pinned via reject_order_with_cluster_orphan test. Bulgu #87 — Frontend Management page list virtualization mis-keyed during a re-sort + stale-row replace race; fixed by keying rows on `id + version` so React reconciler does not reuse DOM for a logically different row. Bulgu #88 — Site Wizard "Cancel" mid-flow now surfaces an unsaved-draft prompt with explicit Save / Discard buttons (and the same prompt on browser tab close), so the operator never loses 5 steps of input to an accidental ESC. Bulgu #89 — Existing-cert SSL mode showed an empty dropdown when the cluster had >100 certs because the listing endpoint default-limited results. Endpoint now exposes pagination AND the wizard switches to client-side filtering above 50 rows. Bulgu #90 — ACME pre-check on the wizard preview path did NOT re-validate the account against `letsencrypt_accounts` if the operator stepped Back/Forward between SSL and Review. Added a debounced re-validation on Review entry. Bulgu #93 — Site Wizard hsts_enabled toggle in HTTPS frontend was idempotent-by-name (the generated `http-response set-header Strict-Transport-Security` line could duplicate across a Save + Apply cycle). The renderer now upserts the header in place. Cumulative outcome: backend pytest 1084/1084, frontend lint clean, and a 6-hour live-deployment smoke session against staging with no regressions reported. |
||
|
|
02b1cb2bca |
feat: v1.5.0 — Site Wizard (Issue #14) + ACME Diagnostic Panel (Issue #13)
Closes #13, Closes #14. This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0 shipped the ACME stability & enterprise audit (Issues #10/#11/#12). v1.5.0 builds on that foundation with two co-equal headline features plus a 22-round audit campaign hardening the prior configuration surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0 lands in v1.5.2). ------------------------------------------------------------------ HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13) ------------------------------------------------------------------ A live pre-flight + post-failure diagnostic surface for every ACME order, reachable from the ACME Automation page. The panel exists to make ACME failures legible to operators who do NOT have shell access to the API host. Endpoints (`backend/routers/acme_diagnostics.py`): POST /api/letsencrypt/orders/{order_id}/diagnostics Run the full 5-check suite (DNS / port-80 / routing / account / agents) and humanize the order's `error_detail` (>=11 RFC-8555 problem types, backwards compatible with legacy plain-string failures). POST /api/letsencrypt/orders/{order_id}/diagnostics/ {check_id}/rerun Re-run a single check in place — used by the "Re-run" button on every row of the modal's pre-flight table. GET /api/letsencrypt/orders/{order_id}/events Merged event timeline combining the typed `acme_order_events` rows with correlated `user_activity_logs` entries (resource_type = 'letsencrypt_order' AND resource_id = order_id). The diagnostic modal auto-tails this timeline every 5 seconds while open. Service-level checks (`backend/services/acme_diagnostics.py`): * DNS resolution via stdlib socket.gethostbyname_ex through run_in_executor (intentionally avoiding an aiodns runtime dep for v1.5.0). * Port-80 HEAD probe, target locked to the order's domains, success on HTTP 200 OR 404, warns on egress timeout (corp egress policies routinely blackhole outbound 80 — fail-hard would be too noisy). * SSRF guard: probe refuses non-public IPs and surfaces the skip in the diagnostic result; IPv4-mapped IPv6 normalisation closes the `::ffff:169.254.169.254` cloud-metadata vector. * HAProxy routing presence check: matches the order's cluster_ids to a port-80 HTTP frontend. * ACME account validity check against `letsencrypt_accounts`. * Agent presence check (>=1 active agent in target cluster). * Every sub-check wrapped in a wall-clock timeout to bound impact on the API event loop. RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate limit on both run and rerun, backed by the (user_id, action, created_at DESC) composite index. Frontend (`frontend/src/components/ACMEAutomation.js`): * "Diagnose" button on every order row + the existing "stuck order" warning row. * Modal with two tabs: - Pre-flight Checks (Antd Table with status pills + Re-run buttons + humanized error banner) - Event Log (Antd Timeline with auto-tail polling, scroll- to-bottom, pause-on-hover) * Correlation IDs surfaced in error banners and individual check fail details for backend-log lookup. ------------------------------------------------------------------ HEADLINE FEATURE B — Site Setup Wizard (Issue #14) ------------------------------------------------------------------ A single guided flow that creates a Backend + Servers + HTTP Frontend (and optional HTTPS Frontend) in one atomic transaction. Endpoints (`backend/routers/site_wizard.py`): POST /api/site-wizard/preview — diff-preview the changeset POST /api/site-wizard/create — atomic execute POST /api/site-wizard/reject — clean rollback (including any wizard_staged ACME orders) GET /api/site-wizard/drafts — draft persistence PUT /api/site-wizard/drafts/{id} — save/update DELETE /api/site-wizard/drafts/{id} Feature surface: * One screen captures both backend (mode + servers) AND frontend (http + optional https + SSL mode) inputs. * SSL modes: ACME (new order, HTTP-01 only for v1.5.0), Upload (existing PEM), Existing (link to a stored cert), or None. * ACME-staged path: wizard_staged_until watermark on the `letsencrypt_orders` row defers finalisation until agent confirmation; per-mode reject cleanly cancels and rolls back the staged order. * Live diff preview against the cluster's current generated config (renderer-evolution noise stripped — track-sc<N> dedup, per-server cookie strip, defaults-cookie inheritance, listen-block flattening). * Draft persistence with PEM stripped at save time (private keys never round-trip through the drafts table). * Per-cluster multi-tenancy: drafts and wizard_staged orders are isolated to the creating user's cluster scope. Frontend (`frontend/src/components/SiteWizard.js`): * 4-step Antd Steps flow: Backend → Frontend → SSL → Review. * Render the live diff preview inline before commit. * Antd Form-level validation mirrors backend Pydantic validators (numeric bounds, HAProxy reserved keywords, ALPN consistency, IPv6 scope-id, domain regex, server name dedup). ------------------------------------------------------------------ AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1 → #82) ------------------------------------------------------------------ v1.5.0 includes 22 adversarial review passes. Each round produced its own commit set in the corporate development line; this squash collapses those into the v1.5.0 release artefact. Highlights: Round 1-4 Site Wizard core: dry-run parity, single-line value injection guard, ACL -f pattern-file block, SSL parity, timeout regex, form-state pin. Round 5-7 defaults-cookie inheritance, server-named-cookie guard, fe/be mode mismatch, duplicate server names, health_check_uri + server_address validators. Round 8-10 cookie_name / cookie_options newline-injection guard, dry-run parity (round 9), TCP-mode HTTP-only feature blockers. Round 11 SSL name path traversal + health-check >= 1. Round 12-13 SSL & ACME deep dive (Bulgu #23-#32). Round 14 single-line value injection (Bulgu #33). Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy reserved keywords, ALPN/TLS consistency, all-backup, multi-domain & multi-user enterprise edges, drain/HSTS/post-completion (Bulgu #34-#53). Round 18-21 concurrency, agent state, TCP-mode HTTP-only, list size caps, IPv6 scope-id, preview account validation, TCP backend + balance uri reject (Bulgu #54-#61). Round 22 FE error visibility + 3x stale-data lockouts, referential integrity + cascade safety, authentication & authorization, multi-cluster isolation, apply_pending_changes concurrency, script injection + bulk import multi-tenancy, prefix-stripped signature comparison (Bulgu #62-#82). ------------------------------------------------------------------ NO CORPORATE-SPECIFIC ARTIFACTS ------------------------------------------------------------------ This squash deliberately sanitises corporate hostnames, container registry references, and TLS secret names into generic placeholders (`your-registry.example.com/your-org`, `haproxy-openmanager*.example.com`, `wildcard-tls`, `taylanbakircioglu/haproxy-openmanager-*`) so the public artefact contains no internal infrastructure detail. Pilot / development history that retained those values stays in the corporate fork and is NOT part of this commit. |
||
|
|
07942a82e8 |
feat: ACME stability & enterprise audit (v1.4.0) — fixes #10 #11 #12
Issue #10 — Silent timezone failure on ACME certificate save - Normalize tz-aware expiry_date to UTC tz-naive before INSERT/UPDATE in ssl_certificates (TIMESTAMP WITHOUT TIME ZONE) — restores ACME download path Issue #11 — Duplicate _acme_challenge_backend in generated config - Generator guard: skip auto-append when backend already rendered - Agent-sync filter: _should_sync_backend() drops system-managed backends and detaches their server rows to prevent orphans - Restore filter: IGNORED_BACKENDS skips reserved names during cluster restore - Parser warning: reserved_backend_names blocks accidental manual import - Cleanup migration: removes orphan rows + cascading server entries (idempotent) Issue #12 — Validated orders required manual completion - New 60s background task complete_pending_acme_orders, flag-independent, multi-replica safe via FOR UPDATE SKIP LOCKED + 30s updated_at watermark - Per-order pg_advisory_lock(0x41434D45, order_id) serializes UI-Complete and auto-task races; idempotency guard returns existing certificate cleanly - retry_order endpoint reports in_progress: true within 30s window so the UI surfaces an info toast instead of duplicating CA requests - ACMEAutomation surfaces stuck orders (status=valid && !ssl_certificate_id) with a one-click Complete action and Cancel fallback; conditional 30s polling Other hardening - ACME state machine error_detail persisted as structured JSON across challenge, finalize, download stages for actionable post-mortems - CertificateRequest Pydantic model: domain regex + min_length/max_length and cluster_ids defaulting to all ACME-enabled clusters when "global" is selected - Renewal cluster fallback now requires acme_enabled=TRUE in addition to active - _complete_certificate preserves manual cluster assignments on renewal, surfaces cluster_errors, raises explicit error on missing private key - Audit logging covers acme_certificate_requested/revoked, ca_chain_imported, account created/deactivated/purged, order retried/cancelled - Settings UI exposes acme.staging_url_override for private test CAs (Pebble) - Schema additions: acme_challenges.attempts (default 0) and last_attempt_at, index idx_letsencrypt_orders_status_updated; all migrations idempotent Tests (41/41 passing) - test_acme_expiry_normalize, test_acme_duplicate_backend, test_acme_state_machine, test_acme_pydantic_validation, test_acme_audit_logging, test_acme_concurrency CI / packaging - docker-build.yml reads version.json and pushes additional product-version tag (e.g. 1.4.0) alongside latest and timestamp build id Closes #10 Closes #11 Closes #12 |
||
|
|
f37f3afd71 |
feat: bulk import change detection, multi-select delete, auto-content stripping
- Server-level change detection in bulk import (field-by-field comparison for 17 server attributes with UPDATE/NO CHANGES status and tooltip) - Multi-select delete for backends and frontends with dependency checks - Dashboard "Backends Summary" address column for servers - Fix unique constraint violation on bulk-create for existing servers (natural key lookup matching DB constraint instead of backend_id FK) - ORDER BY is_active DESC on all entity lookups to prefer active records - Strip auto-generated content (ACME, rate-limit, WAF) from bulk import comparison to eliminate false positive changes on re-import - Fix toolbar overflow with Space wrap prop - Frontend bulk delete modal clarity (selected vs deletable count) - Version bump to 1.3.0 Made-with: Cursor |
||
|
|
d4b58d62c7 |
fix: upgrade react-scripts to 5.0.1 for frontend Docker build compatibility
- react-scripts 4.0.3 -> 5.0.1 (Webpack 5, supports modern JS in node_modules) - @monaco-editor/react 3.7.5 -> ^4.6.0 (resolves peer dep conflict with monaco-editor) - Removed --legacy-peer-deps from Dockerfile (no longer needed, was causing ajv resolution issues) - Added @babel/plugin-proposal-private-property-in-object as explicit devDependency - Removed overrides section (not needed with react-scripts 5) Made-with: Cursor |
||
|
|
68a4d1ac23 |
fix: resolve Babel version conflict in frontend Docker build
Add npm overrides to pin @babel/core and @babel/plugin-proposal-private-property-in-object to versions compatible with react-scripts 4.0.3. Fixes build failure caused by transitive dependency resolution pulling incompatible Babel 7.12.3. Made-with: Cursor |
||
|
|
6aae0f4309 | Initial commit |