Follow-up to #57. The contributed tests render the whole ACMEAutomation tree
(antd Steps + Form + Select) and drive it through all three wizard steps, which
takes 4-9 seconds per test. Under jest's default 5s per-test limit two of them
failed, so `npm test` did not pass as shipped:
✕ an explicitly picked HTTP-01 account survives the step change and is what
gets submitted -> Exceeded timeout of 5000 ms
✕ the wildcard guard still applies on the Review step, where Submit lives
-> Exceeded timeout of 5000 ms
The PR's reported 5/5 holds only when the runner is invoked with an explicit
--testTimeout. Setting it in the file instead means the suite passes however it
is invoked, which matters because the frontend image build runs `npm run build`
and never the tests, so nothing else would have caught this.
Verified with the default runner (no flags) after the change: 5/5 pass.
Test-only. No production code touched.
Multi-account ACME: the certificate wizard honours the selected account (v1.10.3).
Verified before merge. The contributed regression tests were run against the PRE-FIX component to confirm they actually catch the bug: 4 failed / 1 passed, reproducing the reported symptom exactly (Expected "http-01" / Received "dns-01", Review rendering the default account's email instead of the picked one, and the wildcard warning absent). With the fix: 5/5 pass. The three root causes were each confirmed against the tree — rc-field-form's useWatch honours options.preserve (es/useWatch.js:66), the backend default is ORDER BY created_at DESC (routers/letsencrypt.py:505) while the list is served ORDER BY id (:213), and the null-dns_provider normalization matches the backend's own at :473. Frontend production build succeeds; backend suite unchanged at 1243 passed.
Renders the real component and walks it through all three wizard steps, because
the bug these cover was invisible to any unit test: it only appeared once the
wizard advanced PAST the step that owns the account Select, since Form.useWatch
reports only currently-rendered fields.
Assertions describe behaviour rather than markup - the primary evidence is the
POST body (account_id paired with challenge_type), compared against a fixture
whose DNS-01 account is deliberately both the lower id and the older account,
which is the exact shape that made the UI default and the backend default
disagree.
Verified by running the suite against the pre-fix component: the picked-account
test reports challenge_type "dns-01" where "http-01" is expected, the Review
test shows "Active (dns@example.com) / Challenge Method: DNS-01 / DNS provider:
godaddy" for a chosen HTTP-01 account - the reported symptom reproduced - and
the wildcard-guard test finds Submit enabled. Four fail, one passes: the DNS-01
selection path, kept as a positive control because it worked before (the
default the wizard fell back to happened to be the DNS-01 account) and must
keep working after.
Adds src/setupTests.js with the ResizeObserver and matchMedia polyfills jsdom
lacks and Ant Design 5 needs before any Select can open.
Release note covering the three compounding faults, and upgrade notes stating
that this is frontend-only with nothing to do on upgrade. Two points are called
out for operators rather than glossed: installations that never picked an
account explicitly were already using the newest valid account, so only the
preview was wrong; and the wildcard guard that stopped applying on Review was a
lost warning, not a correctness hole, since the backend still rejected those
requests.
With more than one account registered, picking an HTTP-01 account in Request
ACME Certificate still submitted a DNS-01 request, which the API rejected with
"The selected ACME account has no DNS provider configured for DNS-01."
Three faults compounded:
1. Form.useWatch reports only fields that are currently rendered. The account
Select lives on the Configuration step, so the moment the wizard advanced to
Review the watch read undefined and the wizard fell back to the default
account - even though the value was still in the form store. Both wizard
watches now pass preserve: true. The same fault silently disabled the
wildcard guard on Review, the one step where Submit lives.
2. The UI and the backend disagreed on which account is the default. The
backend takes the newest valid account (ORDER BY created_at DESC); the UI
took the first valid entry of a list ordered by id, i.e. the oldest - the
opposite account whenever the two differ. The wizard now resolves the same
account and sends account_id explicitly, so there is no guess left to
disagree about.
3. account_id was read from the form store while challenge_type came from the
reverted account object. Both are now derived from one resolved account, so
the pair can no longer describe two different accounts.
Also: the Review step showed the default account's address instead of the
chosen one, and Submit stayed enabled when the resolved account was
deactivated (the fallback can land on a non-valid account, and the Select
lists deactivated accounts).
Single-account installations are unaffected.
Three dark-mode defects reported on the Apply Management page, all the same
class of bug: light-mode colour literals hardcoded where theme tokens belong.
1. The "Pending Changes" box was painted background #fffbe6 with border
#ffe58f. In dark mode the text on top is light, so the version name,
timestamp and "View Change" link sat on a cream panel and were unreadable.
Measured contrast was 1.03:1; it is now 11.50:1 (secondary text 1.03:1 ->
7.02:1).
2. The added/removed rows in the View Change diff used #f6ffed/#52c41a and
#fff2f0/#ff4d4f, which stayed near-white inside the otherwise dark diff
panel. Now 5.49:1 (added) and 4.01:1 (removed), from 2.21:1 and 2.99:1.
3. The "Apply All Configuration Changes" confirm dialog came up white. This one
is not a colour literal: in Ant Design 5 the STATIC Modal.confirm / message /
notification APIs render into their own detached root and never see the app's
ConfigProvider, so they always fall back to the light algorithm. Registering
ConfigProvider.config({ holderRender }) once at the app root wraps that
detached root in the same ConfigProvider. Verified against the installed antd
5.29.3 source rather than assumed: config-provider/index.js sets
globalHolderRender, and modal/confirm.js wraps the dialog with it. This fixes
EVERY static dialog in the application — 12 components call Modal.confirm —
not just this page.
While in the file, six more instances of the same bug were fixed: the error
Alert border, the VIP pending-delete row, two ACME/pending version panels, the
applied-version panel and the agent-error recommendation box.
Light mode is byte-identical. Each token resolves under the default algorithm
to exactly the literal it replaced (colorWarningBg -> #fffbe6, colorSuccessBg ->
#f6ffed, colorErrorBg -> #fff2f0, colorInfoBg, colorErrorBorder, ...), so this
release can only change dark mode. Contrast was measured by resolving the real
design tokens under both algorithms and computing WCAG ratios, not by eye.
Note the diff rows were already low-contrast in LIGHT mode (2.21:1 and 2.99:1)
and remain so; that is the design system's own success/error pair and changing
it would alter the established light-mode appearance, so it is left alone.
holderRender is registered in an effect rather than during render, since
ConfigProvider.config() mutates antd module state; effects still run long
before a user can click anything that opens a static dialog.
Frontend only: no schema, no SCHEMA_VERSION bump, no API change, no environment
variable, zero agent impact. Backend suite unchanged at 1243 passed.
Closes the follow-up filed during the v1.9.0 CSR review. The private key of a
PENDING CSR is now Fernet-encrypted in the database instead of being stored as
a raw PEM.
Why this key specifically: it is the one key in the system that sits idle. It
is generated at CSR creation, waits for an external CA to sign the request
(days to weeks), and is destroyed the moment the signed certificate is
imported. It is never transmitted to an agent and never leaves the server.
ssl_certificates.private_key_content and the ACME order keys are deliberately
NOT covered, because agents must receive those in plaintext on every poll, so
encrypting them at rest buys nothing without an end-to-end redesign.
Implementation follows the pattern already used for the VRRP secret, TOTP
secrets and DNS provider credentials: a new utils/csr_key_crypto.py with its
own CSR_ENCRYPTION_KEY env var and its own HKDF info string
("csr-private-key-v1"), so rotating one secret class never affects another.
No schema change and deliberately NO SCHEMA_VERSION bump: the Fernet token
replaces the PEM inside the existing ssl_csrs.private_key_pem TEXT column. A
bump would re-run the migration sequence and re-seed the four built-in roles to
their defaults, which is a needless side effect for a storage-format change.
Backward compatible with no data migration. Rows written before this release
hold a raw PEM and are still read unchanged; the discriminator is exact rather
than a heuristic, since a Fernet token is base64url and can never contain the
"-----BEGIN" marker. Legacy rows drain naturally because a CSR's key copy is
NULLed on import.
A key that cannot be decrypted (SECRET_KEY rotated while CSR_ENCRYPTION_KEY was
unset) now fails with an explicit "delete this CSR and create a new one" error.
Previously that situation would have surfaced as the far more confusing
"certificate does not match this CSR's private key".
Also documents all four per-purpose encryption keys in .env.template. Only
VIP_ENCRYPTION_KEY was listed; MFA_ENCRYPTION_KEY and
DNS_PROVIDER_ENCRYPTION_KEY had been missing since v1.6.0 and v1.8.0.
Verified before release, on a corporate pre-production environment and locally:
- Full backend suite 1234 -> 1243 passed (+9 new tests), 0 failed.
- Against a real Postgres: a CSR created through the API stores a Fernet token
with no PEM header in the column, and imports successfully.
- Full 1.10.0 -> 1.10.1 -> 1.10.0 drill on one database volume. The upgrade
logs "Schema already at version 10 (>= 10); skipping migration run", so no
migration executes and the built-in roles are not re-seeded. A CSR created on
1.10.0 with a plaintext key imports successfully after the upgrade, which is
the backward-compatibility guarantee proven against a real row rather than a
mock.
- rsa-2048, rsa-4096 and ecdsa-p384 all round-trip through create, encrypt,
decrypt and import.
- Key derivation is stable across processes: two independent containers sharing
SECRET_KEY decrypt each other's tokens (required for UVICORN_WORKERS > 1 and
multi-replica deployments), while a different SECRET_KEY yields None rather
than a wrong key or an exception.
- Downgrade behaviour was measured, not assumed: 1.10.0 cannot parse the token
and fails with HTTP 500 "key parse failed (encrypted?)" rather than pairing a
wrong key. The rollback note states the measured behaviour.
- No CSR endpoint returns the key in any form: list and detail responses
contain neither a PEM nor a Fernet token.
Not changed here, from the issue's "worth folding in" list: the create rate
limit is not a concurrency guard, create_csr holds a pooled connection across
RSA key generation, detail=str(e) echoes internal error text (a repo-wide
convention), and is_global skips cluster validation in both routers/ssl.py and
routers/csr.py. None are storage concerns and each is a separate change.
GoDaddy DNS-01 provider for ACME (v1.10.0).
Validated on a corporate pre-production environment before merge: backend suite 1221 to 1234 passed (+13, exactly the new GoDaddy tests) with 0 failures; a wire-level harness against a fake GoDaddy API confirmed the apex+wildcard pair coexists, removing the last value uses DELETE rather than PUT [], an unreadable read fails closed with no write attempted, and across every scenario not one request reached a zone-wide endpoint (SPF, DKIM and DMARC survived untouched). The path-guard premise was measured directly: with yarl 1.24.5 a '.' segment normalizes onto the zone-wide TXT endpoint and '..' onto the whole-zone endpoint, so the guard in _rrset_path is load-bearing. No schema, environment, frontend or agent change.
Closes#55
README: add GoDaddy to the two feature bullets and to the DNS-01 provider
catalog, spelling out that the API Key must be a Production key (the first key
the developer dashboard issues is an OTE/test key and is rejected), that the
zone must be in the same account, that the account needs a registered domain
before GoDaddy permits DNS API access, and that a Personal Access Token works
with the Secret left blank. Note that publishing is automatic for GoDaddy as
well as Cloudflare, and add the release-notes entry.
UPGRADE_GUIDE: new section stating there is no SCHEMA_VERSION bump, so the
built-in-role re-seed warning from v1.9.0 does not apply, and no new
environment variable, API-shape or agent change. Two limits are stated
explicitly rather than glossed: the credential check is a read, so a token
with read but not write scope saves successfully and only fails at the first
publish; and downgrading after adopting GoDaddy is not a no-op, because an
unknown provider name degrades DNS-01 orders to the manual-confirm path and
leaves published TXT records marked cleaned without being removed.
Twelve tests, no network and no database, in the existing pure-logic style.
The merge helpers are covered directly (additive add, idempotent republish,
tombstone filtering, remove-one-of-many, remove-the-last-value signalling
DELETE), but helper math alone would stay green if the write path stopped
using it, so add_txt_record and remove_txt_record are also driven against a
recording stub: the assertions pin that a sibling value survives a publish,
that an already-published value issues no write, that an unreadable read
raises instead of replacing the set, that removing the last value emits DELETE
and never an empty PUT, and that no call is ever aimed at a zone-wide path.
_request is exercised through a fake response for the cases that only appear
against the real API: an empty 204 body must not raise, a 3xx must not read as
success (redirects are not followed), a transport failure mid-read must not be
mistaken for an empty body, and each error status must produce a message
naming what the operator has to fix.
Also covers the auth header in both forms, that the sanitizer strips
credentials from composed error text, that the module does not log at all, the
credential-field schema against the upsert validator's own key and length
rules, and the two-key encryption round trip.
Verified by mutation: nine deliberate breakages of the provider - single-value
PUT, empty PUT instead of DELETE, coercing an unreadable read to empty,
treating 3xx as success, following redirects, swallowing transport errors,
dropping the dot-segment guard, lowering the TTL below the API floor, and
removing sanitization - are each caught by at least one test.
Registers a third pluggable DNS provider for ACME DNS-01 alongside Manual and
Cloudflare. Credentials are an API Key + Secret pair; leaving the Secret blank
sends the Key as a Personal Access Token (Bearer), which is the migration path
as GoDaddy retires the sso-key scheme.
GoDaddy's Domains API v1 has no per-value TXT write: PUT on a record set
replaces every value at that name. A certificate covering example.com and
*.example.com publishes two different TXT values at the same
_acme-challenge.example.com, so add/remove are read-modify-write - read the
current set, merge, put the whole list back - with empty-data tombstone rows
filtered out (they are rejected on echo) and DELETE used for the last value,
since PUT with an empty array is rejected.
The zone-wide sibling endpoints (.../records/TXT and .../records) would wipe
SPF/DKIM/DMARC and the whole zone respectively, so the record path is built in
one place that refuses an empty or dot segment. An unreadable record-set read
fails closed rather than being treated as an empty set, because the PUT that
follows would otherwise destroy the coexisting values.
Zone lookup walks name suffixes probing the records API rather than the domain
listing, so zones delegated to GoDaddy nameservers resolve and accounts that
are rejected from the domain-details endpoint still work. Credential and
eligibility failures during the walk surface instead of being reported as
"no managed domain".
Provider errors are sanitized at the single point where GoDaddy-supplied text
enters a message, since those strings are persisted to order events and shown
in the UI. No new dependency, no schema change, no frontend change - the
credential form is rendered from the provider schema.
Found during a v1.8.10 to v1.9.0 upgrade drill on a populated database
(schema v9 to v10) before releasing. Documentation only, no code change.
1. The v1.9.0 upgrade notes said "custom roles need no changes", which reads
as "role data is untouched". It is not: because the SCHEMA_VERSION bump
re-runs the whole idempotent sequence, update_system_roles_to_enterprise_rbac()
issues an unconditional UPDATE roles SET ... permissions = <defaults> for
the four BUILT-IN roles. In the drill an `operator` role that had been
narrowed by removing apply.execute and config.bulk_import came back with
both restored (57 to 59 permissions). Operator-created roles are NOT
affected; the re-seed matches the four built-in names only.
This is pre-existing behaviour of every SCHEMA_VERSION bump and is
documented as intentional in migrations.py, so it is not introduced by the
CSR feature. The v1.7.0 upgrade notes carried this caveat and it was not
carried forward. Restored, with an export and re-apply procedure.
Also clarified why the admin password is safe: the default-user seeding is
guarded by an existence check ("safer than ON CONFLICT"), not an upsert,
so an operator-changed password survives.
2. README release notes jumped from v1.8.7 straight to v1.9.0 because
v1.8.8, v1.8.9 and v1.8.10 were never backfilled. Added all three.
CSR creation (v1.9.0): in-app key + CSR generation and signed-certificate import.
Validated on a corporate pre-production environment before merge: backend suite 1145 to 1221 passed (+76), config generator output byte-identical to 1.8.10 against the same database, real haproxy 2.8 -c accepts a CSR-issued certificate, API surface additive only (+5 endpoints), populated v9 to v10 upgrade drill preserved all data, rollback to 1.8.10 starts cleanly, zero diff in the agent scripts.
Closes#49
New CSRManagement component as a third tab (deep-linkable via ?tab=csr):
- Create modal: name (path-traversal-safe client rules mirroring the
server), CN with wildcard support, SAN tag input, key algorithm select,
optional subject fields in a collapse panel. On success the view modal
opens immediately with the CSR PEM.
- View modal: subject/SAN summary, read-only CSR PEM with copy and a
Download .csr button (Blob download).
- Import modal: paste signed certificate + optional chain, usage type,
global/cluster scope with cluster multi-select, and an optional
certificate-name override for collisions that appeared after CSR
creation; SAN-drift warnings surface in a warning dialog.
- Duplicate action pre-fills the create modal (forceRender so the form
accepts values before first open); delete confirm spells out that a
pending CSR key is destroyed permanently.
- SSL certificate table now renders a distinct CSR source tag next to
the existing Manual / Auto (ACME) tags.
New /api/ssl/csrs endpoint group: generate a private key + CSR server-side
(RSA 2048/4096, ECDSA P-256/P-384; full subject + DNS SANs with wildcard
support), list/detail/delete CSRs, and import the CA-signed certificate.
- New ssl_csrs table (SCHEMA_VERSION 9 -> 10, additive + idempotent); the
migration re-raises on failure so a failed run is retried instead of being
stamped as applied.
- Import verifies the certificate against the stored key as a hard gate
(match=None is treated as an integrity error, not a lenient pass), rejects
malformed and expired certificates with 400, warns on SAN drift, and
creates a normal ssl_certificates row (source=csr, cluster_id=NULL,
last_config_status=PENDING) so it flows through the standard
Apply Management -> agent pull pipeline.
- Concurrency: FOR UPDATE row lock serialises double-import and
delete-during-import; a partial unique index reserves pending CSR names;
soft-deleted same-name certs are reactivated preserving the row id.
- Security: no CSR endpoint ever returns the private key (explicit column
lists, enforced by a static test); the key copy on the CSR row is NULLed
after import; ssl.create/read/delete permissions enforced on every
endpoint incl. reads; per-user rate limit on key generation, which runs
in a worker thread; csr_id and cluster_ids are int32-guarded.
- ssl_service: extract _prepare_cert_fields from create_cert_row (behaviour
unchanged, extraction tests untouched) and add stage_ssl_config_versions
reusing the exact ssl-{id}-create-{ts} version-name scheme.
- Tests: crypto round-trip for all four algorithms, model validation,
import-flow unit tests, endpoint auth/permission pinning, migration and
key-non-exposure static assertions.
Single-source version bump (backend/version.json) with frontend/package.json and
package-lock kept in sync (test_version_consistency). Marks the release that ships
the GHSA-7rhv / GHSA-3p5c / GHSA-3vh4 fixes.
Follow-up to the RCE/missing-auth/SSRF remediation, from a thorough multi-lens
review (3 agents + a black-box audit of all 201 routes). Backend-only; no
agent-script changes.
Regression fix (introduced by the previous commit):
- GET /api/agents was made JWT-only, but deployed agents call it WITH X-API-Key
(not a JWT) to read their applied_config_version and avoid re-applying config on
restart. It now accepts EITHER a valid operator JWT OR a valid agent X-API-Key,
so agents no longer get 401 (which caused a spurious HAProxy reload every restart).
Completeness (GHSA-3p5c siblings the first pass missed — same data class, now JWT):
- dashboard.py: GET /api/haproxy-cluster-pools/{id}/agents (full agent inventory —
a direct anonymous bypass of the GET /api/agents lockdown), /api/pools,
/api/haproxy-cluster-pools, /api/dashboard/stats, /api/dashboard/overview
(auth was optional -> leaked stats/names/health/alerts anonymously),
/api/haproxy/stats.
- waf.py: GET /api/waf/rules. health.py: GET /api/health/errors.
- agent.py: GET /api/agents/generate-uninstall-script/{platform} (agent-management
endpoint; was anonymous) now requires JWT or agent key, like generate-install-script.
- config.py: POST /api/config/{validate,optimize,templates/{id}/generate} were
optional-auth (logging only) and run a HAProxy validator on caller input; now
require a JWT. (bulk-create, parse-bulk, diff and configuration/request were
already mandatory-auth — verified.)
All newly-gated endpoints are frontend-only (axios sends the JWT) or unused;
agents never call them.
SSRF (GHSA-3vh4) gap:
- acme_service._get_nonce fetched directory['newNonce'] (from the attacker-
influenceable directory JSON) with a bare session, http allowed, dual-stack, and
BEFORE the guarded _signed_request POST. Now guarded (assert_public_url +
safe_connector + no redirects + timeout), matching the other ACME sinks.
Correctness:
- Three agent webhooks (config-applied, config-validation-failed, config-sync)
swallowed their auth 401 into a 200 error body via a bare `except Exception`.
Added `except HTTPException: raise` so the 401/403 propagates.
Audit result (live black-box, all 201 routes probed unauthenticated): no data
leak and no unauthenticated mutation anywhere; every sensitive route returns
401/403 (a pre-existing group of read handlers wraps the 401 into a 500 via a
broad except — no data is exposed; left as-is, documented as cosmetic).
Verified: full pytest tests/ (1145 passed, 0 failed; +16 regression tests) + live
localtest stack smoke — agent-key GET /api/agents=200, anonymous=401, all newly
gated endpoints reject anonymous and admit JWT, the 3 webhooks return 401.
Addresses three reported advisories, all verified against the code. Fixes are
entirely server-side — deployed agents already send a valid X-API-Key on every
call, so enforcing it does not require any agent-script change or upgrade.
GHSA-7rhv-c5pc-69r8 (CRITICAL RCE — agent script-template poisoning):
- POST/GET /api/agents/script-templates/{platform} now require the agents.version
permission (was authentication-only), matching POST /versions. Blocks a viewer
JWT from overwriting the root install/upgrade script.
GHSA-3p5c-m5m4-mjpx (missing authentication):
- Agent data-plane endpoints now REQUIRE a valid X-API-Key (was optional/skipped
when the header was absent), checked before any DB access: config,
ssl-certificates (private keys!), upgrade-status, heartbeat (by-name and the
previously auth-less by-id), configuration pending-requests. Removes keyless
heartbeat spoofing and keyless rogue-agent auto-registration.
- Operator/UI endpoints now require a JWT: GET /api/agents, the entire
/api/dashboard-stats router, /api/health/{deep,agents,clusters}, and
/api/ssl/certificates/{id}/config-versions. The simple /api/health liveness
probe stays public. Adds shared auth_middleware.require_authenticated_user.
GHSA-3vh4-gvxx-wm2p (SSRF via ACME directory_url):
- New utils/ssrf_guard.py (https-only + public-IP-only, IPv4-pinned, no redirects),
applied to settings test-connection, acme_service.get_directory and
_signed_request, and validated at Let's Encrypt account creation. The
test-connection response no longer reflects arbitrary upstream JSON keys
(information-disclosure oracle) — only fixed ACME field names.
Verified: full pytest tests/ (1128 passed, 0 failed) + live localtest stack smoke
(valid JWT/key paths return 200/404 as expected; anonymous requests 401; SSRF to
metadata/private/loopback refused). No changes to backend/utils/agent_scripts/*.
Follow-up to the consolidated security bump:
- websocket-driver -> 0.7.5 (CRITICAL, message corruption; dev/build tooling)
- resolve-url-loader -> 5.0.0 (pulls postcss ^8), resolving the last postcss<8.5.10
instance (#12) and #2 at the source. Project uses no SASS, so resolve-url-loader
v4->v5 is inert at build time.
Verified: full production docker build succeeds on node 18; postcss now resolves
to a single 8.5.10 across the tree; deferred dev-only deps unchanged.
Infra-only change: image tag bump in docker-compose and k8s manifest.
Backend redis-py client (redis>=5.0.0, resolves to 8.x) verified compatible
against Redis 8.8.0 for all commands in use (get/setex/incr/expire/delete/ping).
No application code or version change.
The manual Frontend editor, wizard and visual ACL builder hard-rejected the ACL
`-f <file>` flag while bulk import accepted it. Worse, a frontend imported with
an `-f` ACL could not be edited at all (422) until the ACL was dropped.
The original guard predated the fail-safe apply flow: the agent runs `haproxy -c`
before every reload, so a missing pattern file is rejected safely and the previous
config keeps running. Pattern files are operator-managed host files — the same
policy adopted for SPOE filter configs in v1.8.8.
- models: remove the 5 `-f` hard rejects (frontend acl/redirect/use_backend
validators + wizard string/dict-redirect guards); `$(`/backtick and X!X
contradiction guards unchanged
- routers/frontend: `_pattern_file_warnings` helper; non-blocking warning on
create + update responses listing referenced pattern files (empty when no
rule uses `-f` — zero noise)
- routers/config: bulk-import preview advisory listing pattern files per
frontend (cluster config-dir aware, next to the SPOE advisories)
- React: remove the FrontendManagement submit gate and SiteWizard step gate;
ACLRuleBuilder renders informational notes instead of errors and re-adds
`-f (pattern file on host)` to the flag dropdown; create path now renders
server warnings like update
- tests: 4 reject-pins inverted to accept-pins; new test_acl_pattern_file_allow.py
(accept/guards-kept/zero-noise/advisory); full suite green (1094 passed)
Bulk import / manual edit silently dropped `filter spoe engine ...` (Coraza WAF)
and frontend `log-format` because the parser recognised only a fixed directive
set. The regenerated config then missed the SPOE engine, so HAProxy failed with
"unable to find SPOE engine 'coraza' used by the send-spoe-group".
- parser: capture `filter` + `log-format`/`log-format-sd` into new ParsedFrontend fields
- db: additive nullable `log_format` + `filters` TEXT columns on frontends (SCHEMA_VERSION 8->9)
- generator: new `filter` bucket flushed before http-request rules so `filter` precedes
`send-spoe-group`; `log-format` kept in prelude
- bulk import: preview dict, change-detection, persist (create + merge-update); cluster-aware
SPOE pre-flight advisories (missing-filter + host-prerequisite) surfaced in the UI
- manual CRUD: full round-trip (get/create/update) incl. React form fields (no null-wipe)
- reject/rollback: restore the new columns; restore path + wizard helper kept in parity
- backend `option spop-check` recognised (suppresses spurious warning for coraza-spoa)
- tests: test_spoe_filter_import.py; full suite green (1079 passed)
The version.json single-source move (v1.8.7) updated the version READ step but
missed the create-release step, which still ran jq against the deleted repo-root
version.json and failed the workflow. Point it at backend/version.json. This
release step is public-only (GitHub Releases), so it has no corporate counterpart.
The version shown in the UI (backend-sourced via /api/version) could lag the
real release. The canonical version lived at repo-root version.json, but the
backend image is built with context ./backend, so that file did not reach the
container unless a pipeline staged it; the backend then fell back to a hardcoded
constant in main.py that had to be bumped by hand and had drifted (it reported
1.8.4 after 1.8.5 and 1.8.6 shipped).
- version.json moves to backend/version.json (single source of truth), now
committed and co-located with main.py, so COPY . . bakes it into every image
directly - correct version in every deployment, no pipeline staging required.
- backend/main.py reads the co-located file; its in-code fallback is no longer a
real version ("unknown") so it can never silently drift again.
- docker-build.yml reads backend/version.json and drops the staging step;
docker-compose.localtest drops the stale root mount.
- backend/tests/test_version_consistency.py fails the build if main.py hardcodes
a real version, if the canonical file is missing/invalid, or if
frontend/package.json drifts from it.
Bumped to 1.8.7. No functional, schema, API, or agent change. Full suite green.
A user running the API on a 2-core/4GB host reported slow-feeling API
responses (Issue #35 follow-up). Review of the hot paths found no
pathological defect; the dominant factors are the single uvicorn worker
(one core serves all requests) and the constant agent-poll baseline
(4 requests per agent every 30s). Two zero-risk improvements:
- backend/Dockerfile: CMD now honors UVICORN_WORKERS, falling back to
WEB_CONCURRENCY and then 1. Flagless uvicorn natively honors
WEB_CONCURRENCY, so the fallback keeps any deployment that relied on it
byte-for-byte compatible; with the final default of 1 worker uvicorn runs
in-process exactly as before. >1 enables the multiprocess supervisor so
multi-core hosts can use all cores. Background tasks are already
multi-replica safe (FOR UPDATE SKIP LOCKED / advisory locks), as exercised
by the k8s HPA deployment (2-10 replicas). `exec` keeps uvicorn as PID 1
(clean SIGTERM, verified ~1s docker stop with 2 workers).
docker-compose.yml passes UVICORN_WORKERS through as empty-when-unset so a
user-set WEB_CONCURRENCY is never overridden; .env.template documents it.
- routers/agent.py heartbeat (by-name endpoint): the agent's
status/version/upgrade_status were read with three separate single-column
SELECTs against the same row; now one SELECT. Identical values and None
semantics (single consistent snapshot instead of three reads); saves two
round-trips per heartbeat per agent every 30s. The legacy by-id heartbeat
endpoint is untouched; the heartbeat API contract is unchanged for agents
of every version.
- README: new "Performance Tuning" section (worker/replica scaling, and how
to use the X-Response-Time header plus "Slow request detected" logs to
pinpoint slow endpoints).
Verification: full backend suite in docker green (1063 passed, 151 skipped;
also re-run by the runtime image build); worker-count expansion matrix
(unset->1, UVICORN_WORKERS=2->2, WEB_CONCURRENCY=3->3, both->UVICORN_WORKERS,
empty->fallback) all correct; default run confirmed single-process with
uvicorn as PID 1 and healthy API; UVICORN_WORKERS=2 confirmed parent + 2
workers, healthy API, clean shutdown; live heartbeats verified for
register + existing-agent paths AND degraded agents (no stats socket /
haproxy stopped / garbage stats CSV / unknown backend in server_statuses):
all return 200, agent row updates correctly, zero backend errors.
No schema, API, or agent changes.
The background order-completion task (complete_pending_acme_orders, 60s
cycle) failed on EVERY cycle since v1.8.0 with:
[ACME-COMPLETE] Error in completion task: syntax error at or near ")"
Root cause: the bounded DNS-01 retry OR-arm added to the atomic order-claim
query in v1.8.0 (13b65d9) carried one extra closing parenthesis, making the
whole SELECT invalid PostgreSQL. The claim is the task's first statement, so
the generic except swallowed it each minute and NO background ACME work ever
ran on v1.8.0-v1.8.4:
- orders were never claimed for finalize -> download -> save (http-01 too);
- advance_dns01_order never ran, so the DNS-01 TXT record was never
published - DNS-01 with an automated provider (e.g. Cloudflare) could
never validate (exactly the report in Issue #35);
- wizard-staged orders never left wizard_staged (same try block);
- retry_invalid_dns01 / reconcile_dns01_cleanup never executed;
- hourly-created renewal orders could never complete in the background.
Fix: drop the stray ')' (one line). Query semantics are unchanged.
Why the suite missed it: the unit tests mock asyncpg, so raw SQL never
reaches a real parser. Added a regression test that AST-scans the ACME
modules' SQL string literals (comments/quoted literals stripped) and fails
on unbalanced parentheses - it is red on the pre-fix tree and would have
caught the v1.8.0 regression at commit time. Scanned all six ACME modules:
this query was the only unbalanced SQL.
Verification: full backend suite in docker green (1062 passed, 151
skipped); the fixed query EXPLAINs cleanly on postgres:15; live localtest
run shows zero completion-task errors and a seeded pending order is claimed
("[ACME-COMPLETE] Claimed 1 order(s)"). Backend-only, no schema/API/agent
changes; fully backward compatible.
Reported-by: @tkkost (GitHub Issue #35)
The Linux/macOS agent installer could abort during "pre-installation cleanup"
(terminal showed `Killing processes matching: haproxy-agent` then `Killed`,
returning to the prompt) when the install script's own command line contained
"haproxy-agent". The cleanup killed processes via `pgrep -f "$pattern"` starting
with the bare string "haproxy-agent", which also matched the running installer's
own command line and a sudo/PAM ancestor that the $$/$PPID self-exclusion did not
cover, so the installer terminated itself before installing.
- linux_install.sh / macos_install.sh: the cleanup kill loop now targets ONLY
the installed agent - "$INSTALL_DIR/haproxy-agent" (the daemon binary path) and
the agent service/label ("haproxy-agent.service" / "com.haproxy.agent") - never
the bare "haproxy-agent" substring. Neither pattern can match the installer's
own command line. The redundant bare pattern is dropped (the service is stopped
separately, and the binary-path pattern still catches a running daemon).
- frontend (AgentManagement.js): name the downloaded scripts
install-agent-<platform>.sh / uninstall-agent-<platform>.sh (matching the
backend's suggested filename) - defense in depth so this cannot resurface.
Installer-only change. The running agent and its privilege model are unchanged
(it runs as root for HAProxy reload, config writes, keepalived, and self-upgrade).
The cleanup runs only on a full interactive install (gated by SKIP_TO_DAEMON), so
daemon mode, self-upgrade, and config/version apply are unaffected. Both agent
scripts kept in sync. Scripts parse on bash 4.2-5.2; full backend suite green.
Addresses #31.
A self-hosted agent could fail every heartbeat with HTTP 400
`Invalid JSON: Expecting property name enclosed in double quotes` when the
system-info block it collects came back empty on an unusual host. The agent
builds the heartbeat JSON as text, so an empty `$system_info` collapsed the
`$system_info,` line to a bare comma and broke the payload.
- Agent script (linux + macos, kept in sync): guard the fragment-form
register_agent and send_heartbeat builders so an empty system_info falls back
to a valid key and can never emit a bare comma. Uses the most portable bash
glob test (no POSIX class / pattern substitution; verified on bash 3.2-5.2 and
on Ubuntu/Debian/Rocky/Alpine/Amazon Linux). True no-op for healthy agents.
- Backend heartbeat endpoint: parse the body as-is first and only run the
malformed-JSON repair when parsing fails, so a valid heartbeat from any agent
version is byte-for-byte untouched. The repair (now a testable helper) recovers
a leading or doubled comma (the empty-system_info artifact) in addition to the
existing empty-value / trailing-comma fixes.
No agent version bump; self-upgrade and daemon mode are unaffected. Healthy
agents of every version behave identically. Full backend suite green.
Addresses #31.
ZeroSSL/Google account registration failed with
`malformed: The Replay Nonce could not be base64url-decoded`: the ACME client
(a process-wide singleton) kept a single anti-replay nonce shared across
certificate authorities, so a nonce issued by one CA could be sent to another,
and the auto-retry only covered `badNonce`.
- Scope the nonce per CA (self._nonce_by_dir keyed by directory_url): a nonce
from one CA is never sent to another; account registration always uses a fresh
nonce from the target CA.
- Broaden the 400 retry to also recover from the nonce-malformed rejection.
- Fix _b64url_decode padding (used for the EAB HMAC key).
Backend-only; HTTP-01 and Let's Encrypt are unaffected.
Addresses #35.
Follow-up fixes for the DNS-01 feature reported on #35:
- Cloudflare: sanitize the API token (strip surrounding quotes + any non
token68 chars) so a pasted token with quotes/spaces no longer fails with
"Invalid request headers"; verify-on-save shows a precise hint when it
cleaned the input. Covers the automated orchestrator path too.
- ZeroSSL/Google EAB: enter the EAB Key ID and HMAC Key per-account in the
Register Account dialog (falls back to the global Settings value when blank);
base64-validate the HMAC key; humanize the externalAccountRequired failure;
and preserve the deliberate 409/422 instead of downgrading them to 400.
- Apply Management: cluster ACME enable/disable changes now show under a
dedicated "ACME Challenge Routing" section, are counted in the Apply/Reject
dialogs, and Apply/Reject All process them (previously "Rejected 0 HA/VIP
change(s)") - consistent with every other entity. Reject rolls acme_enabled
back to the original via ORDER BY created_at ASC over the snapshot chain.
- getErrorMsg surfaces field-level validation messages.
Backward compatible (additive / strict superset; HTTP-01 unchanged).
Addresses #35.
Add ACME DNS-01 (TXT-record) validation alongside the existing HTTP-01,
for internal/isolated clusters with no public port 80 and for wildcard
certificates. Opt-in via a global kill-switch (default off); HTTP-01 is
byte-for-byte unchanged, with zero agent or rendered-config changes.
- Pluggable DNS provider interface (Manual + Cloudflare). Per-account
credentials are Fernet-encrypted at rest, verified on save, and never
returned by the API or written to logs/events/error_detail.
- Non-blocking per-cycle orchestrator: publish (CAS) -> propagation grace
(across cycles, no in-loop sleep) -> respond -> finalize/download, with a
bounded fresh-order retry chain (1 original + 3 retries) on propagation lag.
- Manual flow: user publishes the TXT record and confirms; manual DNS-01
cannot auto-renew unattended (auto-renew forced off and surfaced in the UI).
- Migration v8: additive, idempotent columns on letsencrypt_accounts/orders
and acme_challenges, plus a new letsencrypt_account_dns_credentials table.
- Challenge-type-aware diagnostics (port80/routing/DNS checks skipped for
DNS-01) and a DNS-01 event timeline in the order detail.
- Frontend: DNS-01 account + credentials management, cert wizard adaptation,
order-detail TXT records + verify, orders/renewal Method columns, and a
Settings kill-switch. README, release notes, and API docs updated.
Implements #35.
Two friction points surfaced in issue #31: (1) the default-login note said
admin/admin but the seeded password is admin123; (2) a reporter logged in at
:3000 (the raw static frontend, no /api behind it) instead of :8080 (nginx,
which serves the UI and proxies /api). Fix the credentials note and add a
clear pointer that :8080 is the single entry point, with the Quick Reference
table annotated accordingly.
Patches the critical shell-quote advisory (quote() does not escape
newlines in object .op values, CVSS 8.1). shell-quote is a dev/build-only
transitive dependency (react-dev-utils / launch-editor) and is not present
in the production image or the browser bundle, so there is no runtime
exposure; this clears the alert and the dev-time risk. Lockfile-only change
verified to install with shell-quote resolving to 1.8.4.
The build workflow tagged the Docker images with the product version but
never created the matching git tag, so the repo Tags/Releases drifted
behind (stuck at the last manual tag, v1.6.0) while Docker Hub had 1.7.8.
Add a step that, after the images are pushed, creates a Release (and its
tag) for the current version.json when one does not already exist, and
grant the job contents:write so it can do so.
Manage highly-available virtual IPs backed by Keepalived (VRRP) directly from the OpenManager
UI — no more SSHing into nodes to install/configure Keepalived by hand. Builds on the agent
pull-architecture: define the VIP centrally, click Apply, and the agents converge.
Highlights:
- New "HA / VIP" tab: create a virtual IP, pick a per-node interface, select which pool nodes
participate (MASTER/BACKUP roles + priorities); live MASTER/BACKUP/FAULT per node.
- On Apply, agents install & configure Keepalived (unicast VRRP, cloud-safe default) across the
major distros (Debian/Ubuntu, RHEL/CentOS/Alma/Rocky, Fedora, SUSE/openSUSE, Alpine) with a
HAProxy health-check, so the VIP fails over automatically when HAProxy drops.
- Single-node (a managed floating IP without failover) and multi-node VRRP failover both work.
- VIP changes ride the standard Apply Management flow with the standard "View Change" diff.
- Approval-gated deletion (safety): deleting a running VIP is staged for approval and the VIP
keeps running, untouched, until you approve it — an agent never tears a VIP down without an
explicit human approval. Per-VIP Diagnostics view; opt-in package uninstall (only on nodes
where OpenManager installed it). A node already running a hand-managed Keepalived is detected
and never overwritten ("externally managed").
- Fully opt-in and backward compatible: nodes/clusters without a VIP are unaffected. Adds
vip_instances + vip_members tables (idempotent SCHEMA_VERSION bump; existing data and
passwords unaffected) and a `vip` RBAC permission group.
- Also includes a HAProxy config-generator robustness fix: auto-inject a stick-table when a
frontend uses a stick counter (track-sc / sc_*_rate) but declares none.
On-prem / L2 (VRRP) scope; the UI notes the cloud caveat.
Supersedes the v1.6.4 stable pin (1.30.2-alpine) with the mainline
patched release. No config, schema, or behavior changes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The bundled nginx reverse proxy was flagged for the nginx 'poolslip' advisory
(affected: mainline <=1.31.0; fixed: stable 1.30.2+ / mainline 1.31.1+). The
config-level mitigation (named capture groups instead of $1/$2 in rewrite) does
not apply — the product's nginx config (nginx/nginx.conf and the k8s configmap)
has no rewrite capture-group directives, only prefix locations + proxy_pass. So
the fix is the version: pin nginx:alpine -> nginx:1.30.2-alpine in
docker-compose.yml and k8s/manifests/10-nginx.yaml.
No config/schema/behavior change (nginx only reverse-proxies). Version bumped to
1.6.4 across all layers. Verified in Docker: nginx -v=1.30.2; nginx -t OK on both
the compose and production configmap configs; all proxied routes work through
nginx; no nginx errors.
A backend server toggled OFF (is_active=false) vanished from the UI with no way
to reactivate it: GET /api/backends honored include_inactive for backends but the
server sub-queries hardcoded 'AND is_active = TRUE'.
- get_backends: server sub-queries now honor include_inactive (default callers
unchanged); added last_config_status to the server payload so the UI can tell a
DISABLED server (re-enableable) from a DELETION (pending delete).
- toggle_server: persists an entity snapshot so an Apply-Management Reject rolls
back is_active (previously left the server stuck disabled).
- BackendServers.js: requests include_inactive, shows disabled servers with the
ON/OFF switch + an 'Inactive' tag, hides only DELETION-pending servers, and
keeps soft-deleted BACKENDS hidden (so include_inactive doesn't resurface them).
- Config generation unchanged: disabled servers stay '# DISABLED:' comments and
convert back to live lines when re-enabled.
Startup migration hardening (multi-replica / rolling-deploy safety): create_essential_tables
fails fast on lock contention and retries; run_all_migrations is serialized by a
session advisory lock and gated by a schema_migrations version marker, so an
already-current schema is skipped instead of issuing lock-heavy DDL that a serving
replica's traffic could block at startup. Idempotent and fail-open.
Version reported consistently across all layers (version.json, backend fallback,
frontend package) -> 1.6.3.
Assigning a HAProxy agent failed with '401: Authorization header missing'
on GET /api/clusters. A cluster-read hardening had made GET /api/clusters and
GET /api/clusters/{id} accept only a user JWT in the Authorization header;
agents authenticate with their agent token in the X-API-Key header, so the
token was never read.
Both endpoints now accept either a user JWT (Authorization) or an agent token
(X-API-Key via validate_agent_api_key), mirroring the existing dual-auth on
POST /api/agents/generate-install-script. Anonymous access is still rejected,
so the original hardening is preserved. The auth guard is placed before the
try block so the failure surfaces as a clean 401 (not the 500-wrapped-401 in
the report). Agent install scripts now consistently send the token via
X-API-Key (pre-flight cluster check on linux/macos, and macOS get_cluster_paths
which previously used the wrong Authorization: Bearer header).
Also normalizes the platform in the uninstall-script generator so macOS agents
(which report platform 'darwin') no longer get a 400 from
GET /api/agents/generate-uninstall-script/darwin.
version 1.6.0 -> 1.6.2.