Bulk import / manual edit silently dropped `filter spoe engine ...` (Coraza WAF)
and frontend `log-format` because the parser recognised only a fixed directive
set. The regenerated config then missed the SPOE engine, so HAProxy failed with
"unable to find SPOE engine 'coraza' used by the send-spoe-group".
- parser: capture `filter` + `log-format`/`log-format-sd` into new ParsedFrontend fields
- db: additive nullable `log_format` + `filters` TEXT columns on frontends (SCHEMA_VERSION 8->9)
- generator: new `filter` bucket flushed before http-request rules so `filter` precedes
`send-spoe-group`; `log-format` kept in prelude
- bulk import: preview dict, change-detection, persist (create + merge-update); cluster-aware
SPOE pre-flight advisories (missing-filter + host-prerequisite) surfaced in the UI
- manual CRUD: full round-trip (get/create/update) incl. React form fields (no null-wipe)
- reject/rollback: restore the new columns; restore path + wizard helper kept in parity
- backend `option spop-check` recognised (suppresses spurious warning for coraza-spoa)
- tests: test_spoe_filter_import.py; full suite green (1079 passed)
Closes#13, Closes#14.
This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0
shipped the ACME stability & enterprise audit (Issues #10/#11/#12).
v1.5.0 builds on that foundation with two co-equal headline features
plus a 22-round audit campaign hardening the prior configuration
surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0
lands in v1.5.2).
------------------------------------------------------------------
HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13)
------------------------------------------------------------------
A live pre-flight + post-failure diagnostic surface for every ACME
order, reachable from the ACME Automation page. The panel exists
to make ACME failures legible to operators who do NOT have shell
access to the API host.
Endpoints (`backend/routers/acme_diagnostics.py`):
POST /api/letsencrypt/orders/{order_id}/diagnostics
Run the full 5-check suite (DNS / port-80 / routing /
account / agents) and humanize the order's `error_detail`
(>=11 RFC-8555 problem types, backwards compatible with
legacy plain-string failures).
POST /api/letsencrypt/orders/{order_id}/diagnostics/
{check_id}/rerun
Re-run a single check in place — used by the "Re-run"
button on every row of the modal's pre-flight table.
GET /api/letsencrypt/orders/{order_id}/events
Merged event timeline combining the typed
`acme_order_events` rows with correlated
`user_activity_logs` entries (resource_type =
'letsencrypt_order' AND resource_id = order_id). The
diagnostic modal auto-tails this timeline every 5 seconds
while open.
Service-level checks (`backend/services/acme_diagnostics.py`):
* DNS resolution via stdlib socket.gethostbyname_ex through
run_in_executor (intentionally avoiding an aiodns runtime
dep for v1.5.0).
* Port-80 HEAD probe, target locked to the order's domains,
success on HTTP 200 OR 404, warns on egress timeout
(corp egress policies routinely blackhole outbound 80 —
fail-hard would be too noisy).
* SSRF guard: probe refuses non-public IPs and surfaces the
skip in the diagnostic result; IPv4-mapped IPv6 normalisation
closes the `::ffff:169.254.169.254` cloud-metadata vector.
* HAProxy routing presence check: matches the order's
cluster_ids to a port-80 HTTP frontend.
* ACME account validity check against `letsencrypt_accounts`.
* Agent presence check (>=1 active agent in target cluster).
* Every sub-check wrapped in a wall-clock timeout to bound
impact on the API event loop.
RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate
limit on both run and rerun, backed by the (user_id, action,
created_at DESC) composite index.
Frontend (`frontend/src/components/ACMEAutomation.js`):
* "Diagnose" button on every order row + the existing
"stuck order" warning row.
* Modal with two tabs:
- Pre-flight Checks (Antd Table with status pills + Re-run
buttons + humanized error banner)
- Event Log (Antd Timeline with auto-tail polling, scroll-
to-bottom, pause-on-hover)
* Correlation IDs surfaced in error banners and individual
check fail details for backend-log lookup.
------------------------------------------------------------------
HEADLINE FEATURE B — Site Setup Wizard (Issue #14)
------------------------------------------------------------------
A single guided flow that creates a Backend + Servers + HTTP
Frontend (and optional HTTPS Frontend) in one atomic transaction.
Endpoints (`backend/routers/site_wizard.py`):
POST /api/site-wizard/preview — diff-preview the changeset
POST /api/site-wizard/create — atomic execute
POST /api/site-wizard/reject — clean rollback (including
any wizard_staged ACME
orders)
GET /api/site-wizard/drafts — draft persistence
PUT /api/site-wizard/drafts/{id} — save/update
DELETE /api/site-wizard/drafts/{id}
Feature surface:
* One screen captures both backend (mode + servers) AND
frontend (http + optional https + SSL mode) inputs.
* SSL modes: ACME (new order, HTTP-01 only for v1.5.0),
Upload (existing PEM), Existing (link to a stored cert),
or None.
* ACME-staged path: wizard_staged_until watermark on the
`letsencrypt_orders` row defers finalisation until agent
confirmation; per-mode reject cleanly cancels and rolls
back the staged order.
* Live diff preview against the cluster's current generated
config (renderer-evolution noise stripped — track-sc<N>
dedup, per-server cookie strip, defaults-cookie
inheritance, listen-block flattening).
* Draft persistence with PEM stripped at save time (private
keys never round-trip through the drafts table).
* Per-cluster multi-tenancy: drafts and wizard_staged orders
are isolated to the creating user's cluster scope.
Frontend (`frontend/src/components/SiteWizard.js`):
* 4-step Antd Steps flow: Backend → Frontend → SSL → Review.
* Render the live diff preview inline before commit.
* Antd Form-level validation mirrors backend Pydantic
validators (numeric bounds, HAProxy reserved keywords, ALPN
consistency, IPv6 scope-id, domain regex, server name
dedup).
------------------------------------------------------------------
AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1 → #82)
------------------------------------------------------------------
v1.5.0 includes 22 adversarial review passes. Each round produced
its own commit set in the corporate development line; this squash
collapses those into the v1.5.0 release artefact. Highlights:
Round 1-4 Site Wizard core: dry-run parity, single-line
value injection guard, ACL -f pattern-file block,
SSL parity, timeout regex, form-state pin.
Round 5-7 defaults-cookie inheritance, server-named-cookie
guard, fe/be mode mismatch, duplicate server
names, health_check_uri + server_address
validators.
Round 8-10 cookie_name / cookie_options newline-injection
guard, dry-run parity (round 9), TCP-mode HTTP-only
feature blockers.
Round 11 SSL name path traversal + health-check >= 1.
Round 12-13 SSL & ACME deep dive (Bulgu #23-#32).
Round 14 single-line value injection (Bulgu #33).
Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy
reserved keywords, ALPN/TLS consistency,
all-backup, multi-domain & multi-user enterprise
edges, drain/HSTS/post-completion (Bulgu
#34-#53).
Round 18-21 concurrency, agent state, TCP-mode HTTP-only,
list size caps, IPv6 scope-id, preview account
validation, TCP backend + balance uri reject
(Bulgu #54-#61).
Round 22 FE error visibility + 3x stale-data lockouts,
referential integrity + cascade safety,
authentication & authorization, multi-cluster
isolation, apply_pending_changes concurrency,
script injection + bulk import multi-tenancy,
prefix-stripped signature comparison
(Bulgu #62-#82).
------------------------------------------------------------------
NO CORPORATE-SPECIFIC ARTIFACTS
------------------------------------------------------------------
This squash deliberately sanitises corporate hostnames, container
registry references, and TLS secret names into generic
placeholders (`your-registry.example.com/your-org`,
`haproxy-openmanager*.example.com`, `wildcard-tls`,
`taylanbakircioglu/haproxy-openmanager-*`) so the public artefact
contains no internal infrastructure detail. Pilot / development
history that retained those values stays in the corporate fork
and is NOT part of this commit.
- Step 3 (Enable ACME on Cluster) now shows a process icon instead of
a misleading green checkmark when ACME is enabled but not yet applied.
Per-cluster "(pending apply)" annotation for multi-cluster setups.
- Step 4 button and all /apply-management navigation buttons now say
"Apply Changes" instead of "Configure" for clearer guidance.
- Setup Guide auto-selects the correct cluster before navigating to
Apply Management, showing pending cluster names in alerts.
- Pending ACME disable changes are now correctly detected in Step 4
even when acme_enabled is already FALSE in the database.
- Entity snapshot rollback for cluster ACME settings: reject correctly
restores acme_enabled/acme_backend_url to pre-change values.
- Deduplication logic prevents "last wins" bug when multiple ACME
toggles are rejected in sequence.
- Connection leak prevention with try/finally around conn2 in ACME
config version creation.
- Step 4 branching uses boolean has_enabled instead of fragile string
truthiness check.
Made-with: Cursor
Root cause: SSL update passes expiry_date as Python datetime object in
new_values. json.dumps() fails on datetime, causing save_entity_snapshot
to return {} (empty). Entity snapshot is never saved in config_version
metadata, so reject/rollback can never find it to restore old values.
Fix: Apply same JSON serialization to new_values as old_values. Only
affects SSL certificates (only entity with datetime in new_values).
All other entity types (frontend, backend, server, waf) are unaffected
as their new_values contain only JSON-serializable types.
Co-authored-by: Cursor <cursoragent@cursor.com>
Problem 1: Apply affects all entities (should only affect changed ones)
Problem 2: Reject rollback not working (entity stays at new value)
Added detailed logging:
- Snapshot creation: JSON test result, field count
- Frontend update: metadata keys, entity_snapshot presence
- Reject: metadata parsing, entity_snapshot detection
- Rollback: entity data, operation type, old_values
- _rollback_update: Before/after values, UPDATE query result
- Verify: Post-rollback database state
This will help identify:
- Is snapshot being created?
- Is metadata being saved to database?
- Is metadata being parsed during reject?
- Is rollback function being called?
- Is UPDATE query executing?
- What are the actual values being restored?
Log locations to check:
kubectl logs deployment/haproxy-openmanager-backend -n haproxy-openmanager | grep 'SNAPSHOT\|ROLLBACK\|REJECT'
Problem: metadata still null, datetime conversion issue
Root cause: asyncpg returns datetime objects that don't serialize properly with isoformat()
Solution: Test each field with json.dumps(), convert non-serializable to str()
Approach:
- Try json.dumps() for each value
- If serializable: use as-is (int, str, bool, list, dict)
- If not serializable: convert to str()
- datetime: use str() (simpler, safer)
- No timezone manipulation (pod is UTC, keep it simple)
This ensures:
- All fields are JSON-safe
- No exceptions during metadata creation
- metadata will be populated (not null)
- Rollback will work
Problem: Frontend update was falling back to old behavior (status=APPLIED)
Cause: old_values contained datetime fields (created_at, updated_at) which are not JSON serializable
Solution: Convert datetime to ISO string before storing in metadata
Changed:
- Convert datetime -> isoformat() + 'Z'
- Keep JSONB/list as-is (already serializable)
- Handle None values
- Ensure all old_values are JSON-safe
This fixes:
- Config version INSERT failure (exception in try block)
- Fallback to old behavior (APPLIED instead of PENDING)
- metadata serialization error
- Entity update now creates PENDING version with snapshot
Changed ENTITY_SNAPSHOT_ENABLED default from false to true.
Reasoning:
- Code is tested and deployed to production
- Backward compatibility verified
- No need for gradual rollout with feature flag
- Entity rollback should work by default
- Users expect reject to rollback entities (not just status change)
Feature flag still exists for emergency disable if needed:
- Set ENTITY_SNAPSHOT_ENABLED=false to disable
- Useful for troubleshooting or rollback scenarios
Default behavior (ENTITY_SNAPSHOT_ENABLED=true):
- Entity update creates snapshot in metadata
- Reject operation rolls back entities to old values
- Bulk import reject deletes new entities, restores updated ones
- Restore reject returns to pre-restore state