Closes#13, Closes#14.
This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0
shipped the ACME stability & enterprise audit (Issues #10/#11/#12).
v1.5.0 builds on that foundation with two co-equal headline features
plus a 22-round audit campaign hardening the prior configuration
surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0
lands in v1.5.2).
------------------------------------------------------------------
HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13)
------------------------------------------------------------------
A live pre-flight + post-failure diagnostic surface for every ACME
order, reachable from the ACME Automation page. The panel exists
to make ACME failures legible to operators who do NOT have shell
access to the API host.
Endpoints (`backend/routers/acme_diagnostics.py`):
POST /api/letsencrypt/orders/{order_id}/diagnostics
Run the full 5-check suite (DNS / port-80 / routing /
account / agents) and humanize the order's `error_detail`
(>=11 RFC-8555 problem types, backwards compatible with
legacy plain-string failures).
POST /api/letsencrypt/orders/{order_id}/diagnostics/
{check_id}/rerun
Re-run a single check in place — used by the "Re-run"
button on every row of the modal's pre-flight table.
GET /api/letsencrypt/orders/{order_id}/events
Merged event timeline combining the typed
`acme_order_events` rows with correlated
`user_activity_logs` entries (resource_type =
'letsencrypt_order' AND resource_id = order_id). The
diagnostic modal auto-tails this timeline every 5 seconds
while open.
Service-level checks (`backend/services/acme_diagnostics.py`):
* DNS resolution via stdlib socket.gethostbyname_ex through
run_in_executor (intentionally avoiding an aiodns runtime
dep for v1.5.0).
* Port-80 HEAD probe, target locked to the order's domains,
success on HTTP 200 OR 404, warns on egress timeout
(corp egress policies routinely blackhole outbound 80 —
fail-hard would be too noisy).
* SSRF guard: probe refuses non-public IPs and surfaces the
skip in the diagnostic result; IPv4-mapped IPv6 normalisation
closes the `::ffff:169.254.169.254` cloud-metadata vector.
* HAProxy routing presence check: matches the order's
cluster_ids to a port-80 HTTP frontend.
* ACME account validity check against `letsencrypt_accounts`.
* Agent presence check (>=1 active agent in target cluster).
* Every sub-check wrapped in a wall-clock timeout to bound
impact on the API event loop.
RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate
limit on both run and rerun, backed by the (user_id, action,
created_at DESC) composite index.
Frontend (`frontend/src/components/ACMEAutomation.js`):
* "Diagnose" button on every order row + the existing
"stuck order" warning row.
* Modal with two tabs:
- Pre-flight Checks (Antd Table with status pills + Re-run
buttons + humanized error banner)
- Event Log (Antd Timeline with auto-tail polling, scroll-
to-bottom, pause-on-hover)
* Correlation IDs surfaced in error banners and individual
check fail details for backend-log lookup.
------------------------------------------------------------------
HEADLINE FEATURE B — Site Setup Wizard (Issue #14)
------------------------------------------------------------------
A single guided flow that creates a Backend + Servers + HTTP
Frontend (and optional HTTPS Frontend) in one atomic transaction.
Endpoints (`backend/routers/site_wizard.py`):
POST /api/site-wizard/preview — diff-preview the changeset
POST /api/site-wizard/create — atomic execute
POST /api/site-wizard/reject — clean rollback (including
any wizard_staged ACME
orders)
GET /api/site-wizard/drafts — draft persistence
PUT /api/site-wizard/drafts/{id} — save/update
DELETE /api/site-wizard/drafts/{id}
Feature surface:
* One screen captures both backend (mode + servers) AND
frontend (http + optional https + SSL mode) inputs.
* SSL modes: ACME (new order, HTTP-01 only for v1.5.0),
Upload (existing PEM), Existing (link to a stored cert),
or None.
* ACME-staged path: wizard_staged_until watermark on the
`letsencrypt_orders` row defers finalisation until agent
confirmation; per-mode reject cleanly cancels and rolls
back the staged order.
* Live diff preview against the cluster's current generated
config (renderer-evolution noise stripped — track-sc<N>
dedup, per-server cookie strip, defaults-cookie
inheritance, listen-block flattening).
* Draft persistence with PEM stripped at save time (private
keys never round-trip through the drafts table).
* Per-cluster multi-tenancy: drafts and wizard_staged orders
are isolated to the creating user's cluster scope.
Frontend (`frontend/src/components/SiteWizard.js`):
* 4-step Antd Steps flow: Backend → Frontend → SSL → Review.
* Render the live diff preview inline before commit.
* Antd Form-level validation mirrors backend Pydantic
validators (numeric bounds, HAProxy reserved keywords, ALPN
consistency, IPv6 scope-id, domain regex, server name
dedup).
------------------------------------------------------------------
AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1 → #82)
------------------------------------------------------------------
v1.5.0 includes 22 adversarial review passes. Each round produced
its own commit set in the corporate development line; this squash
collapses those into the v1.5.0 release artefact. Highlights:
Round 1-4 Site Wizard core: dry-run parity, single-line
value injection guard, ACL -f pattern-file block,
SSL parity, timeout regex, form-state pin.
Round 5-7 defaults-cookie inheritance, server-named-cookie
guard, fe/be mode mismatch, duplicate server
names, health_check_uri + server_address
validators.
Round 8-10 cookie_name / cookie_options newline-injection
guard, dry-run parity (round 9), TCP-mode HTTP-only
feature blockers.
Round 11 SSL name path traversal + health-check >= 1.
Round 12-13 SSL & ACME deep dive (Bulgu #23-#32).
Round 14 single-line value injection (Bulgu #33).
Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy
reserved keywords, ALPN/TLS consistency,
all-backup, multi-domain & multi-user enterprise
edges, drain/HSTS/post-completion (Bulgu
#34-#53).
Round 18-21 concurrency, agent state, TCP-mode HTTP-only,
list size caps, IPv6 scope-id, preview account
validation, TCP backend + balance uri reject
(Bulgu #54-#61).
Round 22 FE error visibility + 3x stale-data lockouts,
referential integrity + cascade safety,
authentication & authorization, multi-cluster
isolation, apply_pending_changes concurrency,
script injection + bulk import multi-tenancy,
prefix-stripped signature comparison
(Bulgu #62-#82).
------------------------------------------------------------------
NO CORPORATE-SPECIFIC ARTIFACTS
------------------------------------------------------------------
This squash deliberately sanitises corporate hostnames, container
registry references, and TLS secret names into generic
placeholders (`your-registry.example.com/your-org`,
`haproxy-openmanager*.example.com`, `wildcard-tls`,
`taylanbakircioglu/haproxy-openmanager-*`) so the public artefact
contains no internal infrastructure detail. Pilot / development
history that retained those values stays in the corporate fork
and is NOT part of this commit.
- Agent scripts now detect and send ip_address in DAEMON heartbeat (Linux: ip route, macOS: ifconfig)
- Backend validates agent-reported IPs via ipaddress stdlib, COALESCE preserves existing on NULL
- IP/VIP change logging (non-critical, try/except wrapped) for operational visibility
- New source_file_hash column on agent_script_templates for reliable update detection
- Migration changed to ON CONFLICT DO NOTHING to prevent overwriting UI-customized scripts on restart
- GET /versions returns script_update_available flag (disk hash vs DB hash comparison with fallback)
- Frontend Alert banner warns users of new agent script versions and directs to Reset to Defaults
- Reset to Defaults and Popconfirm modals explicitly warn about custom script edit loss
- Full backward compatibility: old agents without ip_address field continue working unchanged
Made-with: Cursor
Root cause: config_status enum was created with only PENDING and APPLIED
values. The REJECTED value was never added due to a silent duplicate_object
exception in create_essential_tables(). This caused SSL certificate listing
to crash with "invalid input value for enum config_status: REJECTED" on
fresh installations.
Also adds scrollable containers to Apply Management page to prevent
Agent Sync Status card from being pushed off-screen.
Closes#7
Made-with: Cursor
Fixes#6
- Fix NameError in soft-deleted certificate reactivation path by
reordering variable extraction before DB operations
- Replace silent empty-array returns with HTTP 500 on SQL errors,
making failures visible in both API responses and server logs
- Add primary_domain migration for schema consistency across fresh
and upgraded installations (backfill from legacy domain column)
- Use primary_domain in non-cluster SSL query branch for schema
compatibility
- Surface SSL fetch errors in frontend via toast notifications
- Harden connection cleanup in error handlers with try/except
- Add QEMU + Buildx for linux/amd64,linux/arm64 multi-platform
Docker image builds
- Update GitHub Actions to latest versions (checkout v4, login v3,
build-push v6)
Made-with: Cursor
- Add keepalive_state and keepalive_ip columns to agents table (migration + schema)
- Add keepalive fields to AgentHeartbeat Pydantic model (backward compatible)
- Update heartbeat endpoint to persist keepalive data to DB and cache in Redis
- Add multi-method keepalived detection in agent scripts (journalctl, log files, VIP check)
- Update dashboard-stats agents/status API with Redis-first keepalive lookup
- Update GET /api/agents to include keepalive_state and keepalive_ip
- Show MASTER/BACKUP tag in Dashboard AgentStatusCard
- Show keepalive info in Agent Management registered agents table
- Add Keepalive column to Cluster Management table with VIP search support
Co-authored-by: Cursor <cursoragent@cursor.com>
Implemented comprehensive HAProxy options field support for both backend and frontend entities to enable standard HAProxy directives like 'option http-keep-alive', 'option httplog', 'option forwardfor', etc.
Changes:
- Database: Added 'options' TEXT column to backends and frontends tables
- Models: Added options field to BackendConfig, BackendConfigUpdate, and FrontendConfig
- API Endpoints: Updated CREATE, UPDATE, and GET endpoints to handle options field
* Backend: CREATE/UPDATE/GET with options support
* Frontend: CREATE/UPDATE/GET with options support (fixed 5 SELECT queries)
- Config Generator: Added options block generation for both backends and frontends
- Bulk Import Parser:
* Added options field to ParsedBackend and ParsedFrontend dataclasses
* Implemented option directive parsing with validation
* Added unknown option warnings
* Fixed bulk parse response to include options field
- Bulk Import Merge: Added options field comparison in UPDATE logic
- UI Components:
* BackendServers.js: Added options TextArea form field
* FrontendManagement.js: Added options TextArea form field
Features:
- Multi-line options support (newline-separated format)
- Option validation with known HAProxy options list
- Backward compatible (NULL options for existing entities)
- Bulk import support with merge strategy
- Full CRUD support for both manual and bulk operations
Technical Details:
- Format: Newline-separated TEXT field for multiple options
- Validation: Warns about unknown options but allows them
- Config Generation: Each option written as separate directive
- Agent: Standard HAProxy config validation applies
Total: 10 files modified, ~195 lines added, 26 integration points verified
This is a comprehensive update that adds SSL certificate differentiation
for frontend (HAProxy bind) and server (backend verification) use cases.
FEATURES:
- SSL certificates can be marked as 'frontend' or 'server' usage type
- Frontend SSL: Private key REQUIRED (for HAProxy bind ssl crt)
- Server SSL: Private key OPTIONAL (CA cert only for backend verification)
- UI dropdown for usage type selection
- Dynamic form validation based on usage type
- Filtering: Frontends see only Frontend SSL, Backends see only Server SSL
DATABASE:
- Added usage_type column to ssl_certificates (default: 'frontend')
- Made private_key_content nullable for server SSL support
- Migration automatically runs on pod restart
BACKEND:
- Pydantic v2 compatibility (@field_validator, @model_validator)
- SSL router: usage_type filtering support
- Agent endpoint: usage_type field included
- Improved migration robustness with better error handling
- Fixed duplicate ensure_agents_table() function
- Fixed JSONB permissions insert with json.dumps()
- Fixed ON CONFLICT constraints with explicit checks
FRONTEND:
- SSL Management: Usage Type dropdown with visual feedback
- Frontend Management: Filters only Frontend SSL certificates
- Backend Servers: Filters only Server SSL certificates
- Dynamic private key validation (required for Frontend, optional for Server)
- Improved form UX with color-coded hints
AGENT SCRIPTS (Linux & macOS):
- Support for Server SSL without private key
- Conditional PEM file creation (cert+key vs cert-only)
- usage_type awareness in SSL deployment
- Backward compatible with existing Frontend SSL certificates
DOCKER:
- Increased npm timeout for slow networks (300s → 600s)
- Increased fetch-retries (5 → 10)
- Reduced maxsockets for stability (3 → 1)
All changes are backward compatible. Existing SSL certificates
default to 'frontend' type and continue working unchanged.
Tested with: HAProxy 2.8+, PostgreSQL 15, React 18
Database Migration:
- Added ssl_certificate_id column to backend_servers table
- Added FK constraint to ssl_certificates table
- ON DELETE SET NULL behavior
- Idempotent migration (safe to run multiple times)
Column Details:
Name: ssl_certificate_id
Type: INTEGER
Nullable: YES
Foreign Key: ssl_certificates(id)
On Delete: SET NULL
Migration Function:
add_ssl_certificate_id_to_backend_servers()
Called in run_migrations() at Line 1523
Code Cleanup:
- Removed emojis from migration logs
- Removed emojis from SSL dropdown status icons
- Changed to text: Valid, Expiring, Expired
- Changed to text: Global, Cluster
Error Fixed:
GET /api/backends - 500
column "ssl_certificate_id" does not exist
After migration runs on startup, column will exist and API will work
- Add cluster_id column to agent_config_requests table
- Update config request endpoint to accept and validate cluster_id
- Frontend now sends cluster_id with config requests
- Fixes issue where agents in pools with multiple clusters get wrong config requests
Technical Details:
- Database migration adds cluster_id as nullable foreign key for backward compatibility
- Backend validates that agent belongs to requested cluster via pool_id check
- Improved logging includes cluster name for better traceability
- No impact on existing features (apply management, sync status, entity CRUD)