mirror of
https://github.com/taylanbakircioglu/haproxy-openmanager.git
synced 2026-09-16 23:55:13 +00:00
main
27 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
47cc79dcf7 |
fix(acme): address review findings on the challenge-backend hardening
Five confirmed findings from an adversarial review of the branch, four of them regressions introduced by it. Fall through to the next source when a stored URL cannot be resolved. Stopping at the first non-empty candidate emitted a backend section with no `server` line: the section exists so `haproxy -c` passes and Apply succeeds, then every challenge request 503s from an empty backend with nothing to show for it. Scheme-less values are common — the settings field was free text until this branch — so this was reachable on real installs. Selection moved into `select_acme_backend_source()` so it is testable and the skipped candidates are logged rather than silently dropped. Report a challenge backend with no server line. `extract_acme_backend_target` returns None for that section, and the loopback filter skipped falsy targets, so the case above would have been reported as "challenge route present in applied config" — the new check confirming the very state it exists to catch. Do not narrow the row set feeding the routing check's `fail` branch. Adding a mode filter to the WHERE clause turned a tcp-only port-80 cluster from "ok" into "fail", and the site wizard blocks submit on any failing check, so those installs would have been locked on upgrade day. Mode is now examined in Python and only downgrades to `warn`, using an expression that is character-for- character the renderer's normalisation. Match the agent's config selector. The applied-config lookup omitted `is_active = TRUE`, so it could read a superseded row and report on a config the nodes never received. Extraction now happens in SQL rather than pulling whole configs — these run to hundreds of KB. Select `acme_backend_url` when loading the existing cluster. It was absent, so the entity snapshot recorded old_values as NULL unconditionally and rejecting the pending version wiped the operator's per-cluster URL back to the global loopback default — re-creating the exact failure this branch removes. Also carry `acme_enabled` and `acme_backend_url` through cluster creation. The create model declared neither and the INSERT wrote neither, so a cluster created with ACME switched on came back switched off with no error shown. Refuted and deliberately not changed: settings PUT re-validating a stored loopback value (it validates only what is submitted), an apply-path connection leak (the 422 propagates to a handler that closes it), and the modal discarding backend rejection reasons (the envelope matches). |
||
|
|
bb774141d4 |
fix(acme): make the challenge backend fixable from the panel
Correcting a wrong ACME challenge backend was impossible without a shell, and
even with one the correction did not reach the nodes.
The mint gate only fired when `acme_enabled` flipped. `acme_backend_url` was
written to the DB and minted nothing, so Apply answered "No pending changes to
apply" and the nodes kept the old address forever. It is now decided by
comparing the rendered `server _acme_mgmt` line against the active version —
the one line that answers "would the nodes talk to a different address?".
Comparing whole configs would flag every unrelated pending edit.
The field had no UI at all. Added to the cluster form with validation that
mirrors the backend rules, and keyed on `model_fields_set` so clearing it
reverts to the global setting — with a plain `is not None` test an empty box is
indistinguishable from "not submitted", so a value could never be removed.
Validation is asymmetric on purpose (utils/acme_backend_url):
- at the write boundary, reject what cannot express a reachable target —
including the two silent traps: a scheme-less value became `localhost`, and
an out-of-range port raised inside the generator and destroyed the config
- at render time, never reject. The shipped defaults are themselves loopback,
so refusing to render would make every acme_enabled cluster unappliable,
including for changes unrelated to ACME. Problems are logged and surfaced.
The port-less default stays 8080 rather than moving to HTTP's 80: the bundled
compose publishes nginx on 8080, so installs relying on it work today and the
first sign of breaking them would be the unattended renewal loop months later.
The omission is warned about instead.
RFC1918 is allowed and is usually the right answer here, and no DNS resolution
is performed — both deliberate departures from utils/ssrf_guard, whose policy
is the opposite of what this address needs. What the management host can
resolve says nothing about what the HAProxy node can reach.
Diagnostics stop reporting success on a dead path:
- check_port80 uses GET instead of HEAD and classifies the body. A proxy that
has lost its /.well-known/acme-challenge/ location serves its SPA with HTTP
200, which `status in (200, 404)` accepted as healthy. Warnings also surface
when other domains pass, which previously hid the most diagnostic outcome.
- check_routing filters `mode`, joins `acme_enabled` and reads the APPLIED
config instead of counting database rows, and reports a loopback target.
- every new condition is `warn`, never `fail`: the site wizard blocks submit on
any fail, so a new failing condition would lock every install on upgrade day.
Also: normalise `frontends.mode` once per frontend. It is nullable, and the
raw value was interpolated into `mode {}`, emitting a literal `mode None` that
HAProxy rejects — taking down the whole cluster config. The ACME gate and the
backend-mode check now read the same normalised value.
And stop hardcoding PUBLIC_URL / MANAGEMENT_BASE_URL in docker-compose, which
silently ignored the operator's .env and made the wrong default load-bearing.
|
||
|
|
acfd32dd63 |
fix(acme): stop shipping config-generation failures as applied config
`generate_haproxy_config_for_cluster` reports failure by RETURNING a one-line
comment ("# Error generating configuration: ...") rather than raising. Nothing
in the backend checked for it, so the apply path hashed that comment, stored it
as an APPLIED config version and pushed it to every agent — silently replacing
a cluster's entire haproxy.cfg.
Any exception inside the generator triggers this. An out-of-range port in
`acme_backend_url` is enough: `urlparse('http://h:99999').port` raises
ValueError, the outer `except` swallows it, and the cluster loses its config.
- add `is_config_generation_error()` next to the sentinel definitions so call
sites stop matching the string by hand
- guard the apply path: refuse with 422 and leave the running config in force
- guard the ACME-toggle PENDING mint, and let HTTPException through the local
`except Exception`, which would otherwise report success while no pending
version exists
Also make the challenge backend diagnosable without shell access, since this
block previously emitted no log line at all:
- log the rendered `host:port` and WHICH source chose it (cluster override,
system setting, or the MANAGEMENT_BASE_URL fallback) under the greppable
`ACME-BACKEND` keyword
- warn when the rendered address is loopback: HAProxy resolves it on the node,
not on the management host, so it can only work on an all-in-one install —
and it is exactly what the shipped defaults produce
- log peer, X-Forwarded-For and Host on the challenge endpoint, which answers
"did the request arrive at all?" — the question that separates a wrong
backend address from a wrong response body
Refs the HTTP-01 investigation: a split deployment rendered
`server _acme_mgmt <mgmt>:8080` against a port with no listener, and every
existing check reported success.
|
||
|
|
9e2ea04777 |
feat(haproxy): preserve SPOE filter + frontend log-format on import/edit (v1.8.8, Issue #38)
Bulk import / manual edit silently dropped `filter spoe engine ...` (Coraza WAF) and frontend `log-format` because the parser recognised only a fixed directive set. The regenerated config then missed the SPOE engine, so HAProxy failed with "unable to find SPOE engine 'coraza' used by the send-spoe-group". - parser: capture `filter` + `log-format`/`log-format-sd` into new ParsedFrontend fields - db: additive nullable `log_format` + `filters` TEXT columns on frontends (SCHEMA_VERSION 8->9) - generator: new `filter` bucket flushed before http-request rules so `filter` precedes `send-spoe-group`; `log-format` kept in prelude - bulk import: preview dict, change-detection, persist (create + merge-update); cluster-aware SPOE pre-flight advisories (missing-filter + host-prerequisite) surfaced in the UI - manual CRUD: full round-trip (get/create/update) incl. React form fields (no null-wipe) - reject/rollback: restore the new columns; restore path + wizard helper kept in parity - backend `option spop-check` recognised (suppresses spurious warning for coraza-spoa) - tests: test_spoe_filter_import.py; full suite green (1079 passed) |
||
|
|
a1192e602d |
feat: HA / VIP (Keepalived) management from the UI (#27)
Manage highly-available virtual IPs backed by Keepalived (VRRP) directly from the OpenManager
UI — no more SSHing into nodes to install/configure Keepalived by hand. Builds on the agent
pull-architecture: define the VIP centrally, click Apply, and the agents converge.
Highlights:
- New "HA / VIP" tab: create a virtual IP, pick a per-node interface, select which pool nodes
participate (MASTER/BACKUP roles + priorities); live MASTER/BACKUP/FAULT per node.
- On Apply, agents install & configure Keepalived (unicast VRRP, cloud-safe default) across the
major distros (Debian/Ubuntu, RHEL/CentOS/Alma/Rocky, Fedora, SUSE/openSUSE, Alpine) with a
HAProxy health-check, so the VIP fails over automatically when HAProxy drops.
- Single-node (a managed floating IP without failover) and multi-node VRRP failover both work.
- VIP changes ride the standard Apply Management flow with the standard "View Change" diff.
- Approval-gated deletion (safety): deleting a running VIP is staged for approval and the VIP
keeps running, untouched, until you approve it — an agent never tears a VIP down without an
explicit human approval. Per-VIP Diagnostics view; opt-in package uninstall (only on nodes
where OpenManager installed it). A node already running a hand-managed Keepalived is detected
and never overwritten ("externally managed").
- Fully opt-in and backward compatible: nodes/clusters without a VIP are unaffected. Adds
vip_instances + vip_members tables (idempotent SCHEMA_VERSION bump; existing data and
passwords unaffected) and a `vip` RBAC permission group.
- Also includes a HAProxy config-generator robustness fix: auto-inject a stick-table when a
frontend uses a stick counter (track-sc / sc_*_rate) but declares none.
On-prem / L2 (VRRP) scope; the UI notes the cloud caveat.
|
||
|
|
02b1cb2bca |
feat: v1.5.0 — Site Wizard (Issue #14) + ACME Diagnostic Panel (Issue #13)
Closes #13, Closes #14. This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0 shipped the ACME stability & enterprise audit (Issues #10/#11/#12). v1.5.0 builds on that foundation with two co-equal headline features plus a 22-round audit campaign hardening the prior configuration surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0 lands in v1.5.2). ------------------------------------------------------------------ HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13) ------------------------------------------------------------------ A live pre-flight + post-failure diagnostic surface for every ACME order, reachable from the ACME Automation page. The panel exists to make ACME failures legible to operators who do NOT have shell access to the API host. Endpoints (`backend/routers/acme_diagnostics.py`): POST /api/letsencrypt/orders/{order_id}/diagnostics Run the full 5-check suite (DNS / port-80 / routing / account / agents) and humanize the order's `error_detail` (>=11 RFC-8555 problem types, backwards compatible with legacy plain-string failures). POST /api/letsencrypt/orders/{order_id}/diagnostics/ {check_id}/rerun Re-run a single check in place — used by the "Re-run" button on every row of the modal's pre-flight table. GET /api/letsencrypt/orders/{order_id}/events Merged event timeline combining the typed `acme_order_events` rows with correlated `user_activity_logs` entries (resource_type = 'letsencrypt_order' AND resource_id = order_id). The diagnostic modal auto-tails this timeline every 5 seconds while open. Service-level checks (`backend/services/acme_diagnostics.py`): * DNS resolution via stdlib socket.gethostbyname_ex through run_in_executor (intentionally avoiding an aiodns runtime dep for v1.5.0). * Port-80 HEAD probe, target locked to the order's domains, success on HTTP 200 OR 404, warns on egress timeout (corp egress policies routinely blackhole outbound 80 — fail-hard would be too noisy). * SSRF guard: probe refuses non-public IPs and surfaces the skip in the diagnostic result; IPv4-mapped IPv6 normalisation closes the `::ffff:169.254.169.254` cloud-metadata vector. * HAProxy routing presence check: matches the order's cluster_ids to a port-80 HTTP frontend. * ACME account validity check against `letsencrypt_accounts`. * Agent presence check (>=1 active agent in target cluster). * Every sub-check wrapped in a wall-clock timeout to bound impact on the API event loop. RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate limit on both run and rerun, backed by the (user_id, action, created_at DESC) composite index. Frontend (`frontend/src/components/ACMEAutomation.js`): * "Diagnose" button on every order row + the existing "stuck order" warning row. * Modal with two tabs: - Pre-flight Checks (Antd Table with status pills + Re-run buttons + humanized error banner) - Event Log (Antd Timeline with auto-tail polling, scroll- to-bottom, pause-on-hover) * Correlation IDs surfaced in error banners and individual check fail details for backend-log lookup. ------------------------------------------------------------------ HEADLINE FEATURE B — Site Setup Wizard (Issue #14) ------------------------------------------------------------------ A single guided flow that creates a Backend + Servers + HTTP Frontend (and optional HTTPS Frontend) in one atomic transaction. Endpoints (`backend/routers/site_wizard.py`): POST /api/site-wizard/preview — diff-preview the changeset POST /api/site-wizard/create — atomic execute POST /api/site-wizard/reject — clean rollback (including any wizard_staged ACME orders) GET /api/site-wizard/drafts — draft persistence PUT /api/site-wizard/drafts/{id} — save/update DELETE /api/site-wizard/drafts/{id} Feature surface: * One screen captures both backend (mode + servers) AND frontend (http + optional https + SSL mode) inputs. * SSL modes: ACME (new order, HTTP-01 only for v1.5.0), Upload (existing PEM), Existing (link to a stored cert), or None. * ACME-staged path: wizard_staged_until watermark on the `letsencrypt_orders` row defers finalisation until agent confirmation; per-mode reject cleanly cancels and rolls back the staged order. * Live diff preview against the cluster's current generated config (renderer-evolution noise stripped — track-sc<N> dedup, per-server cookie strip, defaults-cookie inheritance, listen-block flattening). * Draft persistence with PEM stripped at save time (private keys never round-trip through the drafts table). * Per-cluster multi-tenancy: drafts and wizard_staged orders are isolated to the creating user's cluster scope. Frontend (`frontend/src/components/SiteWizard.js`): * 4-step Antd Steps flow: Backend → Frontend → SSL → Review. * Render the live diff preview inline before commit. * Antd Form-level validation mirrors backend Pydantic validators (numeric bounds, HAProxy reserved keywords, ALPN consistency, IPv6 scope-id, domain regex, server name dedup). ------------------------------------------------------------------ AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1 → #82) ------------------------------------------------------------------ v1.5.0 includes 22 adversarial review passes. Each round produced its own commit set in the corporate development line; this squash collapses those into the v1.5.0 release artefact. Highlights: Round 1-4 Site Wizard core: dry-run parity, single-line value injection guard, ACL -f pattern-file block, SSL parity, timeout regex, form-state pin. Round 5-7 defaults-cookie inheritance, server-named-cookie guard, fe/be mode mismatch, duplicate server names, health_check_uri + server_address validators. Round 8-10 cookie_name / cookie_options newline-injection guard, dry-run parity (round 9), TCP-mode HTTP-only feature blockers. Round 11 SSL name path traversal + health-check >= 1. Round 12-13 SSL & ACME deep dive (Bulgu #23-#32). Round 14 single-line value injection (Bulgu #33). Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy reserved keywords, ALPN/TLS consistency, all-backup, multi-domain & multi-user enterprise edges, drain/HSTS/post-completion (Bulgu #34-#53). Round 18-21 concurrency, agent state, TCP-mode HTTP-only, list size caps, IPv6 scope-id, preview account validation, TCP backend + balance uri reject (Bulgu #54-#61). Round 22 FE error visibility + 3x stale-data lockouts, referential integrity + cascade safety, authentication & authorization, multi-cluster isolation, apply_pending_changes concurrency, script injection + bulk import multi-tenancy, prefix-stripped signature comparison (Bulgu #62-#82). ------------------------------------------------------------------ NO CORPORATE-SPECIFIC ARTIFACTS ------------------------------------------------------------------ This squash deliberately sanitises corporate hostnames, container registry references, and TLS secret names into generic placeholders (`your-registry.example.com/your-org`, `haproxy-openmanager*.example.com`, `wildcard-tls`, `taylanbakircioglu/haproxy-openmanager-*`) so the public artefact contains no internal infrastructure detail. Pilot / development history that retained those values stays in the corporate fork and is NOT part of this commit. |
||
|
|
07942a82e8 |
feat: ACME stability & enterprise audit (v1.4.0) — fixes #10 #11 #12
Issue #10 — Silent timezone failure on ACME certificate save - Normalize tz-aware expiry_date to UTC tz-naive before INSERT/UPDATE in ssl_certificates (TIMESTAMP WITHOUT TIME ZONE) — restores ACME download path Issue #11 — Duplicate _acme_challenge_backend in generated config - Generator guard: skip auto-append when backend already rendered - Agent-sync filter: _should_sync_backend() drops system-managed backends and detaches their server rows to prevent orphans - Restore filter: IGNORED_BACKENDS skips reserved names during cluster restore - Parser warning: reserved_backend_names blocks accidental manual import - Cleanup migration: removes orphan rows + cascading server entries (idempotent) Issue #12 — Validated orders required manual completion - New 60s background task complete_pending_acme_orders, flag-independent, multi-replica safe via FOR UPDATE SKIP LOCKED + 30s updated_at watermark - Per-order pg_advisory_lock(0x41434D45, order_id) serializes UI-Complete and auto-task races; idempotency guard returns existing certificate cleanly - retry_order endpoint reports in_progress: true within 30s window so the UI surfaces an info toast instead of duplicating CA requests - ACMEAutomation surfaces stuck orders (status=valid && !ssl_certificate_id) with a one-click Complete action and Cancel fallback; conditional 30s polling Other hardening - ACME state machine error_detail persisted as structured JSON across challenge, finalize, download stages for actionable post-mortems - CertificateRequest Pydantic model: domain regex + min_length/max_length and cluster_ids defaulting to all ACME-enabled clusters when "global" is selected - Renewal cluster fallback now requires acme_enabled=TRUE in addition to active - _complete_certificate preserves manual cluster assignments on renewal, surfaces cluster_errors, raises explicit error on missing private key - Audit logging covers acme_certificate_requested/revoked, ca_chain_imported, account created/deactivated/purged, order retried/cancelled - Settings UI exposes acme.staging_url_override for private test CAs (Pebble) - Schema additions: acme_challenges.attempts (default 0) and last_attempt_at, index idx_letsencrypt_orders_status_updated; all migrations idempotent Tests (41/41 passing) - test_acme_expiry_normalize, test_acme_duplicate_backend, test_acme_state_machine, test_acme_pydantic_validation, test_acme_audit_logging, test_acme_concurrency CI / packaging - docker-build.yml reads version.json and pushes additional product-version tag (e.g. 1.4.0) alongside latest and timestamp build id Closes #10 Closes #11 Closes #12 |
||
|
|
93d7ad8fdb |
feat: add ACME Auto SSL with Let's Encrypt integration (v1.1.0)
Add automated SSL certificate management via ACME protocol (RFC 8555): - Full ACME client implementation (account registration, HTTP-01 challenges, certificate issuance/renewal) - Configurable ACME providers (Let's Encrypt, ZeroSSL, custom CA) via Settings UI - Auto-renewal scheduler with PENDING -> Apply -> APPLIED flow alignment - ACME account management (register, deactivate) from UI - Certificate request wizard with domain validation and cluster targeting - Zero changes to HAProxy agent scripts - challenges routed through existing architecture - Comprehensive security hardening (no private key exposure in API responses) - Full backward compatibility with existing SSL, Apply, Rollback, and Restore workflows - Updated README, API documentation, and Kubernetes deployment notes - Version management embedded in code (v1.1.0) - UI messaging improvements for agent-pull architecture accuracy Made-with: Cursor |
||
|
|
1dcf45b1bb |
fix: dynamic collision detection in config generation
Query agent preserved_listen_blocks dynamically and skip conflicting entities to prevent HAProxy validation errors. Falls back to static reserved list if no agent data available. |
||
|
|
d40c701272 |
fix: resolve Python scoping error in config generation
Remove redundant local 'import json' statements that caused: "cannot access local variable 'json' where it is not associated with a value" Root cause: Local import inside try block made 'json' a local variable, but except clause referenced json.JSONDecodeError before assignment. Changes: - Line 197: Remove local import, use global json (line 5) - Line 514: Remove 'import json as _json', use global json - Use ValueError instead of json.JSONDecodeError (equivalent, JSONDecodeError is a subclass of ValueError) This was causing config generation to fail completely, returning error message as config content instead of actual HAProxy configuration. |
||
|
|
851377aedf |
feat: Add HAProxy proxy name collision prevention system
- Add preserved_listen_blocks column to agents table for storing agent's local listen block names - Implement reserved names check (stats, monitoring, admin, etc.) for frontend/backend creation - Add dynamic collision detection against agent's preserved listen blocks - Apply collision checks to CREATE, UPDATE endpoints and bulk import - Add debug mode for failed config validation (saves to /tmp/haproxy-failed-*.cfg) - Fix JSON character stripping for ACL and use_backend rules - Remove collision protection from agent scripts (now handled by backend) - All collision checks wrapped in try-except for backwards compatibility |
||
|
|
183b4fa9e2 |
fix: Auto-add 'verify none' for SSL backend servers without CA file
When ssl_enabled is true for a backend server but: - No ssl_verify option is explicitly set AND - No CA file (ssl_certificate_id) is specified HAProxy 2.8+ defaults to 'verify required' which fails without a CA. Now auto-adding 'verify none' in this case with a warning log. Users can override this in UI by setting SSL Verification explicitly. This fix is safe for: - Bulk Import: Parser already sets ssl_verify='none' when removing ca-file - Version Diff: Generated config correctly shows 'verify none' - Restore: DB values unchanged, config regeneration applies same fix Fixes: 'verify is enabled by default but no CA file specified' error |
||
|
|
378710db3f |
fix(haproxy-config): ACL rules must come before http-request directives
CRITICAL HAProxy Validation Error Fix Problem: Generated HAProxy config failed validation with error: [ALERT] error detected while parsing an 'http-request deny' condition: no such ACL: 'waf_demo-rule1_path'. Root Cause: Config generator was writing http-request directives BEFORE ACL definitions. HAProxy requires ACLs to be defined before they are referenced. Generated Config (WRONG ORDER): http-request deny if waf_demo-rule1_path ❌ ACL not defined yet! acl waf_demo-rule1_path path_reg ^/admin ❌ Too late! Fix: Reordered frontend config generation: 1. ACL Rules (Line 339-362) - Define ACLs FIRST 2. HTTP Request Headers (Line 364-371) - Use ACLs AFTER Generated Config (CORRECT ORDER): acl waf_demo-rule1_path path_reg ^/admin ✅ Define first acl waf_demo-rule1_method method POST ✅ Define first http-request deny if waf_demo-rule1_path waf_demo-rule1_method ✅ Use after Impact: - Parsing logic: UNCHANGED (no breaking changes) - Database storage: UNCHANGED (no schema changes) - Config generation: FIXED (correct HAProxy syntax) - Bulk import: Works correctly now - Agent config apply: Validation passes Testing: 1. Bulk import config with ACLs and http-request rules 2. Agent applies config successfully 3. HAProxy validation passes Files Changed: - backend/services/haproxy_config.py: Reordered ACL and request_headers generation |
||
|
|
8e1a93f873 |
fix(ssl): add backend server SSL advanced options to haproxy config generation
CRITICAL BUG FIX: Backend server SSL advanced options not in generated HAProxy config Problem: - Database has ssl_sni, ssl_min_ver, ssl_max_ver, ssl_ciphers columns ✅ - Config generation code tries to write them to HAProxy config (line 680-687) ✅ - BUT SELECT statement did NOT include these fields ❌ - Result: server.get('ssl_sni') always returned None Impact: - User sets 'TLS Min Version: TLSv1.2' for backend server in UI - Value saved to database correctly - BUT generated HAProxy config missing 'ssl-min-ver TLSv1.2' - Agent deploys incomplete config, SSL settings lost! Solution: - Added 4 server SSL advanced options to SELECT query (line 616): - ssl_sni (for SNI hostname) - ssl_min_ver (minimum TLS version) - ssl_max_ver (maximum TLS version) - ssl_ciphers (cipher suite override) Config Generation Logic: - Line 680-687: Code already writes these fields to HAProxy config - Line 616: Now SELECT actually retrieves the values from database - Example output: 'server backend1 10.0.0.1:443 ssl sni example.com ssl-min-ver TLSv1.2' Testing: - After deployment, bulk import a config with backend server SSL options - Check generated HAProxy config on agent - Should now see: 'server X ssl ssl-min-ver TLSv1.2 ciphers ...' Files Changed: - backend/services/haproxy_config.py: Line 616 SELECT statement |
||
|
|
2dcdaeeba1 |
fix(config): robust use_backend_rules parsing with comprehensive type handling
PROBLEM: - Config generation was producing invalid syntax: use_backend ["..."] - HAProxy validation failing on agents - Old code had insufficient type checking for JSONB fields ROOT CAUSE: - Missing ELSE branch when use_backend_rules wasn't a list - No handling for unexpected types (tuple, Record, etc.) - No debug logging to track type issues FIX: - Added comprehensive type checking (str, list, tuple, other) - Added debug logging to track type and value - Added fallback parsing for string representation of lists - Enhanced validation to skip invalid entries - Prevents generation of invalid HAProxy syntax IMPACT: - Fixes validation failure for cluster 7 (demo-cluster) - Enables proper parsing of JSONB use_backend_rules - Backward compatible with legacy string format TESTED: - Database validation: use_backend_rules is proper JSONB array - String elements correctly extracted and written - Invalid types caught and logged |
||
|
|
b212fb92bc |
Add SSL Advanced Options support (Backend) - Part 1
FEATURE: Complete SSL Advanced Options implementation for frontend and backend server SSL ✅ DATABASE: - Added SSL parameter columns to frontends table: * ssl_alpn, ssl_npn, ssl_ciphers, ssl_ciphersuites * ssl_min_ver, ssl_max_ver, ssl_strict_sni - Added SSL parameter columns to backend_servers table: * ssl_sni, ssl_min_ver, ssl_max_ver, ssl_ciphers - Migration functions: add_ssl_advanced_options_to_frontends() and add_ssl_advanced_options_to_servers() ✅ MODELS: - FrontendConfig: Added 7 new SSL fields for bind parameters - ServerConfig: Added 4 new SSL fields for server parameters - AgentHeartbeat: Added system_info field (fixes HTTP 422 validation error) ✅ BULK IMPORT PARSER: - Parse alpn, npn, ciphers, ciphersuites, ssl-min-ver, ssl-max-ver, strict-sni from bind lines - Parse sni, ssl-min-ver, ssl-max-ver, ciphers from server lines - Store parsed values in frontend/server objects - User-friendly warnings about imported SSL parameters ✅ CONFIG GENERATOR: - Generate bind lines with SSL advanced options: 'bind :443 ssl crt file.pem alpn h2,http/1.1 ciphers ...' - Generate server lines with SSL advanced options: 'server s1 addr:port ssl sni hostname ssl-min-ver TLSv1.2' - Support both NEW MODE (multiple certs) and OLD MODE (single cert) USER IMPACT: - Bulk import now correctly parses SSL configs with alpn/npn/ciphers - SSL parameters preserved during import (not lost anymore) - Agent heartbeat fixed (no more offline agents) - Ready for UI implementation (next commit) EXAMPLE USAGE: Frontend: bind 0.0.0.0:8443 ssl crt cert1.pem crt cert2.pem alpn h2,http/1.1 Server: server s1 10.1.1.1:443 ssl verify required sni backend.example.com ssl-min-ver TLSv1.2 NEXT: Frontend UI components for editing these SSL options |
||
|
|
f1a3826334 |
fix(backend): Allow backends without servers to be deployed
MINIMAL FIX: Enable backend-first workflow (add servers later)
Problem:
1. User creates backend 'deneme-sil' without servers
2. Config generator SKIPs backend (no servers = skip)
3. User adds server 'server1-sil'
4. Config shows: 'server server1-sil 1.1.1.1:1233' WITHOUT backend block
5. HAProxy validation FAILS (server without backend = syntax error)
User Requirement:
- Create backend first (without servers)
- Assign backend to frontend (default_backend)
- Add servers later
- Standard HAProxy workflow
Solution (MINIMAL - 2 small changes):
1. haproxy_config.py (line 453-454):
OLD: Skip backend if no servers (continue)
NEW: Write backend block anyway (remove continue)
Result:
backend deneme-sil
balance roundrobin
mode http
# (no servers yet - backend will show as DOWN)
Valid HAProxy syntax - check
2. cluster.py (line 1600-1607):
OLD: Only mark backends with servers as APPLIED
NEW: Mark ALL backends as APPLIED (servers optional)
Reason: ALL backends are now in config (even without servers)
Why This Is Better Than Previous Approach:
- Only 2 lines changed (vs 50+ lines)
- No complex logic added
- No risk to existing functionality
- HAProxy naturally handles backends without servers (shows as DOWN)
- Aligns with standard HAProxy usage patterns
Test Scenarios:
- Create backend without servers -> Backend block written to config
- Apply -> HAProxy accepts config (backend DOWN)
- Frontend can use backend (use_backend, default_backend)
- Add server later -> Server added to existing backend block
- Apply -> HAProxy accepts, backend goes UP
HAProxy Behavior:
- Backend without servers: DOWN (no available servers)
- Backend with disabled servers: DOWN (all servers disabled)
- Backend with active servers: UP (servers available)
Related: caa21f0 (has_pending_config fix)
Refs: #backend-workflow #server-optional #haproxy-syntax
|
||
|
|
73ac554add |
fix(backend): Skip backends with no servers in HAProxy config generation
CRITICAL FIX: Backends without servers were causing HAProxy validation failures Problem: - Backend 'silbeni' (ID: 118) oluşturuldu ama server eklenmedi - Config generator backend'i haproxy.cfg'ye yazdı ama server satırı olmadan - HAProxy validation FAIL: 'backend has no servers' - Agent config'i apply etmedi - Frontend'de backend görünmedi (validation fail nedeniyle) Root Cause: - Config generator server olup olmadığını kontrol etmiyordu - HAProxy en az 1 server gerektirir, yoksa validation fail olur - Validation fail = agent apply etmez = backend haproxy.cfg'de görünmez Solution: - Backend loop başında server pre-check eklendi - Server yoksa backend config'e yazılmaz ve WARNING log'lanır - HAProxy validation her zaman başarılı olur (sadece valid backend'ler yazılır) Impact: ✅ Server olmayan backend'ler artık config'e yazılmayacak ✅ HAProxy validation artık fail olmayacak ✅ Agent successfully apply edecek ✅ Kullanıcı frontend'de backend'i görecek (0/0 servers ⚠️ Empty tag ile) ✅ Kullanıcı server ekledikten sonra Apply yapınca backend haproxy.cfg'ye yazılacak Testing: - Server olmayan backend oluştur → Config'e yazılmaz (log: SKIPPING) - Server ekle → Config'e yazılır - Apply → Başarılı Related: a39a5d6 (frontend null/undefined check) Refs: #backend-validation #haproxy-config-generator |
||
|
|
8595656803 |
fix: comprehensive validation for all single-value text fields
Phase 3 - Single Value Field Validation:
- Frontend.default_backend: Added validation to prevent '[]' causing ALERT
- Frontend.monitor_uri: Added validation for monitor endpoint
- Backend.health_check_uri: Added validation for health check path
- Backend.cookie_name: Added validation for cookie persistence
- Server.server_name: Added validation with fallback to server_id
- Server.server_address: Added validation (critical field, skip if invalid)
All single-value text fields now validate against:
- Empty strings
- '[]', '{}', 'null', 'None' invalid values
- Proper error logging and skipping
Additional improvements:
- Removed all emojis from log messages per user request
- Fixed server_address variable usage consistency
- Added proper error messages for debugging
Comprehensive Backend Audit Results:
- Checked all routers (frontend, backend, waf, ssl, config)
- Checked all models (Pydantic validation)
- Checked all services and utils
- Only one config generation file: haproxy_config.py (FULLY FIXED)
- Template files use static strings (no risk)
- Agent scripts use static templates (no risk)
Total fields validated: 30+ across all entity types
Risk level: ZERO - Complete protection against invalid values
|
||
|
|
617e303205 |
fix: additional comprehensive validation for remaining text fields
Phase 2 - Extended Field Validation:
- Frontend.options: Added validation for multiline option directives
- Server.ssl_verify: Added validation for SSL verify parameter
- WAF.redirect_url: Added validation for redirect URL (2 locations)
- WAF.header_name: Added validation with skip on invalid values
- WAF.header_value: Added validation with skip on invalid values
- WAF.path_pattern: Added validation for regex patterns (2 locations)
- WAF.http_method: Added validation for HTTP method filtering
All text fields now validate against invalid values:
- Empty strings, '[]', '{}', 'null', 'None' are skipped
- Warning comments added for debugging invalid WAF rules
- Zero risk of syntax errors in generated HAProxy config
Total fields validated: 24 across Frontend, Backend, Server, and WAF entities
Risk level: ZERO - All string concatenation points secured
|
||
|
|
be4ed94bc4 |
fix: comprehensive validation for all string fields in HAProxy config generation
- Added empty array/null validation for ALL text fields to prevent syntax errors
- Frontend: request_headers, response_headers, tcp_request_rules, acl_rules, use_backend_rules, redirect_rules
- Backend: options, request_headers, response_headers, cookie_options
- Server: cookie_value
- WAF: All 6 custom_condition usage points (IP filter, rate limit, header filter, request filter, geo block, custom rules)
- Prevents invalid syntax like 'redirect []', 'acl []', 'cookie []', 'http-request []'
- All string fields now skip '[]', '{}', 'null', 'None' values before config generation
- Critical fix for bulk imported configs with empty JSON array fields
|
||
|
|
21de3585cb |
fix: prevent invalid empty array syntax in HAProxy config generation
- Skip '[]', '{}', 'null', 'None' strings in redirect_rules, acl_rules, use_backend_rules
- Add validation for request_headers, response_headers, tcp_request_rules
- Prevents 'redirect []' syntax error that causes HAProxy validation failure
- Fixes: parsing [config:70] : error detected in frontend while parsing redirect rule (was '[]')
- All rules now properly filtered before being written to config
|
||
|
|
3e22776b8c |
fix: Improve options field UX and HAProxy config ordering
Fixed three critical issues with options field implementation: 1. Bulk Import Preview UI: - Added options field to frontend expandedRowRender display - Added options field to backend expandedRowRender display - Options now visible in preview before import confirmation 2. HAProxy Config Generator - Best Practice Ordering: Backend: - Moved options to position #2 (after mode/balance, before health checks) - New order: balance → mode → OPTIONS → httpchk → timeouts → cookie → headers Frontend: - Moved options to position #2 (after mode, before default_backend) - New order: bind → mode → OPTIONS → default_backend → timeouts → headers 3. UI Display Improvements: - Options now prominently displayed in bulk import preview - Better visual hierarchy with numbered comments in config generator - Consistent code style with proper whitespace handling Technical Details: - Frontend options placed after mode directive per HAProxy standards - Backend options placed before health checks for better readability - All options rendered as separate lines in preview - hasDetails check updated to include options field Files Modified: - backend/services/haproxy_config.py: Config generation order optimized - frontend/src/components/BulkConfigImport.js: Preview display enhanced |
||
|
|
0fc18fde38 |
feat: Add HAProxy options support for backends and frontends
Implemented comprehensive HAProxy options field support for both backend and frontend entities to enable standard HAProxy directives like 'option http-keep-alive', 'option httplog', 'option forwardfor', etc. Changes: - Database: Added 'options' TEXT column to backends and frontends tables - Models: Added options field to BackendConfig, BackendConfigUpdate, and FrontendConfig - API Endpoints: Updated CREATE, UPDATE, and GET endpoints to handle options field * Backend: CREATE/UPDATE/GET with options support * Frontend: CREATE/UPDATE/GET with options support (fixed 5 SELECT queries) - Config Generator: Added options block generation for both backends and frontends - Bulk Import Parser: * Added options field to ParsedBackend and ParsedFrontend dataclasses * Implemented option directive parsing with validation * Added unknown option warnings * Fixed bulk parse response to include options field - Bulk Import Merge: Added options field comparison in UPDATE logic - UI Components: * BackendServers.js: Added options TextArea form field * FrontendManagement.js: Added options TextArea form field Features: - Multi-line options support (newline-separated format) - Option validation with known HAProxy options list - Backward compatible (NULL options for existing entities) - Bulk import support with merge strategy - Full CRUD support for both manual and bulk operations Technical Details: - Format: Newline-separated TEXT field for multiple options - Validation: Warns about unknown options but allows them - Config Generation: Each option written as separate directive - Agent: Standard HAProxy config validation applies Total: 10 files modified, ~195 lines added, 26 integration points verified |
||
|
|
a5e281b284 |
Feature: Backend Server SSL certificate support + Frontend SSL dropdown enhancement
✨ Backend Server SSL Certificate - Complete Implementation: 1. Model Update (backend/models/backend.py): - Added ssl_certificate_id field to ServerConfig model - Allows selecting SSL certificate from dropdown 2. API Endpoints (backend/routers/backend.py): - CREATE server: Added ssl_certificate_id to INSERT query - UPDATE server: Added ssl_certificate_id to allowed_fields - GET servers: Added ssl_certificate_id to SELECT queries (2 places) 3. Config Generation (backend/services/haproxy_config.py): - SSL certificate lookup by ID - Auto-generate ca-file path: /etc/ssl/haproxy/{cert_name}.pem - Added to server line in HAProxy config Example Generated Config: Before: server es1 10.0.0.1:9200 ssl verify required After: server es1 10.0.0.1:9200 ssl verify required ca-file /etc/ssl/haproxy/star-burgan-com-tr.pem ✨ Frontend SSL Dropdown Enhancement: - Added Global/Cluster-specific tags to Frontend SSL dropdown - Matches Backend Server SSL dropdown design - Shows: [🌍 Global] or [📍 Cluster] with color coding 🔧 Complete SSL Workflow: 1. User edits Backend Server 2. Enables SSL 3. Selects SSL certificate from dropdown 4. Saves → ssl_certificate_id stored in DB 5. Apply Changes → Config generated with ca-file path 6. Agent downloads SSL cert to /etc/ssl/haproxy/ 7. HAProxy uses ca-file for SSL verification ✅ Database Schema: backend_servers table now includes: - ssl_enabled (bool) - ssl_verify (str: none/required) - ssl_certificate_id (int, FK to ssl_certificates) ✅ HAProxy Config Format: server {name} {addr}:{port} ssl verify required ca-file {path} Impact: Backend Server SSL now fully functional with certificate management |
||
|
|
1158e5f2b1 |
Fix: HAProxy validation - ACL and use_backend parsing improvements
🐛 Critical Bug Fixes: - Fixed duplicate 'acl' prefix in generated config (was: 'acl acl Name ...') - Fixed duplicate 'use_backend' prefix in generated config - Added use_backend directive parsing from bulk import configs - Fixed redirect_rules list handling (was causing .strip() error) 🔧 Parser Improvements: - Added use_backend rules parsing (stored as list like ACL rules) - Changed use_backend_rules field from str to list for consistency - Parser now captures all use_backend directives with conditions 🎯 Config Generation Improvements: - Smart prefix detection: only add 'acl' if not already present - Smart prefix detection: only add 'use_backend' if not already present - Support both legacy (string) and new (list) format for rules - Proper JSON parsing with fallback to newline-separated format ✅ HAProxy Validation: - Generated config now passes HAProxy validation (haproxy -c -f) - ACL and use_backend directives in correct HAProxy format - Routing rules properly linked with ACL conditions Example parsed config: acl Elasticsearch hdr(host) -i baremetal-elastic.burgan.com.tr use_backend Elasticsearch if Elasticsearch Tested with full config including multiple ACLs and routing rules. |
||
|
|
6aae0f4309 | Initial commit |