mirror of
https://github.com/taylanbakircioglu/haproxy-openmanager.git
synced 2026-09-16 15:45:11 +00:00
main
22 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
47cc79dcf7 |
fix(acme): address review findings on the challenge-backend hardening
Five confirmed findings from an adversarial review of the branch, four of them regressions introduced by it. Fall through to the next source when a stored URL cannot be resolved. Stopping at the first non-empty candidate emitted a backend section with no `server` line: the section exists so `haproxy -c` passes and Apply succeeds, then every challenge request 503s from an empty backend with nothing to show for it. Scheme-less values are common — the settings field was free text until this branch — so this was reachable on real installs. Selection moved into `select_acme_backend_source()` so it is testable and the skipped candidates are logged rather than silently dropped. Report a challenge backend with no server line. `extract_acme_backend_target` returns None for that section, and the loopback filter skipped falsy targets, so the case above would have been reported as "challenge route present in applied config" — the new check confirming the very state it exists to catch. Do not narrow the row set feeding the routing check's `fail` branch. Adding a mode filter to the WHERE clause turned a tcp-only port-80 cluster from "ok" into "fail", and the site wizard blocks submit on any failing check, so those installs would have been locked on upgrade day. Mode is now examined in Python and only downgrades to `warn`, using an expression that is character-for- character the renderer's normalisation. Match the agent's config selector. The applied-config lookup omitted `is_active = TRUE`, so it could read a superseded row and report on a config the nodes never received. Extraction now happens in SQL rather than pulling whole configs — these run to hundreds of KB. Select `acme_backend_url` when loading the existing cluster. It was absent, so the entity snapshot recorded old_values as NULL unconditionally and rejecting the pending version wiped the operator's per-cluster URL back to the global loopback default — re-creating the exact failure this branch removes. Also carry `acme_enabled` and `acme_backend_url` through cluster creation. The create model declared neither and the INSERT wrote neither, so a cluster created with ACME switched on came back switched off with no error shown. Refuted and deliberately not changed: settings PUT re-validating a stored loopback value (it validates only what is submitted), an apply-path connection leak (the 422 propagates to a handler that closes it), and the modal discarding backend rejection reasons (the envelope matches). |
||
|
|
bb774141d4 |
fix(acme): make the challenge backend fixable from the panel
Correcting a wrong ACME challenge backend was impossible without a shell, and
even with one the correction did not reach the nodes.
The mint gate only fired when `acme_enabled` flipped. `acme_backend_url` was
written to the DB and minted nothing, so Apply answered "No pending changes to
apply" and the nodes kept the old address forever. It is now decided by
comparing the rendered `server _acme_mgmt` line against the active version —
the one line that answers "would the nodes talk to a different address?".
Comparing whole configs would flag every unrelated pending edit.
The field had no UI at all. Added to the cluster form with validation that
mirrors the backend rules, and keyed on `model_fields_set` so clearing it
reverts to the global setting — with a plain `is not None` test an empty box is
indistinguishable from "not submitted", so a value could never be removed.
Validation is asymmetric on purpose (utils/acme_backend_url):
- at the write boundary, reject what cannot express a reachable target —
including the two silent traps: a scheme-less value became `localhost`, and
an out-of-range port raised inside the generator and destroyed the config
- at render time, never reject. The shipped defaults are themselves loopback,
so refusing to render would make every acme_enabled cluster unappliable,
including for changes unrelated to ACME. Problems are logged and surfaced.
The port-less default stays 8080 rather than moving to HTTP's 80: the bundled
compose publishes nginx on 8080, so installs relying on it work today and the
first sign of breaking them would be the unattended renewal loop months later.
The omission is warned about instead.
RFC1918 is allowed and is usually the right answer here, and no DNS resolution
is performed — both deliberate departures from utils/ssrf_guard, whose policy
is the opposite of what this address needs. What the management host can
resolve says nothing about what the HAProxy node can reach.
Diagnostics stop reporting success on a dead path:
- check_port80 uses GET instead of HEAD and classifies the body. A proxy that
has lost its /.well-known/acme-challenge/ location serves its SPA with HTTP
200, which `status in (200, 404)` accepted as healthy. Warnings also surface
when other domains pass, which previously hid the most diagnostic outcome.
- check_routing filters `mode`, joins `acme_enabled` and reads the APPLIED
config instead of counting database rows, and reports a loopback target.
- every new condition is `warn`, never `fail`: the site wizard blocks submit on
any fail, so a new failing condition would lock every install on upgrade day.
Also: normalise `frontends.mode` once per frontend. It is nullable, and the
raw value was interpolated into `mode {}`, emitting a literal `mode None` that
HAProxy rejects — taking down the whole cluster config. The ACME gate and the
backend-mode check now read the same normalised value.
And stop hardcoding PUBLIC_URL / MANAGEMENT_BASE_URL in docker-compose, which
silently ignored the operator's .env and made the wrong default load-bearing.
|
||
|
|
a6166d11b9 |
feat(ssl): add CSR generation and signed-certificate import (backend)
New /api/ssl/csrs endpoint group: generate a private key + CSR server-side
(RSA 2048/4096, ECDSA P-256/P-384; full subject + DNS SANs with wildcard
support), list/detail/delete CSRs, and import the CA-signed certificate.
- New ssl_csrs table (SCHEMA_VERSION 9 -> 10, additive + idempotent); the
migration re-raises on failure so a failed run is retried instead of being
stamped as applied.
- Import verifies the certificate against the stored key as a hard gate
(match=None is treated as an integrity error, not a lenient pass), rejects
malformed and expired certificates with 400, warns on SAN drift, and
creates a normal ssl_certificates row (source=csr, cluster_id=NULL,
last_config_status=PENDING) so it flows through the standard
Apply Management -> agent pull pipeline.
- Concurrency: FOR UPDATE row lock serialises double-import and
delete-during-import; a partial unique index reserves pending CSR names;
soft-deleted same-name certs are reactivated preserving the row id.
- Security: no CSR endpoint ever returns the private key (explicit column
lists, enforced by a static test); the key copy on the CSR row is NULLed
after import; ssl.create/read/delete permissions enforced on every
endpoint incl. reads; per-user rate limit on key generation, which runs
in a worker thread; csr_id and cluster_ids are int32-guarded.
- ssl_service: extract _prepare_cert_fields from create_cert_row (behaviour
unchanged, extraction tests untouched) and add stage_ssl_config_versions
reusing the exact ssl-{id}-create-{ts} version-name scheme.
- Tests: crypto round-trip for all four algorithms, model validation,
import-flow unit tests, endpoint auth/permission pinning, migration and
key-non-exposure static assertions.
|
||
|
|
c79391cd13 |
feat(acl): accept HAProxy -f pattern-file references with advisory warnings (v1.8.9, Issue #38)
The manual Frontend editor, wizard and visual ACL builder hard-rejected the ACL `-f <file>` flag while bulk import accepted it. Worse, a frontend imported with an `-f` ACL could not be edited at all (422) until the ACL was dropped. The original guard predated the fail-safe apply flow: the agent runs `haproxy -c` before every reload, so a missing pattern file is rejected safely and the previous config keeps running. Pattern files are operator-managed host files — the same policy adopted for SPOE filter configs in v1.8.8. - models: remove the 5 `-f` hard rejects (frontend acl/redirect/use_backend validators + wizard string/dict-redirect guards); `$(`/backtick and X!X contradiction guards unchanged - routers/frontend: `_pattern_file_warnings` helper; non-blocking warning on create + update responses listing referenced pattern files (empty when no rule uses `-f` — zero noise) - routers/config: bulk-import preview advisory listing pattern files per frontend (cluster config-dir aware, next to the SPOE advisories) - React: remove the FrontendManagement submit gate and SiteWizard step gate; ACLRuleBuilder renders informational notes instead of errors and re-adds `-f (pattern file on host)` to the flag dropdown; create path now renders server warnings like update - tests: 4 reject-pins inverted to accept-pins; new test_acl_pattern_file_allow.py (accept/guards-kept/zero-noise/advisory); full suite green (1094 passed) |
||
|
|
9e2ea04777 |
feat(haproxy): preserve SPOE filter + frontend log-format on import/edit (v1.8.8, Issue #38)
Bulk import / manual edit silently dropped `filter spoe engine ...` (Coraza WAF) and frontend `log-format` because the parser recognised only a fixed directive set. The regenerated config then missed the SPOE engine, so HAProxy failed with "unable to find SPOE engine 'coraza' used by the send-spoe-group". - parser: capture `filter` + `log-format`/`log-format-sd` into new ParsedFrontend fields - db: additive nullable `log_format` + `filters` TEXT columns on frontends (SCHEMA_VERSION 8->9) - generator: new `filter` bucket flushed before http-request rules so `filter` precedes `send-spoe-group`; `log-format` kept in prelude - bulk import: preview dict, change-detection, persist (create + merge-update); cluster-aware SPOE pre-flight advisories (missing-filter + host-prerequisite) surfaced in the UI - manual CRUD: full round-trip (get/create/update) incl. React form fields (no null-wipe) - reject/rollback: restore the new columns; restore path + wizard helper kept in parity - backend `option spop-check` recognised (suppresses spurious warning for coraza-spoa) - tests: test_spoe_filter_import.py; full suite green (1079 passed) |
||
|
|
a1192e602d |
feat: HA / VIP (Keepalived) management from the UI (#27)
Manage highly-available virtual IPs backed by Keepalived (VRRP) directly from the OpenManager
UI — no more SSHing into nodes to install/configure Keepalived by hand. Builds on the agent
pull-architecture: define the VIP centrally, click Apply, and the agents converge.
Highlights:
- New "HA / VIP" tab: create a virtual IP, pick a per-node interface, select which pool nodes
participate (MASTER/BACKUP roles + priorities); live MASTER/BACKUP/FAULT per node.
- On Apply, agents install & configure Keepalived (unicast VRRP, cloud-safe default) across the
major distros (Debian/Ubuntu, RHEL/CentOS/Alma/Rocky, Fedora, SUSE/openSUSE, Alpine) with a
HAProxy health-check, so the VIP fails over automatically when HAProxy drops.
- Single-node (a managed floating IP without failover) and multi-node VRRP failover both work.
- VIP changes ride the standard Apply Management flow with the standard "View Change" diff.
- Approval-gated deletion (safety): deleting a running VIP is staged for approval and the VIP
keeps running, untouched, until you approve it — an agent never tears a VIP down without an
explicit human approval. Per-VIP Diagnostics view; opt-in package uninstall (only on nodes
where OpenManager installed it). A node already running a hand-managed Keepalived is detected
and never overwritten ("externally managed").
- Fully opt-in and backward compatible: nodes/clusters without a VIP are unaffected. Adds
vip_instances + vip_members tables (idempotent SCHEMA_VERSION bump; existing data and
passwords unaffected) and a `vip` RBAC permission group.
- Also includes a HAProxy config-generator robustness fix: auto-inject a stick-table when a
frontend uses a stick counter (track-sc / sc_*_rate) but declares none.
On-prem / L2 (VRRP) scope; the UI notes the cloud caveat.
|
||
|
|
bd6a31cb0d |
feat: v1.6.0 — Multi-Factor Authentication (Issue #18)
Adds opt-in TOTP-based Multi-Factor Authentication that is fully
backwards compatible with existing logins. Operators choose to enable
MFA per account; nothing changes for users who do not opt in.
Highlights
==========
* RFC 6238 TOTP (6 digits, 30s period, SHA1) with ±30s skew tolerance,
compatible with Microsoft / Google Authenticator, Authy, Duo, 1Password.
* Per-step replay protection (`mfa_last_used_totp_step`) so a captured
code cannot be reused inside the same window.
* Fernet-encrypted TOTP secrets at rest, key resolution via
`MFA_ENCRYPTION_KEY` env (HKDF-derived from `SECRET_KEY` as fallback).
* 10 single-use, bcrypt-hashed backup codes per user, formatted
`XXXX-YYYY` from a confusion-free alphabet (no 0/O/1/I/L).
* Two-step login flow: `POST /api/auth/login` returns `mfa_required`
+ `mfa_token`, then `POST /api/auth/login/mfa-verify` accepts a TOTP
code OR a backup code. JWT is minted only after MFA succeeds.
* Self-service: users enable / disable MFA from their own row in the
Users page; admins reset (single user or bulk) but never enable on
behalf of someone else (matches AWS IAM / GitHub / Google Workspace).
* Bulk emergency reset CLI: `scripts/admin-mfa-reset-all.sh`.
Security hardening
==================
* Atomic transactions with `SELECT … FOR UPDATE` on `mfa_pending_logins`
and `users` rows so concurrent verify / enroll calls cannot race.
* `/api/mfa/enroll/start` refuses re-enrollment when MFA is already on
(prevents silent secret rotation via a stolen JWT).
* Pydantic `ValidationError` messages are sanitized before reaching the
audit log so request bodies (TOTP / backup codes in flight) never
appear in plaintext.
* Slowapi rate limits are per-USER, not per-IP, with a trusted-proxy
XFF strategy so a single ingress address cannot exhaust the bucket
for thousands of operators (`MFA_TRUSTED_PROXY_CIDRS`,
`MFA_RATE_LIMIT_*` env-overridable).
* Login query now scopes to `is_active = TRUE` so a soft-deleted row
with the same username can no longer occlude the active user
(also closes a small account-enumeration side channel).
Database
========
Additive migrations (idempotent `ADD COLUMN IF NOT EXISTS`,
`CREATE TABLE IF NOT EXISTS`):
- users: mfa_enabled, mfa_method, mfa_secret_encrypted,
mfa_enrolled_at, mfa_last_used_at, mfa_last_used_totp_step
- mfa_backup_codes (user_id ON DELETE CASCADE)
- mfa_pending_logins (user_id ON DELETE CASCADE, challenge_token,
attempts, expires_at)
- mfa_pending_enrollments (user_id ON DELETE CASCADE)
Frontend
========
* Login page becomes a 3-phase state machine
(credentials → MFA → submitting); legacy single-step login is
preserved for users who haven't enrolled.
* New MFAEnrollModal (3-step wizard: QR + secret → verify → backup
codes) using `qrcode.react`.
* Users page shows MFA column + per-row enable/disable/reset actions.
Admins viewing other users with MFA off see a non-actionable info
icon explaining that only the user themselves can enable MFA.
Deployment
==========
* `MFA_ENCRYPTION_KEY` is added to `k8s/manifests/03-secrets.yaml` as
a placeholder; `SECRET_KEY` is also placeholder-ized so both are
injected by the existing pipeline pattern (sed-replace + apply).
* No new build-time env vars are required for the frontend. The SPA
uses `window.location.host` for `/api/*` and is routed by the
existing nginx ingress configuration.
* `frontend/.dockerignore` ensures host `.env*` files cannot bleed
into the production bundle.
Tests
=====
* New unit suites:
- `test_mfa_service.py` (TOTP, encryption, backup codes)
- `test_mfa_backwards_compat.py` (regression — non-MFA flow unchanged)
- `test_mfa_rate_limits.py` (env override + dataclass immutability)
- `test_mfa_rate_limit_key.py` (JWT key, trusted-proxy XFF, fallbacks)
* All existing 1000+ unit tests continue to pass.
Documentation
=============
* README MFA section (overview, day-to-day operations, emergency
reset CLI, env variables, rate-limit tuning).
* `scripts/README.md` documents the bulk reset script.
Issue: #18
|
||
|
|
2e7db4d99f |
fix: v1.5.1 — Round-23 + Round-24 audit follow-ups (Bulgu #83 → #93)
A live-deployment audit pass over the v1.5.0 Site Wizard + ACME Diagnostic Panel surface. Two adversarial review rounds (R23, R24) each capped by an end-to-end smoke test against a multi-cluster staging deployment. Bulgu #83 — Frontend Management page warned about stale data without a clear retry CTA. The toast now carries an in-place "Reload" action and the page-level Empty state surfaces the same recovery affordance, so operators never get stuck on a stale-data view without an obvious way out. Bulgu #84 — ACME diagnostics ran with the wrong "last_heartbeat" column reference against the agents table. Aligned the SELECT with the actual schema column (`last_seen`); pinned by an idempotent regression test in `test_acme_diagnostics.py`. Bulgu #85 — ACME order error_detail rendering could leak the raw asyncpg/SQL exception class name when humanize_error_detail encountered an unhandled CA response shape. Added a backwards- compatible fallback branch that emits an "ACME error (raw)" panel without exposing parse_error class name to the user. Bulgu #86 — Multi-cluster apply with concurrent rejects could leave wizard_staged orders dangling without their parent draft. Pinned via reject_order_with_cluster_orphan test. Bulgu #87 — Frontend Management page list virtualization mis-keyed during a re-sort + stale-row replace race; fixed by keying rows on `id + version` so React reconciler does not reuse DOM for a logically different row. Bulgu #88 — Site Wizard "Cancel" mid-flow now surfaces an unsaved-draft prompt with explicit Save / Discard buttons (and the same prompt on browser tab close), so the operator never loses 5 steps of input to an accidental ESC. Bulgu #89 — Existing-cert SSL mode showed an empty dropdown when the cluster had >100 certs because the listing endpoint default-limited results. Endpoint now exposes pagination AND the wizard switches to client-side filtering above 50 rows. Bulgu #90 — ACME pre-check on the wizard preview path did NOT re-validate the account against `letsencrypt_accounts` if the operator stepped Back/Forward between SSL and Review. Added a debounced re-validation on Review entry. Bulgu #93 — Site Wizard hsts_enabled toggle in HTTPS frontend was idempotent-by-name (the generated `http-response set-header Strict-Transport-Security` line could duplicate across a Save + Apply cycle). The renderer now upserts the header in place. Cumulative outcome: backend pytest 1084/1084, frontend lint clean, and a 6-hour live-deployment smoke session against staging with no regressions reported. |
||
|
|
02b1cb2bca |
feat: v1.5.0 — Site Wizard (Issue #14) + ACME Diagnostic Panel (Issue #13)
Closes #13, Closes #14. This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0 shipped the ACME stability & enterprise audit (Issues #10/#11/#12). v1.5.0 builds on that foundation with two co-equal headline features plus a 22-round audit campaign hardening the prior configuration surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0 lands in v1.5.2). ------------------------------------------------------------------ HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13) ------------------------------------------------------------------ A live pre-flight + post-failure diagnostic surface for every ACME order, reachable from the ACME Automation page. The panel exists to make ACME failures legible to operators who do NOT have shell access to the API host. Endpoints (`backend/routers/acme_diagnostics.py`): POST /api/letsencrypt/orders/{order_id}/diagnostics Run the full 5-check suite (DNS / port-80 / routing / account / agents) and humanize the order's `error_detail` (>=11 RFC-8555 problem types, backwards compatible with legacy plain-string failures). POST /api/letsencrypt/orders/{order_id}/diagnostics/ {check_id}/rerun Re-run a single check in place — used by the "Re-run" button on every row of the modal's pre-flight table. GET /api/letsencrypt/orders/{order_id}/events Merged event timeline combining the typed `acme_order_events` rows with correlated `user_activity_logs` entries (resource_type = 'letsencrypt_order' AND resource_id = order_id). The diagnostic modal auto-tails this timeline every 5 seconds while open. Service-level checks (`backend/services/acme_diagnostics.py`): * DNS resolution via stdlib socket.gethostbyname_ex through run_in_executor (intentionally avoiding an aiodns runtime dep for v1.5.0). * Port-80 HEAD probe, target locked to the order's domains, success on HTTP 200 OR 404, warns on egress timeout (corp egress policies routinely blackhole outbound 80 — fail-hard would be too noisy). * SSRF guard: probe refuses non-public IPs and surfaces the skip in the diagnostic result; IPv4-mapped IPv6 normalisation closes the `::ffff:169.254.169.254` cloud-metadata vector. * HAProxy routing presence check: matches the order's cluster_ids to a port-80 HTTP frontend. * ACME account validity check against `letsencrypt_accounts`. * Agent presence check (>=1 active agent in target cluster). * Every sub-check wrapped in a wall-clock timeout to bound impact on the API event loop. RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate limit on both run and rerun, backed by the (user_id, action, created_at DESC) composite index. Frontend (`frontend/src/components/ACMEAutomation.js`): * "Diagnose" button on every order row + the existing "stuck order" warning row. * Modal with two tabs: - Pre-flight Checks (Antd Table with status pills + Re-run buttons + humanized error banner) - Event Log (Antd Timeline with auto-tail polling, scroll- to-bottom, pause-on-hover) * Correlation IDs surfaced in error banners and individual check fail details for backend-log lookup. ------------------------------------------------------------------ HEADLINE FEATURE B — Site Setup Wizard (Issue #14) ------------------------------------------------------------------ A single guided flow that creates a Backend + Servers + HTTP Frontend (and optional HTTPS Frontend) in one atomic transaction. Endpoints (`backend/routers/site_wizard.py`): POST /api/site-wizard/preview — diff-preview the changeset POST /api/site-wizard/create — atomic execute POST /api/site-wizard/reject — clean rollback (including any wizard_staged ACME orders) GET /api/site-wizard/drafts — draft persistence PUT /api/site-wizard/drafts/{id} — save/update DELETE /api/site-wizard/drafts/{id} Feature surface: * One screen captures both backend (mode + servers) AND frontend (http + optional https + SSL mode) inputs. * SSL modes: ACME (new order, HTTP-01 only for v1.5.0), Upload (existing PEM), Existing (link to a stored cert), or None. * ACME-staged path: wizard_staged_until watermark on the `letsencrypt_orders` row defers finalisation until agent confirmation; per-mode reject cleanly cancels and rolls back the staged order. * Live diff preview against the cluster's current generated config (renderer-evolution noise stripped — track-sc<N> dedup, per-server cookie strip, defaults-cookie inheritance, listen-block flattening). * Draft persistence with PEM stripped at save time (private keys never round-trip through the drafts table). * Per-cluster multi-tenancy: drafts and wizard_staged orders are isolated to the creating user's cluster scope. Frontend (`frontend/src/components/SiteWizard.js`): * 4-step Antd Steps flow: Backend → Frontend → SSL → Review. * Render the live diff preview inline before commit. * Antd Form-level validation mirrors backend Pydantic validators (numeric bounds, HAProxy reserved keywords, ALPN consistency, IPv6 scope-id, domain regex, server name dedup). ------------------------------------------------------------------ AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1 → #82) ------------------------------------------------------------------ v1.5.0 includes 22 adversarial review passes. Each round produced its own commit set in the corporate development line; this squash collapses those into the v1.5.0 release artefact. Highlights: Round 1-4 Site Wizard core: dry-run parity, single-line value injection guard, ACL -f pattern-file block, SSL parity, timeout regex, form-state pin. Round 5-7 defaults-cookie inheritance, server-named-cookie guard, fe/be mode mismatch, duplicate server names, health_check_uri + server_address validators. Round 8-10 cookie_name / cookie_options newline-injection guard, dry-run parity (round 9), TCP-mode HTTP-only feature blockers. Round 11 SSL name path traversal + health-check >= 1. Round 12-13 SSL & ACME deep dive (Bulgu #23-#32). Round 14 single-line value injection (Bulgu #33). Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy reserved keywords, ALPN/TLS consistency, all-backup, multi-domain & multi-user enterprise edges, drain/HSTS/post-completion (Bulgu #34-#53). Round 18-21 concurrency, agent state, TCP-mode HTTP-only, list size caps, IPv6 scope-id, preview account validation, TCP backend + balance uri reject (Bulgu #54-#61). Round 22 FE error visibility + 3x stale-data lockouts, referential integrity + cascade safety, authentication & authorization, multi-cluster isolation, apply_pending_changes concurrency, script injection + bulk import multi-tenancy, prefix-stripped signature comparison (Bulgu #62-#82). ------------------------------------------------------------------ NO CORPORATE-SPECIFIC ARTIFACTS ------------------------------------------------------------------ This squash deliberately sanitises corporate hostnames, container registry references, and TLS secret names into generic placeholders (`your-registry.example.com/your-org`, `haproxy-openmanager*.example.com`, `wildcard-tls`, `taylanbakircioglu/haproxy-openmanager-*`) so the public artefact contains no internal infrastructure detail. Pilot / development history that retained those values stays in the corporate fork and is NOT part of this commit. |
||
|
|
71c717364c |
fix: allow dot character in entity names for UI and backend validation
Bulk import accepted dots in frontend/backend/server names but UI and backend validators rejected them with ^[a-zA-Z0-9_-]+$. After import, entities with dots could not be edited. HAProxy itself allows dots in section names, so the regex is expanded to ^[a-zA-Z0-9_.-]+$ across all 12 validation points (5 React form rules, 1 ACL char-strip, 3 Pydantic validators, 1 WAF validator, 2 config-validator warnings). |
||
|
|
93d7ad8fdb |
feat: add ACME Auto SSL with Let's Encrypt integration (v1.1.0)
Add automated SSL certificate management via ACME protocol (RFC 8555): - Full ACME client implementation (account registration, HTTP-01 challenges, certificate issuance/renewal) - Configurable ACME providers (Let's Encrypt, ZeroSSL, custom CA) via Settings UI - Auto-renewal scheduler with PENDING -> Apply -> APPLIED flow alignment - ACME account management (register, deactivate) from UI - Certificate request wizard with domain validation and cluster targeting - Zero changes to HAProxy agent scripts - challenges routed through existing architecture - Comprehensive security hardening (no private key exposure in API responses) - Full backward compatibility with existing SSL, Apply, Rollback, and Restore workflows - Updated README, API documentation, and Kubernetes deployment notes - Version management embedded in code (v1.1.0) - UI messaging improvements for agent-pull architecture accuracy Made-with: Cursor |
||
|
|
b0feca3c2e |
feat: add keepalived VRRP state (MASTER/BACKUP) detection to agent heartbeat
- Add keepalive_state and keepalive_ip columns to agents table (migration + schema) - Add keepalive fields to AgentHeartbeat Pydantic model (backward compatible) - Update heartbeat endpoint to persist keepalive data to DB and cache in Redis - Add multi-method keepalived detection in agent scripts (journalctl, log files, VIP check) - Update dashboard-stats agents/status API with Redis-first keepalive lookup - Update GET /api/agents to include keepalive_state and keepalive_ip - Show MASTER/BACKUP tag in Dashboard AgentStatusCard - Show keepalive info in Agent Management registered agents table - Add Keepalive column to Cluster Management table with VIP search support Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
589f865f59 |
fix(ssl): rollback support + improved ALPN validation error message
PART 1: Rollback Support for SSL Advanced Options =================================================== Problem: SSL advanced options lost on reject/rollback operations Root Cause: - Snapshot creation uses SELECT * (includes all fields) ✅ - old_values contains SSL advanced options ✅ - Rollback UPDATE did NOT restore SSL fields ❌ Impact - Frontend: - User changes ssl_alpn, then clicks Reject - Rollback skipped: ssl_alpn, ssl_npn, ssl_ciphers, ssl_ciphersuites, ssl_min_ver, ssl_max_ver, ssl_strict_sni - Result: All SSL advanced options lost (set to NULL) Impact - Backend Server: - User changes ssl_min_ver='TLSv1.2', then clicks Reject - Rollback skipped: ssl_sni, ssl_min_ver, ssl_max_ver, ssl_ciphers - Result: Security issue - TLS version constraints removed! Fix: - backend/utils/entity_snapshot.py line 285-286: Added 7 frontend SSL fields to rollback UPDATE - backend/utils/entity_snapshot.py line 458: Added 4 server SSL fields to rollback UPDATE PART 2: Improved ALPN Validation Error Message =============================================== Problem: User enters 'http/2' in ALPN field, gets generic error User feedback: Tried 'h2,http/1.1,http/2' → validation error not clear Root Cause: - ALPN standard uses 'h2' for HTTP/2 (not 'http/2') - Validator rejected 'http/2' but didn't explain the correct format Fix: - backend/models/frontend.py line 255-260: Detect common mistakes (http/2, http2, http-2) - Provide helpful error: 'For HTTP/2, use "h2" (not "http/2")' Before: "Invalid ALPN protocol: http/2. Valid protocols: h2, http/1.1, ..." After: "Invalid ALPN protocol: http/2. For HTTP/2, use 'h2' (not 'http/2'). Valid protocols: ..." Testing: 1. Rollback test: Edit frontend SSL, reject, verify SSL fields restored 2. Validation test: Enter 'http/2', verify friendly error message Files Changed: - backend/utils/entity_snapshot.py: _rollback_update() for frontend and server - backend/models/frontend.py: validate_alpn() with better error messages |
||
|
|
e8690245af |
fix(ssl): add change detection and validators for SSL advanced options
Critical additions: 1. Change Detection (backend/routers/config.py) - Added SSL advanced options to bulk import change detection logic - Detects changes in ssl_alpn, ssl_npn, ssl_ciphers, ssl_ciphersuites - Detects changes in ssl_min_ver, ssl_max_ver, ssl_strict_sni - Changes will now appear in version diff 2. Pydantic Validators (backend/models/frontend.py, backend/models/backend.py) - TLS version validator: Only allows valid versions (SSLv3, TLSv1.0-1.3) - ALPN protocol validator: Only allows h2, http/1.1, http/1.0, h2c, spdy/* - NPN protocol validator: Only allows http/1.1, http/1.0, spdy/* - Prevents invalid values from being saved to database - HAProxy validation will not fail due to invalid SSL options Impact: - Bulk import will correctly detect SSL option changes - Apply Management diff will show SSL changes - User cannot enter invalid TLS versions or protocols - Improved UX with early validation errors Previous fixes in this series: - Frontend GET API: Added SSL fields to response - Bulk Import UPDATE: Added SSL fields to UPDATE statement - Backend Server GET API: Added SSL fields to response Test: Bulk import with ALPN change → Should see change in version diff |
||
|
|
cb645b5ef9 |
CRITICAL FIX: Agent Offline Issue - Backend Tolerance for Legacy Agents
🔴 PRODUCTION CRITICAL FIX - Agent HTTP 422 Validation Error PROBLEM: - Production agents sending heartbeat with flat system_info fields - Backend Pydantic model was strict and rejecting unknown fields - Agents going offline with 'Validation error in request data' (HTTP 422) ROOT CAUSE: - Legacy agents embed system_info as flat key-value pairs in heartbeat JSON - Backend expected only defined fields, rejected extra fields - No backward compatibility for agent format variations SOLUTION - BACKEND ONLY (NO AGENT CHANGES): ✅ Added 'extra = "allow"' to AgentHeartbeat Pydantic Config ✅ Backend now accepts both formats: - Flat format: operating_system, kernel_version, etc. (legacy agents) - Nested format: system_info: {...} (future agents) ✅ Updated comments to clarify backward compatibility IMPACT: - ✅ ZERO CHANGES to production agent scripts - ✅ Existing agents will work immediately after backend deploy - ✅ Forward compatible with future agent upgrades - ✅ Tolerant to agent format variations SAFETY: - Minimal change (3 lines) - Pydantic still validates required fields - Extra fields ignored silently (no breaking changes) - Production agents continue without restart or upgrade DEPLOYMENT: 1. Deploy backend (this commit) 2. Agents come online automatically (no action needed) 3. Agent upgrades can happen later (when convenient) This fix ensures production stability without touching agent scripts. |
||
|
|
b212fb92bc |
Add SSL Advanced Options support (Backend) - Part 1
FEATURE: Complete SSL Advanced Options implementation for frontend and backend server SSL ✅ DATABASE: - Added SSL parameter columns to frontends table: * ssl_alpn, ssl_npn, ssl_ciphers, ssl_ciphersuites * ssl_min_ver, ssl_max_ver, ssl_strict_sni - Added SSL parameter columns to backend_servers table: * ssl_sni, ssl_min_ver, ssl_max_ver, ssl_ciphers - Migration functions: add_ssl_advanced_options_to_frontends() and add_ssl_advanced_options_to_servers() ✅ MODELS: - FrontendConfig: Added 7 new SSL fields for bind parameters - ServerConfig: Added 4 new SSL fields for server parameters - AgentHeartbeat: Added system_info field (fixes HTTP 422 validation error) ✅ BULK IMPORT PARSER: - Parse alpn, npn, ciphers, ciphersuites, ssl-min-ver, ssl-max-ver, strict-sni from bind lines - Parse sni, ssl-min-ver, ssl-max-ver, ciphers from server lines - Store parsed values in frontend/server objects - User-friendly warnings about imported SSL parameters ✅ CONFIG GENERATOR: - Generate bind lines with SSL advanced options: 'bind :443 ssl crt file.pem alpn h2,http/1.1 ciphers ...' - Generate server lines with SSL advanced options: 'server s1 addr:port ssl sni hostname ssl-min-ver TLSv1.2' - Support both NEW MODE (multiple certs) and OLD MODE (single cert) USER IMPACT: - Bulk import now correctly parses SSL configs with alpn/npn/ciphers - SSL parameters preserved during import (not lost anymore) - Agent heartbeat fixed (no more offline agents) - Ready for UI implementation (next commit) EXAMPLE USAGE: Frontend: bind 0.0.0.0:8443 ssl crt cert1.pem crt cert2.pem alpn h2,http/1.1 Server: server s1 10.1.1.1:443 ssl verify required sni backend.example.com ssl-min-ver TLSv1.2 NEXT: Frontend UI components for editing these SSL options |
||
|
|
5c1d2e11f9 |
Fix agent offline issue and bulk import SSL parsing
CRITICAL FIXES: 1. Agent Heartbeat Validation Error (HTTP 422) - Added missing 'system_info' field to AgentHeartbeat model - Agents were sending system_info but backend model didn't accept it - No agent script update needed - agents already send this field 2. Bulk Import SSL Parser Enhancement - Fixed parsing of multiple SSL certificates with alpn/npn parameters - Example: 'bind :443 ssl crt cert1.pem crt cert2.pem crt cert3.pem alpn h2,http/1.1' - Old parser stopped at first whitespace after crt path - New parser extracts all crt paths even with alpn/npn/ciphers after them - Added user-friendly warning for SSL parameters (alpn, npn, ciphers) that won't be imported TECHNICAL DETAILS: - backend/models/agent.py: Added system_info: Optional[Dict[str, Any]] - backend/utils/haproxy_config_parser.py: Enhanced SSL bind parsing logic * Parse bind line by splitting and iterating through parts * Extract all crt paths before hitting SSL parameters * Detect and warn about alpn, npn, ciphers, ciphersuites parameters * Inform user these advanced options should be configured manually USER IMPACT: - Agents will come online after backend deployment (no reinstall needed) - Bulk import will correctly parse configs with multiple SSL certs + alpn - Clear warnings shown in UI about SSL parameters not imported |
||
|
|
0fc18fde38 |
feat: Add HAProxy options support for backends and frontends
Implemented comprehensive HAProxy options field support for both backend and frontend entities to enable standard HAProxy directives like 'option http-keep-alive', 'option httplog', 'option forwardfor', etc. Changes: - Database: Added 'options' TEXT column to backends and frontends tables - Models: Added options field to BackendConfig, BackendConfigUpdate, and FrontendConfig - API Endpoints: Updated CREATE, UPDATE, and GET endpoints to handle options field * Backend: CREATE/UPDATE/GET with options support * Frontend: CREATE/UPDATE/GET with options support (fixed 5 SELECT queries) - Config Generator: Added options block generation for both backends and frontends - Bulk Import Parser: * Added options field to ParsedBackend and ParsedFrontend dataclasses * Implemented option directive parsing with validation * Added unknown option warnings * Fixed bulk parse response to include options field - Bulk Import Merge: Added options field comparison in UPDATE logic - UI Components: * BackendServers.js: Added options TextArea form field * FrontendManagement.js: Added options TextArea form field Features: - Multi-line options support (newline-separated format) - Option validation with known HAProxy options list - Backward compatible (NULL options for existing entities) - Bulk import support with merge strategy - Full CRUD support for both manual and bulk operations Technical Details: - Format: Newline-separated TEXT field for multiple options - Validation: Warns about unknown options but allows them - Config Generation: Each option written as separate directive - Agent: Standard HAProxy config validation applies Total: 10 files modified, ~195 lines added, 26 integration points verified |
||
|
|
281e23ea27 |
feat: Add SSL usage_type (Frontend/Server) with conditional private key requirement
This is a comprehensive update that adds SSL certificate differentiation for frontend (HAProxy bind) and server (backend verification) use cases. FEATURES: - SSL certificates can be marked as 'frontend' or 'server' usage type - Frontend SSL: Private key REQUIRED (for HAProxy bind ssl crt) - Server SSL: Private key OPTIONAL (CA cert only for backend verification) - UI dropdown for usage type selection - Dynamic form validation based on usage type - Filtering: Frontends see only Frontend SSL, Backends see only Server SSL DATABASE: - Added usage_type column to ssl_certificates (default: 'frontend') - Made private_key_content nullable for server SSL support - Migration automatically runs on pod restart BACKEND: - Pydantic v2 compatibility (@field_validator, @model_validator) - SSL router: usage_type filtering support - Agent endpoint: usage_type field included - Improved migration robustness with better error handling - Fixed duplicate ensure_agents_table() function - Fixed JSONB permissions insert with json.dumps() - Fixed ON CONFLICT constraints with explicit checks FRONTEND: - SSL Management: Usage Type dropdown with visual feedback - Frontend Management: Filters only Frontend SSL certificates - Backend Servers: Filters only Server SSL certificates - Dynamic private key validation (required for Frontend, optional for Server) - Improved form UX with color-coded hints AGENT SCRIPTS (Linux & macOS): - Support for Server SSL without private key - Conditional PEM file creation (cert+key vs cert-only) - usage_type awareness in SSL deployment - Backward compatible with existing Frontend SSL certificates DOCKER: - Increased npm timeout for slow networks (300s → 600s) - Increased fetch-retries (5 → 10) - Reduced maxsockets for stability (3 → 1) All changes are backward compatible. Existing SSL certificates default to 'frontend' type and continue working unchanged. Tested with: HAProxy 2.8+, PostgreSQL 15, React 18 |
||
|
|
a5e281b284 |
Feature: Backend Server SSL certificate support + Frontend SSL dropdown enhancement
✨ Backend Server SSL Certificate - Complete Implementation: 1. Model Update (backend/models/backend.py): - Added ssl_certificate_id field to ServerConfig model - Allows selecting SSL certificate from dropdown 2. API Endpoints (backend/routers/backend.py): - CREATE server: Added ssl_certificate_id to INSERT query - UPDATE server: Added ssl_certificate_id to allowed_fields - GET servers: Added ssl_certificate_id to SELECT queries (2 places) 3. Config Generation (backend/services/haproxy_config.py): - SSL certificate lookup by ID - Auto-generate ca-file path: /etc/ssl/haproxy/{cert_name}.pem - Added to server line in HAProxy config Example Generated Config: Before: server es1 10.0.0.1:9200 ssl verify required After: server es1 10.0.0.1:9200 ssl verify required ca-file /etc/ssl/haproxy/star-burgan-com-tr.pem ✨ Frontend SSL Dropdown Enhancement: - Added Global/Cluster-specific tags to Frontend SSL dropdown - Matches Backend Server SSL dropdown design - Shows: [🌍 Global] or [📍 Cluster] with color coding 🔧 Complete SSL Workflow: 1. User edits Backend Server 2. Enables SSL 3. Selects SSL certificate from dropdown 4. Saves → ssl_certificate_id stored in DB 5. Apply Changes → Config generated with ca-file path 6. Agent downloads SSL cert to /etc/ssl/haproxy/ 7. HAProxy uses ca-file for SSL verification ✅ Database Schema: backend_servers table now includes: - ssl_enabled (bool) - ssl_verify (str: none/required) - ssl_certificate_id (int, FK to ssl_certificates) ✅ HAProxy Config Format: server {name} {addr}:{port} ssl verify required ca-file {path} Impact: Backend Server SSL now fully functional with certificate management |
||
|
|
bfa0caa006 |
Feature: Full UI support for use_backend rules editing
✨ New Features: - Added use_backend_rules validator to Frontend model - UI now supports editing use_backend rules from Frontend Management page - Array to string conversion for use_backend rules in edit modal 🔧 Model Improvements: - Changed use_backend_rules field type from Optional[str] to Any (list support) - Added parse_use_backend_rules validator (same logic as ACL/redirect rules) - Handles 3 formats: Array, Textarea string (newline-separated), JSON string 💡 UI Improvements: - Frontend edit modal automatically converts use_backend array to multi-line text - Users can edit routing rules line by line in textarea - Format: 'use_backend BackendName if condition' ✅ Complete Workflow: 1. Bulk Import: Config parsed → ACL + use_backend stored as array 2. Frontend Edit: Arrays converted to multi-line string in textarea 3. User edits ACL/use_backend rules in UI 4. Save: Textarea string → validator → array → database 5. Config Generation: Array → HAProxy config format Example workflow: Parse: ['use_backend API if is_api'] → Edit UI: 'use_backend API if is_api' (textarea) → User edits: 'use_backend API_v2 if is_api_v2' → Save: ['use_backend API_v2 if is_api_v2'] → Generate: 'use_backend API_v2 if is_api_v2' (HAProxy config) |
||
|
|
6aae0f4309 | Initial commit |