mirror of
https://github.com/taylanbakircioglu/haproxy-openmanager.git
synced 2026-09-16 15:45:11 +00:00
02b1cb2bca
Closes #13, Closes #14. This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0 shipped the ACME stability & enterprise audit (Issues #10/#11/#12). v1.5.0 builds on that foundation with two co-equal headline features plus a 22-round audit campaign hardening the prior configuration surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0 lands in v1.5.2). ------------------------------------------------------------------ HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13) ------------------------------------------------------------------ A live pre-flight + post-failure diagnostic surface for every ACME order, reachable from the ACME Automation page. The panel exists to make ACME failures legible to operators who do NOT have shell access to the API host. Endpoints (`backend/routers/acme_diagnostics.py`): POST /api/letsencrypt/orders/{order_id}/diagnostics Run the full 5-check suite (DNS / port-80 / routing / account / agents) and humanize the order's `error_detail` (>=11 RFC-8555 problem types, backwards compatible with legacy plain-string failures). POST /api/letsencrypt/orders/{order_id}/diagnostics/ {check_id}/rerun Re-run a single check in place — used by the "Re-run" button on every row of the modal's pre-flight table. GET /api/letsencrypt/orders/{order_id}/events Merged event timeline combining the typed `acme_order_events` rows with correlated `user_activity_logs` entries (resource_type = 'letsencrypt_order' AND resource_id = order_id). The diagnostic modal auto-tails this timeline every 5 seconds while open. Service-level checks (`backend/services/acme_diagnostics.py`): * DNS resolution via stdlib socket.gethostbyname_ex through run_in_executor (intentionally avoiding an aiodns runtime dep for v1.5.0). * Port-80 HEAD probe, target locked to the order's domains, success on HTTP 200 OR 404, warns on egress timeout (corp egress policies routinely blackhole outbound 80 — fail-hard would be too noisy). * SSRF guard: probe refuses non-public IPs and surfaces the skip in the diagnostic result; IPv4-mapped IPv6 normalisation closes the `::ffff:169.254.169.254` cloud-metadata vector. * HAProxy routing presence check: matches the order's cluster_ids to a port-80 HTTP frontend. * ACME account validity check against `letsencrypt_accounts`. * Agent presence check (>=1 active agent in target cluster). * Every sub-check wrapped in a wall-clock timeout to bound impact on the API event loop. RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate limit on both run and rerun, backed by the (user_id, action, created_at DESC) composite index. Frontend (`frontend/src/components/ACMEAutomation.js`): * "Diagnose" button on every order row + the existing "stuck order" warning row. * Modal with two tabs: - Pre-flight Checks (Antd Table with status pills + Re-run buttons + humanized error banner) - Event Log (Antd Timeline with auto-tail polling, scroll- to-bottom, pause-on-hover) * Correlation IDs surfaced in error banners and individual check fail details for backend-log lookup. ------------------------------------------------------------------ HEADLINE FEATURE B — Site Setup Wizard (Issue #14) ------------------------------------------------------------------ A single guided flow that creates a Backend + Servers + HTTP Frontend (and optional HTTPS Frontend) in one atomic transaction. Endpoints (`backend/routers/site_wizard.py`): POST /api/site-wizard/preview — diff-preview the changeset POST /api/site-wizard/create — atomic execute POST /api/site-wizard/reject — clean rollback (including any wizard_staged ACME orders) GET /api/site-wizard/drafts — draft persistence PUT /api/site-wizard/drafts/{id} — save/update DELETE /api/site-wizard/drafts/{id} Feature surface: * One screen captures both backend (mode + servers) AND frontend (http + optional https + SSL mode) inputs. * SSL modes: ACME (new order, HTTP-01 only for v1.5.0), Upload (existing PEM), Existing (link to a stored cert), or None. * ACME-staged path: wizard_staged_until watermark on the `letsencrypt_orders` row defers finalisation until agent confirmation; per-mode reject cleanly cancels and rolls back the staged order. * Live diff preview against the cluster's current generated config (renderer-evolution noise stripped — track-sc<N> dedup, per-server cookie strip, defaults-cookie inheritance, listen-block flattening). * Draft persistence with PEM stripped at save time (private keys never round-trip through the drafts table). * Per-cluster multi-tenancy: drafts and wizard_staged orders are isolated to the creating user's cluster scope. Frontend (`frontend/src/components/SiteWizard.js`): * 4-step Antd Steps flow: Backend → Frontend → SSL → Review. * Render the live diff preview inline before commit. * Antd Form-level validation mirrors backend Pydantic validators (numeric bounds, HAProxy reserved keywords, ALPN consistency, IPv6 scope-id, domain regex, server name dedup). ------------------------------------------------------------------ AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1 → #82) ------------------------------------------------------------------ v1.5.0 includes 22 adversarial review passes. Each round produced its own commit set in the corporate development line; this squash collapses those into the v1.5.0 release artefact. Highlights: Round 1-4 Site Wizard core: dry-run parity, single-line value injection guard, ACL -f pattern-file block, SSL parity, timeout regex, form-state pin. Round 5-7 defaults-cookie inheritance, server-named-cookie guard, fe/be mode mismatch, duplicate server names, health_check_uri + server_address validators. Round 8-10 cookie_name / cookie_options newline-injection guard, dry-run parity (round 9), TCP-mode HTTP-only feature blockers. Round 11 SSL name path traversal + health-check >= 1. Round 12-13 SSL & ACME deep dive (Bulgu #23-#32). Round 14 single-line value injection (Bulgu #33). Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy reserved keywords, ALPN/TLS consistency, all-backup, multi-domain & multi-user enterprise edges, drain/HSTS/post-completion (Bulgu #34-#53). Round 18-21 concurrency, agent state, TCP-mode HTTP-only, list size caps, IPv6 scope-id, preview account validation, TCP backend + balance uri reject (Bulgu #54-#61). Round 22 FE error visibility + 3x stale-data lockouts, referential integrity + cascade safety, authentication & authorization, multi-cluster isolation, apply_pending_changes concurrency, script injection + bulk import multi-tenancy, prefix-stripped signature comparison (Bulgu #62-#82). ------------------------------------------------------------------ NO CORPORATE-SPECIFIC ARTIFACTS ------------------------------------------------------------------ This squash deliberately sanitises corporate hostnames, container registry references, and TLS secret names into generic placeholders (`your-registry.example.com/your-org`, `haproxy-openmanager*.example.com`, `wildcard-tls`, `taylanbakircioglu/haproxy-openmanager-*`) so the public artefact contains no internal infrastructure detail. Pilot / development history that retained those values stays in the corporate fork and is NOT part of this commit.
198 lines
8.4 KiB
Python
198 lines
8.4 KiB
Python
from pydantic import BaseModel, validator
|
|
from typing import Literal, Optional, List
|
|
|
|
class ServerConfig(BaseModel):
|
|
server_name: str
|
|
server_address: str
|
|
server_port: int
|
|
weight: int = 100
|
|
max_connections: Optional[int] = None
|
|
check_enabled: bool = True
|
|
check_port: Optional[int] = None
|
|
backup_server: bool = False
|
|
ssl_enabled: bool = False
|
|
# PR-2 (R11.B): tighten to strict Literal aligned with HAProxy's
|
|
# `server ... ssl verify <none|required>` semantics. Backend-side
|
|
# `verify optional` is NOT supported by HAProxy (only frontend
|
|
# bind-side accepts it) — pre-PR-2 the field accepted arbitrary
|
|
# strings (`'optional'`, `'true'`, etc.) and the generator
|
|
# rendered them verbatim, producing parser errors. Empty strings
|
|
# from the React form are coerced to None by
|
|
# `coerce_ssl_verify_empty_to_none` below.
|
|
ssl_verify: Optional[Literal["none", "required"]] = None
|
|
ssl_certificate_id: Optional[int] = None # SSL certificate for backend server
|
|
|
|
# SSL Advanced Options (server SSL parameters)
|
|
ssl_sni: Optional[str] = None # SNI hostname for backend SSL connections
|
|
ssl_min_ver: Optional[str] = None # Minimum TLS version (TLSv1.2, TLSv1.3)
|
|
ssl_max_ver: Optional[str] = None # Maximum TLS version
|
|
ssl_ciphers: Optional[str] = None # Cipher suite list for backend connections
|
|
|
|
cookie_value: Optional[str] = None
|
|
inter: Optional[int] = None
|
|
fall: Optional[int] = None
|
|
rise: Optional[int] = None
|
|
is_active: bool = True
|
|
|
|
@validator('ssl_min_ver', 'ssl_max_ver')
|
|
def validate_tls_version(cls, v):
|
|
if v is not None:
|
|
valid_versions = ['SSLv3', 'TLSv1.0', 'TLSv1.1', 'TLSv1.2', 'TLSv1.3']
|
|
if v not in valid_versions:
|
|
raise ValueError(f'Invalid TLS version: {v}. Must be one of: {", ".join(valid_versions)}')
|
|
return v
|
|
|
|
@validator('ssl_verify', pre=True)
|
|
def coerce_ssl_verify_empty_to_none(cls, v):
|
|
"""PR-2 (R11.B): React form Select widgets clear to '' (empty
|
|
string) but the strict Literal would reject that. Coerce the
|
|
empty string and the legacy sentinels written by older
|
|
clients into None. Note: server-side mTLS only accepts
|
|
``none`` or ``required`` (HAProxy's `server ... verify`
|
|
keyword has no `optional` mode); a legacy `'optional'`
|
|
value is also coerced to None to fail-safe rather than
|
|
rendering an invalid directive.
|
|
"""
|
|
if v is None:
|
|
return None
|
|
if isinstance(v, str):
|
|
stripped = v.strip().lower()
|
|
if stripped in ("", "[]", "{}", "null"):
|
|
return None
|
|
if stripped == "none":
|
|
return "none"
|
|
if stripped == "required":
|
|
return "required"
|
|
if stripped == "optional":
|
|
# Server-side `verify optional` is invalid HAProxy.
|
|
# Coerce to None so the generator simply omits the
|
|
# directive instead of producing a parser-fatal line.
|
|
return None
|
|
return v
|
|
|
|
class BackendConfig(BaseModel):
|
|
name: str
|
|
cluster_id: int
|
|
balance_method: str = 'roundrobin'
|
|
mode: str = 'http'
|
|
|
|
@validator('name')
|
|
def reject_system_prefix(cls, v):
|
|
# R18 audit fix (round 3 #5): manual backend create previously
|
|
# accepted leading-underscore names (e.g. `_my_backend`). The
|
|
# agent's `_should_sync_backend` filter then dropped any such
|
|
# row from the agent->backend reverse-sync, producing silent
|
|
# control-plane drift between DB and on-disk haproxy.cfg. The
|
|
# wizard's BackendStep already rejected this; align the manual
|
|
# path so the constraint is uniform across entry points.
|
|
if isinstance(v, str) and v.startswith('_'):
|
|
raise ValueError(
|
|
"Backend name must not start with '_' (reserved for "
|
|
"system-managed entities such as the ACME challenge "
|
|
"backend)."
|
|
)
|
|
return v
|
|
health_check_uri: Optional[str] = None
|
|
health_check_interval: Optional[int] = 2000
|
|
health_check_expected_status: Optional[int] = 200
|
|
fullconn: Optional[int] = None
|
|
cookie_name: Optional[str] = None
|
|
cookie_options: Optional[str] = None
|
|
default_server_inter: Optional[int] = None
|
|
default_server_fall: Optional[int] = None
|
|
default_server_rise: Optional[int] = None
|
|
request_headers: Optional[str] = None
|
|
response_headers: Optional[str] = None
|
|
timeout_connect: Optional[int] = 10000
|
|
timeout_server: Optional[int] = 60000
|
|
timeout_queue: Optional[int] = 60000
|
|
options: Optional[str] = None
|
|
servers: List[ServerConfig] = []
|
|
|
|
# Bulgu #68 (round-22 audit) — wizard's BackendStep declares
|
|
# `fullconn: Optional[int] = Field(default=None, ge=0, ...)`
|
|
# (0 == HAProxy "fullconn disabled" sentinel). The manual
|
|
# BackendConfig pre-fix used a single `check_positive` that
|
|
# rejected `<= 0` for ALL of `health_check_interval`,
|
|
# `timeout_connect`, `timeout_server`, `timeout_queue`, AND
|
|
# `fullconn` — so a wizard-created backend with
|
|
# `fullconn=0` (or one persisted before fullconn was
|
|
# introduced and now defaults to 0) would 422 on every PUT,
|
|
# even when the operator was only changing the balance
|
|
# method or adding a server. Same Bulgu #62 "wizard
|
|
# accepted / manual rejects" lockout pattern.
|
|
#
|
|
# Split into two validators: the four timeout/interval fields
|
|
# keep `> 0` (HAProxy parser hard requirement), while
|
|
# `fullconn` switches to `>= 0` mirroring the wizard.
|
|
@validator('health_check_interval', 'timeout_connect', 'timeout_server', 'timeout_queue')
|
|
def check_positive(cls, v):
|
|
if v is not None and v <= 0:
|
|
raise ValueError('Timeout, interval, and connection values must be positive')
|
|
return v
|
|
|
|
@validator('fullconn')
|
|
def check_fullconn_non_negative(cls, v):
|
|
if v is not None and v < 0:
|
|
raise ValueError('fullconn must be >= 0 (0 disables the directive)')
|
|
return v
|
|
|
|
@validator('health_check_expected_status')
|
|
def check_http_status(cls, v):
|
|
if v is not None and (v < 100 or v > 599):
|
|
raise ValueError('HTTP status code must be between 100 and 599')
|
|
return v
|
|
|
|
class BackendConfigUpdate(BaseModel):
|
|
name: Optional[str] = None
|
|
balance_method: Optional[str] = None
|
|
mode: Optional[str] = None
|
|
|
|
@validator('name')
|
|
def reject_system_prefix_update(cls, v):
|
|
# R18 audit fix (round 3 #5): rename guard. Without it an
|
|
# operator could `PUT` a backend's name to `_anything`, which
|
|
# would then be filtered out by the agent reverse-sync.
|
|
if v is not None and isinstance(v, str) and v.startswith('_'):
|
|
raise ValueError(
|
|
"Backend name must not start with '_' (reserved for "
|
|
"system-managed entities)."
|
|
)
|
|
return v
|
|
health_check_uri: Optional[str] = None
|
|
health_check_interval: Optional[int] = None
|
|
health_check_expected_status: Optional[int] = None
|
|
fullconn: Optional[int] = None
|
|
cookie_name: Optional[str] = None
|
|
cookie_options: Optional[str] = None
|
|
default_server_inter: Optional[int] = None
|
|
default_server_fall: Optional[int] = None
|
|
default_server_rise: Optional[int] = None
|
|
request_headers: Optional[str] = None
|
|
response_headers: Optional[str] = None
|
|
timeout_connect: Optional[int] = None
|
|
timeout_server: Optional[int] = None
|
|
timeout_queue: Optional[int] = None
|
|
options: Optional[str] = None
|
|
servers: Optional[List[ServerConfig]] = None
|
|
|
|
# Bulgu #68 (round-22 audit) — same alignment as BackendConfig
|
|
# above. Update path is where the wizard-created `fullconn=0`
|
|
# row most often blows up.
|
|
@validator('health_check_interval', 'timeout_connect', 'timeout_server', 'timeout_queue')
|
|
def check_positive_update(cls, v):
|
|
if v is not None and v <= 0:
|
|
raise ValueError('Timeout, interval, and connection values must be positive')
|
|
return v
|
|
|
|
@validator('fullconn')
|
|
def check_fullconn_non_negative_update(cls, v):
|
|
if v is not None and v < 0:
|
|
raise ValueError('fullconn must be >= 0 (0 disables the directive)')
|
|
return v
|
|
|
|
@validator('health_check_expected_status')
|
|
def check_http_status_update(cls, v):
|
|
if v is not None and (v < 100 or v > 599):
|
|
raise ValueError('HTTP status code must be between 100 and 599')
|
|
return v |