Files
haproxy-openmanager/backend/tests/test_apply_service_extraction.py
taylanbakircioglu 02b1cb2bca feat: v1.5.0 — Site Wizard (Issue #14) + ACME Diagnostic Panel (Issue #13)
Closes #13, Closes #14.

This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0
shipped the ACME stability & enterprise audit (Issues #10/#11/#12).
v1.5.0 builds on that foundation with two co-equal headline features
plus a 22-round audit campaign hardening the prior configuration
surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0
lands in v1.5.2).

------------------------------------------------------------------
HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13)
------------------------------------------------------------------
A live pre-flight + post-failure diagnostic surface for every ACME
order, reachable from the ACME Automation page. The panel exists
to make ACME failures legible to operators who do NOT have shell
access to the API host.

Endpoints (`backend/routers/acme_diagnostics.py`):
  POST /api/letsencrypt/orders/{order_id}/diagnostics
       Run the full 5-check suite (DNS / port-80 / routing /
       account / agents) and humanize the order's `error_detail`
       (>=11 RFC-8555 problem types, backwards compatible with
       legacy plain-string failures).
  POST /api/letsencrypt/orders/{order_id}/diagnostics/
                                {check_id}/rerun
       Re-run a single check in place — used by the "Re-run"
       button on every row of the modal's pre-flight table.
  GET  /api/letsencrypt/orders/{order_id}/events
       Merged event timeline combining the typed
       `acme_order_events` rows with correlated
       `user_activity_logs` entries (resource_type =
       'letsencrypt_order' AND resource_id = order_id). The
       diagnostic modal auto-tails this timeline every 5 seconds
       while open.

Service-level checks (`backend/services/acme_diagnostics.py`):
  * DNS resolution via stdlib socket.gethostbyname_ex through
    run_in_executor (intentionally avoiding an aiodns runtime
    dep for v1.5.0).
  * Port-80 HEAD probe, target locked to the order's domains,
    success on HTTP 200 OR 404, warns on egress timeout
    (corp egress policies routinely blackhole outbound 80 —
    fail-hard would be too noisy).
  * SSRF guard: probe refuses non-public IPs and surfaces the
    skip in the diagnostic result; IPv4-mapped IPv6 normalisation
    closes the `::ffff:169.254.169.254` cloud-metadata vector.
  * HAProxy routing presence check: matches the order's
    cluster_ids to a port-80 HTTP frontend.
  * ACME account validity check against `letsencrypt_accounts`.
  * Agent presence check (>=1 active agent in target cluster).
  * Every sub-check wrapped in a wall-clock timeout to bound
    impact on the API event loop.

RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate
limit on both run and rerun, backed by the (user_id, action,
created_at DESC) composite index.

Frontend (`frontend/src/components/ACMEAutomation.js`):
  * "Diagnose" button on every order row + the existing
    "stuck order" warning row.
  * Modal with two tabs:
    - Pre-flight Checks (Antd Table with status pills + Re-run
      buttons + humanized error banner)
    - Event Log (Antd Timeline with auto-tail polling, scroll-
      to-bottom, pause-on-hover)
  * Correlation IDs surfaced in error banners and individual
    check fail details for backend-log lookup.

------------------------------------------------------------------
HEADLINE FEATURE B — Site Setup Wizard (Issue #14)
------------------------------------------------------------------
A single guided flow that creates a Backend + Servers + HTTP
Frontend (and optional HTTPS Frontend) in one atomic transaction.

Endpoints (`backend/routers/site_wizard.py`):
  POST /api/site-wizard/preview     — diff-preview the changeset
  POST /api/site-wizard/create      — atomic execute
  POST /api/site-wizard/reject      — clean rollback (including
                                       any wizard_staged ACME
                                       orders)
  GET  /api/site-wizard/drafts      — draft persistence
  PUT  /api/site-wizard/drafts/{id} — save/update
  DELETE /api/site-wizard/drafts/{id}

Feature surface:
  * One screen captures both backend (mode + servers) AND
    frontend (http + optional https + SSL mode) inputs.
  * SSL modes: ACME (new order, HTTP-01 only for v1.5.0),
    Upload (existing PEM), Existing (link to a stored cert),
    or None.
  * ACME-staged path: wizard_staged_until watermark on the
    `letsencrypt_orders` row defers finalisation until agent
    confirmation; per-mode reject cleanly cancels and rolls
    back the staged order.
  * Live diff preview against the cluster's current generated
    config (renderer-evolution noise stripped — track-sc<N>
    dedup, per-server cookie strip, defaults-cookie
    inheritance, listen-block flattening).
  * Draft persistence with PEM stripped at save time (private
    keys never round-trip through the drafts table).
  * Per-cluster multi-tenancy: drafts and wizard_staged orders
    are isolated to the creating user's cluster scope.

Frontend (`frontend/src/components/SiteWizard.js`):
  * 4-step Antd Steps flow: Backend → Frontend → SSL → Review.
  * Render the live diff preview inline before commit.
  * Antd Form-level validation mirrors backend Pydantic
    validators (numeric bounds, HAProxy reserved keywords, ALPN
    consistency, IPv6 scope-id, domain regex, server name
    dedup).

------------------------------------------------------------------
AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1#82)
------------------------------------------------------------------
v1.5.0 includes 22 adversarial review passes. Each round produced
its own commit set in the corporate development line; this squash
collapses those into the v1.5.0 release artefact. Highlights:

  Round 1-4   Site Wizard core: dry-run parity, single-line
              value injection guard, ACL -f pattern-file block,
              SSL parity, timeout regex, form-state pin.
  Round 5-7   defaults-cookie inheritance, server-named-cookie
              guard, fe/be mode mismatch, duplicate server
              names, health_check_uri + server_address
              validators.
  Round 8-10  cookie_name / cookie_options newline-injection
              guard, dry-run parity (round 9), TCP-mode HTTP-only
              feature blockers.
  Round 11    SSL name path traversal + health-check >= 1.
  Round 12-13 SSL & ACME deep dive (Bulgu #23-#32).
  Round 14    single-line value injection (Bulgu #33).
  Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy
              reserved keywords, ALPN/TLS consistency,
              all-backup, multi-domain & multi-user enterprise
              edges, drain/HSTS/post-completion (Bulgu
              #34-#53).
  Round 18-21 concurrency, agent state, TCP-mode HTTP-only,
              list size caps, IPv6 scope-id, preview account
              validation, TCP backend + balance uri reject
              (Bulgu #54-#61).
  Round 22    FE error visibility + 3x stale-data lockouts,
              referential integrity + cascade safety,
              authentication & authorization, multi-cluster
              isolation, apply_pending_changes concurrency,
              script injection + bulk import multi-tenancy,
              prefix-stripped signature comparison
              (Bulgu #62-#82).

------------------------------------------------------------------
NO CORPORATE-SPECIFIC ARTIFACTS
------------------------------------------------------------------
This squash deliberately sanitises corporate hostnames, container
registry references, and TLS secret names into generic
placeholders (`your-registry.example.com/your-org`,
`haproxy-openmanager*.example.com`, `wildcard-tls`,
`taylanbakircioglu/haproxy-openmanager-*`) so the public artefact
contains no internal infrastructure detail. Pilot / development
history that retained those values stays in the corporate fork
and is NOT part of this commit.
2026-05-14 00:04:19 +03:00

184 lines
6.0 KiB
Python

"""
v1.5.0 service extraction parity — apply_service.
Asserts:
- _resolve_user_id falls back to the first active admin user using the
CORRECT schema columns (is_admin, is_active) — NOT the non-existent
is_super_admin (M46/R65).
- _mint_internal_jwt passes user_id under both `sub` and `user_id` claims so
it round-trips through get_current_user_from_token.
- apply_cluster_pending delegates to routers.cluster.apply_pending_changes
with a Bearer header (i.e. NEVER duplicates the ~800-line apply pipeline).
"""
from unittest.mock import AsyncMock, patch
import pytest
from services.apply_service import (
_mint_internal_jwt,
_resolve_user_id,
apply_cluster_pending,
)
# ----------------------------------------------------------------------------
# _resolve_user_id
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_resolve_user_id_passthrough_when_user_still_valid():
"""Bulgu #27: a still-active user_id is returned as-is after re-validation."""
fake_conn = AsyncMock()
# First fetchval validates the requested user, returns its id.
fake_conn.fetchval.return_value = 42
with patch(
"services.apply_service.get_database_connection",
AsyncMock(return_value=fake_conn),
), patch(
"services.apply_service.close_database_connection",
AsyncMock(),
):
out = await _resolve_user_id(42)
assert out == 42
@pytest.mark.asyncio
async def test_resolve_user_id_falls_back_when_requested_user_inactive():
"""Bulgu #27: deleted/deactivated created_by must NOT mint a ghost JWT —
fall back to the admin user instead."""
fake_conn = AsyncMock()
# Validation lookup returns None (user gone/inactive); admin fallback returns 1.
fake_conn.fetchval.side_effect = [None, 1]
with patch(
"services.apply_service.get_database_connection",
AsyncMock(return_value=fake_conn),
), patch(
"services.apply_service.close_database_connection",
AsyncMock(),
):
out = await _resolve_user_id(99)
assert out == 1
assert fake_conn.fetchval.await_count == 2
@pytest.mark.asyncio
async def test_resolve_user_id_falls_back_to_active_admin():
fake_conn = AsyncMock()
fake_conn.fetchval.return_value = 1 # admin id
with patch(
"services.apply_service.get_database_connection",
AsyncMock(return_value=fake_conn),
), patch(
"services.apply_service.close_database_connection",
AsyncMock(),
):
out = await _resolve_user_id(None)
assert out == 1
sql, *_ = fake_conn.fetchval.call_args.args
# Schema accuracy: is_admin AND is_active (NOT is_super_admin)
assert "is_admin" in sql
assert "is_active" in sql
assert "is_super_admin" not in sql
@pytest.mark.asyncio
async def test_resolve_user_id_returns_none_when_no_admin():
fake_conn = AsyncMock()
fake_conn.fetchval.return_value = None
with patch(
"services.apply_service.get_database_connection",
AsyncMock(return_value=fake_conn),
), patch(
"services.apply_service.close_database_connection",
AsyncMock(),
):
out = await _resolve_user_id(None)
assert out is None
# ----------------------------------------------------------------------------
# _mint_internal_jwt
# ----------------------------------------------------------------------------
def test_mint_internal_jwt_includes_both_claim_shapes():
"""sub + user_id ensures compat with get_current_user_from_token."""
captured = {}
def fake_create(payload, expires_delta=None):
captured.update(payload)
captured["__expires"] = expires_delta
return "fake-jwt-token"
with patch("services.apply_service.create_access_token", side_effect=fake_create):
token = _mint_internal_jwt(99)
assert token == "fake-jwt-token"
assert captured["sub"] == "99"
assert captured["user_id"] == 99
assert captured["__expires"] is not None
# ----------------------------------------------------------------------------
# apply_cluster_pending
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_apply_cluster_pending_raises_when_no_admin_available():
with patch(
"services.apply_service._resolve_user_id",
AsyncMock(return_value=None),
):
with pytest.raises(RuntimeError, match="no admin user available"):
await apply_cluster_pending(cluster_id=1)
@pytest.mark.asyncio
async def test_apply_cluster_pending_delegates_to_router_with_bearer():
"""Delegate, don't duplicate. We pass through cluster_id, apply_request,
and a Bearer auth header.
"""
apply_mock = AsyncMock(return_value={"applied_count": 3})
# Patch the local-imported symbol via routers.cluster
with patch(
"services.apply_service._resolve_user_id",
AsyncMock(return_value=42),
), patch(
"services.apply_service._mint_internal_jwt",
return_value="fake.jwt.token",
), patch("routers.cluster.apply_pending_changes", apply_mock):
out = await apply_cluster_pending(
cluster_id=7,
apply_request={"force": True},
)
assert out == {"applied_count": 3}
kwargs = apply_mock.call_args.kwargs
assert kwargs["cluster_id"] == 7
assert kwargs["apply_request"] == {"force": True}
assert kwargs["authorization"].startswith("Bearer ")
assert "fake.jwt.token" in kwargs["authorization"]
@pytest.mark.asyncio
async def test_apply_cluster_pending_default_empty_apply_request():
apply_mock = AsyncMock(return_value={})
with patch(
"services.apply_service._resolve_user_id",
AsyncMock(return_value=1),
), patch(
"services.apply_service._mint_internal_jwt",
return_value="t",
), patch("routers.cluster.apply_pending_changes", apply_mock):
await apply_cluster_pending(cluster_id=1)
kwargs = apply_mock.call_args.kwargs
assert kwargs["apply_request"] == {}