Until now the only record of what happened was `user_activity_logs`, which
stores non-GET 2xx operations with no bodies. When something failed you could
see that a counter went up, never what was sent or what came back.
This adds one queryable timeline covering both directions:
- inbound: every API call, including GETs and including 4xx/5xx, with the
user, client IP, status, duration and — redacted, size-capped — the request
and response bodies.
- outbound: every HTTP call the backend makes, tagged with who it went to
(ACME/Let's Encrypt, Cloudflare, GoDaddy, HAProxy stats, agents, the ACME
diagnostics probe).
Outbound rows inherit the inbound request's id, so one operator action and the
CA/DNS calls it triggered read as a single trace: opening a failed "Request
Certificate" shows the exact POST /acme/new-order and the CA's 429 underneath.
Implementation notes:
- Capture is a pure-ASGI middleware that TEES the request and response streams
rather than draining them. `await request.body()` inside a BaseHTTPMiddleware
would consume the receive channel and break the raw-body agent heartbeat
handler. Registered last so it is outermost: it then sees the final
client-visible response and seeds correlation_id_context before the error
handler reads it.
- Rows are written by a batching background writer with a bounded queue, so the
request path never awaits the database and a saturated logger drops rows
visibly (surfaced on the page) instead of blocking. Redaction runs on the
writer, off the request coroutine.
- Secrets never land: headers are an allowlist with Authorization/Cookie kept
only as a presence marker; body keys and value shapes are redacted
(passwords, tokens, api_token, API keys, private-key PEMs, JWTs); the ACME
JWS request body is never stored, because a stored protected+signature pair
is a replayable credential — a summary is logged instead; DNS-provider errors
record only the exception type; the ACME HTTP-01 challenge endpoint is
excluded so key_authorization is never captured.
- Retention is operator-configurable in Settings -> Request Log: separate day
counts for successful and failed rows (7 / 30) plus a hard row cap (500k),
whichever is reached first. Pruned in batches under a Postgres advisory lock,
with the day counts bound as parameters, never interpolated.
- New permissions requestlog.read / requestlog.manage. super_admin and
security_admin get both, operator gets read, viewer gets neither.
Schema: one new table (request_logs) plus its settings seed, SCHEMA_VERSION
10 -> 11, auto-migrated. No existing table altered, no agent or rendered-config
change. Kill switches: REQUEST_LOG_ENABLED=false (middleware never registered)
or the `enabled` toggle in Settings.
Tests: 245 new (7 backend files + 1 frontend), full suite 1655 backend +
17 frontend passing.